Interactive media data generation method and apparatus, electronic device, and storage medium
By loading reference labels and basic prompt words in the multimodal input box to generate prompt word text, the problems of image quality instability and content deviation in the image generation model are solved, and high-quality and high-accuracy image generation is achieved.
Patent Information
- Application Number
- PCT/CN2025/071546
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-17
AI Technical Summary
In the prior art, the image quality generated by the image generation model is unstable, and content deviations are prone to occur, making it difficult to meet the user's image generation needs.
By loading reference tags in the multimodal input box, combining the reference methods and basic prompt words of the media file, the prompt word text is generated, and the target media data is generated to avoid image quality instability and content deviation caused by inaccurate user input.
It improves the quality of image generation and content accuracy, meets the user's image generation needs, and improves interaction efficiency.
Smart Images

Figure CN2025071546_17072025_PF_FP_ABST
Abstract
Description
Interactive media data generation method, device, electronic device and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202410046071.9 filed on January 11, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The embodiments of the present disclosure relate to a method, device, electronic device, and storage medium for generating interactive media data. Background Art
[0003] Currently, AI-based content generation (AIGC) technology can generate corresponding images through user-input prompts and image generation models, thereby greatly improving the efficiency and quality of image generation and lowering the threshold for users to create content.
[0004] However, the quality of images generated by the image generation model will be directly affected by the prompt words entered by the user, which may cause the generated images to have unstable quality and large deviations in image content, making it difficult to meet the user's image generation needs. Summary of the Invention
[0005] The embodiments of the present disclosure provide a method, device, electronic device, and storage medium for generating interactive media data to overcome the problems of unstable quality and large deviation in image content of generated images.
[0006] In a first aspect, an embodiment of the present disclosure provides a method for generating interactive media data, comprising:
[0007] In response to a first operation, a reference tag is loaded in a multimodal input box, wherein the first operation is used to select at least one media file, and the reference tag is used to represent a reference method for the media file in a process of generating target media data with reference to the media file; a prompt word text is generated based on the reference tag and at least one basic prompt word in the multimodal input box; and the target media data is generated based on the prompt word text.
[0008] In a second aspect, an embodiment of the present disclosure provides an interactive media data generating device, including:
[0009] an interaction module configured to load a reference tag in the multimodal input box in response to a first operation, wherein the first operation is used to select at least one media file, and the reference tag is used to represent a reference method for the media file in a process of generating target media data by referring to the media file;
[0010] a processing module, configured to generate a prompt word text according to the reference label and at least one basic prompt word in the multimodal input box;
[0011] A generating module is used to generate the target media data based on the prompt word text.
[0012] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0013] The memory stores computer-executable instructions;
[0014] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the interactive media data generating method as described in the first aspect and various possible designs of the first aspect.
[0015] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the interactive media data generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0016] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the interactive media data generation method described in the first aspect and various possible designs of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] FIG1 is a diagram illustrating an application scenario of the method for generating interactive media data according to an embodiment of the present disclosure;
[0019] FIG2 is a flow chart of a method for generating interactive media data according to an embodiment of the present disclosure;
[0020] FIG3 is a schematic diagram of a process of loading a reference tag in a multimodal input box according to an embodiment of the present disclosure;
[0021] FIG4 is a schematic diagram of another reference tag provided by an embodiment of the present disclosure:
[0022] FIG5 is a schematic diagram of a process of interaction based on a tab window provided by an embodiment of the present disclosure;
[0023] FIG6 is a flowchart of a specific implementation of step S101 in the embodiment shown in FIG2 ;
[0024] FIG7 is a schematic diagram of a process of loading a reference tag by dragging an operation according to an embodiment of the present disclosure;
[0025] FIG8 is a second flow chart of the method for generating interactive media data according to an embodiment of the present disclosure;
[0026] FIG9 is a flowchart of a specific implementation of step S201 in the embodiment shown in FIG8 ;
[0027] FIG10 is a schematic diagram of a process for setting a target position according to an embodiment of the present disclosure;
[0028] FIG11 is a flowchart of a specific implementation of step S202 in the embodiment shown in FIG8 ;
[0029] FIG12 is a structural block diagram of an interactive media data generating apparatus provided by an embodiment of the present disclosure;
[0030] FIG13 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure; and
[0031] FIG14 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0033] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0034] The following explains the application scenarios of the embodiments of the present disclosure:
[0035] FIG1 is a diagram of an application scenario of the interactive media data generation method provided in an embodiment of the present disclosure. The interactive media data generation method provided in an embodiment of the present disclosure can be applied to an application with a media generation function, and more specifically, can be applied to an application scenario in which media data such as images and videos are generated based on user descriptions. The execution subject of this embodiment can be a terminal device that runs the above-mentioned application with a media generation function, or a server that deploys the service end corresponding to the above-mentioned application, or other electronic devices that perform similar functions. Referring to FIG1 , taking the terminal device as the execution subject of the method of this embodiment and the application scenario of generating images based on user descriptions as an example, after running the above-mentioned application, the terminal device displays an image generation page. In the image generation page, at least an input control for receiving prompt words input by the user is provided. The user guides the image generation model to generate the corresponding image by inputting prompt words for limiting the image content into the input control. Specifically, for example, as shown in the figure, after the user enters the prompt word "dancing, full-body photo, soft light effect" in the input control and clicks the "Generate Now" trigger control, the terminal device converts the prompt word into the corresponding prompt word text adapted to the image generation model based on the above prompt word, and inputs the image generation model to generate an image that meets the above "dancing", "full-body photo", and "soft light effect", and displays it in the preview window of the image generation page, completing the image generation process.
[0036] In the related art, when generating AI images based on the above-mentioned "text-to-image" image generation model, the image quality of the generated AI image is directly affected by the prompt word entered by the user. If the prompt word entered by the user is inaccurate or unreasonable, the image generated by the image generation model will deviate from the user's needs. On the other hand, for some complex image content requirements, it is difficult for users to achieve accurate descriptions through prompt words, which leads to the need for users to repeatedly try and optimize the prompt words to improve the quality of the generated image. Therefore, the solutions in the related art lead to problems such as unstable quality and large deviations in image content in the generated images, making it difficult to meet the user's image generation needs.
[0037] The embodiments of the present disclosure provide a method for generating interactive media data to solve the above problems.
[0038] Referring to FIG2 , FIG2 is a flow chart of a method for generating interactive media data according to an embodiment of the present disclosure. The method of this embodiment can be applied in a terminal device. The method for generating interactive media data includes:
[0039] Step S101: In response to a first operation, a reference tag is loaded in a multimodal input box, wherein the first operation is used to select at least one media file, and the reference tag is used to represent a reference method for the media file in the process of generating target media data by referring to the media file.
[0040] For example, referring to the application scenario diagram shown in FIG1 , the application program run by the terminal device can be an application client or a browser (web terminal). No specific restrictions are made here. A multimodal input box is provided in the image generation interface used by the application to generate AI images, wherein the multimodal input box is an input control that can load and display non-text content. The user applies a first operation to the image generation interface to load a reference tag in the multimodal input box, wherein the first operation is at least used to select a media file, and the media file is used as reference information to generate target media data. The target media data includes, for example, images, videos, music, or media (multimedia data) superimposed by two or more media data generated based on AIGC technology. In subsequent embodiments, the case where the target media data is an image (i.e., an application scenario for generating an image) is used as an example for introduction. Other cases are similar and will not be described one by one. Furthermore, the reference tag is a non-text object that maps the above-mentioned media file. Furthermore, the tag content of the reference tag is used to describe the role when generating target media data based on the media file. For example, the tag content of the reference tag includes "reference target object", "reference character distribution", "reference subject", "reference edge information", etc., which represents the role when generating target media data based on the media file.
[0041] For example, FIG3 is a schematic diagram of a process for loading a reference tag in a multimodal input box provided by an embodiment of the present disclosure. The above steps are further explained in conjunction with FIG3. As shown in FIG3, first, a multimodal input box and a media import control are provided in the image generation interface. By clicking the media import control, a folder interface is displayed above the image generation interface. One or more media files are displayed in the folder interface, such as media file #1, media file #2, media file #3, etc. shown in the figure. For example, the media file can be a picture, video file, audio file, etc., and no specific limitation is given. The media file can be a file stored locally on the terminal device or a file stored on the server or other storage medium. Afterwards, the user selects a media file (such as media file #3) from the folder interface as the target media file by applying a first operation (such as single click or double click). The terminal device processes the target media file, generates a corresponding reference tag, and displays it in the multimodal input box.
[0042] Figure 4 is a schematic diagram of another reference tag provided in an embodiment of the present disclosure. As shown in Figure 4, for example, a media file is, for example, a video or a picture; the reference tag includes a tag text and a tag picture, the tag picture includes a thumbnail of the media file, and the tag text is used to describe a target reference method for the media file, such as the "reference character distribution" shown in the figure, that is, referring to the distribution of characters in the above picture to generate target media data.
[0043] Furthermore, after the reference tag is loaded in the multimodal input box, the terminal device can perform subsequent processing on the reference tag in response to further user operations to display the media file corresponding to the reference tag. Exemplarily, the steps of displaying the media file by operating the reference tag include:
[0044] Step S101A: In response to the third operation on the reference tag, displaying a tag window corresponding to the reference tag.
[0045] Step S101B: Display the media file in the tab window.
[0046] Figure 5 is a schematic diagram of a process of interaction based on a tag window provided by an embodiment of the present disclosure. As shown in Figure 5, after loading a reference tag in a multimodal input box, the user, for example, clicks (the third operation) the reference tag (based on the scheme in the previous embodiment, such as clicking the tag text or the tag image), and a tag window is displayed above the reference tag. The tag window is used to display the related information of the reference tag in more detail, as shown in the figure. Optionally, the tag window displays the media file or the thumbnail of the media file corresponding to the reference tag, that is, the media file is displayed in the tag window.
[0047] Optionally, a first operation control (shown as control P1 in the figure) is provided in the tag window. After being triggered, the first operation control is used to replace the first media file corresponding to the reference tag with a second media file. Specifically, for example, when the terminal device responds to the user operation and triggers the first operation control, a folder interface can be displayed, and the second media file can be selected based on the user operation in the folder interface, thereby achieving the replacement of the media file corresponding to the reference tag. The specific operation method has been introduced in the previous embodiment and will not be repeated here.
[0048] Optionally, a second operation control (shown as control P2 in the figure) is provided in the label window. After being triggered, the second operation control is used to delete the reference label, that is, to restore the content in the multimodal input box to the state before executing step S101 of this embodiment. The specific implementation process steps are omitted.
[0049] Optionally, a third control (shown as control P3 in the figure) is provided within the label window. When triggered, the third control is used to modify the reference method for the media file. Specifically, the third control is used to configure and modify the target reference method for the media file, that is, the content corresponding to the label text. For example, by clicking the reference item control, you can change "reference character appearance" to "reference character layout" or set the target character corresponding to "reference character appearance." The specific implementation of this third control can be configured as needed and will not be detailed here.
[0050] In the steps of this embodiment, by displaying the label window corresponding to the reference label, and based on the label window, further detailed configuration and modification of the reference label are achieved, thereby achieving further control over the generated target media data, improving the image quality of the generated target media data, and making the image content of the target media data better meet the user's design needs, and improving interaction efficiency.
[0051] Furthermore, in another possible implementation, the user can also operate the multimodal input box through a more direct interactive method, thereby further improving the efficiency of image generation. For example, the first operation includes a drag operation, and the multimodal input box is set in the image generation interface. As shown in Figure 6, a possible implementation of step S101 includes:
[0052] Step S1011: In response to the drag operation, an indicator mark is displayed in the image generation interface, where the indicator mark is used to indicate a target area in the image generation interface.
[0053] Step S1012: When the drag operation is released within the target area, a reference label is loaded into the multimodal input box.
[0054] For example, a drag operation is a common interactive operation that can be implemented using an input device such as a mouse or a touch screen. The drag operation generally includes several operation steps: select (button down), move (move), and release (button up). The specific implementation principle is not further described here. In this embodiment, the terminal device responds to the drag operation by first selecting one or more target media files from a folder in the select operation step, and then dragging the target media files into the designated target area by moving and releasing the files, thereby loading the target media files into the multimodal input box.
[0055] FIG7 is a schematic diagram of a process of loading a reference tag by a drag operation provided by an embodiment of the present disclosure. The process of the above steps is described in detail below in conjunction with FIG7. For example, first, the terminal device selects a target media file from a folder of the operating system in response to a drag operation applied by the user and drags it to form a drag icon. At the same time, an identifier indicating the drag destination area of the drag operation is displayed in the image generation interface, that is, an indicator. The area in the image generation interface indicated by the indicator is the target area. When the user moves the drag icon into the target area and releases it (that is, when the drag operation is released in the target area), the terminal device loads the media file corresponding to the drag icon, generates a corresponding reference tag, and displays the reference tag in the multimodal input box (after the basic prompt word). In a possible implementation, the target area is the area where the multimodal input box is located. In the steps of this embodiment, by responding to the drag operation, the media file can be quickly loaded into the multimodal input box by dragging, thereby quickly forming the prompt word required to generate the target media data, further improving the operational efficiency of the image generation process.
[0056] Step S102: Generate prompt word text according to the reference label and at least one basic prompt word in the multimodal input box.
[0057] Step S103: Generate target media data based on the prompt word text.
[0058] Exemplarily, after loading the reference tag in the multimodal input box, the reference method for the media file indicated by the reference tag is used to generate the corresponding target media data based on the media file. Specifically, the multimodal input box also includes at least one basic prompt word, wherein the basic prompt word is used to limit the content of the target media data to guide the image generation model to generate the image required by the user. Specifically, the basic prompt words include, for example: "dancing, full-body photo, soft light effect", etc. In other words, the basic prompt words have clear and unambiguous instructions, which can realize the basic content control of the image generated by the image generation model. The basic prompt words can be text information that the user enters in advance in the multimodal input box as needed, and will not be repeated here. Based on the reference label obtained within the multimodal input frame, combined with at least one basic prompt word within the multimodal input frame, a prompt word text is generated. This prompt word text not only implements basic content control (based on the basic prompt word) for the image generated by the image generation model, but also further implements complex content control (based on the reference label) for the image generated by the image generation model. Subsequently, the prompt word text is input into the image generation model to generate target media data that meets the aforementioned content control (basic content control and complex content control) requirements. The image generation model, for example, includes a Generative Adversarial Network (GAN) model, a Transformer network model, a diffusion network model, etc. The specific implementation principles and methods of the image generation model are not further described here.
[0059] Furthermore, illustratively, the specific implementation steps of step S102 include:
[0060] Step S1021: Generate reference prompt words based on the reference tags.
[0061] Step S1022: Generate prompt word text based on the basic prompt words and the reference prompt words.
[0062] Furthermore, illustratively, the specific implementation steps of step S103 include:
[0063] The prompt word text and media files are input into the image generation model to generate target media data.
[0064] The step of generating a reference prompt word based on the reference tag can be implemented using a corresponding parsing template based on the specific implementation of the reference tag. For example, if the reference tag content is "reference person appearance", then based on the corresponding parsing template, it is parsed into the corresponding reference prompt word "target object in reference image X", where "X" can be the file name of the media file. Alternatively, the reference tag content can be directly used as the reference prompt word.
[0065] Next, the reference prompt words and basic prompt words are input into corresponding prompt word templates to generate prompt word texts comprising multiple text paragraphs. The prompt word templates are used to add fixed sentences adapted to the image generation model, enabling the generated prompt word texts to better guide the image generation model in generating target media data that meets the requirements. The prompt word texts are then input into the image generation model. Based on the constraints represented by the reference prompt words in the prompt word texts and the media files indicated by the reference prompt words, the image generation model implements complex content control over the image, thereby generating target media data that meets the requirements for complex content control. The specific implementation process and principles of how the image generation model generates images based on the prompt word texts are not detailed here.
[0066] In an embodiment of the present disclosure, a reference tag is loaded into a multimodal input frame in response to a first operation, wherein the first operation is used to select at least one media file, and the reference tag is used to represent a reference method for the media file in the process of generating target media data with reference to the media file; a prompt word text is generated based on the reference tag and at least one basic prompt word in the multimodal input frame; and the target media data is generated based on the prompt word text. Through interactive operations, the media file is mapped to the multimodal input frame to form a reference tag, and prompt word text is generated based on the reference tag and basic prompt word corresponding to the media file in the multimodal input frame, thereby generating the target media data. This achieves the generation of prompt word text by replacing descriptive text with the media file, so that the generated prompt word text can accurately describe user needs, avoiding the problems of unstable image quality and large deviation in image content caused by inaccurate prompt words entered by the user that cannot accurately describe user needs, thereby improving image quality and image content accuracy.
[0067] Referring to FIG8 , FIG8 is a second flow chart of the interactive media data generation method provided by an embodiment of the present disclosure. Based on the embodiment shown in FIG2 , this embodiment further refines step S101 and adds a step of adjusting the reference tag. The interactive media data generation method includes:
[0068] Step S201: In response to the second operation, the cursor in the multimodal input box is moved to a target position, where the target position is located before or after the basic prompt word.
[0069] For example, when a basic prompt word has been entered in the multimodal input box, loading a reference label into the multimodal input box is equivalent to inserting content, which involves the issue of how to determine the insertion position of the reference label. In one possible implementation, the insertion position can be set by default after the last character in the multimodal input box, that is, the reference label is inserted after the existing basic prompt word; in another possible implementation, the terminal device can dynamically set the insertion position of the reference label based on user needs.
[0070] Furthermore, in the application scenario involved in this embodiment, since the prompt word text needs to be generated based on the basic prompt word and reference label input in the multimodal input box in the subsequent steps, when the arrangement order of the basic prompt word and the reference label is different, the generated prompt word text may be different, and then, the target media data finally generated may be different. Therefore, the control of the insertion position of the reference label can achieve the effect of controlling the image content of the target media data. In this embodiment, the terminal device first moves the cursor in the multimodal input box to the target position based on the second operation applied by the user. The target position is located before or after the basic prompt word, that is, the basic prompt word is a whole composed of multiple characters, and the target position can only be located outside the basic prompt word, thereby avoiding destroying the original basic prompt word. The second operation, for example, includes clicking the left move button, the right move button, etc.
[0071] Furthermore, illustratively, the second operation includes a first editing sub-operation and a second editing sub-operation. As shown in FIG9 , a specific implementation of step S201 includes:
[0072] Step S2011: In response to the first editing sub-operation, the multimodal input box is set from the first state to the second state, wherein when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of tags;
[0073] Step S2012: In response to at least one second editing sub-operation, move the cursor in the multimodal input box to a target position.
[0074] Exemplarily, the second operation includes two operation links: a first editing sub-operation and a second editing sub-operation. The first editing sub-operation is used to set the state of the multimodal input box, and the second editing sub-operation is used to trigger the cursor movement. Specifically, the multimodal input box includes a first state and a second state. When the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters. At this time, the user can move the cursor (for example, through the third editing sub-operation) to modify the basic prompt word. When the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of tags. At this time, since the cursor cannot be moved inside the basic prompt word, the basic prompt word cannot be modified. However, the user can control the cursor to move between (outside) the basic prompt words by implementing the second editing sub-operation, thereby improving the efficiency and accuracy of setting the target position.
[0075] Figure 10 is a schematic diagram of a process for setting a target position provided by an embodiment of the present disclosure. As shown in Figure 10, exemplarily, in the image generation interface, a switch control (shown as Shift_1 in the figure) for switching the state of the multimodal input box and a movement control (shown as "←" and "→" in the figure) for adjusting the left and right movement of the cursor are provided. The basic prompt words "dance, half-length portrait" have been entered in the multimodal input box. Before the user clicks the switch control, the multimodal input box is in the first state. At this time, by triggering the movement control, the cursor can move inside the basic prompt word, for example, the cursor can move to position d1; after the user clicks the switch control (the first editing sub-operation), the multimodal input box is in the second state. At this time, by triggering the movement control (the second editing sub-operation), the cursor can only move outside the basic prompt word, for example, the cursor can move to position d2, but cannot move to the previous position d1.
[0076] In this embodiment, the multimodal input box and the cursor movement are controlled respectively through the first editing sub-operation and the second editing sub-operation, thereby achieving the technical effect of flexibly modifying the basic prompt word and efficiently and accurately setting the target position.
[0077] Step S202: In response to the first operation, a reference tag is loaded at a target position in the multimodal input box.
[0078] For example, after determining the target position, a reference tag can be inserted based on the target position to achieve accurate loading of the reference tag based on the position in the multimodal input box. In one possible implementation, as shown in FIG11 , the specific implementation of step S202 includes:
[0079] Step S2021: In response to the first operation, a target reference mode corresponding to the reference tag is set.
[0080] Step S2022: Generate a reference tag according to the target reference mode and the type of the media file.
[0081] Step S2023: Load the reference tag in the form of a component into the multimodal input box.
[0082] Exemplarily, first, the terminal responds to the first operation and sets the target reference mode corresponding to the reference tag, wherein the first operation is, for example, an operation of selecting a target media file, more specifically, for example, a click operation of selecting a target media file through a folder interface, or a drag operation of dragging the target media file to a target area in an image generation interface. The specific implementation of the above-mentioned first operation has been described in detail in the previous embodiment and will not be repeated here. Furthermore, while responding to the first operation, the terminal device can further display a setting page for setting the target reference mode, and the setting of the target reference mode is completed through the setting page. More specifically, for example, the user enters a descriptive text representing the target reference mode in the setting page, and the terminal device generates the target reference mode based on the descriptive text. In this case, the target reference mode is represented by text; or, the terminal device performs semantic understanding based on the descriptive text and performs regression based on the knowledge base to obtain the target reference mode. In this case, the target reference mode can be identified by an identifier. The specific implementation process for setting the target reference mode through this settings page can be seen in the embodiment shown in FIG5 , which illustrates the process for setting the target reference mode after triggering the third operation control. This method achieves the purpose of configuring and modifying the target reference mode for a media file. It is understood that the process of setting the target reference mode corresponding to a reference tag can also be achieved through other operation methods. The specific operation method can be set as needed and will not be given as examples here.
[0083] In another possible implementation, the target reference mode is a preset fixed value or a dynamic value determined based on the type of the media file. That is, the terminal device does not need to perform the step of obtaining the target reference mode set by the user. In other words, step S202 can be implemented by simply executing steps S2022-S2023, thereby reducing the number of operational steps in specific scenarios and improving image generation efficiency. This will not be further described here.
[0084] Afterwards, according to the set target reference method, a label component is created and loaded into the multimodal input box. The component name of the label component is the label content of the reference label. Generally, the label component has interactive capabilities. When the user applies an operation to the label component to trigger it, the label component will perform the corresponding function, such as the function of displaying the label window in the embodiment shown in Figure 5. The specific implementation method of the label component and the creation steps such as the initialization process are based on the specific operating system and the required settings, and will not be described in detail here. The implementation process of loading the label component into the multimodal input box can be implemented by calling the specific loading interface of the multimodal input box, which will not be described in detail here.
[0085] Furthermore, in the case where the user does not need to set the target reference mode, the first operation includes a paste operation, and the reference label can be loaded more quickly through the paste operation. For example, in another possible implementation, the implementation steps of step S202 include:
[0086] Step S2024: In response to the paste operation, read the media file from the clipboard;
[0087] Step S2025: Load the media file, generate a reference tag based on the media type of the media file, and display it at the target location.
[0088] Exemplarily, a paste operation refers to writing data from the clipboard to a specified object, while the corresponding copy operation is the operation of writing specified data to the clipboard. Copy and cut operations are related technologies known to those skilled in the art, and are implemented in different operating systems and typically have shortcut keys. These operations will not be described in detail here. Specifically, after the cursor moves to the target location, the terminal device responds to the user's paste operation, reads the media file from the clipboard, identifies the media type of the media file, generates a reference tag, and quickly displays it at the target location. For example, if the media file is a video, a label template M_1 corresponding to the "video type" is obtained to generate a corresponding reference tag. Label template M_1 contains fixed text describing the reference method of the "video type" media file, such as "reference the content of this video" or "reference the content and duration of this video." This generates a reference tag and displays it at the target location.
[0089] In this embodiment, a reference tag can be directly generated at the target position of the multimodal input box through a copy-paste operation, thereby achieving the purpose of quickly adding a reference tag to the multimodal input box. In the scenario where multiple media files are used as references, the interaction efficiency of the target media data can be further improved.
[0090] Optionally, it also includes:
[0091] Step S203: Delete the reference label at the target position in the multimodal input frame.
[0092] Exemplarily, on the other hand, after setting the target position by moving the cursor, in addition to inserting (loading) the reference label as shown in the steps of the above embodiment, the reference label at the target position can also be deleted as needed. The specific operation method includes, for example: the terminal device responds to the deletion operation input by the user (for example, pressing the delete button on the keyboard) to trigger the operation and delete the reference label at the target position. More specifically, for example, the deletion operation includes a first deletion sub-operation and a second deletion sub-operation. The first deletion sub-operation and the second deletion sub-operation can be implemented based on the same operation method. For example, the user presses the delete button on the keyboard for the first time, which is the first deletion sub-operation; the user presses the delete button on the keyboard for the second time, which is the second deletion sub-operation; further, when the terminal device responds to the first deletion sub-operation, the reference label of the current target position, that is, the reference label to be deleted, is highlighted; when the terminal device responds to the second deletion sub-operation, the highlighted reference label is deleted; or, the terminal device responds to the user's trigger operation on the second operation control (as shown in Figure 5) to delete the reference label at the target position. For details, please refer to the detailed introduction in the embodiment shown in Figure 5, and no specific restrictions are made here.
[0093] Step S204: Generate prompt word text according to the reference label and at least one basic prompt word in the multimodal input box.
[0094] Step S205: Generate target media data based on the prompt word text.
[0095] In this embodiment, the implementation of step SS204 to step S205 is the same as the implementation of step S102 to step S103 in the embodiment shown in FIG. 2 of the present disclosure, and will not be described in detail here.
[0096] Corresponding to the interactive media data generation method of the above embodiment, FIG12 is a block diagram of the interactive media data generation device provided by the embodiment of the present disclosure. For ease of illustration, only the parts relevant to the embodiment of the present disclosure are shown. Referring to FIG12, the interactive media data generation device 3 includes:
[0097] An interactive module 31 is configured to load a reference tag into the multimodal input box in response to a first operation, wherein the first operation is used to select at least one media file, and the reference tag is used to represent a reference method for the media file in a process of generating target media data by referring to the media file;
[0098] A processing module 32 is configured to generate a prompt word text based on the reference label and at least one basic prompt word in the multimodal input box;
[0099] The generating module 33 is used to generate target media data based on the prompt word text.
[0100] According to one or more embodiments of the present disclosure, the interaction module 31 is further used to: in response to the second operation, move the cursor in the multimodal input box to a target position, and the target position is located before or after the basic prompt word; when the interaction module 31 loads a reference label in the multimodal input box in response to the first operation, it is specifically used to: in response to the first operation, load the reference label at the target position in the multimodal input box.
[0101] According to one or more embodiments of the present disclosure, the second operation includes a first editing sub-operation and a second editing sub-operation. When the interaction module 31 moves the cursor in the multimodal input box to the target position in response to the second operation, it is specifically used to: in response to the first editing sub-operation, set the multimodal input box from the first state to the second state, wherein, when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels; in response to at least one second editing sub-operation, move the cursor in the multimodal input box to the target position.
[0102] According to one or more embodiments of the present disclosure, the first operation includes a paste operation; when the interaction module 31 loads a reference tag at a target position in the multimodal input box in response to the first operation, it is specifically used to: read the media file from the clipboard in response to the paste operation; load the media file, and generate a reference tag and display it at the target position based on the media type of the media file.
[0103] According to one or more embodiments of the present disclosure, the first operation includes a drag operation, and the multimodal input box is set in the image generation interface; when the interaction module 31 loads the reference label in the multimodal input box in response to the first operation, it is specifically used to: display an indicator mark in the image generation interface in response to the drag operation, and the indicator mark is used to indicate a target area in the image generation interface; when the drag operation is released within the target area, the reference label is loaded in the multimodal input box.
[0104] According to one or more embodiments of the present disclosure, when the interaction module 31 loads a reference tag in a multimodal input box in response to a first operation, it is specifically used to: set a target reference mode corresponding to the reference tag in response to the first operation; generate a reference tag according to the target reference mode and the type of the media file; and load the reference tag in the multimodal input box in the form of a component.
[0105] According to one or more embodiments of the present disclosure, the reference tag includes a tag text and a tag image. The tag image includes a thumbnail of the media file, and the tag text is used to describe a target reference method for the media file.
[0106] According to one or more embodiments of the present disclosure, the interaction module 31 is further configured to: display a tag window corresponding to the reference tag in response to a third operation on the reference tag; and present a media file in the tag window.
[0107] According to one or more embodiments of the present disclosure, at least one of the following items is further included: a first operation control is provided in the tag window, and after being triggered, the first operation control is used to replace the first media file corresponding to the reference tag with the second media file; a second operation control is provided in the tag window, and after being triggered, the second operation control is used to delete the reference tag; and a third operation control is provided in the tag window, and after being triggered, the third operation control is used to modify the reference method for the media file.
[0108] According to one or more embodiments of the present disclosure, the processing module 32 is specifically used to: generate reference prompt words based on reference tags; generate prompt word text based on basic prompt words and reference prompt words; and the generation module 33 is specifically used to: input the prompt word text and media files into the image generation model to generate target media data.
[0109] The interactive module 31, processing module 32 and generating module 33 are connected in sequence. The interactive media data generating device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.
[0110] FIG13 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. As shown in FIG13 , the electronic device 4 includes:
[0111] A processor 41, and a memory 42 communicatively connected to the processor 41;
[0112] Memory 42 stores computer-executable instructions;
[0113] The processor 41 executes the computer-executable instructions stored in the memory 42 to implement the interactive media data generation method in the embodiments shown in FIG. 2 to FIG. 11 .
[0114] Optionally, the processor 41 and the memory 42 are connected via a bus 43 .
[0115] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the embodiments corresponding to Figures 2 to 11, and no further details will be given here.
[0116] An embodiment of the present disclosure provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they are used to implement the interactive media data generation method provided in any of the embodiments corresponding to Figures 2 to 11 of the present disclosure.
[0117] An embodiment of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, the interactive media data generation method provided in any one of the embodiments corresponding to FIG. 2 to FIG. 11 of the present disclosure is implemented.
[0118] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.
[0119] Referring to FIG14 , there is shown a schematic diagram of the structure of an electronic device 900 suitable for implementing an embodiment of the present disclosure. The electronic device 900 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG14 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0120] As shown in FIG14 , the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0121] Typically, the following devices can be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although FIG14 shows an electronic device 900 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0122] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0123] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0124] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0125] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0126] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0128] The units or modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit or module does not, in some cases, limit the unit itself.
[0129] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0130] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0131] In a first aspect, according to one or more embodiments of the present disclosure, a method for generating interactive media data is provided, comprising:
[0132] In response to a first operation, a reference tag is loaded in a multimodal input box, wherein the first operation is used to select at least one media file, and the reference tag is used to represent a reference method for the media file in a process of generating target media data with reference to the media file; a prompt word text is generated based on the reference tag and at least one basic prompt word in the multimodal input box; and the target media data is generated based on the prompt word text.
[0133] According to one or more embodiments of the present disclosure, the method further includes: in response to a second operation, moving the cursor in the multimodal input box to a target position, wherein the target position is located before or after the basic prompt word; and in response to the first operation, loading a reference label in the multimodal input box, including: in response to the first operation, loading a reference label at the target position in the multimodal input box.
[0134] According to one or more embodiments of the present disclosure, the second operation includes a first editing sub-operation and a second editing sub-operation, and the moving the cursor in the multimodal input box to the target position in response to the second operation includes: in response to the first editing sub-operation, setting the multimodal input box from the first state to the second state, wherein, when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels; in response to at least one of the second editing sub-operations, moving the cursor in the multimodal input box to the target position.
[0135] According to one or more embodiments of the present disclosure, the first operation includes a paste operation; in response to the first operation, a reference tag is loaded at the target position in the multimodal input box, including: in response to the paste operation, the media file is read from the clipboard; the media file is loaded, and according to the media type of the media file, a reference tag is generated and displayed at the target position.
[0136] According to one or more embodiments of the present disclosure, the first operation includes a drag operation, and the multimodal input box is set in the image generation interface; in response to the first operation, a reference label is loaded in the multimodal input box, including: in response to the drag operation, an indicator mark is displayed in the image generation interface, and the indicator mark is used to indicate a target area in the image generation interface; when the drag operation is released in the target area, the reference label is loaded in the multimodal input box.
[0137] According to one or more embodiments of the present disclosure, loading a reference tag in a multimodal input box in response to a first operation includes: setting a target reference mode corresponding to the reference tag in response to the first operation; generating the reference tag based on the target reference mode and the type of the media file; and loading the reference tag in the multimodal input box as a component.
[0138] According to one or more embodiments of the present disclosure, the reference tag includes a tag text and a tag image, the tag image includes a thumbnail of the media file, and the tag text is used to describe a target reference method for the media file.
[0139] According to one or more embodiments of the present disclosure, the method further includes: in response to a third operation on the reference tag, displaying a tag window corresponding to the reference tag; and presenting the media file in the tag window.
[0140] According to one or more embodiments of the present disclosure, at least one of the following items is further included: a first operation control is provided in the tag window, and after being triggered, the first operation control is used to replace the first media file corresponding to the reference tag with a second media file; a second operation control is provided in the tag window, and after being triggered, the second operation control is used to delete the reference tag; and a third operation control is provided in the tag window, and after being triggered, the third operation control is used to modify the reference method for the media file.
[0141] According to one or more embodiments of the present disclosure, generating a prompt word text based on the reference tag and at least one basic prompt word in the multimodal input box includes: generating a reference prompt word based on the reference tag; generating a prompt word text based on the basic prompt word and the reference prompt word; and generating the target media data based on the prompt word text includes: inputting the prompt word text and the media file into an image generation model to generate the target media data.
[0142] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an apparatus for generating interactive media data, comprising:
[0143] an interaction module configured to load a reference tag in the multimodal input box in response to a first operation, wherein the first operation is used to select at least one media file, and the reference tag is used to represent a reference method for the media file in a process of generating target media data by referring to the media file;
[0144] a processing module, configured to generate a prompt word text according to the reference label and at least one basic prompt word in the multimodal input box;
[0145] A generating module is used to generate the target media data based on the prompt word text.
[0146] According to one or more embodiments of the present disclosure, the interaction module is further used to: move the cursor in the multimodal input box to a target position in response to a second operation, and the target position is located before or after the basic prompt word; when the interaction module loads a reference label in the multimodal input box in response to a first operation, the interaction module is specifically used to: load the reference label at the target position in the multimodal input box in response to the first operation.
[0147] According to one or more embodiments of the present disclosure, the second operation includes a first editing sub-operation and a second editing sub-operation. When the interaction module moves the cursor in the multimodal input box to the target position in response to the second operation, it is specifically used to: in response to the first editing sub-operation, set the multimodal input box from the first state to the second state, wherein when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels; in response to at least one second editing sub-operation, move the cursor in the multimodal input box to the target position.
[0148] According to one or more embodiments of the present disclosure, the first operation includes a paste operation; when the interaction module loads a reference tag at the target position in the multimodal input box in response to the first operation, it is specifically used to: read the media file from the clipboard in response to the paste operation; load the media file, and generate a reference tag and display it at the target position based on the media type of the media file.
[0149] According to one or more embodiments of the present disclosure, the first operation includes a drag operation, and the multimodal input box is set in the image generation interface; when the interaction module loads a reference label in the multimodal input box in response to the first operation, it is specifically used to: display an indicator mark in the image generation interface in response to the drag operation, and the indicator mark is used to indicate a target area in the image generation interface; when the drag operation is released in the target area, the reference label is loaded in the multimodal input box.
[0150] According to one or more embodiments of the present disclosure, when the interaction module loads a reference tag in a multimodal input box in response to a first operation, it is specifically used to: set a target reference mode corresponding to the reference tag in response to the first operation; generate the reference tag according to the target reference mode and the type of the media file; and load the reference tag in the multimodal input box in the form of a component.
[0151] According to one or more embodiments of the present disclosure, the reference tag includes a tag text and a tag image, the tag image includes a thumbnail of the media file, and the tag text is used to describe a target reference method for the media file.
[0152] According to one or more embodiments of the present disclosure, the interaction module is further configured to: display a tag window corresponding to the reference tag in response to a third operation on the reference tag; and present the media file in the tag window.
[0153] According to one or more embodiments of the present disclosure, at least one of the following items is further included: a first operation control is provided in the tag window, and after being triggered, the first operation control is used to replace the first media file corresponding to the reference tag with a second media file; a second operation control is provided in the tag window, and after being triggered, the second operation control is used to delete the reference tag; and a third operation control is provided in the tag window, and after being triggered, the third operation control is used to modify the reference method for the media file.
[0154] According to one or more embodiments of the present disclosure, the processing module is specifically configured to: generate a reference prompt word based on the reference tag; generate a prompt word text based on the basic prompt word and the reference prompt word; and the generation module is specifically configured to: input the prompt word text and the media file into an image generation model to generate the target media data.
[0155] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;
[0156] The memory stores computer-executable instructions;
[0157] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the interactive media data generating method as described in the first aspect and various possible designs of the first aspect.
[0158] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the interactive media data generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0159] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the interactive media data generation method as described in the first aspect and various possible designs of the first aspect.
[0160] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0161] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0162] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An interactive media data generation method, comprising: In response to a first operation, loading a reference label in a multimodal input box, wherein the first operation is at least used to select a media file, and the reference label is used to characterize the reference manner for the media file in the process of generating target media data with reference to the media file; Generating a prompt text according to the reference label and at least one basic prompt word in the multimodal input box; Generating the target media data based on the prompt text.
2. The method according to claim 1, further comprising: In response to a second operation, moving the cursor in the multimodal input box to a target position, where the target position is before or after the basic prompt word; The loading the reference label in the multimodal input box in response to the first operation includes: In response to the first operation, loading a reference label at the target position in the multimodal input box.
3. The method according to claim 2, wherein The second operation includes a first editing sub-operation and a second editing sub-operation. The moving the cursor in the multimodal input box to the target position in response to the second operation includes: In response to the first editing sub-operation, setting the multimodal input box from a first state to a second state, wherein when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels; In response to at least one second editing sub-operation, moving the cursor in the multimodal input box to the target position.
4. The method according to claim 2, wherein The first operation includes a paste operation. The loading a reference label at the target position in the multimodal input box in response to the first operation includes: In response to the paste operation, reading the media file from the clipboard; Loading the media file, and generating a reference label according to the media type of the media file and displaying it at the target position.
5. The method according to claim 1, wherein, The first operation includes a drag operation, and the multimodal input box is set in an image generation interface; The loading the reference label in the multimodal input box in response to the first operation includes: In response to the drag operation, displaying an indication mark in the image generation interface, where the indication mark is used to indicate a target area in the image generation interface; When the drag operation is released within the target area, loading a reference label in the multimodal input box.
6. The method according to claim 1, wherein The loading the reference label in the multimodal input box in response to the first operation includes: In response to the first operation, setting a target reference manner corresponding to the reference label; Generating the reference label according to the target reference manner and the type of the media file; Loading the reference label in the multimodal input box in the form of a component.
7. The method according to any one of claims 1-6, wherein, The reference label includes a label text and a label picture, and the label picture includes a thumbnail of the media file, and the label text is used to describe the target reference manner for the media file.
8. The method according to claim 7, further comprising: In response to a third operation on the reference label, displaying a label window corresponding to the reference label; Within the label window, display the media file.
9. The method according to claim 8 further comprises at least one of the following: A first operation control is provided within the label window, and after being triggered, the first operation control is configured to replace the first media file corresponding to the reference label with a second media file; A second operation control is provided within the label window, and after being triggered, the second operation control is configured to delete the reference label; A third operation control is provided within the label window, and after being triggered, the third operation control is configured to modify the reference manner for the media file.
10. The method according to any one of claims 1-9, wherein, Generating a prompt text according to the reference label and at least one basic prompt word within the multimodal input box includes: Generating a reference prompt word according to the reference label; Generating a prompt text according to the basic prompt word and the reference prompt word; Generating the target media data based on the prompt text includes: Inputting the prompt text and the media file into an image generation model to generate the target media data.
11. An interactive media data generation device, comprising: An interaction module configured to load a reference label within a multimodal input box in response to a first operation, wherein the first operation is at least used to select a media file, and the reference label is used to characterize the reference manner for the media file during the process of generating target media data with reference to the media file; A processing module configured to generate a prompt text according to the reference label and at least one basic prompt word within the multimodal input box; A generation module configured to generate the target media data based on the prompt text.
12. An electronic device, comprising a processor and a memory, wherein The memory is configured to store computer execution instructions; The processor is configured to execute the computer execution instructions stored in the memory, such that the processor executes the interactive media data generation method according to any one of claims 1 to 10.
13. A computer-readable storage medium stores computer-executable instructions, wherein, When the processor executes the computer execution instructions, the interactive media data generation method according to any one of claims 1 to 10 is implemented.
14. A computer program product comprising a computer program, wherein, When the computer program is executed by the processor, the interactive media data generation method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Interactive media data generation method and device, electronic equipment and storage medium
CN120296183A
Intelligent cue word optimization method and system for generating images through characters
CN116012492A
Video generation method and server
CN116233491A
Method and system for content generation, computing device and storage medium
CN117115303A
Model training method, media information synthesizing method, and related apparatus
WO2021098338A1
Cited By
Multimedia resource generation method and device, electronic equipment and storage medium
CN120264087A