Interactive media data generation method and device, electronic equipment and storage medium

By loading reference labels and basic prompt words in the multimodal input box to generate prompt word text, the problems of unstable image quality and content deviation in the image generation model are solved, and higher quality and more accurate image generation are achieved.

CN120296183APending Publication Date: 2025-07-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410046071.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the image quality generated by the image generation model is unstable and is easily affected by the user's input prompt words, resulting in large deviations in the image content and it is difficult to meet user needs.

Method used

By loading reference tags in the multimodal input box, using media files to generate reference tags and basic prompt words, generate prompt word text, and then generate target media data to replace the description text entered by the user, ensuring that the prompt word text accurately describes user needs.

Benefits of technology

The quality of image generation and content accuracy are improved, the instability of image quality and content deviation caused by inaccurate user input is avoided, and the efficiency and effect of image generation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296183A_ABST
    Figure CN120296183A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an interactive media data generation method and device, electronic equipment and a storage medium, a reference tag is loaded in a multi-modal input box in response to a first operation, the first operation is at least used for selecting one media file, and the first operation is used for selecting the media file; the reference label is used for representing a reference mode for the media file in the process of generating the target media data by reference media files; generating a cue word text according to the reference tag and at least one basic cue word in the multi-modal input box; and generating target media data based on the prompt word text. The method comprises the steps of mapping a media file into a multi-modal input box in an interactive operation mode to form a reference label, generating a prompt word text based on the reference label and a basic prompt word corresponding to the media file in the multi-modal input box, and further generating target media data, the prompt word text is generated in a mode of replacing a description text with a media file, so that the generated prompt word text can accurately describe user requirements, the problems of unstable image quality and large image content deviation caused by inaccurate prompt words input by a user and incapability of accurately describing the user requirements are avoided, and the user experience is improved. And the image quality and the image content accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular, to an interactive media data generation method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, based on artificial intelligence-based content generation (AIGC) technology, corresponding images can be generated through the prompt words (prompts) input by users and an image generation model, thereby greatly improving the efficiency and quality of image generation and reducing the threshold for users to create content.

[0003] However, in the prior art, the quality of the images generated by the image generation model is directly affected by the prompt words input by the user, resulting in problems such as unstable image quality and large deviations in image content, making it difficult to meet the user's image generation requirements. Summary of the Invention

[0004] Embodiments of the present disclosure provide an interactive media data generation method, apparatus, electronic device, and storage medium to overcome the problems of unstable image quality and large deviations in image content in the generated images.

[0005] In a first aspect, embodiments of the present disclosure provide an interactive media data generation method, including:

[0006] In response to a first operation, loading reference tags in a multimodal input box, where the first operation is at least used to select one media file, and the reference tags are used to represent the reference manner for the media file in the process of generating target media data with reference to the media file; generating a prompt text according to the reference tags and at least one basic prompt word in the multimodal input box; and generating the target media data based on the prompt text.

[0007] In a second aspect, embodiments of the present disclosure provide an interactive media data generation apparatus, including:

[0008] An interaction module, configured to load reference tags in a multimodal input box in response to a first operation, where the first operation is at least used to select one media file, and the reference tags are used to represent the reference manner for the media file in the process of generating target media data with reference to the media file;

[0009] A processing module, configured to generate a prompt text according to the reference tags and at least one basic prompt word in the multimodal input box;

[0010] A generation module, configured to generate the target media data based on the prompt text.

[0011] In a third aspect, embodiments of the present disclosure provide an electronic device, including: a processor and a memory;

[0012] The memory stores computer-executable instructions;

[0013] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the interactive media data generation method described in the first aspect above and various possible designs of the first aspect.

[0014] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the interactive media data generation method described in the first aspect above and various possible designs of the first aspect is implemented.

[0015] In a fifth aspect, embodiments of the present disclosure provide a computer program product, including a computer program, which implements the interactive media data generation method described in the first aspect above and various possible designs of the first aspect when executed by a processor.

[0016] For the interactive media data generation method, device, electronic device, and storage medium provided in this embodiment, by responding to a first operation, reference tags are loaded into a multimodal input box, where the first operation is at least used to select one media file, and the reference tags are used to characterize the reference manner for the media file in the process of generating target media data with reference to the media file; a prompt text is generated according to the reference tags and at least one basic prompt word in the multimodal input box; and the target media data is generated based on the prompt text. Through an interactive operation method, the media file is mapped into the multimodal input box to form reference tags, and based on the reference tags corresponding to the media file in the multimodal input box and the basic prompt words, a prompt text is generated, and then the target media data is generated, realizing the generation of the prompt text in a way of using the media file instead of the description text, so that the generated prompt text can accurately describe the user's needs, avoiding the problems of unstable image quality and large deviation of image content caused by inaccurate prompt words input by the user and inability to accurately describe the user's needs, and improving the image quality and image content accuracy. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a diagram of an application scenario of the interactive media data generation method provided by an embodiment of the present disclosure;

[0019] Figure 2 It is a flowchart of the interactive media data generation method provided by an embodiment of the present disclosure Figure 1 ;

[0020] Figure 3 It is a schematic diagram of the process of loading reference labels in a multimodal input box provided by an embodiment of the present disclosure;

[0021] Figure 4 It is another schematic diagram of a reference label provided by an embodiment of the present disclosure:

[0022] Figure 5 It is a schematic diagram of the process of interacting based on a label window provided by an embodiment of the present disclosure;

[0023] Figure 6 It is Figure 2 a flowchart of the specific implementation manner of step S101 in the illustrated embodiment;

[0024] Figure 7 It is a schematic diagram of the process of loading reference labels through a drag operation provided by an embodiment of the present disclosure;

[0025] Figure 8 It is a flowchart of the interactive media data generation method provided by an embodiment of the present disclosure Figure 2 ;

[0026] Figure 9 It is Figure 8 a flowchart of the specific implementation manner of step S201 in the illustrated embodiment;

[0027] Figure 10 It is a schematic diagram of the process of setting a target position provided by an embodiment of the present disclosure;

[0028] Figure 11 It is Figure 8 a flowchart of the specific implementation manner of step S202 in the illustrated embodiment;

[0029] Figure 12 It is a structural block diagram of the interactive media data generation device provided by an embodiment of the present disclosure;

[0030] Figure 13 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure;

[0031] Figure 14 It is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0033] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation interfaces are provided for users to choose to authorize or reject.

[0034] The application scenarios of the embodiments of the present disclosure will be explained below:

[0035] Figure 1 FIG. is an application scenario diagram of the interactive media data generation method provided in the embodiments of the present disclosure. The interactive media data generation method provided in the embodiments of the present disclosure can be applied to an application program with a media generation function. More specifically, it can be applied to an application scenario of generating media data such as images and videos based on user descriptions. The execution entity of this embodiment can be a terminal device running the above-mentioned application program with a media generation function, or a server of the service end corresponding to the deployment of the above application program, or other electronic devices with similar functions. Refer to Figure 1 As shown in, taking the terminal device as the execution entity of the method in this embodiment and the application scenario of generating an image based on user descriptions as an example, after the terminal device runs the above application program, an image generation page is displayed. In the image generation page, there is at least an input control for receiving a prompt word input by the user. The user guides the image generation model to generate a corresponding image by inputting a prompt word for restricting the image content into the input control. Specifically, for example, as shown in the figure, when the user inputs the prompt word "dancing, full body photo, soft light effect" into the input control and clicks the trigger control of "generate immediately", the terminal device converts the above prompt word into a prompt word text adapted to the image generation model and inputs it into the image generation model, thereby generating an image that meets the above "dancing", "full body photo", and "soft light effect" and displaying it in the preview window of the image generation page, completing the process of image generation.

[0036] In the related art, in the process of generating AI images based on the above "text-to-image" form of image generation models, the image quality of the generated AI images is directly affected by the prompt words input by the user. If the prompt words input by the user are inaccurate or unreasonable, it will cause a deviation between the images generated by the image generation model and the user's needs. On the other hand, for some complex image content requirements, it is difficult for users to accurately describe them through prompt words. Therefore, the existing solutions in the prior art lead to problems such as unstable image quality and large deviation in image content, making it difficult to meet the user's image generation requirements.

[0037] The embodiments of the present disclosure provide an interactive media data generation method to solve the above problems.

[0038] Refer to Figure 2 , Figure 2 is a flowchart of the interactive media data generation method provided by the embodiments of the present disclosure. Figure 1 The method of this embodiment can be applied to a terminal device. The interactive media data generation method includes:

[0039] Step S101: In response to a first operation, load reference tags in a multimodal input box, where the first operation is at least used to select one media file, and the reference tags are used to characterize the reference manner for the media file in the process of generating target media data with reference to the reference media file.

[0040] Exemplarily, refer to Figure 1The schematic diagram of the application scenario shown, the application program running on the terminal device can be an application client or a browser (Web side), and no specific limitation is made here. In the image generation interface of the application program for generating AI images, there are multiple multi-modal input boxes. Among them, the multi-modal input box is an input control that can load and display non-text content. By applying a first operation to the image generation interface, the user can load a reference label in the multi-modal input box. Among them, the first operation is at least used to select a media file, and the media file is used as reference information to generate target media data. The target media data includes, for example, images, videos, music generated based on AIGC technology, or media (multimedia data) after superimposing two or more media data. In the following embodiments, the case where the target media data is an image (i.e., the application scenario of generating an image) is taken as an example for introduction, and other cases are similar and will not be elaborated one by one. Further, the reference label is a non-text object that maps the above media file. Further, the label content of the reference label is used to describe the role when generating the target media data based on the media file. For example, the label content of the reference label includes "reference target object", "reference person distribution", "reference subject", "reference edge information", etc., that is, it characterizes the role when generating the target media data based on the media file.

[0041] Exemplarily, Figure 3 This is a schematic diagram of the process of loading a reference label in a multi-modal input box provided by an embodiment of the present disclosure. In combination with Figure 3 the above steps are further described as follows. As Figure 3 shown, first, there is a multi-modal input box and a media import control in the image generation interface. By clicking the media import control, a folder interface is displayed above the image generation interface, and one or more media files are displayed in the folder interface, such as media file #1, media file #2, media file #3, etc. shown in the figure. Among them, exemplarily, the media file can be a picture, a video file, an audio file, etc., and no specific limitation is made. The media file can be a file stored locally on the terminal device or a file stored on the server or other storage media. After that, the user selects a media file (such as media file #3) as the target media file from the folder interface by applying a first operation (such as single-clicking or double-clicking), and the terminal device processes the target media file to generate a corresponding reference label and displays it in the multi-modal input box.

[0042] Figure 4 This is another schematic diagram of a reference label provided by an embodiment of the present disclosure. As Figure 4As shown, by way of example, the media file is, for example, a video or a picture; the reference label includes label text and a label picture, and the label picture includes a thumbnail of the media file. The label text is used to describe the target reference manner for the media file. For example, "reference person distribution" shown in the figure, that is, referring to the person distribution in the above picture to generate the target media data.

[0043] Further, after loading the reference label in the multimodal input box, the terminal device can perform subsequent processing on the reference label by responding to further operations of the user to display the media file corresponding to the reference label. Exemplarily, the steps of displaying the media file by operating the reference label include:

[0044] Step S101A: In response to a third operation on the reference label, display a label window corresponding to the reference label.

[0045] Step S101B: In the label window, display the media file.

[0046] Figure 5 It is a schematic diagram of a process for interaction based on a label window provided by an embodiment of the present disclosure. As Figure 5 shown, after loading the reference label in the multimodal input box, the user, for example, clicks (the third operation) on the reference label (based on the solution in the previous embodiment, such as clicking on the label text or the label picture). Above the reference label, a label window is displayed. The label window is used to display more details related to the reference label. As shown in the figure, optionally, the media file corresponding to the reference label or a thumbnail of the media file is displayed in the label window, that is, in the label window, the media file is displayed.

[0047] Optionally, a first operation control (shown as control P1 in the figure) is set in the label window. After being triggered, the first operation control is used to replace the first media file corresponding to the reference label with a second media file. Specifically, for example, when the terminal device responds to the user's operation and triggers the first operation control, a folder interface can be displayed, and in the folder interface, a second media file is selected based on the user's operation, so as to achieve the replacement of the media file corresponding to the reference label. The specific operation method has been introduced in the previous embodiment and will not be elaborated here.

[0048] Optionally, a second operation control (shown as control P2 in the figure) is set in the label window. After being triggered, the second operation control is used to delete the reference label, that is, to restore the content in the multimodal input box to the state before performing step S101 of this embodiment. The specific implementation process steps are not elaborated.

[0049] Optionally, a third operation control (shown as control P3 in the figure) is provided within the label window. After being triggered, the third operation control is used to modify the reference method for the media file. Specifically, the third operation control is used to configure and modify the target reference method for the media file, that is, the content corresponding to the label text. For example, by clicking on the reference item control, "reference by person's appearance" can be modified to "reference by person's layout", or the target person corresponding to "reference by person's appearance" can be set. The specific implementation method of this third operation control can be set as needed and will not be elaborated here.

[0050] In the steps of this embodiment, by displaying the label window corresponding to the reference label and based on the label window, further refinement configuration and modification of the reference label are realized, so as to further control the generated target media data, improve the image quality of the generated target media data, make the image content of the target media data better meet the user's design needs, and improve the interaction efficiency.

[0051] Furthermore, in another possible implementation manner, the user can also operate on the multi-modal input box through a more direct interaction method, so as to further improve the efficiency of image generation. Exemplarily, the first operation includes a drag-and-drop operation. The multi-modal input box is set within the image generation interface, as Figure 6 shown, a possible implementation manner of step S101 includes:

[0052] Step S1011: In response to the drag-and-drop operation, display an indication mark within the image generation interface. The indication mark is used to indicate a target area within the image generation interface.

[0053] Step S1012: When the drag-and-drop operation is released within the target area, load the reference label into the multi-modal input box.

[0054] Exemplarily, the drag-and-drop operation is a common interaction operation, which can be implemented through input devices such as a mouse or through a touch screen. The drag-and-drop operation generally includes several operation links such as selection (button down), movement (move), and release (button up). The specific implementation principle will not be elaborated here. Among them, in this embodiment, the terminal device responds to the drag-and-drop operation. First, in the selection operation link, one or more target media files are selected from the folder, and then through movement and release, the target media files are dragged into the specified target area, so as to implement the step of loading the above target media files into the multi-modal input box.

[0055] Figure 7 FIG. is a schematic diagram of a process for loading a reference label through a drag-and-drop operation provided by an embodiment of the present disclosure. Next, in combination with Figure 7The process of the above steps is introduced in detail. Exemplarily, first, in response to a drag operation applied by the user, the terminal device selects a target media file from the operating system's folder and drags it to form a drag icon. At the same time, an identifier indicating the drag destination area of the drag operation, that is, an indication identifier, is displayed in the image generation interface. The area in the image generation interface indicated by the indication identifier is the target area. When the user moves the drag icon into the target area and releases it (i.e., when the above drag operation is released in the target area), the terminal device loads the media file corresponding to the drag icon, generates a corresponding reference tag, and displays the reference tag in the multimodal input box (after the basic prompt). In a possible implementation manner, the target area is the area where the multimodal input box is located. In the steps of this embodiment, by responding to the drag operation, it is possible to quickly load the media file into the multimodal input box in a drag-and-drop manner, thereby quickly forming the prompt required to generate the target media data, and further improving the operation efficiency of the image generation process.

[0056] Step S102: Generate a prompt text based on the reference tag and at least one basic prompt in the multimodal input box.

[0057] Step S103: Generate target media data based on the prompt text.

[0058] Exemplarily, after loading the reference tags in the multimodal input box, the reference method for the media file indicated by the above reference tags is used to generate the corresponding target media data based on the media file. Specifically, in the multimodal input box, there is also at least one basic prompt. Among them, the basic prompt is used to restrict the content of the target media data to guide the image generation model to generate the image required by the user. Specifically, the basic prompt includes, for example: "dancing, full body photo, soft light effect", etc. Relatively speaking, the basic prompt has a clear and definite indication, which can achieve the basic content control of the image generated by the image generation model. The basic prompt can be the text information pre-entered by the user in the multimodal input box as needed, which will not be elaborated here. On the basis of obtaining the reference tags in the multimodal input box, combined with at least one basic prompt in the multimodal input box, a prompt text is jointly generated, so that the prompt text can not only achieve the basic content control of the image generated by the image generation model (based on the basic prompt), but also further achieve the complex content control of the image generated by the image generation model (based on the reference tag). After that, the prompt text is input into the image generation model, and the target media data that meets the above content control (basic content control and complex content control) requirements can be generated. Among them, the image generation model includes, for example, a Generative Adversarial Network (GAN) model, a Transformer network model, a diffusion network model, etc. The specific implementation principle and method of the image generation model will not be elaborated here.

[0059] Further, exemplarily, the specific implementation steps of step S102 include:

[0060] Step S1021: Generate a reference prompt according to the reference tag.

[0061] Step S1022: Generate a prompt text according to the basic prompt and the reference prompt.

[0062] Further, exemplarily, the specific implementation steps of step S103 include:

[0063] Input the prompt text and the media file into the image generation model to generate the target media data.

[0064] Among them, the step of generating a reference prompt according to the reference tag can be implemented through a corresponding parsing template based on the specific implementation method of the reference tag. For example, if the tag content of the reference tag is "reference the appearance of a person", then based on the corresponding parsing template, it is parsed into the corresponding reference prompt as "the target object in reference picture X", where "X" can be the file name corresponding to the media file. Or, the tag content of the reference tag can also be directly used as the reference prompt.

[0065] After that, referring to the reference prompt and the basic prompt, the corresponding prompt template is input to obtain a prompt text including multiple text paragraphs. The prompt template is used to add fixed statements for adapting to the image generation model, so that the generated prompt text can better guide the image generation model to generate target media data that meets the requirements. Then, the prompt text is input into the image generation model. The image generation model realizes complex content control of the image according to the constraint conditions represented by the reference prompt in the prompt text and the media file indicated by the reference prompt, so as to obtain target media data that meets the complex content control requirements. The specific implementation process and principle of the image generation model generating images based on the prompt text are prior art and will not be elaborated here.

[0066] In the embodiments of the present disclosure, by responding to a first operation, a reference label is loaded into the multimodal input box, where the first operation is at least used to select a media file, and the reference label is used to represent the reference manner for the media file in the process of generating target media data with reference to the reference media file; a prompt text is generated according to the reference label and at least one basic prompt in the multimodal input box; and target media data is generated based on the prompt text. By means of an interactive operation, the media file is mapped into the multimodal input box to form a reference label, and a prompt text is generated based on the reference label corresponding to the media file and the basic prompt in the multimodal input box, and then target media data is generated, realizing the generation of the prompt text in the way of using the media file instead of the descriptive text, so that the generated prompt text can accurately describe the user's needs, avoiding the problems of unstable image quality and large deviation of image content caused by inaccurate prompts input by the user and inability to accurately describe the user's needs, and improving the image quality and the accuracy of image content.

[0067] Reference Figure 8 , Figure 8 is a flowchart of the interactive media data generation method provided by the embodiments of the present disclosure Figure 2 . On the basis of the embodiment shown in Figure 2 , step S101 is further refined, and a step of adjusting the reference label is added. The interactive media data generation method includes:

[0068] Step S201: In response to a second operation, move the cursor in the multimodal input box to a target position, where the target position is before or after the basic prompt.

[0069] Exemplarily, when a basic prompt has been entered in the multimodal input box, loading a reference label into the multimodal input box is equivalent to inserting content. Therefore, there is a problem of how to determine the insertion position of the reference label. In one possible implementation, the insertion position can be default set after the last character in the multimodal input box, that is, after the existing basic prompt, the reference label is inserted; in another possible implementation, the terminal device can dynamically set the insertion position of the reference label based on user needs.

[0070] Furthermore, in the application scenario involved in this embodiment, since the subsequent steps need to generate a prompt text based on the basic prompt and the reference label entered in the multimodal input box, when the arrangement order of the basic prompt and the reference label is different, it may cause the generated prompt text to be different, and further, the final generated target media data to be different. Therefore, controlling the insertion position of the reference label can achieve the effect of controlling the image content of the target media data. In this embodiment, the terminal device first moves the cursor in the multimodal input box to the target position based on the second operation applied by the user. The target position is before or after the basic prompt, that is, the basic prompt is an overall composed of multiple characters, and the target position can only be outside the basic prompt, so as to avoid destroying the original basic prompt. Among them, the second operation includes, for example, clicking the left movement button, the right movement button, etc.

[0071] Furthermore, exemplarily, the second operation includes a first editing sub-operation and a second editing sub-operation, as Figure 9 shown, the specific implementation manner of step S201 includes:

[0072] Step S2011: In response to the first editing sub-operation, set the multimodal input box from the first state to the second state. Among them, when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels;

[0073] Step S2012: In response to at least one second editing sub-operation, move the cursor in the multimodal input box to the target position.

[0074] Exemplarily, the second operation includes two operation links, namely, a first editing sub-operation and a second editing sub-operation. The first editing sub-operation is used to set the state of the multimodal input box, and the second editing sub-operation is used to trigger the cursor movement. Specifically, the multimodal input box includes a first state and a second state. When the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters. At this time, the user can modify the basic prompt word by moving the cursor (for example, by the third editing sub-operation). When the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of tags. At this time, since the cursor cannot be moved inside the basic prompt word, the basic prompt word cannot be modified. However, the user can control the cursor to move between (outside) the basic prompt words by implementing the second editing sub-operation, thereby improving the efficiency and accuracy of setting the target position.

[0075] Figure 10 A schematic diagram of a process of setting a target position provided by an embodiment of the present disclosure is shown in FIG. Figure 10 As shown, exemplarily, in the image generation interface, a switch control (shown as Shift_1 in the figure) for switching the state of the multimodal input box and a mobile control (shown as "←" and "→" in the figure) for adjusting the left and right movement of the cursor are provided. The basic prompt word "dance, half-length portrait" has been entered in the multimodal input box. Before the user clicks the switch control, the multimodal input box is in the first state. At this time, by triggering the mobile control, the cursor can move inside the basic prompt word, for example, the cursor can move to position d1; after the user clicks the switch control (the first editing sub-operation), the multimodal input box is in the second state. At this time, by triggering the mobile control (the second editing sub-operation), the cursor can only move outside the basic prompt word, for example, the cursor can move to position d2, but cannot move to the previous position d1.

[0076] In this embodiment, the multimodal input box and the cursor movement are controlled respectively through the two operation links of the first editing sub-operation and the second editing sub-operation, thereby achieving the technical effect of flexibly modifying the basic prompt words and efficiently and accurately setting the target position.

[0077] Step S202: In response to the first operation, a reference tag is loaded at a target position in the multimodal input box.

[0078] For example, after determining the target position, a reference tag may be inserted based on the target position to achieve accurate loading of the reference tag based on the position in the multimodal input box. Figure 11 As shown, the specific implementation of step S202 includes:

[0079] Step S2021: In response to the first operation, a target reference mode corresponding to the reference tag is set.

[0080] Step S2022: Generate a reference label according to the target reference method and the type of the media file.

[0081] Step S2023: Load the reference label into the multimodal input box in the form of a component.

[0082] Exemplarily, first, in response to a first operation, the terminal sets the target reference method corresponding to the reference label. Herein, the first operation is, for example, an operation of selecting a target media file. More specifically, for example, a click operation of selecting the target media file through a folder interface, or a drag operation of dragging the target media file to a target area within an image generation interface. The specific implementation manners of the above first operation have been introduced in detail in the previous embodiments and will not be elaborated herein. Further, while responding to the first operation, the terminal device may further display a setting page for setting the target reference method. Through this setting page, the setting of the target reference method is completed. More specifically, for example, the user inputs a description text representing the target reference method within the setting page, and the terminal device generates the target reference method according to this description text. In this case, the target reference method is represented by text; or, the terminal device performs semantic understanding based on the description text and performs regression based on a knowledge base to obtain the target reference method. In this case, the target reference method may be identified by an identifier. For the specific implementation process of setting the target reference method through this setting page, reference may be made to Figure 5 In the illustrated embodiment, for the implementation process of setting the target reference method after triggering a third operation control, through the above manner, the purpose of configuring and modifying the target reference method for the media file is achieved. Currently, it can be understood that the process of setting the target reference method corresponding to the reference label may also be implemented through other operation manners, and the specific operation manners may be set according to needs and will not be exemplified one by one herein.

[0083] In another possible implementation manner, the above target reference method is a preset fixed value or a dynamic value determined according to the type of the media file. That is, the terminal device does not need to perform the step of obtaining the target reference method set by the user. That is, steps S2022 - S2023 may be only executed to implement step S202, so as to reduce the operation steps in specific scenarios and improve the efficiency of image generation, which will not be elaborated herein.

[0084] After that, according to the set target reference method, a label component is created and loaded into the multimodal input box. The component name of this label component is the label content of the reference label. Generally, the label component has interaction capabilities. When the user performs an operation on this label component to trigger it, this label component will execute corresponding functions, such as Figure 5The function of the label window is shown in the illustrated embodiment. Regarding the specific implementation of the label component and the creation steps such as the initialization process, they are based on the specific operating system and requirements and will not be elaborated here. The implementation process of loading the label component into the multi-modal input box can be achieved by calling a specific loading interface of the multi-modal input box, which will not be elaborated here.

[0085] Further, in the case where the user does not need to set the target reference method, the first operation includes a paste operation, and the reference label can be loaded more quickly through the paste operation. Exemplarily, in another possible implementation, the implementation steps of step S202 include:

[0086] Step S2024: In response to the paste operation, read the media file from the clipboard;

[0087] Step S2025: Load the media file, and generate a reference label according to the media type of the media file and display it at the target position.

[0088] Exemplarily, the paste operation refers to the operation of writing the data in the clipboard to a specified object, and the corresponding operation is the copy operation, that is, writing to the clipboard for the specified data. The copy operation and the cut operation are prior arts known to those skilled in the art and have corresponding implementations under different operating systems, and usually have operation shortcut keys, which will not be elaborated here. Specifically, after the cursor moves to the target position, the terminal device reads the media file from the clipboard in response to the user's paste operation, identifies the media type of the media file, generates a reference label, and quickly displays it at the target position. For example, when the media file is a video, the label template M_1 corresponding to "video type" is obtained to generate the corresponding reference label. The label template M_1 carries fixed text describing the reference method of the media file of "video type", such as "Refer to the content of this video", "Refer to the content and duration of this video", etc., so as to generate a reference label and display it at the target position.

[0089] In this embodiment, by means of the copy operation - paste operation, at the target position of the multi-modal input box, a reference label can be directly generated, achieving the purpose of quickly adding a reference label to the multi-modal input box. In the scenario of using multiple media files as references, the interaction efficiency of the target media data can be further improved.

[0090] Optionally, it further includes:

[0091] Step S203: Delete the reference label at the target position within the multi-modal input box.

[0092] Exemplarily, on the other hand, after setting the target position by moving the cursor, in addition to inserting (loading) the reference label as shown in the steps of the above embodiments, the reference label at the target position can also be deleted as needed. The specific operation methods include, for example: the terminal device responds to the deletion operation input by the user (such as pressing the delete key on the keyboard) to trigger the operation to delete the reference label at the target position. More specifically, for example, the deletion operation includes a first deletion sub-operation and a second deletion sub-operation. The first deletion sub-operation and the second deletion sub-operation can be implemented based on the same operation method. For example, when the user presses the delete key on the keyboard for the first time, it is the first deletion sub-operation; when the user presses the delete key on the keyboard for the second time, it is the second deletion sub-operation. Further, when the terminal device responds to the first deletion sub-operation, the reference label at the current target position, that is, the reference label to be deleted, is highlighted; when the terminal device responds to the second deletion sub-operation, the highlighted reference label is deleted; or, the terminal device responds to the triggering operation of the user for the second operation control (refer to Figure 5 as shown), and deletes the reference label at the target position. For specific reference, please refer to the detailed introduction in the embodiment shown in Figure 5 , and no specific limitation is made here.

[0093] Step S204: Generate a prompt text based on the reference label and at least one basic prompt word in the multi-modal input box.

[0094] Step S205: Generate target media data based on the prompt text.

[0095] In this embodiment, the implementation manners of step SS204-step S205 are the same as those of step S102-step S103 in the embodiment shown in this disclosure Figure 2 , and will not be elaborated here one by one.

[0096] Corresponding to the interactive media data generation method in the above embodiment, Figure 12 is the structural block diagram of the interactive media data generation device provided by the embodiments of this disclosure. For the convenience of description, only the parts related to the embodiments of this disclosure are shown. Refer to Figure 12 , the interactive media data generation device 3 includes:

[0097] An interaction module 31, configured to load a reference label in the multi-modal input box in response to a first operation, where the first operation is at least used to select a media file, and the reference label is used to characterize the reference manner for the media file in the process of generating target media data with reference to the reference media file;

[0098] A processing module 32, configured to generate a prompt text based on the reference label and at least one basic prompt word in the multi-modal input box;

[0099] A generation module 33, configured to generate target media data based on the prompt text.

[0100] According to one or more embodiments of the present disclosure, the interaction module 31 is further configured to: in response to a second operation, move the cursor in the multimodal input box to a target position, where the target position is before or after the basic prompt; when the interaction module 31 loads a reference label in the multimodal input box in response to a first operation, specifically: in response to the first operation, load the reference label at the target position in the multimodal input box.

[0101] According to one or more embodiments of the present disclosure, the second operation includes a first editing sub-operation and a second editing sub-operation. When the interaction module 31 moves the cursor in the multimodal input box to the target position in response to the second operation, specifically: in response to the first editing sub-operation, set the multimodal input box from a first state to a second state, where when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels; in response to at least one second editing sub-operation, move the cursor in the multimodal input box to the target position.

[0102] According to one or more embodiments of the present disclosure, the first operation includes a paste operation; when the interaction module 31 loads a reference label at the target position in the multimodal input box in response to the first operation, specifically: in response to the paste operation, read a media file from the clipboard; load the media file, and generate and display a reference label at the target position according to the media type of the media file.

[0103] According to one or more embodiments of the present disclosure, the first operation includes a drag operation, and the multimodal input box is set in an image generation interface; when the interaction module 31 loads a reference label in the multimodal input box in response to the first operation, specifically: in response to the drag operation, display an indication identifier in the image generation interface, where the indication identifier is used to indicate a target area in the image generation interface; when the drag operation is released within the target area, load a reference label in the multimodal input box.

[0104] According to one or more embodiments of the present disclosure, when the interaction module 31 loads a reference label in the multimodal input box in response to the first operation, specifically: in response to the first operation, set a target reference method corresponding to the reference label; generate a reference label according to the target reference method and the type of the media file; load the reference label in the multimodal input box in the form of a component.

[0105] According to one or more embodiments of the present disclosure, the reference label includes a label text and a label picture, where the label picture includes a thumbnail of the media file, and the label text is used to describe the target reference method for the media file.

[0106] According to one or more embodiments of the present disclosure, the interaction module 31 is further configured to: in response to a third operation on a reference tag, display a tag window corresponding to the reference tag; and display a media file within the tag window.

[0107] According to one or more embodiments of the present disclosure, at least one of the following is further included: a first operation control is provided within the tag window, and after being triggered, the first operation control is configured to replace a first media file corresponding to the reference tag with a second media file; a second operation control is provided within the tag window, and after being triggered, the second operation control is configured to delete the reference tag; a third operation control is provided within the tag window, and after being triggered, the third operation control is configured to modify the reference manner for the media file.

[0108] According to one or more embodiments of the present disclosure, the processing module 32 is specifically configured to: generate a reference prompt word according to the reference tag; generate a prompt word text according to the basic prompt word and the reference prompt word; the generation module 33 is specifically configured to: input the prompt word text and the media file into an image generation model to generate target media data.

[0109] Among them, the interaction module 31, the processing module 32, and the generation module 33 are connected in sequence. The interactive media data generation device 3 provided in this embodiment can execute the technical solutions of the above method embodiments, and the implementation principles and technical effects are similar, which will not be elaborated here.

[0110] Figure 13 The following is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure, as Figure 13 shown, the electronic device 4 includes:

[0111] a processor 41, and a memory 42 communicatively connected to the processor 41;

[0112] The memory 42 stores computer-executable instructions;

[0113] The processor 41 executes the computer-executable instructions stored in the memory 42 to implement the interactive media data generation method in the embodiment as Figure 2 - Figure 11 shown.

[0114] Optionally, the processor 41 and the memory 42 are connected through a bus 43.

[0115] For relevant descriptions, reference may be made to Figure 2 - Figure 11 the relevant descriptions and effects corresponding to the steps in the corresponding embodiments, and details will not be elaborated here.

[0116] Embodiments of the present disclosure provide a computer-readable storage medium storing computer-executable instructions, which are used to implement the present disclosure when executed by a processor. Figure 2 - Figure 11 The interactive media data generation method provided in any one of the corresponding embodiments.

[0117] Embodiments of the present disclosure provide a computer program product including a computer program, which implements the interactive media data generation method provided in any one of the corresponding embodiments of the present disclosure when executed by a processor. Figure 2 - Figure 11 The interactive media data generation method provided in any one of the corresponding embodiments.

[0118] To implement the above embodiments, embodiments of the present disclosure further provide an electronic device.

[0119] Referring to Figure 14 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 may be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 14 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0120] As Figure 14 shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0121] Typically, the following devices can be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 14 the electronic device 900 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.

[0122] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0123] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0124] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.

[0125] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.

[0126] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0128] The units or modules involved in the embodiments described in this disclosure can be implemented in software or in hardware. Among them, the name of the unit or module does not constitute a limitation on the unit itself in some cases.

[0129] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0130] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0131] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an interactive media data generation method, including:

[0132] In response to a first operation, loading a reference label into a multimodal input box, where the first operation is at least used to select a media file, and the reference label is used to characterize the reference manner for the media file in the process of generating target media data with reference to the media file; generating a prompt text according to the reference label and at least one basic prompt word in the multimodal input box; and generating the target media data based on the prompt text.

[0133] According to one or more embodiments of the present disclosure, the method further includes: in response to a second operation, moving the cursor in the multimodal input box to a target position, where the target position is before or after the basic prompt word; and the loading the reference label into the multimodal input box in response to the first operation includes: in response to the first operation, loading the reference label at the target position in the multimodal input box.

[0134] According to one or more embodiments of the present disclosure, the second operation includes a first editing sub-operation and a second editing sub-operation, and the moving the cursor in the multimodal input box to the target position in response to the second operation includes: in response to the first editing sub-operation, setting the multimodal input box from a first state to a second state, where when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels; and in response to at least one of the second editing sub-operations, moving the cursor in the multimodal input box to the target position.

[0135] According to one or more embodiments of the present disclosure, the first operation includes a paste operation; in response to the first operation, a reference tag is loaded at the target position in the multimodal input box, including: in response to the paste operation, the media file is read from the clipboard; the media file is loaded, and according to the media type of the media file, a reference tag is generated and displayed at the target position.

[0136] According to one or more embodiments of the present disclosure, the first operation includes a drag operation, and the multimodal input box is set in an image generation interface; in response to the first operation, a reference label is loaded in the multimodal input box, including: in response to the drag operation, an indicator mark is displayed in the image generation interface, and the indicator mark is used to indicate a target area in the image generation interface; when the drag operation is released in the target area, the reference label is loaded in the multimodal input box.

[0137] According to one or more embodiments of the present disclosure, loading a reference tag in a multimodal input box in response to a first operation includes: setting a target reference mode corresponding to the reference tag in response to the first operation; generating the reference tag according to the target reference mode and the type of the media file; and loading the reference tag in the multimodal input box in the form of a component.

[0138] According to one or more embodiments of the present disclosure, the reference tag includes a tag text and a tag image, the tag image includes a thumbnail of the media file, and the tag text is used to describe a target reference method for the media file.

[0139] According to one or more embodiments of the present disclosure, the method further includes: in response to a third operation on the reference tag, displaying a tag window corresponding to the reference tag; and displaying the media file in the tag window.

[0140] According to one or more embodiments of the present disclosure, at least one of the following items is also included: a first operation control is provided in the tag window, and the first operation control is used to replace the first media file corresponding to the reference tag with a second media file after being triggered; a second operation control is provided in the tag window, and the second operation control is used to delete the reference tag after being triggered; a third operation control is provided in the tag window, and the third operation control is used to modify the reference method for the media file after being triggered.

[0141] According to one or more embodiments of the present disclosure, generating a prompt text based on the reference label and at least one basic prompt word in the multimodal input box includes: generating a reference prompt word according to the reference label; generating a prompt text according to the basic prompt word and the reference prompt word; generating the target media data based on the prompt text, including: inputting the prompt text and the media file into an image generation model to generate the target media data.

[0142] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an interactive media data generation device, including:

[0143] An interaction module, configured to load a reference label in a multimodal input box in response to a first operation, where the first operation is at least used to select one media file, and the reference label is used to characterize the reference manner for the media file in the process of generating target media data with reference to the media file;

[0144] A processing module, configured to generate a prompt text according to the reference label and at least one basic prompt word in the multimodal input box;

[0145] A generation module, configured to generate the target media data based on the prompt text.

[0146] According to one or more embodiments of the present disclosure, the interaction module is further configured to: move the cursor in the multimodal input box to a target position in response to a second operation, where the target position is before or after the basic prompt word; when the interaction module loads the reference label in the multimodal input box in response to the first operation, specifically: in response to the first operation, load the reference label at the target position in the multimodal input box.

[0147] According to one or more embodiments of the present disclosure, the second operation includes a first editing sub-operation and a second editing sub-operation. When the interaction module moves the cursor in the multimodal input box to the target position in response to the second operation, specifically: in response to the first editing sub-operation, set the multimodal input box from a first state to a second state, where when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels; in response to at least one second editing sub-operation, move the cursor in the multimodal input box to the target position.

[0148] According to one or more embodiments of the present disclosure, the first operation includes a paste operation; when the interaction module loads a reference label at the target position in the multimodal input box in response to the first operation, it is specifically configured to: in response to the paste operation, read the media file from the clipboard; load the media file, and generate a reference label according to the media type of the media file and display it at the target position.

[0149] According to one or more embodiments of the present disclosure, the first operation includes a drag operation, and the multimodal input box is set in the image generation interface; when the interaction module loads a reference label in the multimodal input box in response to the first operation, it is specifically configured to: in response to the drag operation, display an indication identifier in the image generation interface, and the indication identifier is used to indicate a target area in the image generation interface; when the drag operation is released within the target area, load a reference label in the multimodal input box.

[0150] According to one or more embodiments of the present disclosure, when the interaction module loads a reference label in the multimodal input box in response to the first operation, it is specifically configured to: in response to the first operation, set the target reference method corresponding to the reference label; generate the reference label according to the target reference method and the type of the media file; load the reference label in the multimodal input box in the form of a component.

[0151] According to one or more embodiments of the present disclosure, the reference label includes a label text and a label picture, the label picture includes a thumbnail of the media file, and the label text is used to describe the target reference method for the media file.

[0152] According to one or more embodiments of the present disclosure, the interaction module is further configured to: in response to a third operation on the reference label, display a label window corresponding to the reference label; and display the media file in the label window.

[0153] According to one or more embodiments of the present disclosure, it further includes at least one of the following: a first operation control is provided in the label window, and after being triggered, the first operation control is used to replace the first media file corresponding to the reference label with a second media file; a second operation control is provided in the label window, and after being triggered, the second operation control is used to delete the reference label; a third operation control is provided in the label window, and after being triggered, the third operation control is used to modify the reference method for the media file.

[0154] According to one or more embodiments of the present disclosure, the processing module is specifically configured to: generate a reference prompt word according to the reference tag; generate a prompt word text according to the basic prompt word and the reference prompt word; the generating module is specifically configured to: input the prompt word text and the media file into an image generation model to generate the target media data.

[0155] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;

[0156] The memory stores computer-executable instructions;

[0157] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the interactive media data generation method described in the first aspect above and various possible designs of the first aspect.

[0158] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the interactive media data generation method described in the first aspect above and various possible designs of the first aspect are implemented.

[0159] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the interactive media data generation method described in the first aspect above and various possible designs of the first aspect are implemented.

[0160] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0161] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0162] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An interactive media data generation method, characterized in that, Including: In response to a first operation, load a reference label in a multimodal input box, where the first operation is at least used to select a media file, and the reference label is used to characterize the reference manner for the media file in the process of generating target media data with reference to the media file; Generate a prompt text according to the reference label and at least one basic prompt word in the multimodal input box; Generate the target media data based on the prompt text.

2. The method according to claim 1, wherein The method further includes: In response to a second operation, move the cursor in the multimodal input box to a target position, where the target position is before or after the basic prompt word; The loading the reference label in the multimodal input box in response to the first operation includes: In response to the first operation, load the reference label at the target position in the multimodal input box.

3. The method according to claim 2, characterized in that, The second operation includes a first editing sub-operation and a second editing sub-operation. The moving the cursor in the multimodal input box to the target position in response to the second operation includes: In response to the first editing sub-operation, set the multimodal input box from a first state to a second state, where when the multimodal input box is in the first state, the cursor in the multimodal input box moves in units of characters; when the multimodal input box is in the second state, the cursor in the multimodal input box moves in units of labels; In response to at least one second editing sub-operation, move the cursor in the multimodal input box to the target position.

4. The method according to claim 2, characterized in that, The first operation includes a paste operation. The loading the reference label at the target position in the multimodal input box in response to the first operation includes: In response to the paste operation, read the media file from the clipboard; Load the media file, and generate and display a reference label at the target position according to the media type of the media file.

5. The method according to claim 1, wherein The first operation includes a drag operation, and the multimodal input box is set in an image generation interface; The loading the reference label in the multimodal input box in response to the first operation includes: In response to the drag operation, display an indication mark in the image generation interface, where the indication mark is used to indicate a target area in the image generation interface; When the drag operation is released within the target area, load a reference label in the multimodal input box.

6. The method according to claim 1, characterized in that, The loading the reference label in the multimodal input box in response to the first operation includes: In response to the first operation, set the target reference manner corresponding to the reference label; Generate the reference label according to the target reference manner and the type of the media file; Load the reference label in the multimodal input box in the form of a component.

7. The method according to claim 1, wherein The reference label includes a label text and a label picture, where the label picture includes a thumbnail of the media file, and the label text is used to describe the target reference manner for the media file.

8. The method according to claim 7, wherein The method further includes: In response to a third operation on the reference label, display a label window corresponding to the reference label; In the label window, display the media file.

9. The method according to claim 8, wherein Also includes at least one of the following: A first operation control is arranged in the label window. After being triggered, the first operation control is used to replace the first media file corresponding to the reference label with a second media file; A second operation control is arranged in the label window. After being triggered, the second operation control is used to delete the reference label; A third operation control is arranged in the label window. After being triggered, the third operation control is used to modify the reference mode for the media file.

10. The method according to claim 1, characterized in that, Generating a prompt text according to the reference label and at least one basic prompt word in the multimodal input box includes: Generating a reference prompt word according to the reference label; Generating a prompt text according to the basic prompt word and the reference prompt word; Generating the target media data based on the prompt text includes: Inputting the prompt text and the media file into an image generation model to generate the target media data.

11. An interactive media data generation device, characterized in that Including: An interaction module, configured to load a reference label in the multimodal input box in response to a first operation, where the first operation is at least used to select one media file, and the reference label is used to represent the reference mode for the media file in the process of generating target media data by referring to the media file; A processing module, configured to generate a prompt text according to the reference label and at least one basic prompt word in the multimodal input box; A generation module, configured to generate the target media data based on the prompt text.

12. An electronic device, characterized in that, Including: A processor and a memory; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the interactive media data generation method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium. When the processor executes the computer execution instructions, the interactive media data generation method according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the interactive media data generation method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Interaction method and device, electronic equipment and storage medium

    CN120897092A

  • Interactive media data generation method and apparatus, electronic device, and storage medium

    WO2025148986A1