Multimodal content generation methods and apparatuses, intelligent agents, and electronic devices

CN121541807BActive Publication Date: 2026-08-14BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,这种“输入-输出”的黑箱模式,用户意图与生成结果之间存在偏差,反复试错,制作效率低,过程可控性差

Benefits of technology

[0009]根据本公开的第六方面,提供了一种计算机程序产品,包括计算机程序,该计算机程序在被处理器执行时实现根据本公开实施例中任一的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121541807B_ABST
    Figure CN121541807B_ABST
Patent Text Reader

Abstract

This disclosure provides a multimodal content generation method, apparatus, intelligent agent, and electronic device. This disclosure relates to the field of artificial intelligence technology, specifically natural language processing, computer vision, human-computer interaction, and deep learning, and can be applied to scenarios such as intelligent document generation. The specific solution is as follows: receiving user input, including user-uploaded materials and user instructions; analyzing the user input to determine the user's generation intent; based on the generation intent, outputting at least one content generation option; and responding to the user's selection of the target content generation option, generating multimodal content based at least on the generation intent and the materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to the fields of natural language processing, computer vision, human-computer interaction and deep learning, and can be applied to scenarios such as intelligent document generation. Specifically, it relates to a multimodal content generation method and apparatus, intelligent agent and electronic device. Background Technology

[0002] Currently, AI-based content generation tools generally adopt a linear "input-output" processing model. After users submit text instructions and materials, the system directly calls the model to generate results. However, this "input-output" black box model results in a discrepancy between user intent and the generated result, leading to repeated trial and error, low production efficiency, and poor process controllability. Summary of the Invention

[0003] This disclosure provides a method, apparatus, intelligent agent, and electronic device for generating multimodal content.

[0004] According to a first aspect of this disclosure, a multimodal content generation method is provided, comprising: receiving user input, the user input including user-uploaded materials and user instructions; analyzing the user input to determine the user's generation intent; outputting at least one content generation option based on the generation intent; and generating multimodal content based at least on the generation intent and materials in response to the user's selection of the target content generation option.

[0005] According to a second aspect of this disclosure, a multimodal content generation apparatus is provided, comprising: an information receiving module for receiving user input, the user input including user-uploaded materials and user instructions; an intent determination module for analyzing the user input to determine the user's generation intent; an option providing module for outputting at least one content generation option based on the generation intent; and a content generation module for generating multimodal content in response to the user's selection of a target content generation option, based at least on the generation intent and materials.

[0006] According to a third aspect of this disclosure, an intelligent agent is provided, comprising: an input module for receiving input information; a processing module for performing steps in the multimodal content generation method of the first aspect based on the input information to determine a target task and perform the target task to obtain output information; and an output module for outputting the output information obtained by the processing module.

[0007] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the methods described in the embodiments of this disclosure.

[0008] According to a fifth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.

[0009] According to a sixth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.

[0010] By adopting the solution disclosed herein, user intent is accurately identified through comprehensive analysis of materials and user instructions, and content generation options are dynamically provided based on the identification results. This allows users to participate in key decisions through selection operations, thereby avoiding misunderstanding biases from the source. Ultimately, while improving the controllability of the generation process, the efficiency and accuracy of content generation are significantly improved.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a flowchart illustrating a multimodal content generation method according to an embodiment of the present disclosure; Figure 2 This is a schematic diagram of a canvas interface for drawing according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the multimodal content generation task triggering and processing flow according to embodiments of this disclosure; Figure 4 This is a schematic diagram of the asynchronous callback and result processing flow of the generation task according to an embodiment of this disclosure; Figure 5 This is a schematic diagram illustrating the triggering and execution process of the intelligent generation function based on image materials according to an embodiment of this disclosure; Figure 6 This is a schematic diagram of the first and last frame video generation and intelligent optimization process according to an embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of a multimodal content generation apparatus according to an embodiment of the present disclosure; Figure 8 This is a schematic diagram of the structure of an intelligent agent according to an embodiment of the present disclosure; Figure 9 This is a schematic diagram of a scenario based on the multimodal content generation method according to an embodiment of the present disclosure; Figure 10This is a structural diagram of an electronic device used to implement the multimodal content generation method of the embodiments of this disclosure. Detailed Implementation

[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0014] Before introducing the technical solutions of the embodiments of this disclosure, the technical terms that may be used in this disclosure will be further explained: Doodling refers to the freehand drawing operations performed by users on a canvas interface using input devices such as styluses, mice, or touchscreens, targeting specific materials. These drawing marks, as unstructured visual input, are captured, parsed, and given explicit semantics by the system in real time, and then transformed into executable operation instructions.

[0015] In related technologies, after a user submits text instructions and materials, the system directly calls the model to generate the result. This model has significant drawbacks: First, the system only superficially processes the input content, lacking in-depth analysis and strategy matching of the user's true generation intent, leading to a serious deviation between the output and expectations. Second, the process lacks necessary user confirmation and configuration steps, making the generation process uncontrollable and forcing users to repeatedly try and fail to approximate the ideal effect. Finally, the system lacks an effective status feedback mechanism, and the front-end interface cannot reflect the back-end generation progress in real time, resulting in a fragmented user experience. These problems collectively lead to the technical dilemma of low content generation efficiency, unstable quality, and poor user experience.

[0016] To at least partially address one or more of the aforementioned problems and other potential issues, this disclosure proposes a multimodal content generation method. By introducing an intent-driven interaction mechanism and real-time status feedback, the method enhances the controllability of the content generation process, improves the overall execution efficiency from instruction input to result output, and enhances the matching degree between the output content and the expected goals and the final generation quality through accurate understanding of user intent and process intervention.

[0017] This disclosure provides a method for generating multimodal content. Figure 1This is a flowchart illustrating a multimodal content generation method according to an embodiment of the present disclosure. This method can be applied to a multimodal content generation device. The multimodal content generation device is located in an electronic device. The electronic device includes, but is not limited to, fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or ordinary servers. Mobile devices include, but are not limited to, mobile phones, tablets, vehicle terminals, etc. In some possible implementations, the multimodal content generation method can also be implemented by a processor calling computer-readable instructions stored in memory. Figure 1 As shown, the multimodal content generation method includes: S101. Receive user input, which includes user-uploaded materials and user instructions.

[0018] S102. Analyze user input to determine the user's intention to generate the data.

[0019] S103. Based on the generation intent, output at least one content generation option.

[0020] S104. In response to the user's selection of the target content generation option, generate at least multimodal content based on the generation intent and materials.

[0021] In this embodiment of the disclosure, the materials are the basic data for content generation, such as images, video clips, text, data tables, etc.

[0022] In this embodiment of the disclosure, user input refers to all information and instructions submitted to the system by the user in any way (clicking, dragging, drawing, typing, uploading files, etc.). For example, a user uploads an image to the interactive interface, draws a red circle on the image, and simultaneously enters the text instruction "turn this part red".

[0023] In this embodiment of the disclosure, user instructions are operations or text commands issued by the user to express generation needs. These can be graffiti or selection operations on materials on a canvas, or natural language text entered in a dialog box.

[0024] In some implementations, the system listens for and captures user interaction events through different interactive areas (such as canvas and dialog boxes) in the graphical user interface (GUI), including file uploads, drawing on materials, selection operations, and text input, integrating these diverse inputs into a complete processing request. This provides a more natural and accurate way to express intent, reduces user operation and communication costs, and lays the foundation for subsequent accurate intent recognition.

[0025] In this embodiment of the disclosure, the generation intent is the result of the system's understanding of the user's needs. For example, combining a user-uploaded image with the text instruction "anime style", the system identifies the generation intent as "image-to-image style conversion"; if the user selects multiple images, the system might identify the generation intent as "multi-image composite video".

[0026] In some implementations, the system invokes an intent recognition module to perform multimodal comprehensive analysis of the type and quantity of materials and the semantics of user commands. Using a pre-trained deep learning model or rule-based strategy library, the analysis results are mapped to a predefined generation strategy, thereby outputting a structured generation intent. In this way, ambiguous and diverse user input is transformed into a clear and executable generation task, fundamentally avoiding the problem of unexpected generation results due to biased intent understanding, and providing a core guarantee for achieving accurate generation.

[0027] In this embodiment of the disclosure, the content generation options are specific execution paths or configurations recommended to the user by the system based on the generation intent. For example, when the generation intent is "text to image", the system may provide options such as "realistic style" or "cartoon style"; when the generation intent is "multi-image composite video", the options may be "slideshow" or "dynamic scene transition".

[0028] In some implementations, the system retrieves one or more feasible generation paths associated with a pre-configured strategy library based on the determined generation intent, and presents them on the user interface in the form of buttons, lists, or cards for the user to choose from. In this way, control is returned to the user before final generation, allowing them to confirm and select at key decision points, effectively eliminating the uncontrollability of traditional "black box" models and improving the predictability of the generated results.

[0029] In this embodiment of the disclosure, the target content generation option is a content generation option selected by the user from multiple options provided by the system. Multimodal content is the final generated result, and its form may be the same as or different from the input modality. For example, it may be a video generated from text and images, or a new image generated after modification according to doodle instructions.

[0030] In some implementations, the system packages the user's final selection (target option), the identified generation intent, and the original materials into a structured configuration task. It then invokes a background agent (such as an Artificial Intelligence (AI) model) to schedule computing resources to execute the generation task and ultimately output the result. In this way, by driving the background agent through a standardized process that integrates the user's explicit choices, the system ensures the efficient and reliable execution of the generation task, ultimately outputting high-quality, multimodal content that meets the user's expectations.

[0031] Figure 2 A schematic diagram of doodling on the canvas interface is shown, such as... Figure 2 As shown, the user first uploads two images to the canvas interface, adds an arrow and the text "Place here," and then selects these images. The interface then presents guided options: Create a new image (like a Rubik's Cube), Generate a document, Generate a presentation (PowerPoint, PPT), Smart Summary, and Inspiration Stimulation. After selecting "Create a new image," the progress of the generation is displayed on the right side of the canvas interface. Users can engage in deeper communication with the AI ​​assistant through natural language, achieving more precise communication of needs and collaborative creation.

[0032] The technical solution of this disclosure improves the accuracy of intent recognition by comprehensively analyzing materials and user instructions; it dynamically provides content generation options based on the user intent recognition results, allowing users to participate in key decisions through selection operations, thereby avoiding misunderstanding biases from the source, and ultimately improving the controllability of the generation process while improving content generation efficiency and content quality.

[0033] In some embodiments, user instructions include drawing instructions for materials on the canvas interface and / or text instructions entered in the dialog interface.

[0034] Here, the canvas interface is a graphical interface that allows users to directly visualize and edit materials. Doodle commands are freehand marks that users make on the canvas for specific materials. For example, drawing a red circle pointing to the collar on a clothing image indicates a desire to modify that area; or marking "zoom in" next to text indicates a desire to enlarge the font.

[0035] Here, the dialog interface is an area that typically exists in the form of a chat window and receives text commands.

[0036] Here, text commands are natural language commands entered by the user in the dialog interface. For example, entering "change the background to a starry sky" or "combine these three images into a video".

[0037] In some implementations, the system uses its GUI event listening mechanism to capture the user's scribbles on the material in the canvas interface, bind and record them with the material element being applied; at the same time, it captures the text string entered by the user in the dialog interface; finally, the system integrates the instruction information from these two different modalities into unified input data that can be processed by the subsequent intent recognition module.

[0038] Thus, by supporting two complementary instruction methods—doodles and text—the system enriches the means for users to express their creative intentions: doodle instructions provide intuitive and precise spatial location indications, while text instructions offer abstract and flexible semantic description capabilities. The combination of the two enables the system to acquire more comprehensive and accurate user needs information, fundamentally overcoming the limitations of a single instruction mode (text only or doodle only), and laying the foundation for subsequently generating content that highly meets user expectations.

[0039] In some embodiments, after outputting at least one content generation option based on the generation intent, and before generating multimodal content, the method further includes: in response to determining that supplementary information needs to be input for the generation intent, outputting a configuration interface to obtain the supplementary information input for the configuration interface. Specifically, generating multimodal content based at least on the generation intent and materials in response to the user's selection of the target content generation option includes: generating multimodal content based on the generation intent, materials, and supplementary information in response to the user's selection of the target content generation option. That is, determining whether the generation intent requires supplementary information from the user; if so, presenting a configuration interface to the user to obtain the user's configuration input.

[0040] In this embodiment of the disclosure, supplementary information refers to information that the user needs to clarify or refine when the intent may be ambiguous or incomplete. For example, confirming the playback order of multiple images, or selecting a generation style from multiple styles such as "realistic style" and "anime style".

[0041] In this embodiment of the disclosure, the configuration interface is an interactive interface presented by the system to the user for collecting supplementary information. For example, a pop-up configuration card includes an image drag-and-drop sorting area, a style selection button, etc.

[0042] In this embodiment of the disclosure, the configuration input is the final decision parameter provided by the user through the configuration interface. For example, the image order list determined by the user, or the selected "ink painting style" option.

[0043] In some implementations, the system has a built-in intent analyzer that evaluates the complexity and clarity of the identified generation intent based on preset rules (such as the number of materials, the ambiguity of the instruction, and the number of optional parameters). When it is determined that the intent needs to be further clarified, the system dynamically generates and renders a matching configuration interface (such as a configuration card). This configuration interface provides necessary interactive components (such as sorters and selectors) and captures all user interaction results on this interface as structured configuration input. Finally, this input, along with the initial generation intent and materials, is sent to the downstream content generation module.

[0044] Thus, by introducing an intelligent, condition-triggered decision-making layer, an optimal balance is achieved between automated processes and user control. It avoids unnecessary interactive interference when the intent is clear, ensuring the efficiency of the basic process and simplifying operational complexity; at the same time, it proactively guides users to provide key decisions when the intent is ambiguous or multiple possibilities exist, improving the accuracy of generated content and fundamentally solving the problem of repeated trial and error caused by misunderstandings in existing technologies.

[0045] In some embodiments, supplementary information is determined to be required for the generation intent if at least one of the following conditions is met: the explicitness of the generation intent is below a preset threshold; there are multiple materials, and their generation logic depends on a specific arrangement order; the user input contains a combination of material types associated with different generation intents; or there are multiple optional styles or parameters associated with the generation intent.

[0046] In this embodiment of the disclosure, the clarity of the generated intent is a measure of the system's certainty regarding the user intent it identifies. For example, if a user inputs "optimize this image," the clarity of the intent is low (the direction of optimization is unclear); while inputting "increase the resolution of this image to 4K," the clarity of the intent is high.

[0047] In this embodiment of the disclosure, the preset threshold is a pre-set confidence threshold value used to trigger the user confirmation process.

[0048] In this embodiment of the disclosure, the generation logic depends on a specific arrangement order, meaning that the order of the materials directly affects the final generation result. For example, when generating a video from three images A, B, and C, the order A->B->C and C->B->A will tell completely different stories.

[0049] In this embodiment of the disclosure, the combination of material types is associated with different generation intentions, meaning that multiple types of materials provided by the user may correspond to multiple interpretations. For example, if a user uploads an image and a piece of text at the same time, this combination may be associated with the intention of "modifying the image based on the text" or the intention of "generating a text description for the image".

[0050] In this embodiment of the disclosure, the optional styles or parameters are different visualization or execution methods to achieve the same generation intent. For example, for the intent of "text-to-image", optional styles include "realistic", "cartoon", "ink painting", etc.; parameters may include "image ratio" and "generation quantity".

[0051] In some implementations, the system uses a rule engine or lightweight classification model to match the features of the current generation task (such as the semantic ambiguity of user instructions, the number and type of materials, and the number of available style options) with preset triggering conditions in real time. If any condition is met, the user confirmation process is automatically triggered, and the corresponding configuration interface is dynamically generated and presented to collect key decision information.

[0052] Thus, by introducing a multi-dimensional intelligent judgment mechanism, the accuracy and reliability of content generation are improved. By assessing the clarity of the generation intent, the system can effectively intercept ambiguous or low-confidence generation requests, preventing low-quality output due to insufficient understanding at the source and ensuring basic generation quality. For scenarios with multiple material dependencies, the system can identify and ensure that materials with temporal or logical relationships are correctly arranged, avoiding the core defect of illogical generation results due to disordered order. When the input contains a combination of multiple types of materials, the system can accurately identify semantic ambiguities and guide the user to clarify the primary and secondary relationships through proactive questioning, improving the accuracy of intent recognition. By providing multiple selectable styles and parameters, the system can precisely control the generation process based on the user's explicit choices, optimizing the matching degree between the output content and the target scene, thereby improving content quality.

[0053] In some embodiments, the configuration interface is displayed in the form of configuration cards, and its display position is adaptively adjusted according to the current operation interface context, wherein the operation interface context includes a full-screen canvas interface or an embedded dialog interface.

[0054] In this embodiment, the configuration card is a specific configuration interface presented as a floating card, with clear boundaries and a structured layout. For example, it could be a pop-up user interface (UI) component that includes an image sorting area, a style selection button, and a confirmation operation.

[0055] In this embodiment, the user interface context is the user's current primary interactive environment. For example, a full-screen canvas interface is the main design area that the user is focusing on editing, such as the full-screen mode of PPT editing. An embedded dialog interface is a chat window attached to the side or bottom of the main interface, such as the expandable / collapseable AI assistant sidebar in design tools.

[0056] In this embodiment, adaptive adjustment means that the system automatically changes the display logic and position of the configuration card according to the context.

[0057] In some implementations, the system monitors user interaction events and interface status to determine the current active context in real time. When a configuration card needs to be displayed, the system selects the corresponding display scheme from predefined layout strategies based on this context information. If it is in a full-screen canvas interface, it calculates and positions the card at a suitable position on the canvas near the manipulated material. If it is in an embedded dialogue interface, it renders the card as a message bubble in the dialogue flow, thereby achieving intelligent matching between the display position and the interaction scenario.

[0058] In this way, by dynamically adjusting the placement of configuration cards to match the user's operating scenario, the pain point of traditional pop-ups being detached from the user's current task context is fundamentally solved. It not only significantly reduces unnecessary mouse movements and eye jumps, but also improves the overall efficiency of complex content generation tasks and the smoothness of the user experience by maintaining the continuity of the interaction context.

[0059] In some embodiments, the display position is adaptively adjusted according to the current operating interface context, including: in response to determining that the current operating interface context is a dialog interface, presenting the configuration card in the dialog interface; in response to determining that the current operating interface context is a canvas interface, presenting the configuration card on the canvas and placing it close to the material selected by the user.

[0060] In this embodiment of the disclosure, the material selected by the user is one or more target objects specified by the user on the canvas through clicking or selecting. For example, three images highlighted with blue borders on the canvas.

[0061] In some implementations, the system uses the context-aware module of the UI framework to obtain the current focus state of the interface in real time. When it is determined that the current focus is on the dialog interface, the configuration card is inserted as a special system message at the end of the dialog flow for rendering. When it is determined that the current focus is on the canvas interface, the screen coordinates of the selected material are obtained by calculation, and the configuration card is rendered in a suitable area near the coordinates (such as the upper right or lower right of the material) using absolute positioning.

[0062] In this way, by implementing differentiated and intuitive layout strategies in different operating scenarios, the configuration cards are precisely aligned with the user's attention, improving the efficiency of complex tasks and the overall smoothness of the user experience.

[0063] In some embodiments, the agent provides the user with at least one of the following configuration functions through configuration cards: the ability to sort multiple materials by dragging and dropping; the ability to eliminate semantic ambiguity in the user's instructions by selecting a button; and the ability to select a content generation style or parameters through an option list. The multimodal content generation method further includes: in response to receiving an input operation for the configuration card, the agent performs at least one of the following operations: adjusting the arrangement order of multiple materials according to the user's drag-and-drop instructions; eliminating semantic ambiguity in the original instructions by the user's selection of explicit options; and determining specific style parameters for content generation based on the user's selection in a preset list.

[0064] In this embodiment of the disclosure, users are allowed to adjust the order of materials by dragging and dropping. For example, a user can drag and drop three images in the list of configuration cards to arrange them in the order of cover image, first content image, and second content image to generate a short video.

[0065] In this embodiment of the disclosure, explicit options are provided to clarify the user's ambiguous intent. For example, if a user uploads both a landscape image and a poem, the configuration card displays two selection buttons: "Match the image with the poem's mood" and "Generate an image to match the poem." The user can click on one of these to eliminate ambiguity.

[0066] In this embodiment of the disclosure, the specific style parameters are generated attributes presented in a list format that are selectable by the user. For example, a drop-down list or button group containing options such as "realistic photography," "cartoon illustration," and "watercolor style" is used to select the visual style of the generated image.

[0067] In some implementations, the system dynamically generates internal components of the configuration card based on the results of the aforementioned judgment conditions: for sequentially dependent tasks, a draggable list of materials is rendered; for tasks with ambiguous intents, a set of mutually exclusive selection buttons is generated, each button corresponding to a clear intent interpretation; for style parameter selection tasks, optional values ​​are pulled from a predefined resource library and rendered as an option list (such as drop-down menus or button groups); finally, the system captures the user's structured configuration input by listening to the user's interaction events on these components.

[0068] In this way, by integrating multiple targeted interactive components into a unified configuration card, an efficient and accurate "intent calibration" hub is built. It concretizes the user's control at key decision points into executable interactive actions, transforming vague user input into precise machine instructions. This not only reduces the generation failure rate caused by misunderstandings or missing information, but also, by giving users a strong sense of control, ensures that the final generated content highly matches user expectations in terms of logic and accuracy, thereby improving the overall reliability, usability, and user satisfaction of content generation.

[0069] In some embodiments, user input is received, triggered by any of the following operations: in response to the user performing a first operation on the canvas interface, the first operation including a selection operation on at least one material, or a drawing operation on a material and a subsequent selection confirmation operation; in response to the user performing a second operation on the dialog interface, the second operation being an input text instruction.

[0070] In this embodiment of the disclosure, the first operation is a physical action in which the user directly interacts with the material on the canvas. For example, the user selects one or more elements such as images or text boxes by clicking or selecting with the mouse. Another example is that the user first draws a circle on an image (a doodle operation), and then clicks the "magic generate" button that appears in the upper right corner of the image (a confirmation operation).

[0071] In this embodiment of the disclosure, the second operation is a user interaction action in the dialog interface. For example, the user types "Help me turn this picture into a watercolor style" in the chat box and presses the Enter key.

[0072] In some implementations, the system captures user mouse clicks, selection areas, or touch strokes on the canvas interface through a graphical event listener. When an operation that meets preset conditions is detected (such as selecting a material or triggering a generate button after drawing on a material), the system packages the operation and the target material into a generate request. At the same time, the system captures the user's plain text commands through the form submission event of the text input box on the dialog interface. The system maintains a unified task trigger to receive and initialize requests from these two independent interaction channels.

[0073] Thus, by supporting two complementary triggering paradigms—direct canvas manipulation and dialog text commands—a unified generation entry point covering needs from concrete editing to abstract conceptualization is constructed. It satisfies users' pursuit of efficient and precise operational intuition in visual editing scenarios while ensuring flexibility in expressing complex intentions in free creation scenarios. This allows it to adapt to the natural working habits of users with different skill backgrounds in various application scenarios, significantly improving the system's usability and applicability.

[0074] In some embodiments, analyzing user input to determine the user's generation intent includes: matching a generation strategy from a predefined strategy library based on the type and quantity of materials and in conjunction with the semantics of the user's instructions; wherein the generation strategy includes at least one of: text-to-image, text-to-video, image-to-image, multi-image composite video, and mixed text-image editing.

[0075] In this embodiment of the disclosure, the type and quantity of materials are basic attributes of user-provided content. For example, the type includes, but is not limited to, images, videos, text, and audio. The quantity includes, but is not limited to, a single image, a sequence of multiple images, and a mixture of images and text.

[0076] In this embodiment of the disclosure, the semantics of the user instruction is the deeper meaning of the natural language expression input by the user. For example, the instruction "make the image more vibrant" semantically refers to "color enhancement" and "contrast enhancement".

[0077] In this embodiment of the disclosure, the predefined strategy library is a collection of generation capabilities and rules built into the system. For example, it may be a database or configuration file containing processing logic for various generation scenarios.

[0078] In this embodiment of the disclosure, the generation strategy is a specific processing scheme selected by the system. For example, text-to-image: generating a corresponding image based on the text "a cat in a suit"; text-to-video: generating a short video based on the text "the process of sunrise"; image-to-image: performing style conversion, repair, or editing based on the original image; multi-image video synthesis: merging multiple travel photos into a dynamic photo album; mixed text and image editing: adding artistic text to images or modifying the text in images.

[0079] In some implementations, the system uses a multimodal understanding engine to simultaneously analyze the metadata of the source files (such as image size and video duration) and the textual features of user instructions. It uses natural language processing technology to parse the semantics of the instructions, combines computer vision technology to identify the content features of the source files, and then fuses this multi-dimensional information. It performs weighted matching in a predefined policy library and outputs the most suitable generation policy identifier and its confidence score.

[0080] In this way, by comprehensively analyzing the deep semantics of material attributes and user instructions, the optimal generation path can be quickly and accurately located from a massive number of possibilities. This effectively solves the problem of inconsistent generation results caused by the misunderstanding of intent in traditional solutions, and improves the accuracy of content generation.

[0081] In some embodiments, in response to the user's selection of the target content generation option, multimodal content is generated based at least on the generation intent and materials, including: providing generation process feedback simultaneously in the canvas interface and the dialog interface; wherein, in the canvas interface, the output is a generation placeholder with a loading animation displayed at the target location; and in the dialog interface, the output is a status card that updates the generation status in real time.

[0082] In this embodiment of the disclosure, the generation process feedback refers to real-time visual information about the task status provided to the user by the system during content generation. Generation placeholders are temporary visual elements reserved on the canvas to represent content that is about to be generated. For example, a gray rectangle with a loading animation, whose size and position are consistent with the final generated image or video. The loading animation represents the dynamic visual effect of the system processing the task. For example, a progress ring, a pulsating effect, or a skeleton screen displayed on the placeholder.

[0083] In this embodiment of the disclosure, the status card is a structured UI component that displays task progress information in the dialog interface. For example, a message bubble displaying text such as "Generating...65%" or "Processing your video" along with a progress bar.

[0084] In some implementations, once the system confirms that a task has started, it broadcasts a status update event to all active UI components (canvas rendering engine and dialog interface manager). Upon receiving the event, the canvas interface creates and renders a placeholder element with a loading animation at the pre-calculated target position. Simultaneously, the dialog interface inserts a new status card component into the message stream and establishes a long connection with the background task status, enabling it to receive progress updates and refresh the displayed content in real time.

[0085] In this way, by providing complementary and context-adaptive visual feedback in different interface scenarios, the spatial stability and layout predictability of the creation scene are maintained in the form of placeholders in the canvas interface, while a traceable task progress timeline is constructed through status cards in the dialogue interface. Together, the "black box" state of the traditional generation process is transformed into a transparent, visible, and perceptible continuous operation experience, effectively eliminating the uncertainty and anxiety of users during the waiting process, and significantly improving the trust and collaboration efficiency of human-computer interaction.

[0086] In some embodiments, the size of the generated placeholder is set based on a pre-calculated size of the multimodal content. The multimodal content size is the physical size of the multimodal content that the system expects to output and display on the canvas. For example, the display area on the canvas for a 1024×768 pixel image or a 1920×1080 resolution, 15-second video.

[0087] In this embodiment of the disclosure, the size estimation process is expected to be performed before the content is actually generated.

[0088] In some implementations, after receiving a generation task, the system quickly calculates the size of the final result in the canvas coordinate system before the content actually begins to be generated, based on the task type (such as image generation, video compositing), user-preset output parameters (such as resolution, aspect ratio), and the current scaling ratio and layout constraints of the canvas, and immediately applies this size to the generation placeholder to be rendered.

[0089] In this way, by pre-calculating and precisely reserving visual space that is completely consistent with the physical size of multimodal content, the page layout jumps and visual fragmentation caused by dynamic content insertion are fundamentally eliminated. This not only ensures the visual stability and professionalism of the creation interface, but also significantly improves the user's focus and operational continuity during the content generation process by establishing accurate visual expectations, thereby improving content generation efficiency.

[0090] In some embodiments, after the task of generating multimodal content is completed, the method further includes: replacing the generated placeholder with multimodal content in the canvas interface; and updating the generated status card to a result preview card for presenting the multimodal content in the dialog interface.

[0091] In this embodiment, the result preview card is a special UI component used to display the final generated result. For example, it may be an interactive card containing thumbnails, a download button, and operation options.

[0092] In some implementations, when the system receives a notification that the task is completed, the canvas interface will locate the corresponding generation placeholder and smoothly replace it with multimodal content obtained from the storage service; at the same time, the dialog interface will update the corresponding generation status card to a result preview card, which will load a thumbnail or preview link of the result and reconfigure the relevant interactive operation buttons.

[0093] In this way, the generated content and the creation environment are seamlessly integrated in the canvas interface, ensuring a zero-disruption experience for the user's workflow. At the same time, an interactive results management center is established in the dialog interface, forming a complete closed loop from process tracking to result application. This effectively solves the core pain point of the disconnect between result delivery and usage scenarios in traditional content generation tools, and improves creation efficiency and operational smoothness.

[0094] In some embodiments, the result preview card association provides at least one follow-up action option, which includes at least one of the following: add to canvas, regenerate, download, copy.

[0095] In this embodiment of the disclosure, the subsequent operation options are further processing functions associated with the generated results.

[0096] In this embodiment of the disclosure, adding to the canvas means placing the generated result material into the canvas workspace. (Adding to the canvas) provides a seamless connection from the dialog interface to the canvas workspace, allowing users to directly input the generated results into subsequent creative processes without manual import, thus improving workflow continuity.

[0097] In this embodiment, regeneration is performed by executing the generation task again with the same parameters. Regeneration allows for rapid iterative optimization. When the user is not satisfied with the initial result, an alternative solution can be obtained without re-entering instructions, improving trial-and-error efficiency and result satisfaction.

[0098] In some implementations, when the system renders the result preview card, it simultaneously generates a set of operation button components that match the current generated task type. These buttons are embedded at the bottom or side of the card in the form of a horizontal layout or drop-down menu, and a corresponding event handling function is bound to each button. When the user clicks, the corresponding business logic is triggered (such as calling the Canvas Application Programming Interface (API) to add elements, resubmitting the generated task, calling the download interface, or operating the clipboard).

[0099] In this way, high-frequency subsequent operations are aggregated into the result preview scenario, creating a closed-loop experience of "generation-preview-application". This effectively solves the problems of scattered result usage paths and lengthy operation links in the traditional process. It achieves seamless connection of the creation process through "add to canvas", provides low-cost iterative optimization capabilities through "regeneration", and covers the basic needs of content export by relying on "download / copy". This significantly reduces the user's cognitive load and the frequency of step jumps, and comprehensively improves the actual application efficiency after content generation.

[0100] In some embodiments, in response to a user's selection of a target content generation option, multimodal content is generated based at least on the generation intent and materials. Specifically, this includes: constructing a configuration task based on the task type, generation intent, and materials defined by the target content generation option; and invoking an agent to perform content generation operations according to the configuration task to obtain multimodal content.

[0101] In this embodiment of the disclosure, the intelligent agent is a core software module responsible for coordinating and managing the lifecycle of the content generation task. The intelligent agent can be an independent, specialized content generation module. For example, a microservice specifically running a large-scale text-to-image model, or a cloud function specifically performing video compositing.

[0102] In this embodiment of the disclosure, the configuration task is a structured task definition data that encapsulates all the necessary parameters required to generate the execution content, and is used to drive the agent to perform specific operations.

[0103] In some implementations, after the user confirms the generation options on the front end, the front end service sends a request to the task processing system. The system first registers with the task management service, obtains a unique task identifier, and immediately creates a record (pre-placement) with a status of "waiting" for the task in the task processing queue. Subsequently, the system assembles a structured configuration task object based on the target content generation options, generation intent, and material information contained in the front end request. Finally, the system submits the configuration task to an independent intelligent agent through an API call and begins to monitor its processing status.

[0104] In this way, by integrating the user-confirmed target options, generation intentions, and materials into a structured configuration task, the generation instructions are standardized and made more precise. Then, the intelligent agent executes the content generation operation according to the configuration task, which improves the efficiency and quality of content generation.

[0105] In some embodiments, invoking an agent to perform content generation operations according to a configuration task includes: invoking at least two agents to execute multiple configuration tasks in parallel, wherein the configuration tasks executed in parallel include at least two different types of tasks among text-to-image, image-to-image, image processing, video generation, and document generation.

[0106] In some implementations, after receiving the target content generation options, the system parses the task type characteristics contained in the options using predefined routing rules or classification models. Based on the parsing results, it then distributes the task request and its contextual information (such as generation intent, materials, etc.) to corresponding dedicated processing branches. Each branch has its own independent parameter verification, resource allocation, and task execution logic. For example, the text-to-image branch specifically handles the text-to-image generation process, the image-to-image branch specifically handles the image-to-image conversion process, and the video generation branch specifically handles the video compositing and editing process.

[0107] In this way, by scheduling multiple agents to execute at least two heterogeneous tasks such as text-to-graph and graph-to-graph in parallel, dynamic optimization of system resources and efficient parallel processing of heterogeneous computing loads are achieved. This not only solves the problems of resource idleness and processing bottlenecks under a single task queue, but also significantly improves the throughput efficiency of multimodal content generation and the overall system performance through specialized task routing.

[0108] In some embodiments, before invoking the agent to execute the configuration task, the method further includes: performing a legality check on the task, which includes at least one of user identity verification or traffic quota verification. Legality verification is the process of verifying the compliance of the request before task execution. User identity verification is a check to verify user permissions. For example, it verifies the validity of the user's API key, checks whether the user account is in a normal state, and confirms whether the user has the right to use specific generation functions.

[0109] In some implementations, the system verifies the submitter's identity credentials and account status by calling the user service interface, and simultaneously queries the quota management system to check the user's remaining usage quota. The verification process is carried out in real time and synchronously, and if any verification fails, the process will be terminated immediately and a clear error message will be returned.

[0110] In this way, by embedding a legality verification mechanism into the critical path of task execution, a comprehensive service guarantee system is built. This system maintains the security boundary of the system through identification verification and ensures the rationality and sustainability of resource allocation through quota control.

[0111] In some embodiments, before invoking the agent to execute the configuration task, the method further includes: persistently storing the materials and configuration task information to a storage center. The storage center is a unified service or component in the system responsible for data persistence. Examples include cloud storage services, distributed databases, and file server clusters.

[0112] In some implementations, after completing the configuration task construction and validity verification, the system treats the configuration task information and associated material files (or their unique index in the content library) as a complete data record and writes them to the specified persistent storage device in an atomic operation by calling the API interface provided by the storage center, ensuring that the record contains key metadata such as task identifier (ID) and timestamp.

[0113] In this way, the persistent storage step establishes a reliable data snapshot and recovery base point for task execution, ensuring that even if a system interruption or service anomaly occurs during subsequent generation, the key information and context of the task will not be lost, thus providing a solid data foundation for realizing task breakpoint resume, status verification and post-event auditing.

[0114] In some embodiments, invoking an agent to perform a content generation operation according to a configured task includes: obtaining the task status returned by the agent performing the content generation operation; generating status update information based on the task status, the status update information including at least the task progress percentage and a result identifier indicating whether the task was successful or failed; and outputting the status update information through an interactive interface.

[0115] In this embodiment of the disclosure, the task status is an identifier of the stage the generation task is in during its lifecycle. Examples include "Processing," "Success," and "Failure." Status update information is a structured message carrying task status data. Examples include "Generating image..." and "Generating...25%."

[0116] In some implementations, after submitting the configuration task, the system starts a state listener. This listener continuously obtains the latest status of the task by polling the state query interface of the intelligent agent or subscribing to the state change events published by it. Whenever a state change is detected, the system converts the raw state data into structured update information containing progress percentage and result identifier according to predefined mapping rules, and sends it to the front-end interactive interface in real time through push technology. After receiving the information, the front-end calls the corresponding UI update function to dynamically adjust the visual elements displayed to the user.

[0117] In this way, by establishing a standardized status data acquisition and feedback link, the internal task status of the system is transformed into structured information containing progress quantification indicators and result identifiers, realizing the full-link observability of the generation process. This provides data support for task scheduling optimization, anomaly diagnosis, and dynamic resource allocation, thereby significantly improving the controllability and stability of system operation.

[0118] Figure 3 A schematic diagram of the multimodal content generation task triggering and processing flow is shown.

[0119] S301: User input detected.

[0120] S302: Perform intent recognition. The system will perform intent recognition for the user's click operation.

[0121] S303: Based on the identified intent, the task is routed. For example, the task is routed to different processing channels such as "magic image" or "magic video".

[0122] S304: Task parameter construction. S304 may include S304a, S304b, S304c, and S304d.

[0123] S304a: Material Processing. Specifically, based on the attachment list passed in from the front end, the system calls the content selection service to obtain the original image Uniform Resource Locator (URL) and populates it into the file_url field.

[0124] S304b: Graffiti Information Processing. Specifically, based on the screenshot object passed from the front end, the system calls the object storage service to obtain the screenshot URL of the canvas graffiti and populates it into the image_url field.

[0125] S304c: Element information population. Specifically, the system populates all element information into the extended field.

[0126] S304d: Scene settings. Specifically, the system sets the scene type (scence_type) to "wk" and sets the scene to "image editing" or "image-to-video" according to the task type.

[0127] S305: Submit the task and return the task ID. After the parameters are constructed, the system submits the image or video generation task and obtains a unique task ID.

[0128] S306: Task Status Tracking. The system continuously polls the task center using a timer (ticker) to check the task status. Branching is performed based on the polled task status. If a task fails, the process proceeds to the "Message Failure" branch. If the task completes, the generated image or video information is retrieved from storage, and the result is returned to the front-end user interface via a message frame (WriteMessageFrame). The session information is saved, and the process ends. It should be noted that before returning the result to the front-end user interface and saving the session information, there is a "User / Video Triggered Parsing?" checkpoint in the process. Under certain conditions, a parsing task needs to be submitted for further analysis.

[0129] S307: Save session information.

[0130] S308: Returns the result to the front-end user interface.

[0131] This process ensures full-link automation and controllability, from user intent recognition, task parameter preparation, background asynchronous processing to final result return.

[0132] Figure 4 The diagram illustrates the asynchronous callback and result processing flow for task generation. After completing the task, the agent notifies the system of the task status and result through a dedicated callback interface.

[0133] S401: Receive task completion signal.

[0134] S402: Parse the content of the callback request and extract key information, such as task ID (taskId), task status (status), temporary access address of the generated result, or error information.

[0135] S403: Perform risk control security verification on the parsed callback data. If the risk control is successful, execute S404; if the risk control is unsuccessful, execute S405. S404: Save the risk control violation information for this task to the database and end the process directly.

[0136] S405: Based on the callback data, distinguish the final content type of the task. The task is split into three processing branches: Task Analysis Content: Process the generated metadata or text analysis results; Image Generation: Process image-related generation results; Video Generation: Process video-related generation results.

[0137] S406: Update the final result of the task (including the new address after transfer) to the database. For image and video results, the system will perform a transfer operation. That is, the result file returned by the agent and stored in a temporary location will be migrated to the business's own stable storage center and a new, persistent access address will be obtained.

[0138] This process ensures that the results of the generated tasks can be received and processed securely, reliably, and asynchronously, and guarantees content compliance through risk control verification and long-term availability of the results through resource transfer.

[0139] Figure 5 This diagram illustrates the triggering and execution flow of the intelligent image generation function. The user selects one or more images on the canvas. The system first determines if the selected images include any images. If not, the magic generation function is not triggered. If images are included, the system triggers and displays a list of available magic generation functions for the user to choose from. If the user selects "Magic Image Generation," the system uses AI services such as Tongyi Visual Understanding to perform deep analysis and understanding of the image. Based on the understanding results, it generates an image and outputs a new image. If the user selects "Magic Video," the system first determines the number of images selected and proceeds to different sub-processes based on the number: If it's a single image: the "Single Image Animation" function is executed, converting the static image into a dynamic video. If it's two images: the "First and Last Frame Video Synthesis" function is executed, using the two images as the beginning and end of a video to generate a transitional video. If it's three or more images: the "Multi-Frame Synthesis" function (which can be understood as splicing multiple first and last frames) is executed, generating a more complex video. For multiple images (two or more), the system will provide a default order (such as the order in which they were uploaded) and allows users to manually adjust the image order. After the user confirms the order, the system will finally execute the video generation task.

[0140] Figure 6 The diagram illustrates the video generation and intelligent optimization process using the first and last frame sequences. The system confirms the use of the first and last frame sequence mode for video generation, with two images as input. The system first checks for any additional user-provided queries or annotations. If no queries / annotations are found, a default completion order (e.g., image upload order) is used as the video's frame sequence, but users can still manually adjust this order. If queries / annotations are found, the system enters a strategy analysis phase, deeply analyzing the semantics of the user's commands to determine a more logical or user-expected first and last frame order. Regardless of whether it's the default or analyzed order, the system presents the proposed order to the user for final confirmation. After user confirmation, the system generates the video according to the (default or analyzed) order. After video generation, the system asks the user if they want to regenerate. If the user selects "Yes": the system performs a crucial optimization operation—swapping the first and last frame order and regenerating. If the user selects "No": the process ends, and the video generation task is complete.

[0141] This disclosure provides a multimodal content generation apparatus, such as... Figure 7As shown, the device may include: an information receiving module 701 for receiving user input, which includes user-uploaded materials and user instructions; an intent determination module 702 for analyzing user input to determine the user's generation intent; an option providing module 703 for outputting at least one content generation option based on the generation intent; and a content generation module 704 for generating multimodal content in response to the user's selection of a target content generation option, based at least on the generation intent and materials.

[0142] In some embodiments, user instructions include drawing instructions for materials on the canvas interface and / or text instructions entered in the dialog interface.

[0143] In some embodiments, the device further includes: an intent determination module 705. Figure 7 (Not shown in the image) is used to output a configuration interface to obtain supplementary information input for the configuration interface after the option providing module 703 outputs at least one content generation option and before the content generation module 704 generates multimodal content, in response to determining that supplementary information is required for the generation intent; wherein, the content generation module 704 is further used to generate multimodal content based on the generation intent, materials and supplementary information in response to the user's selection of the target content generation option.

[0144] In some embodiments, the conditions upon which the intent determination module 705 makes its determination include at least one of the following: the clarity of the generation intent is lower than a preset threshold; there are multiple materials, and their generation logic depends on a specific arrangement order; the user input contains a combination of material types that are associated with different generation intents; and there are multiple optional styles or parameters associated with the generation intent.

[0145] In some embodiments, the configuration interface provided by the configuration interface module 706 is displayed in the form of a configuration card, and its display position is adaptively adjusted according to the current operation interface context, wherein the operation interface context includes a full-screen canvas interface or an embedded dialog interface.

[0146] In some embodiments, the configuration interface module 706 is further configured to: in response to determining that the operation interface context is a dialog interface, present a configuration card in the dialog interface; in response to determining that the operation interface context is a canvas interface, present a configuration card on the canvas and place it close to the material selected by the user.

[0147] In some embodiments, the content generation module 704 is further configured to: in response to receiving an input operation for the configuration card input, perform at least one of the following operations: adjust the arrangement order of multiple materials according to the user's drag-and-drop instructions; eliminate semantic ambiguity in the original instructions by the user's selection of explicit options; and determine specific style parameters for content generation based on the user's selection in a preset list.

[0148] In some embodiments, the receiving operation of the information receiving module 701 is triggered by any of the following operations: in response to the user performing a first operation on the canvas interface, the first operation including a selection operation on at least one material, or a drawing operation on the material and a subsequent selection confirmation operation; or in response to the user performing a second operation in the dialog interface, the second operation being an input text command.

[0149] In some embodiments, the intent determination module 702 is further configured to: match a generation strategy from a predefined strategy library based on the type and quantity of the material and in conjunction with the semantics of the user instruction; wherein the generation strategy includes at least one of: text-to-image, text-to-video, image-to-image, multi-image composite video, and mixed text-image editing.

[0150] In some embodiments, the device further includes: a feedback module 707 ( Figure 7 (Not shown in the image) is used to provide synchronous feedback on the generation process in both the canvas and dialog interfaces during the generation of multimodal content; in the canvas interface, the output is a generation placeholder with a loading animation displayed at the target location; in the dialog interface, the output is a status card that updates the generation status in real time.

[0151] In some embodiments, in the feedback module 707, the size of the generated placeholder is set based on a pre-calculation of the multimodal content size.

[0152] In some embodiments, the feedback module 707 is further configured to, upon completion of the task of generating multimodal content, replace the generated placeholder with multimodal content in the canvas interface; and update the generated status card to a result preview card for presenting the multimodal content in the dialog interface.

[0153] In some embodiments, the result preview card association provides at least one follow-up action option, which includes at least adding to the canvas, regenerating, downloading, or copying.

[0154] In some embodiments, the content generation module 704 is used to: construct a configuration task based on the task type, generation intent, and materials defined by the target content generation options; and call an agent to perform content generation operations according to the configuration task to obtain multimodal content.

[0155] In some embodiments, the content generation module 704 is further configured to: invoke at least two agents to execute multiple configuration tasks in parallel, wherein the configuration tasks executed in parallel include at least two different types of tasks among text-to-image, image-to-image, image processing, video generation, and document generation.

[0156] In some embodiments, the content generation module 704 is further configured to obtain the task status of the agent performing the content generation operation; generate status update information based on the task status, the status update information including at least the task progress percentage and the result identifier of task success or failure; and output the status update information through the interactive interface.

[0157] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0158] The multimodal content generation device in this embodiment can improve the accuracy of intent recognition, enhance the controllability of the generation process, increase generation efficiency, and improve the accuracy of the content.

[0159] This disclosure provides an intelligent agent, such as... Figure 8 As shown, the intelligent agent includes: an input module 801 for receiving input information; a processing module 802 for executing steps in the multimodal content generation method based on the input information received by the input module 801, to determine the target task and execute the target task to obtain output information; and an output module 803 for outputting the output information obtained by the processing module 802. In some embodiments, the processing module 802 is configured to call an external content generation API or model service when executing the target task. In some embodiments, the intelligent agent is encapsulated as an independently deployable software service, plugin, or cloud API.

[0160] This disclosure provides a scenario illustration of a multimodal content generation method, such as... Figure 9 As shown.

[0161] As previously described, the multimodal content generation method provided in this disclosure is applied to electronic devices. These electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.

[0162] Specifically, the electronic device may perform the following operations: receive user input, which includes user-uploaded materials and user instructions; analyze the user input to determine the user's generation intent; output at least one content generation option based on the generation intent; and generate multimodal content in response to the user's selection of the target content generation option, based at least on the generation intent and materials.

[0163] It should be understood that Figure 9 The scene diagrams shown are merely illustrative and not restrictive; those skilled in the art can interpret them based on... Figure 9 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.

[0164] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0165] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0166] Figure 10 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0167] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0168] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0169] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as multimodal content generation methods. For example, in some embodiments, the multimodal content generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the multimodal content generation method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform a multimodal content generation method by any other suitable means (e.g., by means of firmware).

[0170] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0171] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0172] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0173] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0174] The systems and technologies described herein can be implemented in computing systems that include components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such components, middleware components, or front-end components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0175] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0176] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0177] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating multimodal content, comprising: Receive user input, which includes user-uploaded materials and user instructions; Analyze the user input to determine the user's intention to generate the content; Based on the stated generation intent, at least one content generation option is output; In response to the user's selection of the target content generation option, multimodal content is generated based at least on the generation intent and the material; After outputting at least one content generation option based on the generation intent, and before generating multimodal content, the method further includes: In response to determining that supplementary information is required for the generation intent, a configuration interface is output to obtain the supplementary information input for the configuration interface; the configuration interface is displayed in the form of a configuration card, and its display position is adaptively adjusted according to the current operation interface context, wherein the operation interface context includes a full-screen canvas interface or an embedded dialog interface. In response to the user's selection of the target content generation option, multimodal content is generated based at least on the generation intent and the material, including: In response to the user's selection of the target content generation option, the multimodal content is generated based on the generation intent, the source material, and the supplementary information.

2. The method according to claim 1, wherein, The user instructions include drawing instructions for the material on the canvas interface and / or text instructions entered in the dialog interface.

3. The method according to claim 1, wherein, The supplementary information is determined to be required for the generation intent if at least one of the following conditions is met: The clarity of the generated intent is lower than a preset threshold; The materials are multiple, and their generation logic depends on a specific arrangement order; The user input includes combinations of material types, which are associated with different generation intentions; There are several optional styles or parameters associated with the stated generation intent.

4. The method according to claim 1, wherein, The display position is adaptively adjusted according to the current operating interface context, including: In response to determining that the context of the operation interface is a dialog interface, the configuration card is presented in the dialog interface; In response to determining that the context of the operation interface is a canvas interface, the configuration card is presented on the canvas and placed close to the material selected by the user.

5. The method according to claim 1, wherein, The method further includes: In response to receiving an input operation for the configuration card, perform at least one of the following operations: Adjust the order of multiple materials according to the user's drag-and-drop instructions; Eliminate semantic ambiguity in the original instruction by allowing the user to select explicit options; The specific style parameters for content generation are determined based on the user's selection from a preset list.

6. The method according to claim 1, wherein, The receipt of user input is triggered by any of the following operations: In response to the user performing a first operation on the canvas interface, the first operation includes selecting at least one material, or drawing on the material and then confirming the selection. or In response to the user performing a second operation in the dialog interface, the second operation being inputting a text command.

7. The method according to claim 1, wherein, The analysis of the user input to determine the user's generation intent includes: Based on the type and quantity of the material, and in conjunction with the semantics of the user instruction, a generation strategy is matched from a predefined strategy library; The generation strategy includes at least one of the following: text-to-image, text-to-video, image-to-image, multi-image composite video, and mixed text-image editing.

8. The method according to claim 1, wherein, In response to the user's selection of the target content generation option, multimodal content is generated based at least on the generation intent and the material, including: The generation process feedback is output synchronously in both the canvas interface and the dialog interface; In the canvas interface, the output format is to display a generated placeholder with a loading animation at the target location; In the dialog interface, the output format is a status card that updates the generated status in real time.

9. The method according to claim 8, wherein, The size of the generated placeholder is set based on a pre-calculation of the size of the multimodal content.

10. The method according to claim 8, wherein, After generating the multimodal content, the method further includes: In the canvas interface, replace the generated placeholder with the multimodal content; In the dialog interface, the generated status card is updated to a result preview card for presenting the multimodal content.

11. The method according to claim 10, wherein, The result preview card includes at least one follow-up action option, which includes at least one of the following: add to canvas, regenerate, download, or copy.

12. The method according to claim 1, wherein, In response to the user's selection of the target content generation option, generating multimodal content, based at least on the generation intent and the material, includes: Based on the task type defined by the target content generation options, the generation intent, and the materials, a configuration task is constructed; The intelligent agent is invoked to perform content generation operations according to the configured task, thereby obtaining the multimodal content.

13. The method according to claim 12, wherein, The invoking agent performs content generation operations according to the configured task, including: Invoke at least two agents to execute multiple configuration tasks in parallel, wherein the configuration tasks executed in parallel include at least two different types of tasks among text-to-graph, graph-to-graph, image processing, video generation, and document generation.

14. The method according to claim 12, wherein, The invoking agent performs content generation operations according to the configured task, including: Obtain the task status of the intelligent agent performing the content generation operation; Based on the task status, status update information is generated, which includes at least the task progress percentage and a result identifier indicating whether the task was successful or failed. The status update information is output through the interactive interface.

15. A multimodal content generation apparatus, comprising: The information receiving module is used to receive user input, which includes user-uploaded materials and user instructions; An intent determination module is used to analyze the user input to determine the user's generation intent; An option providing module is configured to output at least one content generation option based on the generation intent; A content generation module is used to generate multimodal content in response to the user's selection of a target content generation option, based at least on the generation intent and the material. The device also includes: An intent determination module is configured to, after the option providing module outputs at least one content generation option and before the content generation module generates multimodal content, output a configuration interface in response to determining that supplementary information needs to be input for the generation intent, so as to obtain the supplementary information input for the configuration interface; the configuration interface is displayed in the form of a configuration card, and its display position is adaptively adjusted according to the current operation interface context, wherein the operation interface context includes a full-screen canvas interface or an embedded dialog interface. The content generation module is further used for: In response to the user's selection of the target content generation option, the multimodal content is generated based on the generation intent, the source material, and the supplementary information.

16. An intelligent agent, comprising: The input module is used to receive input information; The processing module is configured to perform the steps of the multimodal content generation method as described in any one of claims 1 to 14 based on the input information, so as to determine the target task and execute the target task to obtain output information; An output module is used to output the output information obtained by the processing module.

17. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform the method of any one of claims 1-14.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-14.

19. A computer program product comprising a computer program stored on a storage medium, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Image generation method and device, computer equipment and storage medium

    CN117112826A

  • Intelligent agent intention understanding interaction system, method and device

    CN118607535A