Media content processing method, apparatus, device, medium, and product

CN122845895APending Publication Date: 2026-09-29BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611132603.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-29

Smart Images

  • Figure CN122845895A_ABST
    Figure CN122845895A_ABST
Patent Text Reader

Abstract

This document relates to media content processing methods, apparatuses, devices, media, and products. The method includes displaying multiple media contents on a first interface. The method also includes receiving a first instruction instructing the execution of a media content processing operation on at least one first media content, the multiple media contents including at least one first media content. The method further includes displaying at least one second media content on the first interface, the at least one second media content being generated based on at least one first media content and the first instruction, the at least one second media content being displayed at the location corresponding to the at least one first media content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This article covers the field of content processing in general, and specifically covers media content processing methods, devices, equipment, media and products. Background Technology

[0002] With the rapid development of mobile internet, smart terminals, and digital content platforms, media content such as images, videos, and GIFs has become an important medium for users to record their lives, express their opinions, showcase their creativity, and engage in social interaction. Users can conveniently shoot, acquire, edit, and publish media content using computing devices such as mobile phones, tablets, and personal computers, and disseminate it through social platforms, short video platforms, instant messaging applications, or content communities. The rich variety of media content formats enhances the intuitiveness and appeal of information expression, enabling users to more efficiently record their lives, share their interests, showcase their work, and communicate socially.

[0003] With the development of artificial intelligence, computer vision, image processing, and mobile computing power, media content editing tools have gradually evolved from basic editing tools into comprehensive creation tools. These tools can be standalone applications or integrated into other types of applications, providing users with diverse functions such as filters, beautification, stickers, text, music, special effects, templates, and cover editing. This helps users generate media works with personalized expression and high visual quality in a short time. The development of media editing technology has lowered the barrier to content creation, enriched users' ways of expression, and promoted the efficient production, sharing, and dissemination of media content in social, entertainment, and digital media scenarios. Summary of the Invention

[0004] This article provides a method, apparatus, device, medium, and product for media content processing.

[0005] According to a first aspect of this document, a media content processing method is provided. The method includes displaying multiple media contents on a first interface. The method further includes receiving a first instruction, which instructs the execution of a media content processing operation on at least one first media content, wherein the multiple media contents include at least one first media content. The method also includes displaying at least one second media content on the first interface, wherein the at least one second media content is generated based on at least one first media content and the first instruction, and the at least one second media content is displayed at the corresponding position of the at least one first media content. This method can improve the editing efficiency of multiple media contents, enhance the consistency of the editing state, and improve the user experience.

[0006] According to a second aspect of this document, a media content processing apparatus is provided. The apparatus includes a first interface display module for displaying multiple media contents on the first interface; a first instruction receiving module for receiving a first instruction, the first instruction instructing the execution of a media content processing operation on at least one first media content, the multiple media contents including at least one first media content; and a media content display module for displaying at least one second media content on the first interface, the at least one second media content being generated based on at least one first media content and the first instruction, and the at least one second media content being displayed corresponding to the position of at least one first media content. This apparatus can improve the interactive efficiency, the continuity of result presentation, and the editing experience of multimedia content editing.

[0007] In a third aspect of this document, an electronic device is provided, including at least one processor and a storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to the first aspect of this document. This electronic device can improve the interactive efficiency of multimedia content editing, the continuity of result presentation, and the editing experience.

[0008] In a fourth aspect of this document, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to the first aspect of this document. The program stored in this computer-readable storage medium can support multimedia content editing and result display based on natural language, thereby improving editing efficiency and user experience.

[0009] In a fifth aspect of this document, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the method according to the first aspect of this document. With this computer program product, when the program is executed by a processor, at least one of multiple media contents can be edited based on natural language input, and the generated result can be displayed, thereby improving the efficiency of multimedia content editing and the effect of result display.

[0010] It should be understood that the content described in this section is not intended to define the key or essential features of this document, nor is it intended to limit the scope of this document. Other features of this document will become readily apparent from the following description. Attached Figure Description

[0011] The above and other objects, features and advantages of this document will become more apparent from a more detailed description taken in conjunction with the accompanying drawings, wherein like reference numerals generally denote like parts.

[0012] Figure 1 The illustration shows a schematic diagram of an example environment in which the device and / or method may be implemented;

[0013] Figure 2 The illustration shows a sample method for media content processing;

[0014] Figure 3 This diagram illustrates an example of a module architecture for collaborative editing of multimedia content.

[0015] Figure 4 The illustration shows a sample process for multi-graph intent understanding;

[0016] Figure 5 The diagram illustrates an example flow of differential parameter inference;

[0017] Figure 6 The illustration shows an example of a multi-image carousel and slot status;

[0018] Figure 7 The illustration shows an example of a multi-image carousel and slot status;

[0019] Figure 8 The diagram illustrates an example flow of slot-level process control.

[0020] Figure 9 The illustration shows an example of a media content pre-editing page.

[0021] Figure 10 This diagram illustrates an example of a media content editing interface.

[0022] Figure 11 The illustration shows a sample interface for media content editing.

[0023] Figure 12 This diagram illustrates another example of the interface in media content editing.

[0024] Figure 13 The illustration shows a sample of the interface after media content editing is completed.

[0025] Figure 14 The illustration shows an example of a media content publishing page.

[0026] Figure 15 The illustration shows an example of an AI-generated content creation dialog page.

[0027] Figure 16 The illustration shows a schematic block diagram of a media content processing device;

[0028] Figure 17 A schematic block diagram is shown for an example device used to implement the contents of this article. Detailed Implementation

[0029] It is understood that the data involved in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0030] It is understandable that before using the technical content disclosed in this article, users should be informed of the type, scope of use, and usage scenarios of the personal information involved in this article in an appropriate manner in accordance with relevant laws and regulations, and their authorization should be obtained.

[0031] For example, upon receiving a user's proactive request, a prompt message can be sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described in this document.

[0032] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0033] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.

[0034] The present invention will now be described in more detail with reference to the accompanying drawings. While the drawings illustrate the contents of this document, it should be understood that this document can be implemented in various forms and should not be construed as limited to the contents set forth herein; rather, these contents are provided to provide a more thorough and complete understanding of the document. It should be understood that the drawings and contents of this document are for illustrative purposes only and are not intended to limit the scope of this document.

[0035] In this description, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0036] In existing media content editing technologies, users typically edit multiple images, videos, and other media content by operating on each image individually or by selecting multiple images and applying the same settings uniformly. For example, a user can first add a filter, adjust the brightness, or apply effects to one image and then copy the same editing parameters to other images; or, a user can select multiple images in the media management interface and apply the same filter, adjustment parameters, or effects to all selected images. However, these methods usually rely on multiple selections and manual configurations by the user in the graphical interface, making it difficult to directly express editing needs for multiple media content using natural language.

[0037] Furthermore, existing batch editing methods typically apply the same editing parameters directly to multiple media content images. However, different media content images may differ in brightness, color temperature, people, scenes, and genres. Using the same parameters may result in inconsistent final visual effects across different media content images. For example, an image that was originally bright may become overexposed after a uniform increase in brightness, while an image that was originally dark may still be underexposed. At the same time, existing multimedia content editing interfaces have limited capabilities in organizing multiple media content images, displaying historical results, showing generation progress, and controlling cancellation. This can easily lead to problems such as discontinuous editing processes for multiple media content images, unintuitive result viewing, and the impact of failed or canceled tasks on the overall editing workflow.

[0038] To address this, this paper proposes a method for media content processing. In this method, a computing device displays multiple media contents on a first interface. The computing device then receives a first instruction, which directs a media content processing operation to be performed on at least one first media content, where the multiple media contents include at least one first media content. The computing device then displays at least one second media content on the first interface. This at least one second media content is generated based on at least one first media content and the first instruction, and is displayed at the corresponding position of the at least one first media content. Through this method, users can process at least one media content from multiple media contents based on the first instruction and view the generated results, thereby improving the efficiency of multimedia content processing and the effect of result display.

[0039] The following will describe this article in further detail with reference to the accompanying drawings. Figure 1 An example environment 100 is shown in which the devices and / or methods described herein may be implemented. Figure 1 As shown, environment 100 may include computing device 102 and application 104 running on computing device 102. Computing device 102 can display a corresponding interface, receive user operations, process interface data, and display processing results by running application 104.

[0040] Examples of computing device 102 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants, media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices. The application 104 running on the computing device may be a content platform application, a social application, a content creation application, a shopping application, an office application, a live streaming application, or a combined application that includes one or more of the above functions.

[0041] The steps performed by the computing device 102 in this article, such as interface display, information presentation, operation reception, data processing, content updating, page switching, and result display, can all be implemented by the application 104 running in the computing device 102, and will not be described in detail hereafter.

[0042] The above combination Figure 1 A schematic diagram illustrating an example environment in which the devices and / or methods described herein may be implemented is provided below. Figure 2 A diagram illustrating an example method for processing media content according to this article. Figure 2 The method in can be derived from Figure 1 The computing device 102 or any suitable device in the system can be used for execution.

[0043] like Figure 2 As shown, in example method 200, at box 202, the computing device can display multiple media content items on a first interface. For example, the multiple media content items may include multiple images, GIFs, animations, videos, or other processable media content uploaded or selected by the user. The multiple media content items can be displayed in thumbnail form or in non-thumbnail form, allowing the user to view multiple media content items on the first interface.

[0044] In one scenario, multiple media content items can be added to the first interface through batch selection or drag-and-drop. The computing device can then determine the display positions of these multiple media content items on the first interface and display the corresponding media content at those positions.

[0045] In one scenario, the first interface may include a first input area and a first media content display area. Multiple media content items may be displayed in the first media content display area. While viewing the media content in the first media content display area, the user can input natural language commands in the first input area, allowing the computing device to receive these commands. In this case, the user can directly input their processing needs for at least one media content, expressed in natural language, in the first input area of ​​the first interface.

[0046] In one scenario, the computing device can receive a first operation triggered on a first interface and display an input component on the first interface. For example, the first interface may have a component for rotating prompts, which can rotate multiple different prompts, each of which can perform different media content processing operations on the media content. If a first operation is received on this component, the input component can be displayed. The input component can display a second input area and a second media content display area. Multiple media content items can be displayed in the second media content display area. The user can enter natural language commands in the second input area and can view the multiple media content items that can be processed currently through the second media content display area in the input component.

[0047] In one scenario, the input component may further include multiple prompts, each of which may indicate a media content processing operation. Additionally, these multiple prompts are the carousel-like prompts described above. For example, the prompts may include recommended effects, recommended filters, recommended final product effects, recommended image quality enhancement operations, or other processing operations applicable to the media content. The computing device may receive a trigger operation on the first prompt and display at least one third media content on the input component. The at least one third media content may be generated based on at least one second media content and the first prompt, and the at least one third media content may be displayed correspondingly at the position of the at least one second media content. For example, after generating at least one second media content, clicking the first prompt can continue to perform media content processing operations corresponding to the first prompt on the at least one second media content to generate at least one third media content, and display the at least one third media content at the position of the corresponding second media content.

[0048] In one scenario, the computing device can also receive a trigger operation on a first control. The first control can be located on a first interface or on an input component of the first interface. For example, a first control for switching to a dialog interface can be set on the first interface or in a displayed input component. Upon receiving a trigger operation on the first control, the computing device can display a dialog interface and show at least one second media content within it. The user can continue editing at least one second media content within the dialog interface. Thus, the user can enter the dialog interface from the first interface and further adjust the media content through dialog interaction.

[0049] In one scenario, the dialog interface may include a second control. The computing device may display at least one fifth media content within the dialog interface, which may be obtained by editing at least one second media content. The computing device may receive a trigger operation on the second control and, upon receiving the trigger operation, display a first interface showing at least one fifth media content. For example, the second control may be a close control; after the user clicks the close control, the dialog interface can be closed, returning to the first interface, where the at least one fifth media content edited and processed within the dialog interface is displayed.

[0050] In one scenario, the first control can be located on an input component of the first interface, and the input component can include multiple prompts, with the first control being one of these prompts. For example, the first control could be "Chat with AI," "No ideas? Chat with AI," or other prompts used to enter a dialogue interface.

[0051] At box 204, the computing device can receive a first instruction, which instructs the execution of a media content processing operation on at least one first media content, wherein the multiple media content includes at least one first media content. In one case, the first instruction may be a natural language instruction. For example, a natural language instruction may be "brighten these up a bit," "add a filter to the second image," "beautify it," or "the first two images are warm-toned, the last two are cool-toned," etc. Additionally, the media content processing operation may be a media content editing operation. The computing device can determine, based on the first instruction, at least one first media content to be processed and the corresponding media content processing operation.

[0052] In one scenario, before receiving the first instruction, the computing device may receive a selection operation, which can be used to select at least one first media content from multiple media content. The selection operation may include a selection triggered on at least one media content, or a natural language instruction indicating the selection of at least one media content. For example, a user may first select one or more images from multiple media content and then input the first instruction. Alternatively, the computing device may select at least one first media content from multiple media content based on the first instruction. For example, expressions such as "second image," "image with people," or "first two images" in the first instruction can be used to determine the corresponding first media content. Thus, the target media content can be determined either through a selection operation initiated by the user on the media content or through a natural language instruction indicating the selection.

[0053] In one scenario, at least one first media content may include multiple first media content items, and correspondingly at least one second media content may include multiple second media content items. When generating multiple second media content items, the computing device can establish media content processing tasks for each of the multiple first media content items and execute the multiple media content processing tasks concurrently to improve the generation efficiency of the multiple second media content items. When generating multiple second media content items, the media content processing parameters for each first media content item to its corresponding second media content item may be different, or the editing parameters for each of the multiple second media content items may be different. Additionally, these differences in editing parameters can be determined by viewing the parameters of the editing controls. In one scenario, the computing device can determine the corresponding media content processing parameters based on the visual characteristics of the first media content items, whereby the visual characteristics include at least one of the following: brightness, color temperature, scene, or genre. For example, for media content items with different brightness, color temperature, or scenes, the computing device can use different brightness parameters, filter intensity parameters, color temperature parameters, or other editing parameters. Thus, even if multiple media content items are generated based on the same first instruction, processing results that better match the characteristics of each media content can be obtained through different processing parameters. For example, if a user enters "brighten these pictures a bit", the brightness control parameter for the first media content, which has a brighter original image, can be +5, while the brightness control parameter for the second media content, which has a darker original image, can be +15. By checking the editing control parameters of the corresponding second media content, the user can determine that different editing parameters were used when the different second media content was generated.

[0054] At box 206, the computing device can display at least one second media content on the first interface. The at least one second media content is generated based on at least one first media content and a first instruction, and the at least one second media content is displayed in the corresponding position of the at least one first media content. For example, when the first instruction instructs to apply the same style to all media content, the computing device can display the corresponding second media content in the original positions of multiple first media content; when the first instruction only instructs to process one of the media content, the computing device can display the second media content in the original position of the first media content.

[0055] In one scenario, displaying multiple media contents on the first interface may include: displaying multiple slots on the first interface, each slot displaying multiple media contents; displaying at least one second media content on the first interface may include: displaying at least one second media content in a slot corresponding to at least one media content. For example, after a user performs a media content processing operation on the media content in the second slot, the processed second media content can be displayed in the second slot.

[0056] In one scenario, the computing device can receive a second operation targeting a first slot on a first interface, where at least one slot may include the first slot. The computing device can display at least one fourth media content corresponding to the first slot. The at least one fourth media content may be obtained by performing media content processing operations on the media content in the first slot at different times. For example, the media content in the same slot may undergo different filters, effects, final effects, or parameter adjustments at different times to obtain multiple processing results, which can be displayed in a display area. When the user clicks on different slots, the display area can show multiple processing results corresponding to that slot. In this way, the user can view multiple historical versions or processing results corresponding to the same slot.

[0057] In one scenario, the computing device can receive a third operation, which may be triggered in response to one of at least one fourth media content. Upon receiving the third operation, the computing device can display the fourth media content corresponding to the third operation in the media content display area of ​​the first interface. For example, after a user clicks on a historical version, the media content display area can switch to displaying the media content corresponding to that historical version.

[0058] In one scenario, the computing device can also display multiple slots in a carousel, with the selected slot highlighted. For example, the selected slot can be marked with a highlighted border, magnified display, a pause indicator, or other visual styles. If a pause button is received, the computing device can stop the carousel.

[0059] In one scenario, the computing device may display a notification on at least one piece of second media content. This notification may indicate that the at least one piece of second media content was updated via a first instruction. For example, after the second media content is generated in the first interface, the computing device may display a red dot, update marker, new content marker, corner mark, or other visual cue at the corner, edge, or adjacent position of the second media content to indicate to the user that the media content at that location has been updated from the first media content to the second media content.

[0060] This method improves the interactive and processing efficiency of multimedia content processing, reducing the operational costs for users who need to operate on multiple media content items one by one and configure parameters item by item. Simultaneously, it enhances the seamless integration between the initial command, natural language processing requirements, prompts, dialogue interaction, location-based display, and slot-based presentation, enabling the continuous and intuitive presentation of processing results for multiple media content items. Furthermore, by supporting differentiated processing parameters for different media content, displaying historical processing results, and slot-based carousel display, it improves the adaptability, interpretability, viewability, and user experience of multimedia content processing results.

[0061] The above combination Figure 2 A diagram illustrating the example methods for media content processing described in this paper is provided below. Figure 3 A schematic diagram illustrating an example of the modular architecture of a dual-mode image editing system. Figure 3 Example 300 in the example can be derived from Figure 1 The computing device 102 or any suitable device in the process is processed by running the application.

[0062] like Figure 3 As shown, Example 300 illustrates a computing device architecture for collaborative editing of multiple media contents. This architecture may include a multi-image messaging module 302, an editing intent understanding module 304, an image referencing resolution module 306, a current-round target image set 308, an image individual feature analysis module 310, a differential parameter inference engine 312, a concurrent editing execution engine 314, a transaction manager 316, a slot-level process controller 318, a multi-image carousel and slot management module 320, and an editor carousel and dialogue stream write-back module 322.

[0063] The multi-image messaging module 302 can receive multiple images uploaded by users in batches and present them in a structured format of multiple image message cards in the dialog interface or editing interface. These images can be replaced with other media content, including videos. Each image can have a corresponding material identifier, thumbnail content, media type, and display order, facilitating subsequent individual identification, selection, and editing of multiple media contents.

[0064] In one scenario, the multi-image messaging module 302 can provide multiple image selection options in the dialog input area. Users can add multiple images by batch selecting from their album or by dragging and dropping. The computing device can support mixed-genre selection; for example, still images, animated photos, and videos can be uploaded in the same batch. The multiple images uploaded by the user can be presented as a batch image message in the dialog stream. This batch image message can include fields such as message identifier, message type, information about the multiple images, layout method, and total number of images.

[0065] In one scenario, the multi-image message module 302 can automatically determine the layout based on the number of images or the shooting relationship between them. For example, when the number of images is greater than or equal to 3, a grid layout can be used to display multiple thumbnails in a multi-column grid manner; when the number of images is 2 to 4, a horizontal scrolling layout can be used to support users to swipe left and right to browse multiple thumbnails; when multiple images belong to a group of consecutive shots, a stacked layout can be used, and the front-to-back relationship between multiple thumbnails can be displayed by slight rotation or offset.

[0066] In one scenario, the multi-image message module 302 can also support image focusing operations. After a user clicks on a thumbnail, the image corresponding to that thumbnail can be set as the current editing main image. The current editing main image can be marked with a highlight border, a selection icon, or other visual styles. If subsequent natural language input does not explicitly specify the target media content, the editing operation can default to the current editing main image. The material identifier of the current editing main image can be associated with the original frame information in the intelligent editing service input protocol.

[0067] The editing intent understanding module 304 can be used to analyze the natural language input by the user to identify the editing scope and editing operation indicated by the natural language. The editing scope can include full batch processing, specified media content processing, group processing, or conditional filtering. For example, when the natural language input is "brighten these a bit," the editing intent understanding module 304 can determine that the editing operation applies to multiple uploaded images; when the natural language input is "the first two are warm-toned, the last two are cool-toned," the editing intent understanding module 304 can determine that multiple media contents are divided into different groups, and different groups correspond to different editing operations.

[0068] The image reference resolution module 306 can be used to resolve media content references in the user's natural language into specific material identifiers. For example, the image reference resolution module 306 can resolve expressions such as "second image," "this one," "previous image," and "brightest one" into one or more corresponding material identifiers. Thus, the editor intent understanding module 304 and the image reference resolution module 306 can jointly determine the target image set 308 for this round, that is, one or more target media contents that need to be edited in this round.

[0069] The image individual feature analysis module 310 can be used to extract features from each target image in the target image set 308. The extracted individual visual features may include brightness, color temperature, and scene. For example, the image individual feature analysis module 310 can extract the brightness distribution, dominant color temperature, saturation, indoor or outdoor scene labels, and genre information such as still images, dynamic photos, or videos of the target images. By extracting the above individual visual features, the computing device can distinguish the original visual differences between different media content during batch editing.

[0070] The differential parameter inference engine 312 can generate image-by-image editing instructions for each target image in the target image set 308 based on the user's natural language input and the individual visual features extracted by the image individual feature analysis module 310. For example, the differential parameter inference engine 312 can first obtain basic editing parameters based on the natural language input, and then adjust the basic editing parameters differentially based on the visual features of each target image, such as brightness, color temperature, scene, or genre, thereby obtaining editing instructions applicable to each target image. Thus, multiple images can achieve visually consistent or user-initiated editing results even with different parameters.

[0071] The concurrent editing execution engine 314 can be used to concurrently perform editing processing on multiple target images in the current target image set 308 based on the editing operation instructions generated by the differential parameter inference engine 312. The transaction manager 316 can be used for transaction management of multimedia content editing processing, including committing the editing results and updating effect frames and history when editing processing is successful, and executing failure handling strategies such as partial commit, full rollback, or retrying failed tasks when editing processing fails. The transaction manager improves the reliability and recoverability of the batch editing process of multimedia content.

[0072] The slot-level process controller 318 can maintain a separate generation process for each of multiple slots, and provide slot-level progress display, consumption display, pause or cancellation control. A cancellation operation on a slot can only affect the generation process corresponding to that slot, without affecting the generation processes of other slots. The multiple image carousel and slot management module 320 can organize multiple images in the editor by slot and carousel the slots, maintaining the original image, historical editing results, and current editing results for each slot. The editor carousel and dialog stream write-back module 322 can write back the editing results, historical results, or current display results of each slot to the editor carousel area and dialog stream for users to view and continue editing.

[0073] The above combination Figure 3 A schematic diagram illustrating an example of a module architecture for collaborative editing of multiple media contents is provided below; Figure 4 A diagram illustrating an example process for understanding the intent of multiple media contents.

[0074] like Figure 4As shown in Example 400, this illustrates a schematic diagram of an example flow for understanding multiple media content intents. After receiving user natural language input 402, the intelligent editing service can perform edit scope identification 404. Edit scope identification 404 can determine the scope type corresponding to the editing operation indicated by the natural language input through keyword matching and semantic pattern analysis. Scope type 406 can include full batch processing, specified media content processing, group processing, or conditional filtering. In one scenario, the edit scope type, triggering mode, example, and mapping relationship can be shown in the following table:

[0075] For full batch processing, it applies to all images 412, thus the intelligent editing service can determine that the editing operation applies to all uploaded images and uses all uploaded media content as the target media image set 420. For specified media content processing, the intelligent editing service can perform image reference resolution 408 to parse the image reference in natural language into specific material identifiers. Then, it can determine the sequence number, indication, attribute, time, and content reference 414. For group processing, the intelligent editing service can perform group logic resolution 410 to determine different groups and their corresponding editing intentions. For conditional filtering processing, the intelligent editing service can filter a subset of images that meet the conditions from multiple images based on the conditional expressions in natural language.

[0076] In the 408 image referencing resolution process, the intelligent editing service can determine the target image based on sequence reference, indicator reference, attribute reference, time reference, or content reference. Sequence reference can be used to parse expressions in the form of "the Nth image" and obtain the corresponding index; indicator reference can map "this image" to the current main image being edited, and map "the previous image" or "the next image" to the image adjacent to the current main image being edited; attribute reference can determine "the brightest image" or "the darkest image" based on attributes such as brightness; time reference can determine "the previous image" based on the operation history; and content reference can determine "the image with corresponding features" based on the content recognition results.

[0077] In the group logic parsing 410, the intelligent editing service can identify grouping quantifiers in natural language and determine the image range corresponding to each group based on these quantifiers. Subsequently, the intelligent editing service can identify the editing intent corresponding to each group and generate a grouping mapping table 416. Additionally, subsets can be filtered using conditional filtering.

[0078] After determining the full set of images, specified images, grouped images, or conditionally filtered images, the intelligent editing service can identify the target image set 420 and further analyze the target operation type and parameters 422. The target operation type can include filters, beautification, basic adjustments, special effects, stickers, or intelligent generation, etc. Subsequently, the intelligent editing service can output the target image set, target operation type, and editing parameters to the differential parameter inference engine 424 to generate editing operation instructions corresponding to each target image.

[0079] The above combination Figure 4 A diagram illustrating an example flow for understanding the intent of multiple media contents is provided below. Figure 5 A schematic diagram illustrating an example flow of differential parameter inference.

[0080] like Figure 5 As shown in Example 500, the differential parameter inference process is illustrated. Upon receiving the user's editing intent 502, the intelligent editing service can perform basic parameter inference 506. Basic parameter inference 506 can be used to map the user's natural language description into specific editing operation instructions and obtain basic editing parameters. For example, when the user inputs "Japanese fresh style," it can be mapped to a filter-type editing operation instruction, where the filter is enabled and the filter resource identifier is a preset filter identifier; when the user inputs "brighten it a bit," it can be mapped to a basic adjustment-type editing operation instruction; when the user inputs "warm tone," it can be mapped to a basic adjustment-type editing operation instruction. These basic editing parameters can serve as the basis for subsequent image-by-image differential adjustments.

[0081] The target image set 504 may include one or more target images determined based on natural language input. For each target image in the target image set 504, the intelligent editing service can perform image-by-image individual visual feature extraction 508 to obtain the individual visual feature vector corresponding to that target image. The individual visual feature vector can be used to characterize information such as brightness, color temperature, saturation, scene, and genre of the target image.

[0082] After obtaining the basic editing parameters and the individual visual feature vectors of the target images, the intelligent editing service can calculate an adjustment amount of 510 for each image. For the i-th target image, the adjustment amount can be calculated by a fine-tuning function based on the difference between the individual visual feature vector of that image and the target effect feature vector. Subsequently, the intelligent editing service can add the basic parameters and the adjustment amount to obtain the personalized editing parameters corresponding to that image.

[0083] In one scenario, the adjustment calculation 512 may include a brightness adjustment fine-tuning strategy. For brightness adjustment, the target brightness can be determined first based on the base brightness and the brightness adjustment amount corresponding to the user's intent. Then, the difference between the target brightness and the average brightness of the current image can be calculated, and the difference can be limited 514 to prevent over-adjustment.

[0084] In one scenario, the adjustment calculation 512 may also include a filter intensity fine-tuning strategy. For example, when a target is detected in the target image, the filter intensity 516 can be reduced to decrease the filter's impact on the target area; when the difference between the dominant color temperature of the target image and the preset color temperature of the filter is large, the filter intensity can be increased or corresponding compensation can be made to ensure that different images have a more consistent visual appearance after applying the same filter intent.

[0085] In one scenario, the adjustment calculation 512 may also include a color temperature fine-tuning strategy. For example, the target color temperature can be set as the target red-blue channel ratio, and the color temperature adjustment amount can be calculated based on the difference between the target red-blue channel ratio and the dominant color temperature of the current image, so that the color temperature of each target image gradually converges to the target color temperature 518.

[0086] In one scenario, the differential parameter inference process can also include effect frame difference analysis 520. For images with existing effect frames, the intelligent editing service can compare the differences between the original frame and the effect frame to analyze the editing effects already applied to the image. For example, when it is detected that an image has already applied editing effects such as increasing brightness by 5 or contrast by 10, the intelligent editing service can avoid repeatedly overlaying or conflicting with existing editing effects when generating new editing operation instructions, or it can intelligently overlay them based on existing editing effects.

[0087] Therefore, the intelligent editing service can generate personalized editing parameters 522 and output image-by-image editing instructions 524. The core of this process is that batch editing pursues visual consistency, rather than the consistency of parameter values. By overlaying image-by-image differential adjustments on a unified base parameter, images with different exposures, color temperatures, or genres can present a consistent or harmonious style after applying the same user's editing intent.

[0088] The above combination Figure 5 A schematic diagram illustrating an example flow of differential parameter inference is provided below. Figure 6 A schematic diagram illustrating an example process of concurrent scheduling and transaction processing.

[0089] like Figure 6As shown in Example 600, a sample flow of concurrent scheduling and transaction processing is illustrated. After receiving a batch editing request 602 for N images, the intelligent editing service can enter the transaction start phase 604. In the transaction start phase 604, the intelligent editing service can record a snapshot of the pre-edit state of each image. The pre-edit state snapshot can include original frame information and editor capability states, such as the original frame and editing capabilities. Therefore, in the event of subsequent editing failure or user rollback, the pre-edit state of the image can be restored based on this pre-edit state snapshot.

[0090] Subsequently, the intelligent editing service can dynamically determine the concurrency level C based on device performance, as shown in step 606. Device performance may include the number of central processing unit (CPU) cores, available memory, and current load. The concurrency level C can be used to represent the number of image editing tasks that can be executed in parallel at the same time. By dynamically determining the concurrency level C based on the device status, batch editing efficiency can be improved while avoiding device lag, insufficient memory, or editing failures caused by processing multiple media content pieces simultaneously.

[0091] After determining the concurrency level C, the intelligent editing service can divide the editing tasks corresponding to N images into batches, taking the result of N divided by C and rounding it up to the nearest whole batch, as shown in step 608. Each batch can include no more than C image editing tasks. The main image can always be assigned to the first batch and processed first, so that users can see the main editing results as soon as possible.

[0092] In step 610, the intelligent editing service can concurrently execute the editing operation instructions corresponding to each image. For the main image, since it is prioritized in the first batch, the preview can be updated immediately after processing, as shown in step 612, without waiting for all images to be processed. This improves the WYSIWYG (What You See Is What You Get) experience, allowing users to more quickly confirm whether the editing results meet their expectations.

[0093] During concurrent execution, all image editing operations in a single batch editing request can be executed as a single transaction. Specifically, during the transaction execution phase, the editing operation instructions corresponding to each image can be executed concurrently; at step 614, it is determined whether all images have been successfully edited. If all images have been successfully edited, the process can proceed to transaction commit 616 and update the effect frame and operation history, for example, updating the effect frame and operation history. Thus, the editing results of multiple images can be submitted uniformly and used as edited images in subsequent editing processes.

[0094] In step 614, if it is determined that not all images were edited successfully, the process can proceed to failure handling strategy 618. The failure handling strategy can be determined based on a user-preset strategy or by real-time selection. As shown in step 620, strategy A can be a partial submission strategy, i.e., skipping failed images and submitting the remaining images normally. As shown in step 622, strategy B can be a full rollback strategy, i.e., rolling all images back to their pre-editing state. As shown in step 624, strategy C can be a failed task retry strategy, i.e., re-executing the editing process only for failed images.

[0095] In one scenario, this process can also be combined with front-end performance optimization strategies. These strategies can include thumbnail pre-generation and Least Recently Used (LRU) caching, list virtualization rendering, and progressive loading. Thumbnail pre-generation and LRU caching reduce redundant decoding and loading overhead; list virtualization rendering renders only thumbnails within the visible area, reducing UI rendering pressure; and progressive loading prioritizes loading the main image preview and loads other images on demand based on user browsing or computing device scheduling needs. This improves UI responsiveness and overall editing smoothness in batch editing scenarios involving multiple media contents.

[0096] The above combination Figure 6 A schematic diagram illustrating an example process of concurrent scheduling and transaction processing is provided below. Figure 7 A diagram illustrating an example of multiple media content carousels and slot status.

[0097] like Figure 7 As shown in Example 700, this diagram illustrates a carousel of multiple media contents and slot status. In one scenario, a slot status object can be maintained for each image. The slot status object can be used to record information such as the material identifier corresponding to the slot, whether it participates in the current generation, carousel status, original image reference, historical image reference, result image reference, currently displayed image reference, loading status, new result prompt, and source anchor point.

[0098] In one scenario, each slot can have a corresponding slot state 710. Slot state 710 can include a selected state 712, a carousel state 714, a stacked state 716, a currently displayed image reference 718, a loading percentage 720, and a new result red dot 722. Specifically, the selected state 712 indicates whether the media content corresponding to the slot is participating in the current generation round; the carousel state 714 indicates whether the slot is playing, paused, or not yet carouseled; the stacked state 716 organizes the original image reference, historical image reference, and result image reference corresponding to the slot; the currently displayed image reference 718 indicates the version of the media content currently displayed in the editor and can default to the latest result; the loading percentage 720 indicates the loading progress of the generation process corresponding to the slot; and the new result red dot 722 indicates that there are new generated results in the slot that have not yet been viewed.

[0099] During the carousel and display of multiple media content, the editor's multi-image area 702 can automatically carousel multiple slots by default. For example, slot 1 in frame 704, slot 2 in frame 706, and slot 3 in frame 708. During the carousel, the currently carouseled slot can be highlighted, for example, by using a highlighted border, magnification, or other visual styles. The resulting image can always be displayed at the forefront of the corresponding slot, allowing users to prioritize the latest generated results. If a pause button is received, the editor can pause the carousel; if an operation is received targeting a specific slot, the editor can manually locate that slot.

[0100] In one scenario, carousel loading tips can be displayed by slot. If a slot is being generated, its thumbnail can show the loading percentage; if the slot is in the process of carousel loading, the carousel can restart from that slot. Since multiple media content carousel thumbnails occupy the space originally used for the tooltip component in the editor, in multi-media content scenarios, the tooltip component in the single-image scenario can be removed from the input box to avoid obscuring the image display area in the editor.

[0101] In one scenario, historical images and original images in each slot can be collapsed in a stacked manner, as shown in Stack 716. If an operation is received for a stacked image, the editor can enter the history expanded state and display the original and historical images associated with that slot. If an operation is received for a specific historical image or original image, the editor can switch the display to the media content corresponding to that historical image or original image. The history expanded and collapsed states can be linked across multiple slots, meaning multiple slots can be expanded or collapsed simultaneously, and this expanded or collapsed state can be consistent with the history expanded or collapsed state on the input box page.

[0102] In one scenario, users can add or remove images for generation using the plus sign control, and can select multiple slots for this round of generation. The input field can automatically populate all selected slot images from the editor's carousel. When a user attempts to deselect all, or if only one image remains and the user continues to deselect it, the computing device can display a prompt, such as "Select at least one," and prevent deselecting all slots. This ensures that each generation operation has at least one target image.

[0103] In one scenario, slot backfilling rule 724 can include three cases: one-to-one, one-to-many, and many-to-one. For the one-to-one case, a single input image generates a single result, which can be backfilled into the result reference of the slot corresponding to that input image, and this single result can be displayed first in the slot, as shown in single result display 726. For the one-to-many case, a single input image generates multiple results. The model link can support up to 9 results, and the conversational link currently supports 9 results. Multiple result images can be added to the corresponding slot, and the last result is selected by default. The history of other slots can remain unchanged, as shown in multiple results with the last one selected by default 728. For the many-to-one case, multiple input images collaboratively generate a single result, which can be backfilled into the aggregate slot, as shown in aggregate slot backfilling single result 730.

[0104] The above combination Figure 7 This diagram illustrates an example of multiple media content carousels and slot statuses. The following section combines... Figure 8 A schematic diagram illustrating an example flow of slot-level process control.

[0105] like Figure 8 As shown in Example 800, a sample flow of slot-level process control is illustrated. After initiating multi-graph batch generation 802, the concurrent editing execution engine can create a separate generation process 804 for each slot participating in the generation, and each generation process can be bound one-to-one with the slot identifier of the corresponding slot. The generation processes are isolated from each other, so that if the generation process corresponding to any slot is canceled, fails, or completes, it will not affect the generation processes corresponding to other slots.

[0106] Each generation process can report its corresponding slot's progress percentage, generation cost, and generation status (806). Specifically, each generation process can write its progress percentage, generation cost, and running status into the loading field of its corresponding slot. The running status can include running, completed, or failed. Because the loading information for each slot is maintained separately, different slots can display their own generation progress and status independently, without interference.

[0107] During the generation process, the progress of each slot can be displayed via a mask overlay. The progress information in the mask overlay can include percentage progress, estimated resource consumption, and estimated remaining time. For example, the mask overlay could display "Generating 30%" (814), "Estimated consumption of 30 starlight" (816), or "AI generating, 10 seconds remaining" (818). Slot thumbnails can also simultaneously display the loading percentage, allowing users to directly perceive the generation progress of each slot across multiple media content carousels.

[0108] Additionally, user operation 810 can be identified. In one scenario, the overlay can include an entry point for collapsing generation or collapsing loading. If an operation 820 for collapsing generation or collapsing loading is received, the computing device can collapse the overlay without terminating the generation process of the corresponding slot; the generation task continues to execute in the background. After generation is complete, the overlay of the corresponding slot can automatically disappear. The generation bar at the bottom of the editor can be displayed in a generating state during the generation process, for example, displaying "Image generation in progress...", and providing entry points for canceling or completing the image; when historical results exist, the generation bar can switch from the default state to viewing historical states, and the application process of the effect can be displayed on the screen in an animated manner.

[0109] Users can also perform cancellation or pause operations on individual slots. If a cancellation operation is received for a specific slot, the computing device can terminate only the generation process 822 related to that slot, while the generation processes of other slots continue to execute, and their loading status will not be cancelled. This allows for "cancellation only of processes related to the image in this slot," preventing the cancellation of one image from affecting the generation process of other images. The computing device can also support manual pause 824 for a single slot, pausing the generation process of that slot without affecting other slots.

[0110] The generation process for each slot can be completed asynchronously. If a slot completes generation (812) first, and a new result exists for that slot, the editor can immediately display the latest result (826) for that slot without waiting for all other slots to complete. When entering the history expansion state or dialog page, if there are still slots that have not yet been generated, the computing device can handle them according to fallback logic, such as continuing to display the loading status or original image of that slot, thus not blocking the viewing of completed slots.

[0111] When a new result arrives in a particular slot, a red dot will appear in that slot to indicate that a new result is available for viewing. The red dot will only disappear after the user manually clicks on the slot; if the slot is simply rotated to without the user clicking, the red dot will remain displayed. This method prevents users from missing new results that have been asynchronously generated.

[0112] The above combination Figure 8 A schematic diagram illustrating an example flow of slot-level process control is provided below. Figure 9 A diagram illustrating an example of multiple media content pre-editing pages.

[0113] like Figure 9 As shown in Example 900, a pre-editing page is displayed after the user selects multiple media contents and before entering the formal editing operation. This page may include page 902, effects selection control 904, next step control 906, and multimedia content selection area 908. Page 902 can display material illustrations corresponding to the multiple media contents selected by the user, and provide entry points for functions such as music, settings, sharing, text, topics, stickers, or more.

[0114] The effects selection control 904 allows users to choose recommended effects or final effects to apply to multiple media content. The control displays the name of the currently recommended effect and provides a final effect control, enabling users to process multiple media content as a single image based on the selected effect. In a single-image scenario, the input box can display multiple recommended words in a carousel; however, in a multi-image scenario, continuing to display multiple recommended words in a carousel can easily lead to interface congestion and information confusion. Therefore, in a multi-image scenario, the effects selection control 904 can converge to display a single recommended word, i.e., only one recommended word. This recommended word can come from the intelligent system's recommendation of the currently selected media content and can be re-recommended as the currently selected media content changes. In the event of recommendation anomalies or the inability to obtain recommendation results for the currently selected media content, the effects selection control 904 can display a fallback recommended word.

[0115] The next step control 906 can be used to trigger entry from the pre-editing page to the editing page. If an operation is received on the next step control 906, the computing device can display the editing page so the user can continue editing the selected media content. The slot area 908 can display multiple slots showing thumbnail content corresponding to the selected media content. The user can select the media content to view or edit by clicking on any thumbnail in the slot area 908. The selected thumbnail content can be identified by a border, highlight, magnification, or other visual styles.

[0116] This method allows users to access recommended effects, a finished product entry point, a next step entry point, and a current editing object selection point before entering the formal editing page. This enables users to switch between multiple selected media content and pre-determine the desired finished product effects or editing direction.

[0117] The above combination Figure 9This diagram illustrates an example of a media content pre-editing page. The following section combines... Figure 10 A schematic diagram illustrating an example of a media content editing interface.

[0118] like Figure 10 As shown in Example 1000, multiple media content editing interfaces are illustrated. These multiple media content editing interfaces may include page 1002, a recommended effects area 1004, an AI dialogue entry point 1006, and multiple media content selection areas 1008. Page 1002 can be used to display a schematic diagram of the media content currently being edited, and to display input and selection controls related to the editing of multiple media content after the user pulls up the keyboard.

[0119] The recommended effects area 1004 can be used to display multiple recommended effects to users. Recommended effects can include "transform into a snowy beach," "one-click HD," and "vacation film filter," among others. Users can click on any recommended effect in the recommended effects area 1004 to select the desired editing effect to apply to multiple media content.

[0120] The AI ​​dialogue entry point 1006 can be used to initiate a dialogue with the AI ​​system. If an operation is received on the AI ​​dialogue entry point 1006, the computing device can display a dialogue editing interface.

[0121] The multiple media content selection area 1008 can be used to display thumbnails corresponding to multiple selected media contents, and supports users in selecting multiple media contents to participate in the current round of editing. Users can check or uncheck one or more thumbnails in the multiple media content selection area 1008 to determine the target set of images to be edited together. Selected thumbnails can be marked with a checkmark, a highlighted border, or other visual styles. In addition, users can enter information below the media content selection area 1008. For example, users can enter "adjust these pictures to a Japanese fresh style" or "beautify them." The computing device can determine the editing scope, target media content, and corresponding editing operations based on the natural language input by the user. The computing device can perform collaborative editing on multiple images based on the multiple images selected in the multiple media content selection area 1008 and the natural language input by the user.

[0122] The above combination Figure 10 A schematic diagram illustrating an example of a media content editing interface is provided below. Figure 11 A diagram illustrating an example of editing in the media content editing interface.

[0123] like Figure 11As shown, Example 1100 illustrates the result generation page after a user selects at least one media content and enters natural language to initiate editing. This result generation page may include page 1102, a generation progress indicator 1104, a collapse generation control 1106, and a slot area 1108. Page 1102 can be used to display a schematic diagram of the media content currently generating the result, and can retain entry points for editing functions such as music, settings, sharing, text, topics, stickers, or more.

[0124] The generation progress indicator 1104 can be used to display the current generation progress of the editing results. For example, the generation progress indicator 1104 can display "Generating 30%" and the corresponding progress bar to inform the user that the editing results of the current media content are being generated. The generation progress indicator 1104 can correspond to the currently selected thumbnail content, allowing the user to perceive that the materials displayed on the current page are undergoing editing processing.

[0125] The collapse generation control 1106 can be used to collapse the generation progress prompt. If an operation to collapse the generation control 1106 is received, the computing device can shrink the generation progress prompt and display it in a preset area of ​​the page, such as shrinking it to the corresponding slot, the edge of the page, or other positions that do not obstruct the display of the material. After the generation prompt is collapsed, the generation process can continue without being interrupted by the collapse operation.

[0126] The slot area 1108 can include multiple slots for displaying thumbnails corresponding to multiple media contents selected by the user. The material illustrations displayed on the current result generation page can correspond to the thumbnails previously selected by the user. If the user selected a thumbnail in the thumbnail selection area 1108 before starting the generation process, page 1102 can prioritize displaying the material illustrations of the media content corresponding to that thumbnail and display the generation progress of that media content during the generation process.

[0127] Using this method, the computing device can display the generation status of the currently selected media content on the results generation page, and allow users to reduce the obstruction of the page content by the generation prompts by collapsing the generation controls, while allowing the generation task to continue to execute.

[0128] The above combination Figure 11 This diagram illustrates an example of the interface used in media content editing. The following section combines... Figure 12 A diagram illustrating another example of an interface in media content editing.

[0129] like Figure 12As shown in Example 1200, this illustrates the interface after the progress indicator is collapsed during media content editing. This interface may include page 1202, the collapsed generation panel 1204, and status notification information 1206. Page 1202 can be used to display a schematic diagram of the currently edited media content and retain entry points for functions such as music, settings, sharing, text, topics, stickers, or more.

[0130] During the generation process, if an operation to collapse the generation control is received, the computing device can collapse the generation progress indicator, which was originally located in the middle of the page or near the material display area, and display the collapsed generation indicator in a preset area. For example, the collapsed generation indicator can be displayed in the generation bar 1204. The generation bar 1204 can display the current image generation status and can include a cancel control and a final image control. Therefore, without affecting the continued execution of generation, the obstruction of the material display area by the generation indicator can be reduced.

[0131] If an operation is received targeting the cancel control in the generation section 1204, the computing device can cancel the current media content generation process. After canceling generation, the user can select another operation or re-initiate the generation request. Thus, the user can actively terminate the current generation task during the generation process without waiting for it to finish naturally.

[0132] In one scenario, if a user triggers an operation that conflicts with the current generation process before it is complete—such as adding an image or switching to an editing operation that requires waiting for generation to finish—the computing device can display status message 1206. Status message 1206 can be used to inform the user that the media content is still being generated, for example, displaying a message such as "Artificial Intelligence (AI) generation in progress, 10 seconds remaining." This message can inform the user that the current operation needs to wait for generation to complete, or that the current generation task needs to be canceled before continuing.

[0133] If a conflicting operation is received, the computing device can prevent the corresponding operation from executing immediately before the generation process is complete and before it has been cancelled. The operation can only be allowed to continue after generation is finished, or after the user terminates the current generation process via the cancellation control. This avoids executing incompatible operations while media content is still being generated, thereby improving the consistency of the editing process and the stability of the user interface.

[0134] The above combination Figure 12 This diagram illustrates another example of an interface used in media content editing. The following section combines... Figure 13 A schematic diagram illustrating an example of the interface after media content processing is complete.

[0135] like Figure 13 As shown in Example 1300, this illustrates a schematic diagram of a media content processing completion interface. This interface may include page 1302, a slot area 1304, a result information area 1306, and a return control 1308. Page 1302 can be used to display a schematic diagram of the media content for which media content processing has been completed. In one scenario, when multiple media contents have completed media content processing, the computing device can automatically display the first media content to complete the processing, allowing the user to view the generated results promptly.

[0136] Slot area 1304 can be used to display thumbnail content corresponding to multiple media contents, and to indicate the currently playing or displayed media content. Slots that are rotated to or selected can be marked with borders, highlights, pause indicators, or other visual styles. In one scenario, if a slot contains new results that the user has not yet viewed, that slot can display a red dot. This red dot only disappears after the user manually clicks on the corresponding slot; if the slot is only rotated to but the user does not click on it, the red dot can remain displayed.

[0137] The results information area 1306 can be used to display relevant results information for the media content corresponding to the current slot. If an operation is received on the results information area 1306 or the corresponding slot, the computing device can display at least one fourth media content for that media content. At least one fourth media content can be obtained by performing media content processing operations on the media content in the corresponding slot at different times. Multiple fourth media contents can include the original version of the media content, at least one historical edited version, and the currently processed version. The user can click on any of the fourth media contents to switch the display of the corresponding version of the media content on page 1302.

[0138] The return control 1308 can be used to return to the previous version or the previous editing state. If an operation is received on the return control 1308, the computing device can revert the current media content to the previous version and allow the user to continue making modifications based on that previous version. Thus, the user can view different historical versions of the media content after processing and revert to an earlier version to continue editing when needed.

[0139] In one scenario, the generated result can be linked to a source anchor. For a single image, the generated result can correspond to a single source anchor; for multiple images, the generated result can be linked to the source anchor corresponding to the first image, and aggregated text can be displayed at the anchor, such as "Artificial Intelligence". "etc." If an operation is received targeting this source anchor, the computing device can display an anchor aggregation page. This anchor aggregation page can be used to trace the source of multiple images in a unified manner, such as displaying the generation chain, intelligent effects used, editing sources, or related generation records for each image. The display and navigation logic of this source anchor can be consistent with the online "AI-enabled" anchor logic, thereby achieving cross-image aggregation and unified source tracing in multi-image generation scenarios.

[0140] The above combination Figure 13 This diagram illustrates an example of the interface after media content processing is complete. The following section combines... Figure 14 An illustration depicting an example of a media content publishing page.

[0141] like Figure 14 As shown in Example 1400, this diagram illustrates the process of a user entering a media content publishing page after generating desired media content. This publishing page may include page 1402, a media content preview area 1404, a thumbnail content selection area, a publishing information editing area, and a publishing operation area. Page 1402 can be used to host functions such as previewing the media content to be published, filling in publishing information, and publishing control.

[0142] The media content preview area can be used to display a mock-up of the media content generated by the user. Users can use this area to view the overall effect of the content to be published. In one scenario, the media content preview area may also include cover editing controls. If an operation to edit the cover controls is received, the computing device can display a cover editing interface so that the user can select or adjust the cover of the media content to be published.

[0143] The thumbnail selection area can be used to display thumbnails corresponding to multiple generated or selected media content. Users can switch between the currently previewed media content by clicking on different thumbnails. Selected thumbnails can be marked with borders, highlights, or other visual styles. The thumbnail selection area can also include add controls; if an operation to add controls is received, the computing device can continue adding media content or enter the media selection page.

[0144] If an operation is received for a time-limited daily control, the computing device can configure the media content to be published as content to be displayed for a limited time. If an operation is received for a work to be published control, the computing device can generate a publishing request based on the currently previewed media content and the publishing information filled in by the user, and publish the media content to the target platform.

[0145] Using this method, after a user completes editing multiple images or media content, the computing device can link the generated results to the publishing page, allowing the user to continue editing the cover, previewing the content, filling in the title and description, selecting topics and tags, setting the visibility range, and finally publishing.

[0146] The above combination Figure 14 This diagram illustrates an example of a media content publishing page. The following section combines... Figure 15 An illustration depicting an example of an AI-generated dialogue page.

[0147] like Figure 15 As shown in Example 1500, this illustrates a schematic diagram of an AI-powered content creation dialogue page. This AI-powered content creation dialogue page may include page 1502, a selected media content display area 1504, an intelligent system response message area 1506, and a publishing entry point 1508. Page 1502 can be used to host the dialogue interaction content between the user and the intelligent system, and can display a natural language input area at the bottom so that the user can continue to input creative ideas or editing needs.

[0148] The publishing entry point 1508 can be used to incorporate the edited results selected by the user into the publishing process. For example, ... Figure 15 As shown, the publishing entry 1508 can be displayed as "Publish (2)", where the number can represent the number of currently selected generated results. If an operation is received for the publishing entry 1508, the computing device can bring the multiple result images selected by the user into the publishing process. The publishing page can display the original image of this publishing as well as all generated historical images, thereby ensuring consistent result transmission between the editing page, the AI ​​creation dialogue page, and the publishing page.

[0149] Figure 16 A schematic block diagram of a device for processing media content according to this article is shown. Figure 16 As shown, device 1600 can Figure 1 The device 1600 is implemented in a computing device 102, and the device 1600 includes a first interface display module 1602 for displaying multiple media contents on the first interface; a first instruction receiving module 1604 for receiving a first instruction, the first instruction instructing the execution of a media content processing operation on at least one first media content, the multiple media contents including at least one first media content; and a media content display module 1606 for displaying at least one second media content on the first interface, the at least one second media content being generated based on at least one first media content and the first instruction, the at least one second media content being displayed at the position of at least one first media content.

[0150] In one scenario, the first instruction is a natural language instruction, and the first instruction receiving module 1604 includes: a natural language instruction receiving and determining module, used for receiving a natural language instruction input in the first input area when the first interface includes a first input area and a first media content display area, with multiple media contents displayed in the first media content display area; or a first operation receiving module, used for receiving a first operation triggered on the first interface; and a second input area receiving module, used for displaying an input component on the first interface, displaying a second input area and a second media content display area in the input component, with multiple media contents displayed in the second media content display area, and receiving a natural language instruction input in the second input area.

[0151] In one scenario, the input component further includes multiple prompts, each of which indicates a media content processing operation. The first instruction receiving module 1604 further includes: a first prompt trigger operation receiving module for receiving a trigger operation on the first prompt, the multiple prompts including the first prompt; and a third media content display module for displaying at least one third media content on the input component, the at least one third media content being generated based on at least one second media content and the first prompt, the at least one third media content being displayed at the location of the at least one second media content.

[0152] In one scenario, the first instruction receiving module 1604 further includes: a first control trigger operation receiving module, used to receive a trigger operation on a first control, the first control being on a first interface or on an input component of the first interface; a second media content display module, used to display a dialog interface, displaying at least one second media content in the dialog interface; and a second media content editing module, used to edit at least one second media content in the dialog interface.

[0153] In one scenario, the dialog interface includes a second control, and the first instruction receiving module 1604 further includes: a fifth media content display module for displaying at least one fifth media content on the dialog interface, wherein the at least one media content is obtained by editing at least one second media content; a second control trigger operation receiving module for receiving a trigger operation for the second control; and a first interface display module for displaying a first interface on which at least one fifth media content is displayed.

[0154] In one scenario, the first control is on an input component of a first interface, and the input component also includes multiple prompts, with the first control being one of the multiple prompts.

[0155] In one scenario, prior to receiving the first instruction, the device 1600 further includes a selection operation receiving module for receiving a selection operation, the selection operation being used to select at least one media content from a plurality of media content. The selection operation includes either a selection operation triggered on at least one media content, or a natural language instruction indicating the selection of at least one media content.

[0156] In one scenario, the first interface display module 1602 further includes: a plurality of slot display modules, configured to display a plurality of slots on the first interface, wherein a plurality of media contents are displayed in each of the plurality of slots.

[0157] In one scenario, the media content display module 1606 further includes a second media content slot display module, configured to display at least one second media content in a slot corresponding to at least one media content.

[0158] In one embodiment, the device 1600 further includes: a second operation receiving module for receiving a second operation on a first slot on a first interface, at least one slot including the first slot; and a fourth media content display module for displaying at least one fourth media content corresponding to the first slot, the at least one fourth media content being obtained by performing media content processing operations on the media content in the first slot at different times.

[0159] In one embodiment, the device 1600 further includes: a third operation receiving module for receiving a third operation triggered in response to at least one of the fourth media contents; and a fourth media content display module for displaying the fourth media content corresponding to the third triggered operation in the media content display area of ​​the first interface.

[0160] In one embodiment, the device 1600 further includes: a plurality of slot carousel display module for carousel display of a plurality of slots, wherein the slot being carouseled to is highlighted; and a carousel stop module for stopping the carousel in response to receiving an operation on a pause button.

[0161] In one scenario, at least one first media content includes multiple first media content, and at least one corresponding second media content includes multiple second media content. The media content processing parameters of each first media content to its corresponding second media content are different, or the editing parameters of each of the multiple second media content are different.

[0162] In one embodiment, the device 1600 further includes a prompting display module for displaying a prompting label on the at least one second media content, the prompting label indicating that the at least one second media content is content updated via a first instruction.

[0163] Figure 17 A schematic block diagram of an example device 1700 that can be used to implement this document is shown. Figure 1 The computing device 102 can be implemented using device 1700. As shown, device 1700 includes a central processing unit (CPU) 1701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 1702 or loaded from storage unit 1708 into random access memory (RAM) 1703. Various programs and data required for the operation of device 1700 can also be stored in RAM 1703. CPU 1701, ROM 1702, and RAM 1703 are interconnected via bus 1704. Input / output (I / O) interface 1705 is also connected to bus 1704.

[0164] Multiple components in device 1700 are connected to I / O interface 1705, including: input unit 1706, such as a keyboard, mouse, etc.; output unit 1707, such as various types of displays, speakers, etc.; storage unit 1708, such as a disk, optical disk, etc.; and communication unit 1709, such as a network card, modem, wireless transceiver, etc. Communication unit 1709 allows device 1700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0165] The various processes and handling described above, such as method 200, can be executed by processing unit 1701. For example, in one case, method 200 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1708. In another case, part or all of the computer program can be loaded and / or installed on device 1700 via ROM 1702 and / or communication unit 1709. When the computer program is loaded into RAM 1703 and executed by CPU 1701, one or more actions of example method 200 described above can be performed.

[0166] This document can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this document.

[0167] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0168] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0169] The computer program instructions used to perform the operations described herein may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In one scenario, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays, can execute computer-readable program instructions to implement various aspects of this document by utilizing the state information of the computer-readable program instructions.

[0170] Various aspects of this document are described herein with reference to flowchart illustrations and / or block diagrams of the methods, apparatus (systems), and computer program products thereof. It should be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0171] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0172] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0173] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to this document. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0174] The above description is exemplary and not exhaustive, nor is it limited to the disclosed content. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the description. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to technology in the market, or to enable others skilled in the art to understand the disclosure herein.

Claims

1. A media content processing method, comprising: Multiple media content are displayed on the first screen; Receive a first instruction, the first instruction instructing to perform a media content processing operation on at least one first media content, the plurality of media content including the at least one first media content; as well as At least one second media content is displayed on the first interface. The at least one second media content is generated based on the at least one first media content and the first instruction. The at least one second media content is displayed at the position of the at least one first media content.

2. The method according to claim 1, wherein, The first instruction is a natural language instruction, and receiving the first instruction includes: The first interface includes a first input area and a first media content display area, wherein the plurality of media contents are displayed in the first media content display area, and the natural language instructions input in the first input area are received; or, Receive a first operation, which is triggered in response to the first interface; and An input component is displayed on the first interface, and a second input area and a second media content display area are displayed in the input component. The plurality of media contents are displayed in the second media content display area, and the natural language instructions input in the second input area are received.

3. The method of claim 2, wherein the input component further comprises a plurality of prompts, each of the plurality of prompts indicating a media content processing operation, and the method further comprises: Receive a trigger operation on the first prompt content, wherein the plurality of prompt content includes the first prompt content; as well as The input component displays at least one third media content, which is generated based on the at least one second media content and the first prompt content, and the at least one third media content is displayed at the position of the at least one second media content.

4. The method according to claim 2, further comprising: Receive a trigger operation on a first control, which is on the first interface or on the input component of the first interface; A dialog interface is displayed, in which at least one second media content is displayed; as well as The at least one second media content is edited in the dialog interface.

5. The method of claim 4, wherein the dialog interface includes a second control, and the method further includes: At least one fifth media content is displayed in the dialog interface, and the at least one media content is obtained by editing the at least one second media content; Receive trigger operations for the second control; as well as The first interface is displayed, and the first interface displays at least one fifth media content.

6. The method according to claim 4, wherein, The first control is on the input component of the first interface, and the input component also includes multiple prompts, with the first control being one of the multiple prompts.

7. The method according to claim 1, further comprising, before receiving the first instruction: Receive a selection operation, the selection operation being used to select at least one media content from the plurality of media content; The selection operation includes: a selection operation triggered on the at least one media content, or a natural language instruction indicating the selection of the at least one media content.

8. The method according to claim 1, wherein, The step of displaying multiple media contents on the first interface includes: displaying multiple slots on the first interface, wherein multiple media contents are displayed in each of the multiple slots; Displaying at least one second media content on the first interface includes: displaying the at least one second media content in the slot corresponding to the at least one media content.

9. The method of claim 8, further comprising: Receive a second operation for a first slot on a first interface, wherein the at least one slot includes the first slot; as well as Display at least one fourth media content corresponding to the first slot, wherein the at least one fourth media content is obtained by performing media content processing operations on the media content in the first slot at different times.

10. The method of claim 9, further comprising: Receive a third operation, which is triggered in response to one of the at least one fourth media content; as well as The fourth media content corresponding to the third trigger operation is displayed in the media content display area of ​​the first interface.

11. The method of claim 8, further comprising: The multiple slots are displayed in a carousel, with the slot being featured in the carousel highlighted. as well as In response to receiving an operation on the pause button, the carousel is stopped.

12. The method according to claim 1, wherein the at least one first media content includes a plurality of first media content, and the corresponding at least one second media content includes a plurality of second media content, wherein the media content processing parameters of each first media content to its corresponding second media content are different, or the editing parameters of each of the plurality of second media content are different.

13. The method according to claim 1, further comprising: A prompt icon is displayed on the at least one second media content, indicating that the at least one second media content is content updated via the first instruction.

14. A media content processing apparatus, comprising: The first interface display module is used to display multiple media contents on the first interface. A first instruction receiving module is configured to receive a first instruction, which instructs to perform media content processing operations on at least one first media content, wherein the plurality of media content includes the at least one first media content. as well as A media content display module is used to display at least one second media content on the first interface. The at least one second media content is generated based on the at least one first media content and the first instruction, and the at least one second media content is displayed at the position of the at least one first media content.

15. An electronic device comprising: At least one processor; as well as A memory for storing at least one program, which is executed by the at least one processor to enable the at least one processor to implement the method according to any one of claims 1-13.

16. A computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method according to any one of claims 1-13.

17. A computer program product comprising a computer program, said computer program being executed by a processor to implement the method according to any one of claims 1-13.