Method, apparatus, device, storage medium and product for editing media content
Patent Information
- Application Number
- CN202611141277.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-29
AI Technical Summary
相关的图片编辑方案通常支持对整体画面的编辑处理,但在局部选取与元素级拆分编辑方面,仍面临选取精度和编辑灵活性方面的困难
[0008]以此方式,通过提供区域选取模式和基于选取区域的多元素图层拆分能力,能够实现对媒体内容中局部区域内多个元素的精确拆分与单独编辑,从而提高媒体内容的编辑效率。
Smart Images

Figure CN122845858A_ABST
Abstract
Description
Technical Field
[0001] The examples in this article generally relate to the field of computers, and in particular to methods, apparatus, devices, computer-readable storage media, and products for editing media content. Background Technology
[0002] In the field of media content editing, users have a persistent need for precise editing of specific areas within images. While existing image editing solutions typically support overall image editing, they still face challenges in terms of selection accuracy and editing flexibility when it comes to selecting specific areas and performing element-level segmentation editing. Summary of the Invention
[0003] In a first aspect, a method for editing media content is provided. The method includes: receiving a first operation, the first operation being configured to enable a first selection mode, the first selection mode being configured to receive an operation of a first type; receiving a second operation, the second operation being configured to select at least one region of first media content, the second operation corresponding to the first type; and presenting a plurality of media layers, the plurality of media layers corresponding to different elements within the at least one region.
[0004] In a second aspect, an apparatus for editing media content is provided. The apparatus includes: a first receiving module configured to receive a first operation, the first operation being used to enable a first selection mode, the first selection mode being used to receive an operation of a first type; a second receiving module configured to receive a second operation, the second operation being used to select at least one area of first media content, the second operation corresponding to the first type; and a rendering module configured to render a plurality of media layers, the plurality of media layers corresponding to different elements in the at least one area.
[0005] In a third aspect, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0008] In this way, by providing a region selection mode and the ability to split multiple elements into layers based on the selected region, it is possible to accurately split and edit multiple elements in a local area of media content, thereby improving the editing efficiency of media content.
[0009] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of the example environment is shown; Figures 2A to 2J The diagram illustrates sample interfaces for editing media content using a box selection method in several scenarios; Figures 3A to 3D The diagram illustrates sample interfaces for editing media content using a point-and-click method in several scenarios. Figure 4 The flowcharts show example processes for editing media content in several scenarios; Figure 5 Block diagrams of apparatuses for editing media content in various scenarios are shown; and Figure 6 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation
[0011] The examples in this document will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.
[0012] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.
[0013] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0014] The examples in this document may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and rules. In the examples, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.
[0015] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0016] As used in this article, the term "media content" refers to digital content used for editing or processing, which may include, but is not limited to, images, video frames, composite images, or other visual digital information.
[0017] The term "selection mode" as used in this article refers to an interaction mode used to determine how a user selects a target on media content. Selection modes may include, but are not limited to, box selection mode and point selection mode, with different selection modes corresponding to different types of operations.
[0018] As used in this article, the term "region" refers to a designated portion of media content, the extent of which can be defined by dragging or other means. Regions can have various geometric shapes, including but not limited to rectangles, ellipses, or irregular shapes.
[0019] The term "media layer" as used in this article refers to a separately manipulated unit of content extracted from media content. Each media layer may correspond to one or more elements within the media content and can be moved, scaled, deleted, or otherwise edited independently. In some contexts, a "media layer" may also be referred to as a layer or editable layer.
[0020] As used in this article, the term "element" refers to an identifiable and separable visual object in media content, which may include, but is not limited to, characters, animals, text, graphics, backgrounds, or other visual components.
[0021] As used in this article, the term "content item" refers to a thumbnail representation of at least a portion of a media layer within a layer component. There is a correspondence between content items and media layers; users can select the corresponding media layer by selecting a content item.
[0022] As used in this article, the term "first component" refers to a user interface component used to present content items corresponding to multiple media layers in a hierarchical order. The first component may include, but is not limited to, layer panels, sidebars, or floating windows.
[0023] The term "entity segmentation" as used in this article refers to a technique that uses computer vision models to perform semantic analysis on media content, in order to identify and separate different visual entities within the media content. Entity segmentation can be implemented based on deep learning models, including but not limited to instance segmentation models or semantic segmentation models.
[0024] The term "media material" as used in this article refers to material resources that can be used in the process of editing media content. Media material can be converted from the media layer or exist independently, and is used to add, replace, or combine content in the editing screen.
[0025] As mentioned above, in the field of image editing, the local editing capabilities supported by relevant editing solutions are mainly limited to local removal and modification. Related image editors offer interactive editing capabilities and various on-screen operation functions. However, users have specific editing needs for local areas and specific selections, such as local processing, removal of unwanted objects, local enhancement or color correction, and adjustment or scaling of local elements. When using related tools for local adjustments, the core requirement is precise local adjustments while maintaining consistency across other parts, replacing complex manual adjustment methods using masks. Therefore, providing more flexible and precise capabilities for selecting and splitting local elements for editing has become a problem that needs to be solved.
[0026] A scheme for editing media content is proposed. According to the scheme, an electronic device receives a first operation from a user to enable a first selection mode, which is used to receive a first type of operation (e.g., a drag-and-drop operation). Then, the electronic device receives a second operation to select at least one area of first media content, the second operation corresponding to the first type. Subsequently, the electronic device presents multiple media layers, each corresponding to a different element within the at least one area.
[0027] Using the above approach, users can precisely select a specific area within media content through a selection mode, and then split multiple different elements within that area into different media layers, enabling individual editing of each element. Since there's no need for global splitting of the entire media content, the electronic device only needs to perform element recognition and layer generation calculations on the user-selected local area, thus reducing client-side computational overhead and memory usage. Simultaneously, since only the processing results of a local area need to be rendered to the interface, the number of rendering objects and drawing data that the client needs to update is reduced, thereby lowering interface update overhead.
[0028] The following describes various examples of this scheme in further detail with reference to the accompanying drawings.
[0029] Example Environment Figure 1 A schematic diagram of example environment 100 is shown. (e.g.) Figure 1 As shown, example environment 100 may include electronic device 110.
[0030] In this example environment 100, electronic device 110 may run an application 120 that supports editing media content. Application 120 may be any suitable type of application for editing media content, including but not limited to: image editing applications, video editing applications, multimedia authoring applications, or other suitable applications. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.
[0031] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting the editing of media content.
[0032] In some cases, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, virtual reality (VR) / augmented reality (AR) devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some cases, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0033] Server 130 can be a single physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and basic cloud computing services. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 in electronic devices 110 that support editing media content.
[0034] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some cases, server 130 and electronic device 110 can exchange signaling information through their communication connection.
[0035] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the scheme.
[0036] The following description of the example will continue with reference to the accompanying drawings.
[0037] Example Interaction Figure 2A A schematic diagram of an interface 200A for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 2A This describes the process by which a user enters partial editing mode.
[0038] like Figure 2A As shown, on interface 200A, electronic device 110 displays first media content 202, a point-and-click control 204, a box-selection control 206, and multiple functional controls 208-1 to 208-5. The first media content 202 contains multiple elements. For example, the first media content 202 includes elements 202-1 to 202-4. Elements 202-1 to 202-4 correspond to different objects in the first media content 202. For example, element 202-1 corresponds to the "deer" in the first media content 202, and element 202-2 corresponds to the "kangaroo" in the first media content 202. The point-and-click control 204 is used to enable a second selection mode that receives a click-and-select operation. The box-selection control 206 is used to enable a first selection mode that receives a box-selection operation. The multiple functional controls 208-1 to 208-5 are used to trigger corresponding editing functions.
[0039] In one scenario, in response to receiving a user's trigger operation on the selection control 206 (i.e., the first operation), the electronic device 110 activates the first selection mode and switches the interface 200A from global editing mode to partial editing mode. After switching to partial editing mode, the image presented in the first media content 202 can be changed to the result of algorithmic processing and overlay, such as removing decorative materials and filter effects, but retaining portrait beautification and existing editing effects. Simultaneously, the function control area switches to the function items corresponding to the selection interaction.
[0040] In some cases, after switching to partial editing mode, if the user has not yet selected any area on the first media content 202, the function controls may be in a non-interactive state (e.g., grayed out); once the user selects any area on the first media content 202, the corresponding function controls become interactive. In some cases, the functions supported in partial editing mode may include, but are not limited to, eliminating, removing backgrounds, adjusting, copying, and splitting elements.
[0041] In this way, the number of operation steps and interface jumps required to access the partial editing function is reduced. Furthermore, since the electronic device 110 only needs to load the functional controls and processing logic related to the selection interaction in partial editing mode, rather than maintaining all the resources for global editing and partial editing simultaneously, the client's memory usage and rendering overhead are reduced.
[0042] In some cases, the electronic device 110 presents a first set of controls, which corresponds to a first set of editing functions. In response to the activation of a first selection mode, the electronic device 110 stops presenting the first set of controls and presents a second set of controls, which corresponds to a second set of editing functions. At least some of the functions in the second set of editing functions differ from those in the first set of editing functions, and these at least some functions correspond to the first selection mode.
[0043] As an example, such as Figure 2A As shown, the electronic device 110 can display a first set of controls in the interface 200A. The first set of controls may include function controls 208-1 to 208-5. Function controls 208-1 to 208-5 are used to trigger the first set of editing functions. For example, when function control 208-3 is clicked, the electronic device 110 can enable a cropping function to support the user in cropping the first media content 202. When the user enables the first selection mode through the selection control 206, such as... Figure 2B As shown, the electronic device 110 can replace the first set of controls with the second set of controls. For example, in interface 200B, the electronic device 110 can stop displaying function controls 208-1 to 208-5 and display the eliminate control 212, copy control 214, and split control 216. The eliminate control 212, copy control 214, and split control 216 correspond to the second set of editing functions. For example, the copy control is used to trigger the copying of the area or element selected by the user. Thus, the interface only retains the editing functions applicable to the currently selected area, reducing the occupation of the operation area by irrelevant functions and user accidental touches.
[0044] In some scenarios, when the first selection mode is enabled, the electronic device 110 can add at least one control to present a second set of controls on the interface, while retaining at least some of the controls in the first set. In other words, although the first set of controls stops being presented as the original control combination, at least some of the controls in the first set can continue to be presented as controls in the second set. The newly added at least one control corresponds to an editing function that can be executed in the first selection mode. For example, after the user enables the first selection mode through the selection control 206, the electronic device 110 can retain the original function controls 208-4 and add an elimination control 212, a copy control 214, and a split control 216.
[0045] In other scenarios, when the first selection mode is enabled, the electronic device 110 can remove at least one control from the first group of controls and present the remaining controls from the first group as the second group of controls. The removed control corresponds to an editing function that is unrelated to the first selection mode, unavailable in the first selection mode, or unsuitable for performing on the selected area. For example, the first group of controls may include controls for triggering editing functions such as cropping, filtering, and background settings. When the user enables the first selection mode via the selection control 206, the electronic device 110 can stop presenting controls for triggering cropping, filtering, or background setting functions, while retaining controls for triggering adjustments to the selected area, thus forming the second group of controls.
[0046] In some cases, the first group of controls corresponds to a first number, and the second group of controls corresponds to a second number, where the first number differs from the second number. For example, such as... Figure 2A As shown, the first group of controls may include functional controls 208-1 to 208-5, with a total number of 5. When the user enables the first selection mode through the selection box control 206, as shown... Figure 2B As shown, the electronic device 110 stops displaying function controls 208-1 to 208-5 and displays eliminate control 212, copy control 214, and split control 216, with a second number of 3, which is less than the first number. In other examples, the electronic device 110 may add one or more controls corresponding to the first selection mode while retaining at least some of the controls in the first group of controls, making the second number greater than the first number.
[0047] Figure 2B A schematic diagram of an interface 200B for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 2B This describes the process of selecting an area using a box selection operation.
[0048] like Figure 2B As shown, on interface 200B, electronic device 110 can display first media content 202 and an area 210 obtained through drag-and-drop operations. Area 210 is a rectangular selection area obtained based on the first start position and the first end position of the first drag-and-drop operation, and this area 210 covers elements 202-1 and 202-2. At the associated position of the upper left corner of area 210, electronic device 110 displays a first control 210-1, which is used to delete the current selection box. In the function control area of interface 200B, electronic device 110 displays an erase control 212, a copy control 214, and a split control 216.
[0049] In some cases, the user performs a first drag operation in the preview area (i.e., an implementation of a second operation, the second operation corresponding to a first type of operation, which can be a drag operation). The first start position and the first end position of this first drag operation determine the range of the first area 210. Figure 2B As an example, region 210 has a rectangular geometry obtained through the first drag operation. In this way, users can more easily select the local area they want to edit, reducing the steps required to define the editing scope through multiple clicks or menu navigation. The electronic device 110 only needs to allocate selection rendering resources and subsequent processing calculations for the user-specified local area, without needing to preload processing data for the entire media content, thereby reducing client-side computational overhead and memory usage.
[0050] In some cases, electronic device 110 displays a first control 210-1 at an associated location within the first area 210. In response to receiving a user interaction with the first control 210-1, electronic device 110 stops displaying the first area 210. Figure 2B As an example, the first control 210-1 is a delete button located in the upper left corner of area 210. Clicking this button cancels the current selection. In this way, users can cancel the selection with a single click without performing an additional undo operation, reducing the number of steps required to cancel the selection. After canceling the selection, the electronic device 110 can immediately release the rendering resources and mask data associated with that selection, thus reducing the client's memory usage.
[0051] In some cases, the second operation may also include a second drag operation. After the user completes the first drag operation to form the first region 210 (e.g., the user's finger leaves the screen), the user can perform another drag operation in the preview area to form a second region. The electronic device 110 presents the second region on the first media content, based on the second start position and the second end position. In the case of multiple selection boxes, the user can delete the corresponding selection boxes individually using the delete button on each selection box. In this way, based on the above selection mode scheme, the user can define multiple non-adjacent selection regions simultaneously without having to start a splitting process for each region, reducing the number of repetitive operations in multi-region selection scenarios. The electronic device 110 can merge the processing requests for multiple regions into a single batch processing, reducing the number of processing requests sent to the server 130, thereby reducing the bandwidth consumption of the end-to-cloud communication and the server query load.
[0052] In some cases, electronic device 110 receives a third operation from a user for editing a first attribute associated with at least one region. In response to this operation, electronic device 110 renders at least one region in a first style, which represents the edited first attribute. As an example, such as... Figure 2B As shown, electronic device 110 can receive a third operation via interface 200B. The third operation can be associated with region 210 and can, for example, be used to edit a first attribute of region 210. For instance, the third operation can be used to adjust the display position of region 210 within the first media content 202. As another example, the third operation can be used to adjust the display size of region 210. As an example, the third operation can indicate a first position. After receiving the third operation, electronic device 110 can present region 210 in a first style. This first style can indicate that region 210 has been moved to a first position. As another example, the third operation can also be a scaling operation on region 210. After receiving the scaling operation, electronic device 110 can present region 210 in a first style. The first style can indicate the display size corresponding to region 210 after scaling. In this way, users can adjust the selected region with more flexible operation, thereby effectively improving the efficiency of human-computer interaction.
[0053] In some cases, when the area determined by the selection operation only covers a part of an element, the processing model used by the electronic device 110 is still able to recognize the complete outline of the element and include the part of the element outside the selection box into the corresponding media layer, so as to ensure that the split media layer contains the complete visual representation of the element.
[0054] Figure 2C A schematic diagram of the interface 200C for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 2C This describes the process of splitting the selected area into multiple media layers.
[0055] like Figure 2B As shown, on interface 200B, electronic device 110 can display control 216. Control 216 is used to trigger the splitting of layers in at least one selected area (e.g., area 210). When control 216 is triggered, electronic device 110 can use an appropriate algorithm or model to split multiple elements in area 210 into multiple media layers, where the multiple media layers correspond to different elements in at least one area (i.e., splitting multiple elements in area 210 into different media layers). For example, as... Figure 2CAs shown, electronic device 110 can split element 202-2 in region 210 from the selected area to obtain media layer 218. Media layer 218 is presented as a bounding box, indicating that it is selected. In this state, media layer 218 can be edited (e.g., moved, flipped, scaled, etc.). It should be understood that although multiple elements in region 210 are split into multiple media layers, only the bounding box of one media layer can be presented after splitting to indicate that the media layer is selected by default after splitting. As an example, when multiple media layers are split, electronic device 110 can select the topmost media layer by default according to the hierarchy of the media layers. In some scenarios, electronic device 110 can simultaneously set all the split media layers to the selected state, or determine one or more media layers to be selected by default based on at least one of the following: the stacking order of each media layer, its position in region 210, its area, its prominence, its confidence level, or the user's previous selection.
[0056] In some scenarios, electronic device 110 can also receive selection operations for all elements in the first media content and determine all identified elements as elements to be split. For example, a user can trigger a select all control or select all elements by covering a selection area that covers the entire first media content. In response to a splitting request for all elements (e.g., triggering control 216), electronic device 110 can use an appropriate algorithm or model to split all elements in the first media content into multiple individually editable media layers.
[0057] In some cases, after layer splitting, the electronic device 110 can also present a first component 220, which is a layer panel that presents multiple content items 222, 224, and 226 in hierarchical order. These multiple content items correspond to different media layers. Content item 222 is used to display at least a portion of media layer 218, corresponding to element 202-2. Content item 224 is used to display another media layer (e.g., ...). Figure 2D At least a portion of the media layer 228 shown corresponds to element 202-1. Content item 226 is used to display at least a portion of the background media layer. At the associated location of content item 222, the electronic device 110 also presents a delete control 222-1 for requesting the deletion of the corresponding media layer.
[0058] In some cases, electronic device 110 presents multiple content items 222, 224, and 226 in a first order within first component 220. These multiple content items correspond to multiple media layers obtained by splitting, and the first hierarchical relationship of the multiple media layers is related to the first order of the content items in first component 220. The term "hierarchical relationship" as used herein can refer to the stacking or occlusion relationship of multiple media layers within media content, and can be implemented by the hierarchical parameters of the corresponding media layers, the drawing order, or the order in the rendering queue. Figure 2C As an example, in the first component 220, the content item 222 is located at the top, indicating that the corresponding media layer 218 has the highest layer hierarchy on the canvas. For example, media layer 218 is drawn or overlaid on top of other media layers. In this way, based on the above-mentioned scheme of splitting elements in an area into multiple media layers, users can intuitively view all the split media layers and their hierarchical relationships through the layer panel, without having to select elements on the canvas one by one to confirm the existence and order of each layer, reducing the number of exploratory operations required for layer management. The electronic device 110 reduces the rendering overhead of the interface and the number of interface objects that need to be updated by presenting media layers in the form of thumbnailed content items in the layer panel, rather than rendering the selection status indicators of all media layers at full size on the interface simultaneously.
[0059] Figure 2D A schematic diagram of the interface 200D for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 2D This describes the process of selecting a media layer through the Layers panel.
[0060] like Figure 2D As shown, the electronic device 110 can receive a selection operation on the content item 224. After receiving the selection operation, the electronic device 110 can display the media layer 228 in a bounding box style in the interface 200D to indicate that the media layer is selected and can be edited. In some scenarios, after splitting the media layer, the electronic device 110 can also receive a selection operation on the element 202-1. After receiving the selection operation, the electronic device 110 can also display the media layer 228 in a bounding box style in the interface 200D. Furthermore, the electronic device 110 can also display the content item 224 in a preset style in the first component 220 to indicate that the corresponding media layer 228 has been selected. For example, after the media layer 228 is selected, the electronic device 110 can associate it with the content item 224 and display a delete control. As another example, after the media layer 228 is selected, the electronic device 110 can adjust the background color of the content item 224 to indicate that the content item 224 and its corresponding media layer 228 have been selected.
[0061] In some cases, electronic device 110 receives a seventh operation from the user on a first content item (e.g., content item 224), the seventh operation representing a selection of the first content item. In response to this operation, electronic device 110 presents a second media layer 228 in a third style, the third style representing that the second media layer 228 is selected. Figure 2D As an example, when a user clicks on content item 224 in the first component 220, the media layer 228 in the canvas is presented as a bounding box. In this way, building upon the previous scheme of presenting media layer content items in the layer panel, the user can precisely select the corresponding media layer on the canvas by clicking on the content item in the layer panel, even if the media layer is obscured by other elements or difficult to click directly. This eliminates the need for repeated attempts to select elements on the canvas, reducing the number of attempts required to select obscured layers. The electronic device 110 only needs to update the selection status indicator (such as the bounding box) of the selected media layer, without re-rendering the entire canvas, reducing the local refresh overhead of the canvas area.
[0062] Figure 2E A schematic diagram of the interface 200E for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 2E This describes the process of adjusting the layer order.
[0063] In some cases, electronic device 110 receives an eighth operation from a user, which adjusts multiple content items in a first component from a first order to a second order. In response to this operation, electronic device 110 presents multiple media layers according to a second hierarchical relationship corresponding to the second order. Figures 2D to 2E As an example, electronic device 110 can receive a drag operation on content item 224. Upon receiving the drag operation, electronic device 110 can move content item 224 to the position of content item 222 in the direction indicated by the drag operation. After the position is moved, the vertical order of content item 222 and content item 224 is relative to... Figure 2D The swapping changes the stacking order of media layers on the canvas. In this way, building upon the layer panel scheme described above, users can directly adjust the stacking order of media layers by dragging content items in the layer panel, eliminating the need for multiple forward or backward operations on the canvas to gradually adjust the layers, thus reducing the steps required for layer sorting. The electronic device 110 redetermines the stacking order of each media layer based on the adjusted order information and updates the canvas rendering. It only needs to redraw the affected media layers according to the new layer hierarchy, without recalculating the rendering order of all layers on the entire canvas, reducing the rendering overhead of the canvas.
[0064] Figure 2F A schematic diagram of the interface 200F for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 2FThis describes the process of editing the attributes of the media layer.
[0065] In some cases, electronic device 110 can receive a tenth operation for editing a second attribute of the fifth media layer; and present the fifth media layer in a fifth style, the fifth style representing the edited second attribute. Figure 2F As an example, electronic device 110 can receive a tenth operation. The tenth operation may include, for example, a drag operation on media layer 228. Such a drag operation can indicate a second position. Upon receiving such a drag operation, electronic device 110 can move media layer 228 to the second position (corresponding to the fifth pattern) to indicate that media layer 228 is relative to... Figure 2D The position has changed. In some scenarios, the tenth operation can also be a scaling operation for media layer 228. Upon receiving the scaling operation, electronic device 110 can present media layer 228 at the adjusted display size (corresponding to the fifth style) to indicate that media layer 228 is relative to... Figure 2D The size has changed. In other scenarios, the tenth operation can include not only adjusting the display position and size, but also adjusting the presentation angle, flip state, and color parameters of the fifth media layer. In this way, based on the above selection and splitting scheme, users can directly adjust the position and size of the split media layers on the canvas by dragging and zooming gestures, without having to modify these attributes through an additional attribute panel or parameter input, reducing the number of operation steps and interface jumps required for attribute editing. The electronic device 110 only needs to recalculate the transformation matrix and render the single media layer being edited, without having to re-render all layers in the entire canvas, reducing the rendering overhead in attribute editing scenarios.
[0066] Figure 2G A schematic diagram of a 200G interface for editing media content based on various scenarios is shown below. (See below for reference.) Figure 2G This describes the process of deleting the media layer.
[0067] In some cases, electronic device 110 receives a first request from a user for deleting a third media layer (e.g., via...). Figure 2F (Triggered by the delete control 224-1 in the middle). In response to this request, the electronic device 110 stops rendering the third media layer and the second content item, the second content item being used to display at least a portion of the third media layer. Figures 2F to 2GAs an example, after triggering the delete control 224-1, both media layer 228 and content item 224 stop appearing on the interface, and only other content items remain in the first component 220. In this way, based on the layer panel solution described above, users can delete unwanted media layers with a single click using the nearest delete control. Simultaneously, the corresponding displays in the canvas and layer panel are updated synchronously, eliminating the need for users to perform deletion operations separately in the canvas and layer panel, thus reducing the steps required for layer deletion. After deleting the media layer, the electronic device 110 releases the rendering resources and memory space occupied by that media layer, correspondingly reducing the client's memory usage.
[0068] Figure 2H A schematic diagram of an interface 200H for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 2H This describes the process of copying area elements by selecting them using a box.
[0069] like Figure 2H As shown, after selecting region 210, electronic device 110 can receive a trigger operation on control 214. When control 214 is triggered, electronic device 110 can present media layer 230 on interface 200H. Elements in region 210 can form media layer 230 with the same visual representation, and this combined media layer 230 is presented as an editable new media layer with a bounding box. In some cases, when copying by selecting a box, electronic device 110 directly copies the content of the entire area selected by the user, which is equivalent to cropping and copying the area.
[0070] In some cases, electronic device 110 receives a third request related to at least one of a plurality of media layers. In response to this request, electronic device 110 presents at least one media element corresponding to the at least one media layer. Figure 2C and Figure 2HAs an example, elements that have been split or copied can be converted into newly added image assets in the draft of the current media content. As an example, electronic device 110 can receive a third request. This third request, referred to as an export request, is used to export at least one split media layer as media assets (image assets). After exporting, electronic device 110 can present the exported media assets in an appropriate area or component (such as a media management component). Users can then perform further editing or related operations on the media assets. As an example, users can edit the exported media assets in another interface, such as moving, scaling, cropping, color correcting, adding effects, or combining them with other media components. As another example, users can also apply the media assets to video generation, graphic design, or other content creation tasks. In this way, the split elements can be reused in different editing tasks or creation projects, reducing the need for users to repeatedly perform element recognition, segmentation, and extraction operations, and improving the efficiency of media content creation.
[0071] In some cases, such as Figure 2I As shown, after splitting the media layers, the electronic device 110 can present a first component 220. In the first component 220, the electronic device 110 can be associated with the split media layers, presenting controls for adjusting the hierarchical relationship of the media layers. For example, the electronic device 110 can present a first adjustment control 220-1 and a second adjustment control 220-2 in the first component 220. The first adjustment control 220-1 and the second adjustment control 220-2 correspond to content items 222 and 224, respectively; that is, controls 220-1 and 220-2 correspond to different media layers and are used to adjust the hierarchical relationship of different media layers. In some scenarios, the electronic device 110 can also present content item 226 in the first component 220, which corresponds to the background media layer. However, in this example, the electronic device 110 may not present the adjustment controls corresponding to the background media layer; that is, the background media layer can be configured not to support hierarchical order adjustment and is presented in a locked state at the bottom of the content item sequence. In other scenarios, to ensure interface integrity, the electronic device 110 may also present a third adjustment control corresponding to the background media layer in the first component 220. However, in this example, the third adjustment control can be configured to be non-interactive. For example, the third adjustment control can be presented in a grayed-out style.
[0072] In some cases, such as Figure 2JAs shown, after splitting the media layer, the electronic device 110 can present only the content items corresponding to the split media layer in the first component 220, without presenting the content items corresponding to the background media layer, to indicate that the background media layer is configured not to support the adjustment of the hierarchical order. For example, the electronic device 110 can present only the content items 222 and 224 corresponding to the split media layer in the first component 220.
[0073] Figure 3A A schematic diagram of an interface 300A for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 3A This describes the process of selecting elements using the point selection mode.
[0074] like Figure 3A As shown, on interface 300A, electronic device 110 displays first media content 302, candidate elements 302-1 and 302-2, a selection control 304, and multiple functional controls 306-1 to 306-5. The selection control 304 is selected, indicating that a second selection mode is enabled. When the second selection mode is activated, candidate elements 302-1 and 302-2 are presented in a second style (e.g., color highlighting), indicating that these candidate elements support click selection. For example, when the second selection mode is activated, electronic device 110 can display appropriate animation effects, such as using specific colors to represent selectable candidate elements in the first media content 302.
[0075] In some cases, when the electronic device 110 enters the point-and-click mode, it can use an entity segmentation model to pre-identify and segment the first media content 302 to obtain multiple candidate elements. Further, the electronic device 110 can trigger the determination of a second set of elements from the identified candidate elements. For example, such as... Figure 3AAs shown, the electronic device 110 can provide the first media content 302 to the first model for segmentation. Further, the first model can output multiple candidate elements, along with the corresponding segmentation region, confidence level, and boundary information for each candidate element. The electronic device 110 can trigger a filtering process for multiple candidate elements based on at least one of the following: confidence level, display area, visible ratio, boundary integrity, and degree of overlap with other candidate elements. For example, the electronic device 110 can trigger the removal of candidate elements with a confidence level below a preset threshold, noise regions with excessively small areas, and candidate elements corresponding to the background, and merge duplicate segmentation regions indicating the same visual object. The electronic device 110 can also trigger a sorting of the filtered candidate elements based on their salience and editability in the first media content 302, and determine a predetermined number of the top-ranked candidate elements as a second group of elements. For example, the second group of elements may include elements suitable for user selection and editing, such as characters, animals, text, or foreground objects. The electronic device 110 presents the second group of elements in a second style to indicate to the user that the corresponding element is selectable. In this way, invalid or duplicate candidate elements in the interface can be reduced, the number of candidate elements rendered can be decreased, and the accuracy of the user selecting the target element can be improved.
[0076] In some scenarios, initial recognition can be completed beforehand after the user imports the image. In other cases, cached recognition results can be used directly for already uploaded images, eliminating the need for repeated processing. Users can select candidate elements in the preview area by clicking, and deselect by double-clicking the same element or using other appropriate interactions. Deselected elements are no longer highlighted. Users can also select multiple elements in different ways and deselect them one by one.
[0077] In some cases, electronic device 110 receives a fourth operation from the user on the click control 304. This fourth operation enables a second selection mode, which receives a second type of operation (i.e., a click-to-select operation). In response to enabling the second selection mode, electronic device 110 presents a second set of elements 302-1 and 302-2 in a second style, indicating that the second set of elements supports selection. Subsequently, electronic device 110 receives a fifth operation from the user to select a first set of elements of the first media content 302, corresponding to the second type. Figure 3AAs an example, when the user clicks on candidate elements 302-1 and 302-2, the electronic device 110 can adjust these elements to the selected state. In this way, based on the aforementioned box selection mode scheme, the user can also select a single target element through precise clicking, without having to draw a selection box on the canvas to define the area, reducing the number of steps required to select a single element. The electronic device 110 directly responds to the user's clicking operation with the pre-segmented candidate elements, without having to re-perform the segmentation calculation for each click, reducing the response latency and computational overhead of the clicking interaction.
[0078] In some cases, electronic device 110 receives a sixth operation from the user, which represents a selection of a third group of elements, different from the second group of elements. Further, electronic device 110 may present the third group of elements in a second style. As an example, electronic device 110 also supports user selection of other elements that were not segmented during the pre-segmentation process. For example, such as... Figure 3A As shown, element 302-3 is not segmented during the pre-segmentation process. Electronic device 110 can receive a selection operation for selecting element 302-3. Upon receiving such a selection operation, electronic device 110 can trigger further image segmentation of the first media content 302 to adjust element 302-3 into a selectable element. Furthermore, electronic device 110 can present the selected element 302-3 in a second style (e.g., a highlight style or other appropriate style). In this case, the selected element 302-3 also supports editing events such as copying, adjusting, and splitting layers.
[0079] In some cases, for areas not covered by entity segmentation, when a user clicks on such an area, the electronic device 110 can use a second model to further refine the identification of the user's click location, thereby determining the element corresponding to that location and presenting it in a second style. As an example, when a user clicks on a target location in the first media content 302 that is not identified as a candidate element by the first model (e.g., element 302-3), the electronic device 110 determines the position information of the target location in the image coordinate system of the first media content 302 based on the touch coordinates of the sixth operation. This position information may include point coordinates, a local region centered on those point coordinates, an area covered by the touch trajectory, or a prompt mask generated based on the touch location.
[0080] The electronic device 110 provides the first media content 302 and the aforementioned location information to the second model. Using the location information as a segmentation prompt, the second model identifies and segments visual objects located at or covering that location within the first media content 302, outputting one or more target segmentation regions and their confidence levels. For example, when a target location is simultaneously located within multiple overlapping elements, the second model can determine multiple segmentation regions covering the target location and filter them based on at least one of the following: confidence level, visible area, boundary integrity, hierarchical relationship, and degree of matching with the target location. Elements corresponding to segmentation regions that meet preset conditions are then identified as the third group of elements. In some cases, the electronic device 110 can identify the element with the highest matching degree as the third group of elements; in other cases, it can identify multiple elements covering the target location with a confidence level higher than a preset threshold as the third group of elements for further selection by the user.
[0081] The electronic device 110 presents the third group of elements in a second style, such as highlighting the outline or region corresponding to the third group of elements, thereby indicating to the user the selectable elements determined by the second model. In this way, the user can further refine the selection of objects that are not recognized by the first model, are inaccurately recognized, or are partially obscured by other elements by clicking on the location, without having to manually draw the object boundaries; the electronic device 110 can also narrow down the recognition and segmentation range of the second model based on the location information, reducing the amount of data processed by the model and the computational overhead, and improving the response speed of fine selection.
[0082] In this way, users can further select more refined elements beyond the pre-segmentation results without switching to other tools or manually drawing masks, reducing the number of operation steps and tool switching times in fine selection scenarios. The electronic device 110 only needs to call the second model to perform fine segmentation on the local area of the user's click position, rather than re-performing global segmentation on the entire media content, reducing the computational overhead and response latency in fine selection scenarios.
[0083] Figure 3B A schematic diagram of an interface 300B for editing media content according to certain scenarios is shown below. (See below for reference.) Figure 3B This describes the process of forming the same media layer after selecting multiple elements.
[0084] like Figure 3BAs shown, the user can select elements 302-1 and 302-2 and click on control 306-5. After control 306-5 is clicked, the electronic device 110 can display a media layer 308 formed by the selected elements 302-1 and 302-2 on the interface 300B. The media layer 308 is presented as a movable, scalable, or editable media layer with a bounding box. Furthermore, the electronic device 110 can display the content item 312 and background content item 314 of the media layer 308 in the first component 310. In some cases, when the user selects multiple elements, the electronic device 110 extracts all the selected elements to form a single media layer, which the user can then move and scale.
[0085] Figure 3C and 3D Schematic diagrams of interfaces 300C and 300D for editing media content according to certain scenarios are shown below. (Refer to the following...) Figure 3C and 3D This describes the process of copying elements by clicking to form a new media layer.
[0086] In some cases, electronic device 110 receives a ninth operation from a user, representing a selection of a third region, which includes a first element (e.g., element 302-2). In response to receiving a second request (e.g., a user triggering a copy control 306-4), electronic device 110 presents a fourth media layer 316 corresponding to the second element, which has the same visual representation as the first element. Figures 3C to 3D As an example, element 302-2 forms a new media layer 316 with the same visual representation, which is presented as an editable bounding box.
[0087] In this way, based on the above solution, users can copy selected elements with a single click and directly obtain a new, editable layer on the canvas, without having to export, edit externally, and re-import elements from the screen. This reduces the number of operation steps and application switching required in element copying scenarios. When copying elements, the electronic device 110 can reuse existing element segmentation results to generate a new media layer without re-performing segmentation and extraction calculations, thus reducing the computational overhead and response latency of the copying operation.
[0088] In some cases, the electronic device 110 further receives a fourth request related to the first media content and presents second media content, which includes the first element and the second element. As an example, after a user obtains a second element with the same visual representation as the first element through a copy operation, the user can trigger a redraw control to initiate the fourth request. The electronic device 110 provides the first media content, including the first and second elements, the position and outline information of the first and second elements respectively, and preset redraw instructions to the image generation model. The image generation model performs a fusion process on the first element, the second element, and their surrounding areas based on the position, size, pose, and occlusion relationship of the first and second elements on the canvas to generate the second media content. In the second media content, the first and second elements are redrawn and merged into the same image, making the edges, textures, lighting, shadows, colors, and perspective relationships around the second element match the background and other content in the image.
[0089] In some cases, the electronic device 110 may define only the second element and its surrounding local area as the area to be redrawn, while keeping the content outside the area unchanged; alternatively, the electronic device 110 may also trigger a complete redraw of the first media content. The electronic device 110 may construct a mask to constrain the area to be retained based on the first media content, and construct a mask for the area to be generated based on the contours of the first and second elements, thereby controlling the image generation model to naturally integrate the copied second element into the image while retaining the original image content. After the image generation model completes processing, the electronic device determines the complete image output by the model as the second media content, and uses the second media content to replace or cover the first media content for presentation.
[0090] In this way, the copied elements, which are separate media layers, can be regenerated and combined with the original image to form a new picture. This reduces problems such as abrupt edges, inconsistent lighting, or unnatural occlusion between the second element and the background, and makes it easier for users to continue editing, saving, or exporting the redrawn second media content as a whole.
[0091] In some cases, after the editing function effect is applied, the electronic device 110 can display a comparison control on the interface, which allows the user to switch between the state before and after editing for preview.
[0092] The above solution allows users to precisely split and edit local areas of media content at the element level, achieving precise adjustments while maintaining consistency with other parts, thus replacing the editing method that requires complex masking and manual adjustment.
[0093] In some cases, the electronic device 110 can first use an entity segmentation model to determine the segmentation region and outline of the selected element, and separate the element from the original image into a separate media layer. After the selected element is extracted to form a separate media layer by point selection, the electronic device 110 can use an image generation model to fill in the empty areas in the original image caused by the extraction of the element in real time, so as to generate filled content that blends with the background of the original image, and present the filled complete image as the base image. After selecting an appropriate area of media content by box selection, the electronic device 110 can use a model that supports multi-object segmentation or multi-layer generation to determine multiple elements in the selected area and their corresponding segmentation regions, and generate multiple media layers. Furthermore, after multiple elements in the selected area are split into multiple media layers by box selection, the electronic device 110 can also use an image generation model to fill in the gaps in the selected area. Whether selecting or selecting a media layer by point or frame, the image generation model can generate fill content that matches the original image when filling in holes, based on at least one of the following: texture, color, brightness, lighting direction, shadow, perspective, and style of the image content surrounding the hole. This ensures that the fill content is harmonious with the surrounding image content in terms of lighting and visual effects. Therefore, it reduces obvious holes, abrupt edge changes, or inconsistent lighting caused by element separation, resulting in a more natural and complete visual effect for the filled base image.
[0094] In some situations, the electronic device 110 supports undoing and redoing applied editing effects. In partial editing mode, undo and redo can be applied separately to each editing capability. Even after exiting partial editing mode, the user can still undo for each capability.
[0095] In some scenarios, the element recognition, segmentation, and completion processes involved in the above solutions can be implemented through various methods. In a first implementation, the processing model can be entirely deployed on the electronic device 110, where element recognition, segmentation, and layer generation are performed locally without communication with the server 130. In a second implementation, the processing model can be entirely deployed on the server 130, where the electronic device 110 sends media content and interactive input information to the server 130, which then performs element recognition and segmentation processing and returns the results to the electronic device 110. In a third implementation, the electronic device 110 and the server 130 jointly perform the processing: for example, the electronic device 110 can perform preliminary entity segmentation locally to quickly present candidate elements, while when the user requires more refined selection or higher-quality layer generation, the electronic device 110 can send relevant data to the server 130 for further processing.
[0096] In some cases, the electronic device 110 can also present a global split control in the editing panel of the global editing mode. In response to a user's triggering of the global split control, the electronic device 110 acquires the third media content to be split; the third media content can be an image after removing overlay materials, filters, or added backgrounds, and can retain the basic processing effects already applied to the third media content. The electronic device 110 can use an appropriate model to analyze the aforementioned third media content to identify multiple elements, generate editable media layers for each element, and fill in the empty areas after element removal to obtain a filled base image. After processing, the electronic device 110 expands the layer panel, presenting the content items corresponding to the multiple media layers in stacking order, with the topmost media layer selected by default; after the user deselects, the electronic device 110 returns to the main panel of the global editing mode. The output of the global split can include the filled base image and multiple generated media layers, where the base image can be locked at the bottom layer and its order cannot be adjusted, while other media layers can respond to order adjustment operations. In this way, in scenarios where the entire media content needs to be broken down at once, users can directly obtain multiple media layers that can be edited separately, without having to select individual areas one by one.
[0097] Example process Figure 4 A flowchart of an example process 400 for editing media content is shown, based on several scenarios. Process 400 can be implemented at electronic device 110. See below for reference. Figure 4 To describe process 400.
[0098] like Figure 4 As shown, in box 410, electronic device 110 receives a first operation, the first operation being used to enable a first selection mode, the first selection mode being used to receive a first type of operation.
[0099] In box 420, electronic device 110 receives a second operation for selecting at least one area of first media content, the second operation corresponding to the first type.
[0100] In box 430, electronic device 110 presents multiple media layers, each media layer corresponding to a different element in at least one area.
[0101] In this way, users can enable selection mode and select a local area to split different elements within that area into individual media layers for editing, reducing the steps required to perform global splitting of the entire media content. Since the electronic device 110 only needs to perform element recognition and layer generation calculations on the user-selected local area, rather than performing global segmentation on the entire media content, the client's computational overhead and memory usage are correspondingly reduced. At the same time, since only the splitting results of a local area need to be updated in the interface, the client's rendering overhead and the amount of data that needs to be drawn are reduced, thereby lowering interface update overhead.
[0102] In some cases, at least one region includes a first region, and receiving a second operation includes receiving a second operation, the second operation including a first drag operation in the first media content, the first drag operation including a first start position and a first end position, the first region being obtained based on the first start position and the first end position.
[0103] In this way, based on the selection mode scheme described above, users can intuitively define the position and range of the selection area through drag-and-drop operations, reducing the number of steps required to define the editing area. The electronic device 110 directly determines the area range based on the start and end positions of the drag, eliminating the need for a global scan of the entire media content and reducing the computational overhead of the area determination process.
[0104] In some cases, the electronic device 110 may present a first area in the first media content, the first area having a geometry obtained through a first drag operation.
[0105] In this way, based on the above solution, users can intuitively view the geometry and coverage of the selected area on the canvas, reducing the number of misoperations caused by the selection area being invisible. The rendering of the geometric shape indication of the selection area on the canvas by the electronic device 110 only involves lightweight border drawing, which has a small impact on the overall rendering overhead of the canvas.
[0106] In some cases, the electronic device 110 may present a first control at an associated location in the first area; and in response to receiving an interactive operation, stop presenting the first area, the interactive operation being used to trigger the first control.
[0107] In this way, based on the above solution, users can deselect the area with a single click using the nearest control, reducing the steps required for the deselection operation. After deselecting the area, the electronic device 110 can release associated rendering resources and mask data, reducing the client's memory usage.
[0108] In some cases, the second operation also includes a second drag operation, which includes a second start position and a second end position. The method further includes: in the first media content, presenting a second region, which is obtained based on the second start position and the second end position, and at least one region includes the second region.
[0109] In this way, based on the above solution, users can define multiple non-adjacent selection areas for batch processing, reducing the number of repetitive operations in multi-area scenarios. Electronic device 110 can merge processing requests from multiple areas, reducing the number of requests sent to the server and lowering the bandwidth usage for end-to-cloud communication.
[0110] In some cases, the electronic device 110 may receive a third operation for editing a first attribute, the first attribute being associated with at least one region; and presenting at least one region in a first style, the first style representing the edited first attribute.
[0111] In this way, based on the above solution, users can directly edit the properties of regions or media layers on the canvas, reducing the steps required to set properties through additional panels. The electronic device 110 only needs to re-render the object being edited, reducing rendering overhead in property editing scenarios.
[0112] In some cases, the first attribute includes at least one of the following: the display position of at least one area; the display size of at least one area.
[0113] In this way, based on the above solution, users can flexibly adjust the display position and size of the area, improving editing accuracy. The electronic device 110 only needs to update the transformation parameters of the adjusted area, without recalculating the layout of the entire canvas, thus reducing computational overhead.
[0114] In some cases, the electronic device 110 may receive a fourth operation for enabling a second selection mode, the second selection mode for receiving a second type of operation; receive a fifth operation for selecting a first group of elements of the first media content, the fifth operation corresponding to the second type; and present a first media layer, the first media layer corresponding to the first group of elements.
[0115] In this way, based on the aforementioned box selection mode, users can also accurately select individual elements by clicking, reducing the number of steps required to select a single element. The electronic device 110 directly responds to the clicking operation based on the pre-segmentation results, reducing the response latency of the clicking interaction.
[0116] In some cases, electronic device 110 may respond to enabling a second selection mode by presenting a second group of elements in a second style, the second style indicating that the second group of elements supports selection.
[0117] In this way, building upon the above approach, users can intuitively understand which elements are selectable before operation, reducing the number of exploratory clicks. Electronic device 110 pre-renders the highlighted states of candidate elements, eliminating the need to calculate the highlighted area in real-time every time the user hovers or clicks, thus reducing real-time computational overhead during the interaction process.
[0118] In some cases, the first set of elements is determined based on the following process: segmenting the first media content using a first model to obtain multiple candidate elements; and determining the second set of elements from the multiple candidate elements.
[0119] In this way, based on the above scheme, the electronic device 110 can automatically identify candidate elements in the media content using the first model, without the user needing to manually mark or draw element boundaries. The segmentation calculation of the first model can be completed in advance after the image is imported, so that the user can obtain candidate elements when entering the point-and-click mode, reducing user waiting time and interaction response latency.
[0120] In some cases, electronic device 110 may receive a sixth operation, which represents the selection of a third set of elements, which is different from the second set of elements; and the presentation of the third set of elements in a second style.
[0121] In this way, based on the above solution, users can flexibly switch between different combinations of elements, reducing the number of steps required to re-enter the selection process. Electronic device 110 only needs to update the visual identifier state of the element being switched, reducing the rendering overhead of the selection switching operation.
[0122] In some cases, the third set of elements is determined based on the following process: determining the location information corresponding to the sixth operation; and providing the location information and the first media content to the second model to determine the third set of elements.
[0123] In this way, based on the above scheme, the electronic device 110 can use the second model to perform fine segmentation of the user's click position, realizing the selection of elements in the pre-segmented uncovered area. The electronic device 110 only needs to call the second model for the local area of the click position, instead of re-performing global segmentation, thus reducing the computational overhead of fine selection.
[0124] In some cases, presenting multiple media layers includes: in a first component, presenting multiple content items in a first order, the multiple content items corresponding to multiple media layers, the first hierarchical relationship of the multiple media layers being related to the first order.
[0125] In this way, based on the above solution, users can intuitively view the hierarchical relationship of media layers through layer components, reducing the exploratory operations required to confirm the layer order. Electronic device 110 presents media layers as thumbnail content items, reducing the rendering overhead of the layer management interface.
[0126] In some cases, the electronic device 110 may receive a seventh operation, which represents the selection of a first content item for displaying at least a portion of the second media layer; and present the second media layer in a third style, which represents that the second media layer is in a selected state.
[0127] In this way, building upon the above approach, users can select obscured media layers through the Layers panel, reducing the number of times they need to attempt to select the obscured layer. The electronic device 110 only needs to update the status indicator of the selected media layer, reducing the local refresh overhead of the selection operation.
[0128] In some cases, the electronic device 110 may receive an eighth operation for adjusting multiple content items from a first order to a second order; and for presenting multiple media layers according to a second hierarchical relationship, the second hierarchical relationship corresponding to the second order.
[0129] In this way, based on the above solution, users can adjust the layer order by dragging and dropping content items, reducing the number of operations required to move layers forward or backward. The electronic device 110 only needs to redraw the affected media layers in the new order, reducing the rendering overhead of the sorting operation.
[0130] In some cases, electronic device 110 may receive a first request to delete the third media layer; and stop displaying the third media layer and a second content item, the second content item being used to display at least a portion of the third media layer.
[0131] In this way, based on the above solution, users can delete unnecessary media layers with a single click and simultaneously update the layer panel, reducing the steps required for deletion. After deletion, the electronic device 110 releases the corresponding rendering resources and memory space, reducing client-side resource consumption.
[0132] In some cases, electronic device 110 may receive a ninth operation, which represents the selection of a third region including a first element; and in response to receiving a second request, present a fourth media layer corresponding to the second element, which has the same visual representation as the first element.
[0133] In this way, based on the above solution, users can copy selected elements with a single click and obtain a new, editable layer, reducing the number of steps involved in copying elements. The electronic device 110 can reuse existing segmentation results to generate new media layers, further reducing the computational overhead of the copying operation.
[0134] In some cases, electronic device 110 may receive a fourth request related to the first media content; and present second media content including the first element and the second element.
[0135] In this way, building upon the above approach, users can continue editing media content containing both the original and copied elements, reducing the number of steps required to switch between different editing states. The electronic device 110 merges the copied elements with the original content for rendering, eliminating the need to maintain a separate editing canvas and reducing memory usage.
[0136] In some cases, electronic device 110 may receive a third request, which is related to at least one of a plurality of media layers; and present at least one media material, which corresponds to at least one media layer.
[0137] In this way, based on the above solution, users can convert the split media layers into reusable materials, reducing the number of steps required to acquire materials. Electronic device 110 can directly reuse existing media layer data without recalculation, reducing the overhead of redundant calculations.
[0138] In some cases, the electronic device 110 can receive a tenth operation for editing a second attribute of the fifth media layer; and present the fifth media layer in a fifth style, which represents the edited second attribute.
[0139] In this way, based on the above solution, users can directly edit the properties of the media layer and preview the effect in real time, reducing the number of steps required for property editing. The electronic device 110 only needs to re-render the media layer being edited, reducing the rendering overhead of the editing operation.
[0140] In some cases, the second attribute includes at least one of the following: the display position of the fifth media layer; the display size of the fifth media layer; the rendering angle of the fifth media layer; the flip state of the fifth media layer; and the color parameters of the fifth media layer.
[0141] In some cases, the electronic device 110 may present a first set of controls corresponding to a first set of editing functions; and in response to the activation of a first selection mode, stop presenting the first set of controls and present a second set of controls corresponding to a second set of editing functions, at least some of the functions of the second set of editing functions being different from the first set of editing functions, and at least some of the functions corresponding to the first selection mode.
[0142] In this way, after the first selection mode is enabled, the electronic device 110 can stop displaying the first set of controls that are unrelated to the current selection mode and display the second set of controls that are associated with the first selection mode, so that the user can directly select the editing function applicable to the currently selected object from the second set of controls, reducing the operation steps required for the user to find the target editing function and the probability of accidentally touching unrelated controls.
[0143] In some cases, the first group of controls corresponds to a first number, and the second group of controls corresponds to a second number, wherein the first number is different from the second number.
[0144] Example devices and equipment A corresponding apparatus for implementing the above methods or processes is also provided. Figure 5 Block diagrams of an apparatus 500 for editing media content are shown in several scenarios. The apparatus 500 can be implemented as or included in an electronic device 110. The various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.
[0145] like Figure 5 As shown, the device 500 includes a first receiving module 510 configured to receive a first operation, the first operation being used to enable a first selection mode, the first selection mode being used to receive a first type of operation; a second receiving module 520 configured to receive a second operation, the second operation being used to select at least one area of a first media content, the second operation corresponding to the first type; and a presentation module 530 configured to present multiple media layers, the multiple media layers corresponding to different elements in at least one area.
[0146] In some cases, the second receiving module 520 may also be configured such that at least one region includes the first region, and receiving the second operation includes receiving the second operation, the second operation including a first drag operation in the first media content, the first drag operation including a first start position and a first end position, the first region being obtained based on the first start position and the first end position.
[0147] In some cases, the presentation module 530 may also be configured to present a first area in the first media content, the first area having a geometry obtained through a first drag operation.
[0148] In some cases, the second receiving module 520 may also be configured to present the first control at an associated location in the first region; and to stop presenting the first region in response to receiving an interactive operation, the interactive operation being used to trigger the first control.
[0149] In some cases, the second receiving module 520 may also be configured to include a second drag operation, the second drag operation including a second start position and a second end position, and the method further includes: in the first media content, presenting a second region, the second region being obtained based on the second start position and the second end position, at least one region including the second region.
[0150] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a third operation for editing a first attribute, the first attribute being associated with at least one region; and to present at least one region in a first style, the first style representing the edited first attribute.
[0151] In some cases, the first attribute includes at least one of the following: the display position of at least one area; the display size of at least one area.
[0152] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a fourth operation, which is used to enable a second selection mode, which is used to receive a second type of operation; receive a fifth operation, which is used to select a first group of elements of the first media content, which corresponds to the second type; and present a first media layer, which corresponds to the first group of elements.
[0153] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to present a second group of elements in a second style in response to enabling a second selection mode, the second style indicating that the second group of elements supports selection.
[0154] In some cases, the first set of elements is determined based on the following process: segmenting the first media content using a first model to obtain multiple candidate elements; and determining the second set of elements from the multiple candidate elements.
[0155] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a sixth operation, which represents the selection of a third group of elements, which is different from the second group of elements; and to present the third group of elements in a second style.
[0156] In some cases, the third set of elements is determined based on the following process: determining the location information corresponding to the sixth operation; and providing the location information and the first media content to the second model to determine the third set of elements.
[0157] In some cases, the presentation module 530 may also be configured to present multiple media layers, including: in the first component, presenting multiple content items in a first order, the multiple content items corresponding to multiple media layers, the first hierarchical relationship of the multiple media layers being related to the first order.
[0158] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a seventh operation, which represents the selection of a first content item, the first content item being used to display at least a portion of the second media layer; and to present the second media layer in a third style, the third style representing that the second media layer is in a selected state.
[0159] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive an eighth operation, which is used to adjust multiple content items from a first order to a second order; and to present multiple media layers according to a second hierarchical relationship, the second hierarchical relationship corresponding to the second order.
[0160] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a first request for deleting the third media layer; and to stop rendering the third media layer and a second content item for displaying at least a portion of the third media layer.
[0161] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a ninth operation, the ninth operation representing the selection of a third region, the third region including the first element; and in response to receiving a second request, to present a fourth media layer, the fourth media layer corresponding to the second element, the second element having the same visual representation as the first element.
[0162] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a fourth request related to the first media content; and to present second media content, which includes the first element and the second element.
[0163] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a third request, the third request being related to at least one of a plurality of media layers; and to present at least one media material, the at least one media material corresponding to at least one media layer.
[0164] In some cases, the first receiving module 510 or the second receiving module 520 may also be configured to receive a tenth operation for editing a second attribute of the fifth media layer; and to present the fifth media layer in a fifth style, which represents the edited second attribute.
[0165] In some cases, the second attribute includes at least one of the following: the display position of the fifth media layer; the display size of the fifth media layer; the rendering angle of the fifth media layer; the flip state of the fifth media layer; and the color parameters of the fifth media layer.
[0166] In some cases, the presentation module 530 may also be configured to present a first set of controls, the first set of controls corresponding to a first set of editing functions; and in response to the activation of the first selection mode, to stop presenting the first set of controls and present a second set of controls, the second set of controls corresponding to a second set of editing functions, at least some of the functions of the second set of editing functions being different from the first set of editing functions, the at least some functions corresponding to the first selection mode.
[0167] In some cases, the first group of controls corresponds to a first number, and the second group of controls corresponds to a second number, wherein the first number is different from the second number.
[0168] The modules included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some cases, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0169] Figure 6 A block diagram of an electronic device 600 in which one or more examples may be implemented is shown. It should be understood that... Figure 6 The electronic device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 6 The illustrated electronic device 600 can be used to implement the electronic device 110 discussed above.
[0170] like Figure 6As shown, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processing units or processors 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processor 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 600.
[0171] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof). Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.
[0172] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various examples.
[0173] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.
[0174] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0175] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0176] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0177] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0178] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0179] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0180] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method of editing media content, comprising: receiving a first operation to enable a first selection mode, the first selection mode to receive a first type of operation; receiving a second operation to select at least one region of a first media content, the second operation corresponding to the first type; and presenting a plurality of media layers corresponding to different elements in the at least one region.
2. The method of claim 1, wherein the at least one region comprises a first region, and the receiving a second operation comprises: receiving the second operation comprising a first drag operation in the first media content, the first drag operation comprising a first start position and a first end position, the first region resulting based on the first start position and the first end position.
3. The method of claim 2, further comprising: presenting the first region in the first media content, the first region having a geometry resulting from the first drag operation.
4. The method of claim 3, further comprising: presenting a first control at an associated position of the first region; and stopping presenting the first region in response to receiving an interaction operation to trigger the first control.
5. The method of claim 2, wherein the second operation further comprises a second drag operation, the second drag operation comprising a second start position and a second end position, the method further comprising: presenting a second region in the first media content, the second region resulting based on the second start position and the second end position, the at least one region comprising the second region.
6. The method of claim 1, further comprising: receiving a third operation to edit a first property, the first property related to the at least one region; and presenting the at least one region in a first style, the first style characterizing the edited first property.
7. The method of claim 1, further comprising: receiving a fourth operation to enable a second selection mode, the second selection mode to receive a second type of operation; receiving a fifth operation to select a first group of elements of a first media content, the fifth operation corresponding to the second type; and presenting a first media layer corresponding to the first group of elements.
8. The method of claim 7, further comprising: presenting a second group of elements in a second style in response to enabling the second selection mode, the second style characterizing that the second group of elements supports selection.
9. The method of claim 8, further comprising: receiving a sixth operation characterizing a selection of a third group of elements, the third group of elements different from the second group of elements; and presenting the third group of elements in the second style.
10. The method of claim 1, wherein the presenting a plurality of media layers comprises: In the first component, a plurality of content items are presented in a first order, the plurality of content items corresponding to a plurality of media layers, a first hierarchical relationship of the plurality of media layers being related to the first order.
11. The method of claim 10, further comprising: receiving a seventh operation, the seventh operation characterizing a selection of a first content item, the first content item being for presenting at least part of a second media layer; and presenting the second media layer in a third style, the third style characterizing the second media layer being in a selected state.
12. The method of claim 11, further comprising: receiving an eighth operation, the eighth operation being for adjusting the plurality of content items from the first order to a second order; and presenting the plurality of media layers according to a second hierarchical relationship, the second hierarchical relationship corresponding to the second order.
13. The method of claim 10, further comprising: receiving a first request, the first request being for deleting a third media layer; and stopping presenting the third media layer and a second content item, the second content item being for presenting at least part of the third media layer.
14. The method of claim 1, further comprising: receiving a ninth operation, the ninth operation characterizing a selection of a third region, the third region including a first element; and in response to receiving a second request, presenting a fourth media layer, the fourth media layer corresponding to a second element, the second element having a same visual representation as the first element.
15. The method of claim 1, further comprising: presenting a first set of controls, the first set of controls corresponding to a first set of editing functions; and in response to the first selection mode being enabled, stopping presenting the first set of controls and presenting a second set of controls, the second set of controls corresponding to a second set of editing functions, at least part of the second set of editing functions being different from the first set of editing functions, the at least part corresponding to the first selection mode.
16. The method of claim 1, further comprising: receiving a third request, the third request being related to at least one media layer of the plurality of media layers; and presenting at least one media asset, the at least one media asset corresponding to the at least one media layer.
17. An apparatus for editing media content, comprising: a first receiving module configured to receive a first operation, the first operation being for enabling a first selection mode, the first selection mode being for receiving a first type of operation; a second receiving module configured to receive a second operation, the second operation being for selecting at least one region of a first media content, the second operation corresponding to the first type; and a presenting module configured to present a plurality of media layers, the plurality of media layers corresponding to different elements in the at least one region.
18. An electronic device, comprising: at least one processor; and a memory coupled to the at least one processor, the memory comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform the method of any of claims 1-17. at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-16.
19. A computer-readable storage medium having stored thereon computer- executable instructions executable by a processor to implement a method according to any one of claims 1-16.
20. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1-16.