Media content processing method and apparatus, device and storage medium
By selecting templates in the media content editing interface and utilizing generative artificial intelligence technology, the problem of poor realism in hairstyle replacement in existing tools has been solved, achieving high-quality hairstyle and hair color replacement and improving the realism and efficiency of editing.
Patent Information
- Application Number
- PCT/CN2025/110218
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-24
- Filing Date
- 2025-07-23
- Publication Date
- 2026-01-29
AI Technical Summary
Existing media content editing tools struggle to achieve high-quality, realistic effects when changing a person's hairstyle.
By selecting templates in the editing interface, visual content corresponding to the target template is generated. Combined with generative artificial intelligence technology, hairstyles and hair colors can be replaced, ensuring the relevance and authenticity of the visual content.
It improves the replacement effect of hairstyles and hair colors in media content, enhancing the realism of the generated results and editing efficiency.
Smart Images

Figure CN2025110218_29012026_PF_FP_ABST
Abstract
Description
Method, device, apparatus and storage medium for processing media content
[0001] The present application claims priority to the Chinese patent application No. 202411001299.2, filed on July 24, 2024, entitled “Method, device, apparatus and storage medium for processing media content”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The example embodiments of the present disclosure generally relate to the field of computer, and in particular, to a method, device, apparatus and computer readable storage medium for processing media content. BACKGROUND
[0003] In recent years, editing tools provide people with various capabilities to edit media content. For example, image or video editing tools can help users optimize the image or video to be edited. In some scenarios, people expect to be able to replace the hairstyle, e.g., the hair style and / or color, of a person in the media content with high quality. SUMMARY
[0004] In a first aspect of the present disclosure, a method for processing media content is provided. The method comprises: displaying an editing interface of first media content, the first media content comprising a plurality of face objects; in response to a selection of a first face object in the plurality of face objects, displaying a set of templates in the editing interface, the set of templates being associated with different hair styles; and based on a selection of a target template in the set of templates, displaying second media content in the editing interface, the second media content comprising first visual content corresponding to a target hair style indicated by the target template and second visual content corresponding to the first face object, the first visual content being visually associated with the second visual content, the second media content being generated based on the first media content and a preset prompt item corresponding to the target template.
[0005] In a second aspect of the present disclosure, a device for processing media content is provided. The device comprises: an interface display module configured to display an editing interface of first media content, the first media content comprising a plurality of face objects; a template display module configured to, in response to a selection of a first face object in the plurality of face objects, display a set of templates in the editing interface, the set of templates being associated with different hair styles; and a content display module configured to, based on a selection of a target template in the set of templates, display second media content in the editing interface, the second media content comprising first visual content corresponding to a target hair style indicated by the target template and second visual content corresponding to the first face object, the first visual content being visually associated with the second visual content, the second media content being generated based on the first media content and a preset prompt item corresponding to the target template.
[0006] In a third aspect of the disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of the disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer- executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0009] It is to be understood that the details set forth herein are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which:
[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments according to the present disclosure can be implemented;
[0012] FIGS. 2A-2B show example interfaces according to some embodiments of the present disclosure;
[0013] FIGS. 3A-3B show example interfaces according to some embodiments of the present disclosure;
[0014] FIG. 4 shows a schematic diagram of generating media content according to some embodiments of the present disclosure;
[0015] FIG. 5 shows a flowchart of an example process of processing media content according to some embodiments of the present disclosure;
[0016] FIG. 6 shows a schematic block diagram of an example apparatus for processing media content according to some embodiments of the present disclosure; and
[0017] FIG. 7 shows a block diagram of an electronic device capable of implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of the present disclosure.
[0019] It should be noted that the titles of any sections / sub-sections provided herein are not limiting. Various embodiments are described throughout, and any type of embodiment can be included under any section / sub-section. Moreover, embodiments described in any section / sub-section can be combined with any other embodiments described in the same section / sub-section and / or different section / sub-section in any manner.
[0020] In the description of embodiments of the present disclosure, the term "includes" and its derivatives, such as "including," should be understood in an open, inclusive sense, that is, "including, but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "an embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Additional explicit and implicit definitions can be found below. The terms "first," "second," etc. can refer to different or the same objects. Additional explicit and implicit definitions can be found below.
[0021] Data of users, acquisition and / or use of data, etc. can be involved in embodiments of the present disclosure. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the present disclosure, all collection, acquisition, processing, processing, forwarding, use, etc. of data are performed on the premise that users are aware of and confirm. Accordingly, when implementing embodiments of the present disclosure, the type of data or information that can be involved, the range of use, the scenario of use, etc. should be informed to users and authorized by users in a proper manner according to relevant laws and regulations. The specific informing and / or authorization manner can vary according to actual situations and application scenarios, and the scope of the present disclosure is not limited in this respect.
[0022] In the present specification and embodiments, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed range. Users refuse to process personal information other than the necessary information required for basic functions, which will not affect the user's use of basic functions.
[0023] As mentioned above, in the editing scenario of media content, replacing the hairstyle of a person is a common editing requirement. Some traditional editing tools can achieve the replacement of the hairstyle by superimposing the visual content corresponding to the hairstyle into the media content. However, the replacement result of the hairstyle obtained based on this way has poor reality.
[0024] Embodiments of the present disclosure propose a scheme for processing media content. According to the scheme, a first media content is displayed in an editing interface, the first media content comprising a plurality of face objects; in response to a selection of a first face object in the plurality of face objects, a set of templates is displayed in the editing interface, the set of templates being associated with different hair styles; and in response to a selection of a target template in the set of templates, a second media content is displayed in the editing interface, the second media content comprising first visual content corresponding to a target hair style indicated by the target template and second visual content corresponding to the first face object, the first visual content being visually associated with the second visual content, the second media content being generated based on the first media content and a preset prompt item corresponding to the target template.
[0025] In this way, embodiments of the present disclosure can apply a specified hair style (e.g., hairstyle and / or color) to a person or other object in the media content, and improve the reality of the generated media content.
[0026] Various example implementations of the scheme are described in further detail below in conjunction with the accompanying drawings.
[0027] Example Environment
[0028] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, the example environment 100 can include an electronic device 110.
[0029] In this example environment 100, the electronic device 110 can run an application 120 that supports interface interaction. The application 120 can be any suitable type of application for interface interaction, examples of which can include, but are not limited to, a camera application, a clipping application, or other suitable application. A user 140 can interact with the application 120 via the electronic device 110 and / or its attached devices.
[0030] In the environment 100 of FIG. 1, if the application 120 is in an active state, the electronic device 110 can present an interface 150 for supporting interface interaction through the application 120.
[0031] In some embodiments, the electronic device 110 communicates with the server 130 to enable provisioning of services of the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a palmtop computer, a portable gaming terminal, a VR / AR device, a Personal Communication System (PCS) terminal, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including an accessory or peripheral device for the foregoing, or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface to a user (such as “wearable” circuitry, etc.).
[0032] The server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms. The server 130 may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like. The server 130 can provide background services for the application 120 in the electronic device 110 that supports a virtual scene.
[0033] A communication connection can be established between the server 130 and the electronic device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, and the like, and embodiments of the present disclosure are not limited in this regard. In embodiments of the present disclosure, the server 130 and the electronic device 110 can implement signaling interaction through the communication connection therebetween.
[0034] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.
[0035] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0036] Example Interaction
[0037] FIGS. 2A-2B illustrate example interfaces 200A-200B, in accordance with some embodiments of the present disclosure. The interfaces 200A-200B can be provided by, for example, the electronic device 110 illustrated in FIG. 1.
[0038] As shown in FIG. 2A, the electronic device 110 can provide an editing interface 200A for editing a first media content 205.
[0039] In some embodiments, the electronic device 110 can identify a plurality of facial objects in the first media content 205, e.g., a facial object 210-1 and a facial object 210-2 (individually or collectively referred to as facial objects 210).
[0040] Further, the electronic device 110 can receive a selection of the facial object 210-1 in the interface 200A. In some embodiments, the electronic device 110 may, for example, receive a tap, a long press, or the like of the facial object 210-1 by a user to determine that the user selects the facial object 210-1.
[0041] In some embodiments, the electronic device 110 can also display, in the interface 200A, indication elements corresponding to the plurality of facial objects 210, e.g., an indication element 215-1 and an indication element 215-2.
[0042] In some embodiments, the electronic device 110 can receive a preset operation (e.g., a tap) of the indication element 215-1 by a user, which indicates a selection of the first facial object 210-1.
[0043] Further, as shown in FIG. 2A, the electronic device 110 may, for example, display the indication element 215-1 according to a target style to represent that the corresponding first facial object 210-1 is in a selected state. For example, the electronic device 110 can change a color, a line width, or the like of the indication element 215-1.
[0044] As shown in FIG. 2A, the electronic device 110 can provide, in the interface 200A, a set of templates in response to the selection of the first facial object 210-1, e.g., a template 220-1, a template 220-2, a template 220-3, and a template 220-4 (individually or collectively referred to as templates 220).
[0045] In some embodiments, such templates 220 can correspond to different hairstyles. In some embodiments, different templates 220 can also correspond to different preset hint items. As will be introduced below, such preset hint items can be used to generate an application result of a hairstyle.
[0046] In some embodiments, as shown in FIG. 2A, the electronic device 110 can further provide a label 225 for editing the hairstyle and a label 230 for editing the hair color. The editing process regarding the hair color will be described with reference to FIGS. 3A and 3B.
[0047] Further, the electronic device 110 may, for example, receive a selection of the template 220-3 by the user. Accordingly, as shown in FIG. 2B, the electronic device 110 can display a second media content 235 in the interface 200B. The second media content 235 can be corresponding to the selected template 220-3, i.e., the hairstyle corresponding to the template 220-3 can be applied to the first face object 210-1.
[0048] Specifically, as shown in FIG. 2B, the second media content 235 can include a visual content 240-1 corresponding to the first face object 210-1. The visual content 240-1 can include a first visual content corresponding to the target hairstyle, i.e., the hair portion in the picture. Additionally, the visual content 240-1 can further include a second visual content corresponding to the first face object 210-2, i.e., the face portion in the picture.
[0049] In some embodiments, the second visual content can be consistent with the picture portion corresponding to the first face object 210-1 in the first media content 205. In some embodiments, the second visual content can be further generated based on at least part of the visual content of the first media content. For example, compared with the first media content 205, the second media content 235 can adaptively adjust the face picture of the first face object 210-1.
[0050] In some embodiments, in the second media content 235, the visual content 240-2 corresponding to the second face object 210-2 without the adjusted hairstyle can be consistent with the first media content 205.
[0051] In some embodiments, the electronic device 110 can further receive a request for regenerating the second media content 235. For example, as shown in FIG. 2B, the electronic device 110 can update the template 220-3 to an entry for initiating the request for regenerating.
[0052] Accordingly, the electronic device 110 can receive a preset operation on the template 220-3, and can accordingly trigger the generation of a new third media content based on the first media content 205 and the preset prompt item corresponding to the template 220-3.
[0053] In some embodiments, the second media content and the third media content can be associated with different seed files to ensure that the regenerated third media content is different from the second media content.
[0054] The specific generation process of the second media content 235 will be described in detail below with reference to FIG. 4, which is not described here in detail.
[0055] Through the above-described process, embodiments of the present disclosure can support replacing the hairstyle of a person or other object in media content (e.g., a picture or a video) by using generative artificial intelligence technology, and provide a more realistic generation result.
[0056] In some embodiments, embodiments of the present disclosure can also support editing the color of the hair. An example process of editing the color of the hair will be described below with reference to FIGS. 3A and 3B.
[0057] As shown in FIG. 3A, the electronic device 110 can provide an editing interface 300A for editing the media content 305. Similarly, the media content 305 may, for example, include one or more face objects, such as face object 310-1 and face object 310-2 (individually or collectively referred to as face object 310).
[0058] Further, the electronic device 110 can receive a selection of the face object 310-2 in the interface 300A. In some embodiments, the electronic device 110 may, for example, receive a user operation of tapping, long-pressing, or the like on the corresponding position of the face object 310-2 to determine that the user selects the face object 310-2.
[0059] Alternatively, the electronic device 110 can display the indication element 315-1 and the indication element 315-2. Further, the electronic device 110 may, for example, receive a preset operation of the user on the indication element 315-2, and accordingly determine that the user selects the face object 310-2.
[0060] Further, the electronic device 110 can provide a color configuration control in the interface 300A. Taking FIG. 3A as an example, the electronic device 110 may, for example, receive a selection of the label 330 to provide the color configuration control. As an example, the electronic device 110 may, for example, also receive a selection of the label 325 to provide a set of templates corresponding to different hairstyles as shown in FIG. 2A.
[0061] Further, as shown in FIG. 3A, the electronic device 110 may, for example, receive a configuration operation of the user via the color configuration control, and determine a target color to be applied. For example, the electronic device 110 may, for example, provide a plurality of templates (also referred to as a second set of templates) corresponding to different preset colors, such as template 320-1, template 320-2, template 320-3, and template 320-4, by using the color configuration control. Further, the electronic device 110 may, for example, receive a selection of the template 320-4 corresponding to “color D” by the user. In some embodiments, the color configuration control may, for example, also specify the target color to be applied by other appropriate manners.
[0062] Further, the electronic device 110 can display the editing result of the media content 305. Specifically, the electronic device 110 can apply the selected specified “color D” to the target visual content in the media content 305, which can correspond to the hair associated with the selected face object 310-2.
[0063] In some embodiments, different colors can correspond to different preset prompt items. The electronic device 110 can utilize the generative model to generate new media content based on the preset prompt item corresponding to “color D” and the media content 305.
[0064] Specifically, as shown in FIG. 3B, the electronic device 110 can present the generated media content 335. The media content 335 can include visual content 340-2 corresponding to the face object 310-2. The visual content 340-2 can include first visual content corresponding to the specified “color D”, i.e., the hair portion in the picture. Additionally, the visual content 340-2 can also include second visual content corresponding to the face object 310-2, i.e., the face portion in the picture.
[0065] In some embodiments, the second visual content can remain consistent with the picture portion in the first media content 305 corresponding to the first face object 310-2. In some embodiments, the second visual content can also be generated based on at least part of the visual content of the media content 305. For example, compared with the first media content 305, the second media content 335 can adaptively adjust the face picture of the face object 310-2.
[0066] In some embodiments, in the media content 335, the visual content 340-1 corresponding to the face object 310-1 without adjusted hair color can remain consistent with the media content 305.
[0067] In this way, embodiments of the present disclosure can further support replacement of hair color for a person or other object in media content (e.g., a picture or a video), thereby improving the editing efficiency of the media content.
[0068] Generation of media content
[0069] The following will continue to refer to FIG. 4 to describe an example process of generating media content corresponding to a target hair style. In some scenarios, the media content can be generated by the electronic device 110 and / or the server 130.
[0070] As shown in FIG. 4, the electronic device 110 and / or the server 130 can obtain a first media content 405 to be processed, and can accordingly determine a plurality of regions corresponding to a plurality of face objects in the first media content, thereby generating mask data 410. It should be understood that the regions corresponding to the face objects include both face regions and corresponding hair regions.
[0071] Further, the electronic device 110 and / or the server 130 can determine at least one target region in the plurality of regions, the at least one target region corresponding to at least one face object in the plurality of face objects different from the first face object.
[0072] For example, the electronic device 110 and / or the server 130 can determine a difference between the mask data 410 and mask data corresponding to the face object whose hairstyle is to be edited, and accordingly determine at least one face region corresponding to other face objects.
[0073] Further, the electronic device 110 and / or the server 130 can generate a first intermediate media content 415 by blurring the at least one target region in the first media content.
[0074] Additionally, the electronic device 110 and / or the server 130 can provide the first intermediate media content 415 and a preset prompt item corresponding to a target hairstyle (e.g., a specified hairstyle and / or hair color) to a target model, thereby obtaining a second intermediate media content 420 generated by the target model. In some embodiments, the target model can include any appropriate generative model, and the present disclosure is not intended to be limited in this regard.
[0075] Additionally, the electronic device 110 and / or the server 130 can generate a second media content 425 corresponding to the target hairstyle based on the first media content 405 and the second intermediate media content 420.
[0076] In some embodiments, the electronic device 110 and / or the server 130 can generate a second mask indicating a smoothing region based on a first mask corresponding to the first face object (i.e., the face object to be edited). As an example, the second mask can be determined based on a difference between the mask data 410 and the first mask.
[0077] Further, the electronic device 110 and / or the server 130 can smooth the first media content 405 and the second intermediate media content 420 according to the second mask to generate the second media content 425. Based on such a smoothing operation, the generated second media content 425 can have a picture content corresponding to the first face object closer to the second intermediate media content 420, and other picture contents closer to the first media content 405.
[0078] Thus, the electronic device 110 and / or the server 130 can implement a more smooth picture transition, and improve the reality of the picture.
[0079] Example process
[0080] FIG. 5 shows a flowchart of an example process 500 of processing media content, according to some embodiments of the present disclosure. The process 500 can be implemented at the electronic device 110. The process 500 is described below with reference to FIG. 1.
[0081] As shown in FIG. 5, at block 510, the electronic device 110 displays an editing interface of a first media content, the first media content including a plurality of facial objects.
[0082] At block 520, the electronic device 110 displays, in the editing interface, a set of templates in response to a selection of a first facial object of the plurality of facial objects, the set of templates being associated with different hair styles.
[0083] At block 530, the electronic device 110 displays, in the editing interface, a second media content based on a selection of a target template of the set of templates, the second media content including first visual content corresponding to a target hair style indicated by the target template and second visual content corresponding to the first facial object, the first visual content being visually associated with the second visual content, the second media content being generated based on the first media content and a preset prompt item corresponding to the target template.
[0084] In some embodiments, the second visual content is visual content of the first media content corresponding to the first facial object; or the second visual content is generated based on at least part of visual content of the first media content.
[0085] In some embodiments, the process 500 further includes receiving a regeneration request for the second media content; and displaying, in the editing interface, a third media content, the third media content being generated based on the first media content and the preset prompt item corresponding to the target template, the third media content being different from the second media content.
[0086] In some embodiments, receiving the regeneration request for the second media content includes receiving a first operation for the target template, the first operation indicating the regeneration request for the second media content.
[0087] In some embodiments, the process 500 further includes displaying, in the editing interface, a plurality of indication elements corresponding to the plurality of facial objects; and receiving a second operation for a target indication element of the plurality of indication elements, the second operation indicating a selection of a target facial object corresponding to the target indication element.
[0088] In some embodiments, the process 500 further includes, in response to the second operation on the target indication element, displaying the target indication element in the target style in the editing interface to represent a selected state of the target indication element.
[0089] In some embodiments, the set of templates includes: a first set of templates corresponding to different hairstyles; and / or a second set of templates corresponding to different colors.
[0090] In some embodiments, the second media content is generated based on the following process: determining, from the first media content, a plurality of regions corresponding to a plurality of facial objects; determining at least one target region of the plurality of regions, the at least one target region corresponding to at least one facial object of the plurality of facial objects different from the first facial object; generating first intermediate media content by blurring the at least one target region in the first media content; providing the first intermediate media content and a preset prompt to the target model to generate second intermediate media content; and generating the second media content based on the first media content and the second intermediate media content.
[0091] In some embodiments, generating the second media content based on the first media content and the second intermediate media content includes: generating a second mask indicating a smoothing region based on a first mask corresponding to the first facial object; and smoothing the first media content and the second intermediate media content according to the second mask to generate the second media content.
[0092] Example apparatus and device
[0093] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above-described methods or processes. FIG. 6 shows a schematic structural block diagram of an example apparatus 600 for processing media content according to certain embodiments of the present disclosure. The apparatus 600 can be implemented as or included in the electronic device 110. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.
[0094] As shown in FIG. 6, the apparatus 600 includes: an interface display module 610 configured to display an editing interface of first media content, the first media content including a plurality of facial objects; a template display module 620 configured to display, in response to a selection of a first facial object of the plurality of facial objects, a set of templates in the editing interface, the set of templates being associated with different hairstyles; and a content display module 630 configured to display, based on a selection of a target template of the set of templates, second media content in the editing interface, the second media content including first visual content corresponding to a target hairstyle indicated by the target template and second visual content corresponding to the first facial object, the first visual content being visually associated with the second visual content, the second media content being generated based on the first media content and a preset prompt corresponding to the target template.
[0095] In some embodiments, the second visual content is visual content in the first media content corresponding to the first facial object; or the second visual content is generated based on at least part of visual content of the first media content.
[0096] In some embodiments, the apparatus 600 further includes a request generation module configured to: receive a re-generation request for the second media content; and display, in the editing interface, third media content generated based on the first media content and the preset prompt item corresponding to the target template, the third media content being different from the second media content.
[0097] In some embodiments, the request generation module is further configured to: receive a first operation for the target template, the first operation indicating the re-generation request for the second media content.
[0098] In some embodiments, the apparatus 600 further includes an element display module configured to: display, in the editing interface, a plurality of indication elements corresponding to a plurality of facial objects; and receive a second operation for a target indication element in the plurality of indication elements, the second operation indicating a selection of a target facial object corresponding to the target indication element.
[0099] In some embodiments, the element display module is further configured to: in response to the second operation for the target indication element, display the target indication element in the editing interface in a target style to represent a selected state of the target indication element.
[0100] In some embodiments, the set of templates includes: a first set of templates corresponding to different hairstyles; and / or a second set of templates corresponding to different colors.
[0101] In some embodiments, the second media content is generated based on a process including: determining, from the first media content, a plurality of regions corresponding to a plurality of facial objects; determining at least one target region in the plurality of regions, the at least one target region corresponding to at least one facial object in the plurality of facial objects different from the first facial object; generating first intermediate media content by blurring the at least one target region in the first media content; providing the first intermediate media content and the preset prompt item to the target model to generate second intermediate media content; and generating the second media content based on the first media content and the second intermediate media content.
[0102] In some embodiments, the apparatus 600 further includes a mask generation module configured to: generate a second mask indicating a smoothing region based on a first mask corresponding to the first facial object; and smooth the first media content and the second intermediate media content according to the second mask to generate the second media content.
[0103] As shown in FIG. 7, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 can include, but are not limited to, one or more processors 710 or processing units, memory 720, storage 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processor 710 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in memory 720. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.
[0104] Electronic device 700 typically includes a plurality of computer storage media. Such media can be any available media that is accessible by electronic device 700 and includes both volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. Storage 730 can be removable or non-removable media and can include machine- readable media such as a flash drive, a magnetic disk drive, or any other media that can be used to store information and / or data and that can be accessed by electronic device 700.
[0105] Electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard drive"), and a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM, a DVD, etc.). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0106] Communication unit 740 enables communication with other electronic devices over communication media. Additionally, functionality of components of electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in a distributed computing environment.
[0107] Input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through communication unit 740, as desired, in order to communicate with one or more devices that enable a user to interact with electronic device 700, or to communicate with any device (e.g., a network card, a modem, etc.) that enables electronic device 700 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0108] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0109] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0110] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0111] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0112] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The computer program product can also be an article of manufacture that comprises a computer readable medium. The medium can comprise a hard disk drive, a memory, or a floppy diskette, which can be accessed using a drive unit. Additionally, the medium can comprise a storage device that can store program codes. The storage device can include, but is not limited to, devices needing a platter and a read / write head, optical disk drives such as CD-ROM, DVD, Blu-ray Disc™ drives, memory devices such as flash drives, memory sticks, or any device that stores digital information. Additionally, the medium can include a single storage device or a plurality of storage devices.
[0113] Various implementations of the disclosure have been described in detail above. The foregoing description is exemplary and explanatory only, and not exhaustive, of the disclosed implementations. Many modifications and variations of the implementations described herein are possible and will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the implementations described. The description of the term used herein is chosen for the purpose of explaining the principles of the implementations, practical application, or improvement over technology in the field, or to enable other ordinary skilled persons in the art to understand the various implementations disclosed herein.
Claims
1. A method for processing media content, comprising: The editing interface for the first media content is displayed, which includes multiple facial objects; In response to the selection of a first facial object among the plurality of facial objects, a set of templates is displayed in the editing interface, the set of templates being associated with different hair styles; as well as Based on the selection of a target template from the set of templates, second media content is displayed in the editing interface. The second media content includes first visual content corresponding to the target hair style indicated by the target template and second visual content corresponding to the first facial object. The first visual content is visually associated with the second visual content. The second media content is generated based on the first media content and a preset prompt item corresponding to the target template.
2. The method according to claim 1, wherein: The second visual content is the visual content in the first media content corresponding to the first facial object; or The second visual content is generated based on at least a portion of the visual content of the first media content.
3. The method according to claim 1, further comprising: Receive a request to regenerate the second media content; as well as The editing interface displays third media content, which is generated based on the first media content and the preset prompt item corresponding to the target template. The third media content is different from the second media content.
4. The method of claim 3, wherein receiving a regeneration request for the second media content comprises: Receive a first operation for the target template, the first operation indicating the regeneration request for the second media content.
5. The method according to claim 1, further comprising: The editing interface displays multiple indicator elements corresponding to the multiple facial objects; as well as Receive a second operation for a target indicator element among the plurality of indicator elements, the second operation indicating the selection of the target facial object corresponding to the target indicator element.
6. The method according to claim 5, further comprising: In response to the second operation on the target indicator element, the target indicator element is displayed in the editing interface in a target style to indicate the selected state of the target indicator element.
7. The method according to claim 1, wherein the set of templates comprises: The first set of templates corresponding to different hairstyles; and / or The second set of templates corresponds to different colors.
8. The method of claim 1, wherein the second media content is generated based on the following process: Identify multiple regions corresponding to the multiple facial objects from the first media content; Determine at least one target region among the plurality of regions, the at least one target region corresponding to at least one facial object among the plurality of facial objects that is different from the first facial object; First intermediate media content is generated by blurring at least one target region in the first media content; Provide the first intermediate media content and the preset prompts to the target model to generate the second intermediate media content; as well as The second media content is generated based on the first media content and the second intermediate media content.
9. The method according to claim 8, wherein generating the second media content based on the first media content and the second intermediate media content comprises: Based on the first mask corresponding to the first facial object, a second mask indicating the smooth area is generated; as well as The first media content and the second intermediate media content are smoothed according to the second mask to generate the second media content.
10. An apparatus for processing media content, comprising: The interface display module is configured to display an editing interface for first media content, which includes multiple facial objects; A template display module is configured to display a set of templates associated with different hairstyles in the editing interface in response to the selection of a first facial object among the plurality of facial objects. as well as The content display module is configured to display second media content in the editing interface based on the selection of a target template from the set of templates. The second media content includes first visual content corresponding to the target hair style indicated by the target template and second visual content corresponding to the first facial object. The first visual content is visually associated with the second visual content. The second media content is generated based on the first media content and a preset prompt item corresponding to the target template.
11. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processor.
12. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 9.
13. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Hair style recommending method and system based on image identification
CN108305146A
Electronic device for generating image including 3D avatar reflecting face motion through 3D avatar corresponding to face and method of operating same
CN111742351A
Hair style conversion method and device and storage medium
CN115660948A
Image generation method, service provision method, apparatus, medium, and program product
CN117853604A
Media content processing method and device, equipment and storage medium
CN119002780A