Image processing method and apparatus, and terminal device, storage medium and program product

WO2026194383A1PCT designated stage Publication Date: 2026-09-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/146005
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-21
Filing Date
2025-12-26
Publication Date
2026-09-24

Smart Images

  • Figure CN2025146005_24092026_PF_FP_ABST
    Figure CN2025146005_24092026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are an image processing method and apparatus, and a terminal device, a storage medium and a program product. The image processing method comprises: displaying a first interface, which comprises a first control; in response to a touch performed on the first control, displaying a first image and / or a second image, wherein the first image comprises a first face, and the second image comprises a second face; and in response to the selection of the first image and / or the second image, blending the second face into the first image, wherein the first image after blending comprises a third face, which is different from the first face.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing methods, apparatus, terminal devices, storage media and software products

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510346720.1, filed on March 21, 2025, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to an image processing method, apparatus, terminal device, storage medium, and program product. Background Technology

[0004] In the field of multimedia editing, face blending refers to merging the faces of one or more images into the faces of another image to obtain face-blended videos or images, which are quite interesting.

[0005] Currently, traditional face fusion is pre-packaged and cannot be modified based on the image used. For example, an application can provide a fixed image containing a face, and the user can select an image that can include another face. The terminal device can then fuse the face from the user-selected image into the face in the fixed image. However, the inventors have found that in traditional face fusion, the user cannot select arbitrary images for face fusion, the interaction of face fusion is relatively fixed, and there is a technical problem of poor flexibility in face fusion. Summary of the Invention

[0006] This disclosure provides an image processing method, apparatus, terminal device, storage medium, and program product.

[0007] This disclosure provides an image processing method, which includes:

[0008] Display a first interface, the first interface including a first control;

[0009] In response to a touch of the first control, a first image and / or a second image are displayed, the first image including a first face and the second image including a second face;

[0010] In response to the selection of the first image and / or the second image, the second face is fused into the first image, wherein the fused first image includes a third face that is different from the first face.

[0011] This disclosure also provides another image processing method, which includes:

[0012] Display a first interface, which includes a camera-captured image and a first control;

[0013] In response to a touch of the first control, a first effect is displayed, the first effect including a first image, the first image including a first face;

[0014] In response to a touch of the first special effect, a first multimedia is generated based on the first special effect and a second image captured by the camera. The first multimedia includes an image in which a second face from the second image is fused to a first face from the first image.

[0015] This disclosure also provides an image processing apparatus, which includes a display module and a fusion module, wherein:

[0016] The display module is used to display a first interface, the first interface including a first control;

[0017] The display module is further configured to, in response to touch of the first control, display a first image and / or a second image, wherein the first image includes a first face and the second image includes a second face;

[0018] The fusion module is used to fuse the second face into the first image in response to the selection of the first image and / or the second image, wherein the fused first image includes a third face that is different from the first face.

[0019] This disclosure also provides an image processing apparatus, which includes a display module and a generation module, wherein:

[0020] The display module is used to display a first interface, which includes a camera-captured image and a first control.

[0021] The display module is further configured to, in response to touch of the first control, display a first special effect, the first special effect including a first image, the first image including a first face;

[0022] The generation module is used to generate a first multimedia based on the first special effect and a second image captured by a camera, in response to a touch of the first special effect. The first multimedia includes an image after fusing a second face from the second image to a first face from the first image.

[0023] This disclosure also provides a terminal device, including: a processor and a memory;

[0024] The memory stores computer-executed instructions;

[0025] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the various image processing methods that may be involved above.

[0026] This disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the various image processing methods described above.

[0027] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the various image processing methods described above. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of this disclosure;

[0030] Figure 2 is a schematic flowchart of an image processing method provided in an embodiment of this disclosure;

[0031] Figure 3 is a functional schematic diagram of single-person face fusion provided in an embodiment of this disclosure;

[0032] Figure 4 is a functional schematic diagram of multi-person face fusion provided in an embodiment of this disclosure;

[0033] Figure 5 is a schematic diagram of another multi-person face fusion function provided in an embodiment of this disclosure;

[0034] Figure 6 is a schematic diagram of a process for displaying a first interface according to an embodiment of the present disclosure;

[0035] Figure 7 is a schematic diagram of another process for displaying the first interface provided in an embodiment of this disclosure;

[0036] Figure 8 is a schematic diagram of a process for displaying a first image according to an embodiment of the present disclosure;

[0037] Figure 9 is a schematic diagram of displaying a third image according to an embodiment of this disclosure;

[0038] Figure 10 is a schematic diagram of displaying a fourth image according to an embodiment of this disclosure;

[0039] Figure 11 is a schematic diagram of another display of the first image provided by an embodiment of this disclosure;

[0040] Figure 12 is a schematic diagram of a selection button being grayed out according to an embodiment of this disclosure;

[0041] Figure 13 is a schematic diagram of a process for generating a first template according to an embodiment of this disclosure;

[0042] Figure 14 is a flowchart illustrating another image processing method provided in an embodiment of this disclosure;

[0043] Figure 15 is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this disclosure;

[0044] Figure 16 is a schematic diagram of another image processing apparatus provided in an embodiment of this disclosure; and

[0045] Figure 17 is a schematic diagram of the structure of a terminal device provided in an embodiment of this disclosure. Detailed Implementation

[0046] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0047] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0048] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the terminal device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0049] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the terminal device.

[0050] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0051] For ease of understanding, the concepts involved in the embodiments of this disclosure will be explained below.

[0052] The application scenarios of the embodiments of this disclosure will now be described in detail with reference to Figure 1.

[0053] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of this disclosure. Referring to Figure 1, the device includes a terminal device. The interface displayed on the terminal device includes image A, image B, and a fusion control. Image A includes face 1, and image B includes face 2. Image A can be a predetermined image, and image B can be an image selected by the user. When the user clicks the fusion control, the terminal device can fuse face 2 into image A; that is, the terminal device can fuse the features of face 2 into face 1 in image A to obtain image C. Image C can include face 3, which can be the result of fusing face 1 and face 2. The terminal device can perform face fusion on image A and image B based on any feasible face algorithm; this embodiment of the disclosure does not limit this.

[0054] Optionally, the application scenarios of this disclosure can be template creation scenarios, template usage scenarios, special effects creation scenarios, and special effects usage scenarios. Figure 1 is only an example of the application scenarios of this disclosure and is not a limitation on the application scenarios.

[0055] Face blending refers to the merging of faces. It can merge faces from one or more images into faces from another image, resulting in blended videos or images with high levels of interest and interactivity. Currently, traditional face blending is pre-packaged and cannot be modified. For example, when developing a face blending function, developers can configure an image including face A. When using the function, the user can select another image, which may include face B. The terminal device can then merge face B into face A to obtain the blended face C. However, in this method, face A is fixed, and the style and background of the image containing face A cannot be modified. Users cannot select arbitrary images for face blending, resulting in a relatively fixed interaction and limited flexibility in both the blending process and the user experience.

[0056] This disclosure provides an image processing method. A terminal device can display a first interface, which may include a first control. The first control may include a first sub-control, a second sub-control, and a third sub-control. Touching the first or second sub-control allows the terminal device to display a first image and a second image. In response to touch the third sub-control, the terminal device can display a second image. In response to the selection of the first and / or second image, a second face is merged into the first image to obtain a merged first image. The merged first image may include a third face different from the first face. In this method, when a user touches the first, second, or third sub-control, the terminal device can display the first and / or second image associated with that sub-control. Therefore, the user can quickly and easily select the first and second images for face merging, reducing operational complexity and improving the efficiency of face merging interaction. Furthermore, since the user can freely select the first and second images, the flexibility and fun of face merging interaction can be improved, enhancing the user experience.

[0057] The technical solutions of this disclosure and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this disclosure will now be described with reference to the accompanying drawings.

[0058] Figure 2 is a schematic flowchart of an image processing method provided in an embodiment of this disclosure. Referring to Figure 2, the method may include:

[0059] S201. Display the first interface, which includes the first control.

[0060] The execution entity of this disclosure can be a terminal device or an image processing device installed in the terminal device. The image processing device can be implemented in software, or it can be implemented using a combination of software and hardware; this disclosure does not limit this approach.

[0061] The first control can be associated with a function for face fusion, which can be facial fusion, that is, merging faces from multiple images.

[0062] In some embodiments, this disclosure allows face fusion based on a first image and a second image. The first image may include a first face, and the second image may include a second face. For example, the images required for face fusion may include a facial source image and a fusion base image, wherein the facial source image can be used to provide facial material. For example, the facial source image may include a face that can be used for face fusion, and the fusion base image may be the target image for face fusion, i.e., other faces are fused into the face in the fusion base image.

[0063] In some embodiments, the first image may be a face-blending background image, and the second image may be a facial material image; that is, the image including the first face is the face-blending background image, and the image including the second face is the facial material image.

[0064] For example, image A includes face 1, image B includes face 2, the background of image A (the area other than the face) is background 3, and the background of image B is background 4. If image A is a face blending background image (first image) and image B is a face material image (second image), then when the terminal device performs face blending on image A and image B, it can blend face 2 with image A, and face 2 is blended with face 1 in image A. The background of the new image obtained is background 3 (the background of image A), and the new image face is the face after the blending of face 1 and face 3.

[0065] In some embodiments, the first control may include a first sub-control, a second sub-control, and a third sub-control, wherein the first sub-control may be associated with a single-person face fusion function, the second sub-control may be associated with a multi-person face fusion function, and the third sub-control may be associated with a face prediction function. For example, the first interface may include the first sub-control. For example, the first interface may include the second sub-control. For example, the first interface may also include the third sub-control. For example, the first interface may also include the first sub-control, the second sub-control, and the third sub-control; this disclosure does not limit this aspect.

[0066] In some embodiments, the function of single-person face fusion can refer to fusing a single face. For example, a first image may include a first face, and a second image may also include a second face. The function of single-person face fusion can fuse the second face in the second image into the first face in the first image. For example, the first image may also include multiple first faces, and the second image may include a second face. The function of single-person face fusion can fuse one second face in the second image with any one of the first faces in the first image.

[0067] The following explanation, with reference to Figure 3, will illustrate the function of single-person face blending.

[0068] Figure 3 is a functional schematic diagram of single-person face fusion provided by an embodiment of this disclosure. Referring to Figure 3, it includes a first image and a second image. The first image includes face 1, face 2, and a triangular marker, while the second image includes face 3 and a square marker. A terminal device (not shown in Figure 3) can perform single-person face fusion processing on the first and second images to obtain a fused face result. This fused face result includes the triangular marker (with the same background as the first image) and the face resulting from the fusion of face 1 and face 3. Optionally, the terminal device can also fuse face 2 and face 3; this embodiment of the disclosure does not limit this.

[0069] In some embodiments, the function of multi-face fusion can refer to fusing multiple faces. For example, a first image may include multiple faces, and a second image may also include multiple faces. The multi-face fusion function can fuse multiple faces in the second image into multiple faces in the first image. For example, one face in the first image is fused with one face in the second image. For example, the first image may include multiple faces, and the second image may include one face. If the number of second images is greater than one, then the multi-face fusion function can fuse each face in the second image with a face in the first image.

[0070] The following explanation, in conjunction with Figures 4 and 5, details the function of multi-person face fusion.

[0071] Figure 4 is a functional schematic diagram of multi-person face fusion provided by an embodiment of this disclosure. Referring to Figure 4, it includes: a first image and a second image. The first image includes face 1, face 2, and a triangular marker; the second image includes face 3, face 4, and a square marker. A terminal device (not shown in Figure 4) can perform multi-person face fusion processing on the first image and the second image to obtain the fused face result. This fused face result includes the triangular marker (with the same background as the first image), the face fused from face 1 and face 3, and the face fused from face 2 and face 4.

[0072] Figure 5 is a functional schematic diagram of another multi-person face fusion provided by an embodiment of this disclosure. Referring to Figure 5, it includes: a first image, a second image A, and a second image B. The first image includes face 1, face 2, and a triangular marker; the second image A includes face 3 and a square marker; and the second image B includes face 4 and a square marker. A terminal device (not shown in Figure 5) can perform multi-person face fusion processing on the first image, second image A, and second image B to obtain a face fusion result. This face fusion result includes a triangular marker (with the same background as the first image), the face fused from face 1 and face 3, and the face fused from face 2 and face 4.

[0073] It should be noted that in the single-person face fusion function and the multi-person face fusion function, faces can be paired arbitrarily (i.e., fused after pairing) or paired in a preset order (as shown in Figure 5, the faces of the first image are fused with the faces of the second image from left to right). This disclosure does not limit this.

[0074] In some embodiments, the function of face prediction can refer to predicting new faces. For example, a first image may include a first face, a second image A may include a second face, and a second image B may include another second face. The function of face prediction can be based on fusing the second faces of the second images A and B together into a first face in the first image.

[0075] In some embodiments, this disclosure can be applied to template scenarios, where the template may include a pre-designed video or image production framework. The template may include elements such as text, images, audio, and animation. Users can input videos or images into the template and obtain videos or images that match the template style. Templates may include text / image templates and video templates. When the template is a text / image template, the user can generate an image using the template; when the template is a video template, the user can generate a video using the template.

[0076] In some embodiments, in a template generation scenario, the terminal device can display a first interface based on the following feasible implementation: displaying a second interface, the second interface including a first control and a first slot, and displaying the first interface in response to touch on the first control. This allows users to generate face-blending templates in template application scenarios, improving the flexibility of face-blending.

[0077] The second interface can be used for multimedia editing. For example, the second interface can be a multimedia editing interface, which may include an editing area and a preview area, where the preview area can display the image to be edited in the multimedia editing interface.

[0078] In some embodiments, the first slot includes a second image. The first slot is used to display the second image.

[0079] In some embodiments, the editing area may include a first slot. For example, the editing area may include a main track and at least one sub-track, wherein slots can be added to each track, and the slots can be filled with materials such as images, videos, and audio. Furthermore, the images and videos in the slots can be displayed in the preview area of ​​the second interface.

[0080] In some embodiments, the second interface may include a toolbar, which may include the first control. For example, the bottom of the second interface may include a toolbar, which may include the first control, as well as other controls such as filter controls and sticker controls.

[0081] The process of displaying the first interface will be explained below with reference to Figure 6.

[0082] Figure 6 is a schematic diagram of a process for displaying a first interface according to an embodiment of this disclosure. Referring to Figure 6, a second interface is also included. The second interface may include a preview area, an editing area, and a toolbar. The editing area may include slots for image 1 and image 2, the preview area may include image 1 and image 2 (not shown in Figure 6), and the toolbar may include filter controls, sticker controls, and face blending controls. The face blending control may be a first control. After the user clicks the face blending control, the terminal device (not shown in Figure 6) can jump from the second interface to the first interface. The first interface may include controls for single-person face blending, controls for multi-person face blending, controls for face prediction, and an area for adding images. The area for adding images can be used to display the first and second images selected by the user.

[0083] In some embodiments, in special effects generation scenarios, the terminal device can display a first interface based on the following feasible implementation: displaying a third interface, which is the interface captured by the camera and includes a third control; and displaying the first interface in response to touch on the third control. This allows users to generate face-blending effects in special effects application scenarios, improving the flexibility of face-blending generation.

[0084] In some embodiments, the third interface may include a toolbar, which may include third controls. For example, the third interface may be a camera shooting interface, and the bottom of the interface may include a toolbar, which may include third controls, such as filter controls, sticker controls, and text controls.

[0085] The following, with reference to Figure 7, describes another process for displaying the first interface.

[0086] Figure 7 is a schematic diagram of another process for displaying the first interface provided by an embodiment of this disclosure. Referring to Figure 7, a third interface is also included. The third interface is the interface captured by the camera, and may include the image captured by the camera. The bottom of the third interface may include a toolbar, which includes filter controls, sticker controls, and face blending controls (third controls). After the user clicks the face blending control, the terminal device (not shown in Figure 7) can jump from the third interface to the first interface. The first interface may include controls for single-person face blending, controls for multi-person face blending, controls for face prediction, and an area for adding images.

[0087] S202, in response to touch of the first control, display the first image and / or the second image.

[0088] In some embodiments, the response to a touch of the first control can also be a response to a touch of a first sub-control, a second sub-control, or a third sub-control. For example, the first control displayed by the terminal device on the first interface includes a first sub-control, a second sub-control, and a third sub-control. If the user clicks the first sub-control, the terminal device can determine that the user has selected the single-person face fusion function; if the user clicks the second sub-control, the terminal device can determine that the user has selected the multi-person face fusion function; and if the user clicks the third sub-control, the terminal device can determine that the user has selected the face prediction function.

[0089] In some embodiments, during template generation, the terminal device can display a first image and / or a second image based on the following feasible implementation: displaying the first image and the second image in response to touch on a first sub-control or a second sub-control, and displaying the second image in response to touch on a third sub-control. In this way, the terminal device can display a blended background image and a facial image based on the selected second function, facilitating user selection of the blended image, reducing operational complexity, and improving interaction efficiency.

[0090] In some embodiments, the number of faces in the first image displayed by the terminal device is associated with a sub-control touched by the user. Specifically, displaying the first image and the second image in response to touch of the first or second sub-control can be as follows: The first image includes one first face; the second image includes multiple first faces. For example, the first sub-control may be associated with a single-person face merging function, and the second sub-control may be associated with a multi-person face merging function. If the user clicks the first sub-control, the number of first faces in the first image displayed on the first interface can be 1 (or greater than 1, which is not limited in this embodiment); if the user clicks the second sub-control, the number of faces in the first image displayed on the first interface is N (or greater than N, where N is an integer greater than 1). This allows the user to quickly select the face merging background image associated with the face merging function, improving the efficiency of user interaction.

[0091] In some embodiments, the first interface may include a first area and / or a second area, wherein a first image is displayed in the first area and a second image is displayed in the second area. This facilitates the user's selection of the first and second images, improving the accuracy of user interaction.

[0092] In some embodiments, the first area may include a plurality of preset first images, or it may include a first image selected by the user in the album, and this disclosure does not limit this.

[0093] In some embodiments, the second area may include a second image, which may be a second image to be edited or a second image selected by the user in the album; this disclosure does not limit this. For example, in the second interface, if the editing area includes slots for second image 1 and second image 2, then the second area of ​​the first interface may include second image 1 and second image 2, and the second area may also include slots for second image 1 and second image 2, thus facilitating user selection of the second image.

[0094] In some embodiments, when a user clicks the third sub-control, the terminal device can determine that the user has selected the face prediction function. The terminal device can display the second area on the first interface without displaying the first area (the first image can be a system default image). That is, when the user clicks the third sub-control, the terminal device can display the second image on the first interface, or it can choose not to display the first image. It should be noted that in some embodiments, when the user clicks the third sub-control, the user can also display the first area on the first interface, or display the first image. This disclosure does not limit this.

[0095] The process of displaying the first image on the terminal device will be explained below with reference to Figure 8.

[0096] Figure 8 is a schematic diagram of a process for displaying a first image according to an embodiment of this disclosure. Referring to Figure 8, the interface includes: a first interface. The first interface includes a control for single-person face fusion (first sub-control), a control for multi-person face fusion (second sub-control), a control for face prediction (third sub-control), and an area for adding images. When the user clicks the control for single-person face fusion, the first interface also includes a first region and a second region. The first region may include first image 1, first image 2, first image 3, and first image 4, and each first image in the first region includes a face. The second region includes second image A and second image B.

[0097] Referring to Figure 8, when the user clicks the control for multi-person face fusion, the terminal device can replace the first image displayed in the first region. This first region may include first image 5, first image 6, first image 7, and first image 8, and each first image in the first region includes multiple faces. When the user clicks the control for face prediction, the terminal device can de-display the first region on the first interface, while the second region remains displayed. In this way, when the user selects different sub-controls, the images displayed by the terminal device can be associated with the face fusion function, making it easier for the user to select images for face fusion, reducing operational complexity, and improving the efficiency of face fusion.

[0098] In some embodiments, in a template generation scenario, the image processing method further includes: displaying a third image from a user-authorized photo album in response to a touch associated with the first area; and displaying the touched third image in the first area in response to a touch on the third image. This allows the user to select a first image from the photo album, improving the flexibility of interaction.

[0099] The third image includes one or more first faces. For example, if a user clicks the first sub-control, the third image in the album may include one first face; if the user clicks the second sub-control, the third image in the album may include multiple first faces.

[0100] It should be noted that the album can also display images that do not include faces, but users cannot select images that do not include faces.

[0101] Optionally, the first area may include a control for displaying a photo album, which the terminal device can display when the user clicks the control, and the photo album includes a third image.

[0102] The process of the terminal device displaying the third image in the first area will be explained below with reference to Figure 9.

[0103] Figure 9 is a schematic diagram of displaying a third image according to an embodiment of this disclosure. Referring to Figure 9, it includes: a first interface. The first interface includes controls for single-person face fusion, controls for multi-person face fusion, controls for face prediction, and an area for adding images. When the user clicks the single-person face fusion control, the first interface also includes a first area and a second area. The first area may include a first image 1 (a face) and controls for displaying a photo album. The second area includes a second image A and a second image B.

[0104] Referring to Figure 9, when the user clicks the control for displaying the photo album, the terminal device (not shown in Figure 9) can display the photo album. This album may include third images 5, 6, 7, 8, 9, and 10, each containing one face. When the user clicks on third image 5, the terminal device can redisplay the first interface and display third image 5 in the first area of ​​the first interface. This allows the user to flexibly select the first image, improving the flexibility of the interaction.

[0105] In some embodiments, in a template generation scenario, the image processing method further includes: displaying a fourth image in the album in response to a touch associated with the second area; and displaying the touched fourth image in the second area in response to a touch on the fourth image. This allows the user to select a second image from the album, improving the flexibility of interaction.

[0106] The process of the terminal device displaying the fourth image in the second area will be explained below with reference to Figure 10.

[0107] Figure 10 is a schematic diagram of displaying a fourth image according to an embodiment of this disclosure. Referring to Figure 10, it includes: a first interface. The first interface includes controls for single-person face fusion, controls for multi-person face fusion, controls for face prediction, and an area for adding images. When the user clicks the single-person face fusion control, the first interface also includes a first area and a second area. The first area may include a first image 1 (a face) and controls for displaying a photo album. The second area includes a second image A, a second image B, and controls for displaying a photo album.

[0108] Referring to Figure 10, when the user clicks the control for displaying the photo album in the second area, the terminal device (not shown in Figure 10) can display the photo album. This album may include fourth images 5, 6, 7, 8, 9, and 10, each containing one face. When the user clicks on fourth image 5, the terminal device can redisplay the first interface and display fourth image 5 in the second area of ​​the first interface. This allows the user to flexibly select the second image, improving the flexibility of the interaction.

[0109] In some embodiments, after the user selects a second image, the terminal device may perform face blending processing, and if the second image selected by the user does not include a face, the terminal device may generate a prompt message to prompt the user to select an image that includes a face.

[0110] In some embodiments, in a special effects generation scenario, the terminal device can display a first image based on the following feasible implementation: in response to touch of a first sub-control or a third sub-control, displaying a first image including a first face; in response to touch of a second sub-control, displaying a first image including multiple first faces. In this way, the terminal device can automatically select the first image, and the user does not need to select a second image. Therefore, the terminal device can quickly generate special effects, improving the efficiency of special effects generation.

[0111] In the scene where special effects are generated, the user can select the first image without needing to select the second image, which can be the system default image.

[0112] When the user clicks the third sub-control, the first image on the first interface can be a sample image. In actual application, the terminal device can merge the faces in the two second images into the face in the sample image.

[0113] The process of displaying at least one first image associated with the second function on the terminal device will now be described with reference to Figure 11.

[0114] Figure 11 is a schematic diagram of another method for displaying a first image according to an embodiment of this disclosure. Referring to Figure 11, it includes: a first interface. The first interface includes controls for single-person face fusion, controls for multi-person face fusion, controls for face prediction, a fusion target image, and multiple parameters. When the user clicks the single-person face fusion control, the fusion target image is the first image 1, and the first interface can display this first image 1 (a face). The fusion duration is 3, the blending mode is normal, the stretching mode is automatic adaptation, the color is white, and the transparency is 50%.

[0115] Referring to Figure 11, when the user clicks the control for merging multiple faces, the target image changes from first image 1 to first image 2, and the first interface displays first image 1 as first image 2 (multiple faces). When the user clicks the control for face prediction, the first interface can predict the gender of the face; the user can select male, and the first image 2 displayed on the first interface changes to first image 3, which may include a male face. Thus, in the scene where special effects are generated, the user does not need to select a second image, and the terminal device can assist the user in selecting the first image, improving the flexibility and efficiency of the interaction.

[0116] In some embodiments, in the special effects generation scenario, the first interface further includes a fourth control, and the image processing method further includes: displaying a fifth image in response to touch of the fourth control; and replacing the first image displayed on the first interface with the touched fifth image in response to touch of the fifth image. Thus, in the special effects generation scenario, users can flexibly select the face blending background image, improving the flexibility of interaction.

[0117] The fifth image includes a first image of a first face, or multiple first images of first faces. For example, the fifth image may be associated with a sub-control clicked by the user, which will not be elaborated further in this disclosure.

[0118] Optionally, the fourth control can be used to display a photo album, or it can be used to modify the predicted gender of a face; however, this embodiment of the present disclosure does not limit this.

[0119] Optionally, in the embodiment shown in Figure 11, the first interface may further include a fourth control. When the user clicks the control for single-person face fusion or multi-person face fusion, if the user clicks the fourth control, the terminal device can display a photo album. The user can select an image in the photo album and replace the first image 1 or the first image 2. When the user clicks the control for face prediction, if the user clicks the fourth control, the predicted gender of the face can be switched from male to female, and the first image 3 will also become an image of a female face.

[0120] S203, in response to the selection of the first image and / or the second image, the second face is fused into the first image.

[0121] The fused first image includes a third face, which is different from the first face. For example, in the embodiment shown in Figure 3, face 1 can be the second face, face 3 can be the first face, and face 1 + face 3 can be the third face. The third face is the face obtained by fusing the first face and the second face. Therefore, the third face is still in the first image. However, the third face is different from the first face before fusion. That is, the terminal device processes the first face in the first image based on the second face (fusing the information of the second face) to obtain the third face.

[0122] In some embodiments, after generating the fused first image, the terminal device may also generate a first template or a first effect. The first template or first effect includes the fused first image. For example, the first template can be used to generate a video, which may include the fused first image. For instance, the first template may include a first image and a second image. When a user uses the first template, they can select a second image and replace the second image in the first template. Thus, the user can generate a video based on the first template, which may include the fused first image, determined based on the first image in the first template and the second image selected by the user.

[0123] It should be noted that in the application scenario of the first template, the user can select a second image, and in the application scenario of the first special effect, the user can take a picture of a second image. The function of the first special effect is similar to that of the first template, and will not be described in detail in this embodiment.

[0124] In some embodiments, the first image and the second image further include markers indicating that the image has been touched. In response to selection of the first image and / or the second image, the terminal device may specifically: in response to touch of the first image, highlight the markers of the touched first image; in response to touch of the second image, when the number of touched second images equals the number of first faces in the touched first images, gray out the markers of the other second images. In this way, the terminal device can assist the user in selecting second images, avoiding situations where the user selects too many or too few second images, leading to face merging failure, reducing the complexity of user interaction, improving the flexibility of user interaction, and enhancing the user experience.

[0125] In some embodiments, when the marker of the second image is grayed out, it indicates that the second image cannot be selected, that is, the user cannot select the second image for face blending.

[0126] In some embodiments, the second region may include slots for second images. Therefore, the slots for second images may also include selection buttons. When a user clicks the selection button, it indicates that the user selects the second image in the slot, and the terminal device can perform face blending based on the second image. For example, the upper right corner of the slot may include a selection button, and the second region may include slot 1 and slot 2. After the user clicks the selection button for slot 1, the terminal device can perform face blending based on the second image in slot 1. After the user clicks the selection buttons for both slot 1 and slot 2, the terminal device can perform face blending based on the second images in both slot 1 and slot 2.

[0127] In some embodiments, when the slot selection button is grayed out, the slot cannot be selected, that is, the user cannot select the second image in the slot for face blending.

[0128] Optionally, the terminal device may determine the number of faces in the first image based on any feasible implementation method, and this embodiment of the disclosure does not limit this.

[0129] For example, the first image includes two faces, and the second area includes three slots. If the user selects two slots, the terminal device can gray out the selection button for the last slot, that is, the user cannot continue to select the last slot.

[0130] The process of graying out the selection buttons for other slots in the second area by the terminal device will be explained below with reference to Figure 12.

[0131] Figure 12 is a schematic diagram of a selection button being grayed out according to an embodiment of this disclosure. Referring to Figure 12, it includes: a first interface. The first interface includes controls for single-person face fusion, controls for multi-person face fusion, controls for face prediction, and an area for adding images. When the user clicks the single-person face fusion control, the first interface also includes a first area and a second area. The first area includes a first image, and the second area includes slots for image A and image B. Image A includes a face, and image B also includes a face.

[0132] Referring to Figure 12, the first image includes a face. When the user selects the first image, the terminal device (not shown in Figure 12) can display the second image in the area where the image is added. When the user selects a slot for image A, the selection button for that slot is highlighted. Since the user has already selected a face, the selection button for the slot for image B is grayed out, and the user can no longer select a slot for image B. That is, the terminal device can perform face merging based on the first image and image A.

[0133] In the embodiment shown in Figure 12, if the user needs to merge the first image and image B, the user can click on the slot of image A again to deselect the slot of image A. Therefore, the selection buttons for the slots of image A and image B are both clickable. The user can click on the selection button for the slot of image B, so that the terminal device can merge the face based on the first image and image B.

[0134] In some embodiments, the first interface further includes a second control. After the user selects the first image and / or the second image, the terminal device can generate a first template based on the following feasible implementation: in response to touch of the second control, the second interface is redisplayed, and a second slot is displayed on the second interface, wherein the second slot includes the first image after face fusion; the terminal device can generate the first template, wherein the first template is used to generate the image after face fusion. In this way, after the terminal device generates the first image after face fusion, the user can adjust the display position and duration of the first image after face fusion based on the slot of the first image after face fusion, thereby improving the flexibility of the first template.

[0135] The process of generating the first template will now be explained with reference to Figure 13.

[0136] Figure 13 is a schematic diagram of a process for generating a first template according to an embodiment of this disclosure. Referring to Figure 13, it includes: a first interface. The first interface includes controls for single-person face fusion, controls for multi-person face fusion, controls for face prediction, an area for adding images, and a confirmation control (second control). When the user clicks the single-person face fusion control, the first interface also includes a first area and a second area. The first area includes a first image 1 and a first image 2, and the second area includes slots for image A and image B. The first image 1 includes two faces, and the first image 2, image A, and image B each include one face. When the user selects the first image 1, the terminal device can display the first image 1 in the image addition area.

[0137] Referring to Figure 13, when the user clicks the confirmation control, the terminal device (not shown in Figure 13) can jump from the first interface to the second interface. The second interface includes a preview area, an editing area, a toolbar, and a generation control. The editing area includes slots for image A, image B, and the first image 1 after face merging. The toolbar includes filter controls, sticker controls, and face merging controls. After the user clicks the generation control, the terminal device can generate a first template based on image A, image B, and the first image 1 after face merging.

[0138] Optionally, when the terminal device generates the first template, it can fix the slot in the second area. After fixing the slot, the user does not need to select the second image for that slot when using the first template. For example, if the first template includes two slots, the user would need to reselect the corresponding second image for each slot when using the first template. However, if the second image in one slot is fixed during the generation of the first template, the user only needs to select the second image for the other slot. This increases the application scenarios for using the first template and improves its flexibility.

[0139] In some embodiments, the user can edit the slots in the first image after face merging. For example, the user can delete the slots in the first image after face merging, and the terminal device can stop the face merging process. For example, the user can crop the slots in the first image after face merging, and the user can also modify the second image. After the second image is modified, the terminal device can re-perform the face merging process.

[0140] In some embodiments, since the face fusion process takes a certain amount of time, if the user clicks on controls such as publish or preview during the face fusion process, the terminal device can generate a prompt message to remind the user that the face fusion is not yet finished. When the face fusion is completed, the terminal device can notify the user that the face fusion is complete. If the face fusion fails, the terminal device can also notify the user that the face fusion has failed and display the reason for the failure (e.g., face not detected, network abnormality, etc.).

[0141] In some embodiments, after the terminal device determines the first image after face fusion, it can cache the first image after face fusion locally. When the user uses the first image after face fusion again, the terminal device does not need to regenerate the first image after face fusion.

[0142] In some embodiments, the slot of the first image after face fusion can be at the end of the main track or at the position of the first selected slot within the drawing, and the length of the slot of the first image after face fusion is the same as the length of the first selected slot. This disclosure does not limit this aspect.

[0143] In some embodiments, the image processing method described above may further include: generating a first special effect. The first special effect can be used to generate a face-blended image. For example, in a scenario where the special effect is generated, since the user does not need to select a second image, the terminal device can display a first image. After user confirmation, the terminal device can generate the first special effect. When the first special effect is used, the user can take an image including a face, and based on that image and the first image selected during the creation of the first special effect, generate a face-blended image.

[0144] In some embodiments, during the special effects generation scenario, the first interface may further include a preview control. When the user clicks the preview control, the terminal device can display a photo album, allowing the user to select an image from the album and apply the special effects to that selected image, thus obtaining a preview result. If the second image selected by the user is insufficient for face blending during the preview, the terminal device can prompt the user to select a second image. This allows users to preview the effect while creating special effects, thereby improving the efficiency of special effects creation.

[0145] In some embodiments, when creating templates for single-person or multi-person face fusion, the user can select a first image and a second image; when creating templates for facial prediction, the user can select a second image; and when creating special effects for single-person, multi-person, and facial prediction, the user can select a first image.

[0146] It should be noted that in the embodiments of this disclosure, the first face and the second face do not specifically refer to one face, but rather indicate that the first face is the face in the face blending background image and the second face is the face in the face blending material. In the embodiments of this disclosure, the second face can be blended into the first face.

[0147] This disclosure provides an image processing method. A terminal device can display a first interface. In response to touch on a first sub-control or a second sub-control in the first interface, the terminal device can display a first image and a second image. In response to touch on a third sub-control in the first interface, the terminal device can display a second image. In response to selection of the first image and / or the second image, a second face is merged into the first image, wherein the merged first image includes a third face, which is different from the first face. Thus, when a user clicks on a sub-control, the terminal device can display the image required for the face-merging function associated with that sub-control. Therefore, the user can quickly and easily select the image for face merging, reducing operational complexity and improving the efficiency of face merging. Furthermore, since the user can freely select the image for face merging, the flexibility and fun of face merging can be improved, enhancing the flexibility of face merging interaction and thus improving the user experience.

[0148] Based on the embodiment shown in Figure 2, this disclosure also includes another image processing method. The other image processing method will be described in detail below with reference to Figure 14.

[0149] Figure 14 is a schematic flowchart of another image processing method provided in an embodiment of this disclosure. Referring to Figure 14, it includes:

[0150] S1401. Display the first interface, which includes the image captured by the camera and the first control.

[0151] The image captured by the camera may include one face or multiple faces; this disclosure does not limit this.

[0152] The first control can be used to display a first special effect. For example, the first interface may include the captured content and the first control for displaying the first special effect. When the terminal device displays the first interface, the first interface may include the captured image, and the bottom of the captured image may include the first control, wherein when the user clicks the first control, the bottom of the captured interface may display the application of the first special effect.

[0153] In some embodiments, the main interface of a multimedia editing application may include a plus sign control. When the user clicks the plus sign control, the terminal device may display a first interface. It should be noted that the terminal device may also display the first interface based on any feasible implementation method, and this disclosure does not limit this.

[0154] S1402, In response to touch on the first control, display the first special effect.

[0155] The first special effect includes a first image, and the first image includes a first face. For example, when a user clicks the first control, the terminal device can display a special effects panel, which may include the first special effect.

[0156] S1403, in response to a touch of the first special effect, generate the first multimedia based on the first special effect and a second image captured by the camera.

[0157] The first multimedia includes an image obtained by fusing a second face from a second image to a first face from a first image. Optionally, the first multimedia can be a video or an image; this disclosure does not limit the specific type of multimedia.

[0158] For example, when a terminal device launches a multimedia editing application, it can display the main interface of the application, which includes a plus sign control at the bottom. When the user clicks the plus sign control, the terminal device can display a shooting interface, which may include the captured image and a first control. When the user clicks the first control, an effects panel can be displayed at the bottom of the shooting interface, which may include a first special effect. When the user clicks the first special effect, the captured image can display a face-blending result, combining the face in the image with the face from the first special effect. When the user clicks "shoot," the terminal device can generate a video or image that includes the face-blending result.

[0159] In some embodiments, the first interface may further include a second control, and the image processing method further includes: in response to touch of the second control, displaying a first template, the first template including a third image, the third image including a first face; in response to touch of the first template, displaying a fourth image, the fourth image including a second face; and in response to touch of the fourth image, generating second multimedia. Thus, the face-blending result in the second multimedia is associated with the second face in the user-selected fourth image, improving the display flexibility of the second multimedia.

[0160] The second multimedia includes an image after fusing the second face of the fourth image into the first face of the third image.

[0161] In some embodiments, the terminal device may display a fourth image based on the following possible implementation: playing a video that includes function buttons, and displaying the fourth image in response to touch of the function buttons.

[0162] The function buttons can be used to create videos that apply the template. For example, a function button could be a "Shoot the same style" button, which, when clicked by the user, allows the device to create a new video that applies the template.

[0163] For example, the video played by the terminal device is generated based on a first template. The video may include a "Take the same photo" function button. When the user clicks the function button, the terminal device may display a photo album, which may include a fourth image, which may include a second face.

[0164] For example, the video is created based on a first template, which includes a first image and two second images. When the video is played on a terminal device, the user can click a "shoot the same thing" button in the video. The terminal device can display a photo album, which may include multiple fourth images. The user can select two fourth images, and the terminal device can generate a new video based on the first image in the first template and the two selected fourth images. This new video may include the first image, the two selected fourth images, and a fusion result where the second faces from the two selected fourth images are blended into the first face of the first image.

[0165] This disclosure provides an image processing method. A first interface is displayed. In response to touch on a first control in the first interface, a first special effect is displayed. In response to touch on the first special effect, a terminal device generates first multimedia based on the first special effect and a second image captured by a camera. Furthermore, in response to touch on the second control in the first interface, the terminal device can display a first template. In response to touch on the first template, a fourth image is displayed. In response to touch on the fourth image, the terminal device can generate second multimedia. In the above method, the first multimedia may include an image after fusing a second face from the second image to a first face from the first image. The second multimedia may include an image after fusing a second face from the fourth image to a first face from a third image. Therefore, users can quickly and easily create face-blended videos based on the first template and the first special effect, improving the efficiency of user interaction.

[0166] Figure 15 is a schematic diagram of an image processing apparatus provided in an embodiment of this disclosure. Referring to Figure 15, the image processing apparatus 1500 includes a display module 1501 and a fusion module 1502, wherein:

[0167] The display module 1501 is used to display a first interface, the first interface including a first control;

[0168] The display module 1501 is further configured to, in response to touch of the first control, display a first image and / or a second image, wherein the first image includes a first face and the second image includes a second face;

[0169] The fusion module 1502 is configured to, in response to the selection of the first image and / or the second image, fuse the second face into the first image, wherein the fused first image includes a third face that is different from the first face.

[0170] The image processing apparatus provided in this disclosure allows users to quickly select images for face fusion, reducing operational complexity. Furthermore, since users can freely select images for face fusion, the flexibility and fun of face fusion are improved, as is the flexibility of face fusion interaction, thereby enhancing the user experience.

[0171] According to one or more embodiments of this disclosure, the display module 1501 is specifically used for:

[0172] In response to a touch on the first sub-control or the second sub-control, the first image and the second image are displayed;

[0173] In response to a touch of the third sub-control, the second image is displayed.

[0174] The image processing apparatus provided in this disclosure allows the terminal device to display the image required for the face fusion function associated with the sub-control when the user selects different sub-controls, making it easier for the user to select the face fusion image, reducing operational complexity, and improving interaction efficiency.

[0175] According to one or more embodiments of this disclosure, the display module 1501 is specifically used for:

[0176] In response to a touch of the first sub-control, the first image and the second image are displayed, the first image including a first face;

[0177] In response to a touch of the second sub-control, the first image and the second image are displayed, the first image including a plurality of first faces.

[0178] In this way, the display of the first image can be associated with the first sub-control and the second sub-control, improving the flexibility of image display and reducing the complexity of user interaction.

[0179] According to one or more embodiments of this disclosure, the first interface includes a first region and / or a second region, wherein the first image is displayed in the first region and the second image is displayed in the second region.

[0180] In this way, the terminal device can display the first image and the second image in separate areas, avoiding the user selecting the wrong image for face merging and improving the accuracy of user interaction.

[0181] According to one or more embodiments of this disclosure, the display module 1501 is further configured to:

[0182] In response to a touch associated with the first area, a third image from a user-authorized photo album is displayed, the third image including one or more first faces;

[0183] In response to a touch of the third image, the touched third image is displayed in the first area.

[0184] This allows users to select the first image from their album, increasing the flexibility of the interaction.

[0185] According to one or more embodiments of this disclosure, the display module 1501 is further configured to:

[0186] In response to a touch associated with the second area, a fourth image in the album is displayed, the fourth image including one or more second faces;

[0187] In response to a touch of the fourth image, the touched fourth image is displayed in the second area.

[0188] This allows users to select a second image from their photo album, increasing the flexibility of the interaction.

[0189] According to one or more embodiments of this disclosure, the fusion module 1502 is specifically used for:

[0190] In response to a touch on the first image, the marker of the touched first image is highlighted;

[0191] In response to a touch on a second image, when the number of touched second images is equal to the number of first faces in the touched first image, the markers of the other second images are grayed out.

[0192] In this way, the terminal device can assist the user in selecting the second image, avoiding the situation where the user selects too many or too few second images, which would lead to face fusion failure. This reduces the complexity of user interaction, increases the flexibility of user interaction, and improves the user experience.

[0193] According to one or more embodiments of this disclosure, the display module 1501 is further configured to:

[0194] Display a second interface, the second interface including a first control and a first slot, the first slot including the second image;

[0195] In response to a touch of the first control, the first interface is displayed.

[0196] In this way, users can generate face-blending templates in template application scenarios, improving the flexibility of face blending.

[0197] According to one or more embodiments of this disclosure, the display module 1501 is specifically used for:

[0198] In response to a touch of the second control, the second interface is redisplayed;

[0199] The second interface displays a second slot, which includes the first image after face fusion.

[0200] A first template is generated, which is used to generate an image after facial fusion.

[0201] In this way, after the terminal device generates the face fusion result, the user can adjust the display position and duration of the face fusion result based on the slot of the face fusion result, thereby improving the flexibility of the first template.

[0202] According to one or more embodiments of this disclosure, the display module 1501 is further configured to:

[0203] Display a third interface, which is the interface captured by the camera, and the third interface includes a third control;

[0204] In response to a touch of the third control, the first interface is displayed.

[0205] In this way, users can generate face-blending effects in scenarios where special effects are applied, increasing the flexibility of face-blending generation.

[0206] According to one or more embodiments of this disclosure, the display module 1501 is further configured to:

[0207] In response to a touch on a first child control or a third child control, a first image including a first face is displayed;

[0208] In response to a touch on the second sub-control, a first image including multiple first faces is displayed;

[0209] In this way, users do not need to select a first and second image for face blending, thus enabling terminal devices to quickly generate special effects and improve the efficiency of special effects generation.

[0210] According to one or more embodiments of this disclosure, the display module 1501 is further configured to:

[0211] In response to a touch of the fourth control, a fifth image is displayed, the fifth image including a first image of a first face, or multiple first images of first faces;

[0212] In response to a touch of the fifth image, the first image displayed on the first interface is replaced with the touched fifth image.

[0213] In this way, users can replace the first image in the context of special effects creation, improving the flexibility of face-blending interaction.

[0214] The image processing apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0215] Figure 16 is a schematic diagram of another image processing apparatus provided in an embodiment of this disclosure. Referring to Figure 16, the image processing apparatus 1600 includes a display module 1601 and a generation module 1602, wherein:

[0216] The display module 1601 is used to display a first interface, the first interface including a picture captured by the camera and a first control;

[0217] The display module 1601 is further configured to, in response to touch of the first control, display a first special effect, the first special effect including a first image, the first image including a first face;

[0218] The generation module 1602 is configured to, in response to a touch of the first special effect, generate a first multimedia based on the first special effect and a second image captured by a camera, the first multimedia including an image after fusing a second face from the second image to a first face from the first image.

[0219] In this way, since the face fusion result in the first multimedia is determined based on the first image in the first special effect and the captured second image, users can capture videos or images with the first special effect, which increases the fun and interactivity of shooting and improves the user experience. Furthermore, users can quickly and easily capture multimedia with the first special effect applied in the shooting interface, thus reducing the complexity of user interaction, reducing the complexity of multimedia generation, and improving the efficiency of multimedia generation.

[0220] According to one or more embodiments of this disclosure, the display module 1601 is further configured to:

[0221] In response to a touch of the second control, a first template is displayed, the first template including a third image, the third image including a first face;

[0222] In response to a touch on the first template, a fourth image is displayed, the fourth image including a second face;

[0223] In response to a touch on the fourth image, a second multimedia is generated, the second multimedia comprising an image in which a second face of the fourth image is fused to a first face of the third image.

[0224] In this way, since the face fusion result in the first multimedia is determined based on the first image in the first template and the selected fourth image, users can quickly and easily shoot and play videos with the same effect, reducing the complexity of operation and improving the efficiency of video editing.

[0225] The image processing apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0226] Figure 17 is a schematic diagram of the structure of a terminal device provided in an embodiment of this disclosure. Referring to Figure 17, it shows a schematic diagram of the structure of a terminal device 1700 suitable for implementing an embodiment of this disclosure. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The terminal device shown in Figure 17 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this disclosure.

[0227] As shown in Figure 17, the terminal device 1700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1701, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1702 or a program loaded from storage device 1708 into random access memory (RAM) 1703. The RAM 1703 also stores various programs and data required for the operation of the terminal device 1700. The processing unit 1701, ROM 1702, and RAM 1703 are interconnected via a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.

[0228] Typically, the following devices can be connected to I / O interface 1705: input devices 1706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1709. Communication device 1709 allows terminal device 1700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 17 shows terminal device 1700 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0229] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1709, or installed from storage device 1708, or installed from ROM 1702. When the computer program is executed by processing device 1701, it performs the functions defined in the methods of embodiments of this disclosure.

[0230] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), or any suitable combination thereof.

[0231] The aforementioned computer-readable medium may be included in the aforementioned terminal device; or it may exist independently and not assembled into the terminal device.

[0232] The aforementioned computer-readable medium carries one or more programs, which, when executed by the terminal device, cause the terminal device to perform the method shown in the above embodiments.

[0233] This disclosure provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements various methods that may be involved in the above embodiments.

[0234] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements various methods that may be involved in the above embodiments.

[0235] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0236] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than that indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0237] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0238] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), etc.

[0239] The terminal device, computer-readable storage medium, and computer program product provided in this disclosure embodiment can display a first image and a second image required for the face fusion function associated with a sub-control clicked by the user. Therefore, it is convenient for the user to select the image for face fusion, reducing the complexity of operation. Furthermore, since the user can freely select the image for face fusion, the flexibility and fun of face fusion can be improved, the flexibility of face fusion interaction can be improved, and thus the user experience can be enhanced.

[0240] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0241] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0242] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0243] It is understood that the data involved in this technical solution (including but not limited to the data itself, its acquisition, or its use) shall comply with the requirements of relevant laws, regulations, and provisions. Data may include information, parameters, and messages, such as flow control instructions.

[0244] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0245] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely exemplary forms of implementing the claims.

Claims

1. An image processing method, comprising: Display a first interface, the first interface including a first control; In response to a touch of the first control, a first image and / or a second image are displayed, the first image including a first face and the second image including a second face; In response to the selection of the first image and / or the second image, the second face is fused into the first image, wherein the fused first image includes a third face that is different from the first face.

2. The method according to claim 1, wherein, The first control includes a first sub-control, a second sub-control, and a third sub-control. In response to touch on the first control, it displays a first image and / or a second image, including: In response to a touch on the first sub-control or the second sub-control, the first image and the second image are displayed; In response to a touch of the third sub-control, the second image is displayed.

3. The method according to claim 2, wherein, In response to a touch on the first sub-control or the second sub-control, displaying the first image and the second image includes: In response to a touch of the first sub-control, the first image and the second image are displayed, the first image including a first face; In response to a touch of the second sub-control, the first image and the second image are displayed, the first image including a plurality of first faces.

4. The method according to claim 2 or 3, wherein, The first interface includes a first area and / or a second area, wherein the first image is displayed in the first area and the second image is displayed in the second area.

5. The method according to claim 4, wherein, After displaying the first image and / or the second image, the method further includes: In response to a touch associated with the first area, a third image from a user-authorized photo album is displayed, the third image including one or more first faces; In response to a touch of the third image, the touched third image is displayed in the first area.

6. The method according to claim 4 or 5, wherein, After displaying the first image and / or the second image, the method further includes: In response to a touch associated with the second area, a fourth image in the album is displayed, the fourth image including one or more second faces; In response to a touch of the fourth image, the touched fourth image is displayed in the second area.

7. The method according to any one of claims 1-6, wherein, The first and second images also include markers indicating that the images have been touched; The response to the selection of the first image and / or the second image includes: In response to a touch on the first image, the marker of the touched first image is highlighted; In response to a touch on a second image, when the number of touched second images is equal to the number of first faces in the touched first image, the markers of other second images are grayed out.

8. The method according to any one of claims 1-7, wherein, The first interface for display includes: Display a second interface, the second interface including a first control and a first slot, the first slot including the second image; In response to a touch of the first control, the first interface is displayed.

9. The method according to claim 8, wherein, The first interface also includes a second control, and the method further includes: In response to a touch of the second control, the second interface is redisplayed; The second interface displays a second slot, which includes the first image after face fusion. A first template is generated, which is used to generate an image after facial fusion.

10. The method according to any one of claims 1-9, wherein, The first interface is displayed, including: Display a third interface, which is the interface captured by the camera, and the third interface includes a third control; In response to a touch of the third control, the first interface is displayed.

11. The method according to claim 10, wherein, The step of displaying the first image in response to touch of the first control includes: In response to a touch on a first child control or a third child control, a first image including a first face is displayed; In response to a touch on the second sub-control, a first image including multiple first faces is displayed.

12. The method according to claim 11, wherein, The first interface also includes a fourth control, and the method further includes: In response to a touch of the fourth control, a fifth image is displayed, the fifth image including a first image of a first face, or multiple first images of first faces; In response to a touch of the fifth image, the first image displayed on the first interface is replaced with the touched fifth image.

13. The method according to any one of claims 10-12, further comprising: Generate a first special effect, which is used to generate the image after facial fusion.

14. An image processing method, comprising: Display a first interface, which includes a camera-captured image and a first control; In response to a touch of the first control, a first effect is displayed, the first effect including a first image, the first image including a first face; In response to a touch of the first special effect, a first multimedia is generated based on the first special effect and a second image captured by the camera. The first multimedia includes an image in which a second face from the second image is fused to a first face from the first image.

15. The method according to claim 14, wherein, The first interface also includes a second control, and the method further includes: In response to a touch of the second control, a first template is displayed, the first template including a third image, the third image including a first face; In response to a touch on the first template, a fourth image is displayed, the fourth image including a second face; In response to a touch on the fourth image, a second multimedia is generated, the second multimedia comprising an image in which a second face of the fourth image is fused to a first face of the third image.

16. An image processing apparatus, comprising a display module and a fusion module, wherein: The display module is configured to display a first interface, the first interface including a first control; The display module is further configured to display a first image and / or a second image in response to a touch of the first control, the first image including a first face and the second image including a second face; The fusion module is configured to, in response to the selection of the first image and / or the second image, fuse the second face into the first image, wherein the fused first image includes a third face that is different from the first face.

17. An image processing apparatus, comprising a display module and a generation module, wherein: The display module is configured to display a first interface, the first interface including a camera-captured image and a first control; The display module is further configured to display a first effect in response to touch of the first control, the first effect including a first image, the first image including a first face; The generation module is configured to, in response to a touch of the first special effect, generate a first multimedia based on the first special effect and a second image captured by a camera, the first multimedia including an image after fusing a second face from the second image to a first face from the first image.

18. A terminal device, comprising a processor and a memory, wherein, The memory stores computer-executable instructions; When the processor executes the computer-executable instructions stored in the memory, it implements the image processing method as claimed in any one of claims 1-13, or the image processing method as claimed in claim 14 or 15.

19. A computer-readable storage medium storing computer-executable instructions, wherein, When the processor executes the computer-executable instructions, it implements the image processing method as claimed in any one of claims 1-13, or the image processing method as claimed in claim 14 or 15.

20. A computer program product comprising a computer program, wherein, When the computer program is executed by the processor, it implements the image processing method as claimed in any one of claims 1-13, or the image processing method as claimed in claim 14 or 15.