Image processing method and device, electronic equipment, storage medium and program product
By resizing and stylizing the target object, the problem of unstable stylized effects was solved, thus achieving greater flexibility in image processing and richer display effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2026-03-10
AI Technical Summary
The stylized effects processing in existing technologies is unstable, resulting in significant differences in the effects when processing different images, which affects the user experience.
By acquiring the first image of the target object and redrawing it in the target style, the size of the target object is transformed to generate a second image, thus enhancing the flexibility and stability of special effects processing.
When processing different images under the same target style, it ensures that the generated images can all present the size transformation effect of the target object, enriching the special effects processing methods and improving the display effect of special effects images.
Smart Images

Figure CN121639445A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to computer application technology, and more particularly to an image processing method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] In image processing and video production, special effects are highly favored by users, such as stylized effects. Users can apply selected stylized effects to images or videos to create images that exhibit effects corresponding to the stylized effect.
[0003] In related technologies, when using stylized effects to process images, the resulting special effects images exhibit a certain degree of randomness. Therefore, when the same stylized effect is applied to different images, the resulting special effects images will show significant differences, making the stylized effect processing results unstable and affecting the user experience. Summary of the Invention
[0004] This disclosure provides an image processing method, apparatus, electronic device, storage medium, and program product to achieve the effect of performing targeted local stylization redrawing of target objects in an image while simultaneously resizing the target objects, thereby enhancing the flexibility of the special effects processing.
[0005] In a first aspect, embodiments of this disclosure provide an image processing method, the method comprising:
[0006] In response to an image setting operation, a first image is acquired, wherein the first image includes a target object and the target object is presented at a first size;
[0007] In response to an image processing request, the target object in the first image is redrawn in the target style to obtain a second image, and the second image is displayed; wherein at least a portion of the target object in the second image is presented at a second size and the target object presents the target style.
[0008] Secondly, embodiments of this disclosure also provide an image processing apparatus, the apparatus comprising:
[0009] An image acquisition module is configured to acquire a first image in response to an image setting operation, wherein the first image includes a target object and the target object is presented at a first size;
[0010] An image processing module is configured to, in response to an image processing request, redraw the target object in the first image in a target style to obtain a second image, and display the second image; wherein at least a portion of the target object in the second image is presented at a second size and the target object presents the target style.
[0011] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0012] One or more processors;
[0013] Storage device for storing one or more programs.
[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in any of the embodiments of this disclosure.
[0015] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image processing method as described in any of the embodiments of this disclosure.
[0016] Fifthly, this disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the image processing method described in any embodiment of the present invention.
[0017] The technical solution of this disclosure, in response to an image setting operation, acquires a first image, wherein the first image includes a target object and the target object is presented at a first size. This achieves the effect of supporting multiple methods to acquire the original image to be processed, enabling users to customize the original image settings, and enhancing the flexibility of image acquisition methods. Furthermore, in response to an image processing request, the target object in the first image is redrawn according to a target style to obtain a second image, which is then displayed. In the second image, at least a portion of the target object is presented at a second size and the target object exhibits the target style. This solves the problem of unstable special effects processing in related technologies, ensuring that when different first images are processed based on the same target style, the generated second images corresponding to the first images can all present the specific effect of target object size transformation. Furthermore, it achieves the effect of performing targeted local stylization redrawing of the target object in the image while simultaneously transforming the target object's size, enhancing the flexibility of the special effects processing process, enriching the special effects processing methods, and enriching the special effects display effects of the special effects images. Attached Figure Description
[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0019] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure;
[0020] Figure 2 This is a schematic flowchart of another image processing method provided in an embodiment of the present disclosure;
[0021] Figure 3 This is a schematic flowchart illustrating a method for generating a second mask image using an image processing technique according to an embodiment of this disclosure.
[0022] Figure 4 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure;
[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0030] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0031] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0032] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0033] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0034] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0035] Figure 1This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure. This embodiment is applicable to situations where a target object in an image to be processed undergoes local stylization processing. The method can be executed by an image processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 1 As shown, the method in this embodiment may specifically include:
[0036] S110. In response to an image setting operation, a first image is acquired, wherein the first image includes a target object and the target object is presented at a first size.
[0037] The image setting operation can be understood as an operation used to set the image to be processed with special effects. Optionally, the image setting operation may include at least one of the following: triggering the image setting control; receiving audio information including a trigger setting wake-up word; receiving an image setting instruction; the current limb movement being consistent with a preset setting trigger limb movement. The first image can be understood as the image to be processed. The first image may include a target object, and the target object is presented at a first size. The target object may be any object included in the first image. Optionally, the target object may be a person, animal, plant, or building, etc.; or, the target object may also be any limb part, such as a person's face or limbs, etc. The first size may be any size. Optionally, the first size may be the original size of the target object in the first image. For example, the first image may include images of people and buildings, and a person may be used as the target object, with the original size of the person in the first image used as the first size. It should be noted that the number of first images may be one or more, and regardless of whether it is one or more, the technical solutions provided in the embodiments of this disclosure can be used to process the first image.
[0038] In the embodiments of this disclosure, the image setting operation may include various implementation methods, such as image capture operation or image upload operation. These two implementation methods will be described below.
[0039] Optionally, in response to an image setting operation, acquiring a first image includes: in response to an image capturing operation, capturing an image using a target capturing device to obtain the first image.
[0040] The image capture operation can be understood as the operation of acquiring an image to be processed through image capture. Optionally, the image capture operation may include at least one of the following: triggering an image capture control; receiving audio information including a trigger wake-up word; receiving an image capture command; and the current limb movement being consistent with a preset capture trigger limb movement. The target capture device may be a capture device installed on a terminal device, such as a smartphone camera or tablet camera; or it may be a capture device externally connected to the terminal device, such as a camera.
[0041] As an optional implementation of this embodiment, a shooting start control can be pre-set in the display interface. Upon detecting a trigger operation on the shooting start control, the shooting interface can be entered. At this time, the image displayed in the shooting interface is the image within the field of view of the target shooting device. The content of the image displayed in the shooting interface can be updated by moving the target shooting device. Furthermore, upon detecting a user triggering a shooting operation (e.g., triggering the shooting control or clicking the terminal device screen), the target shooting device can capture the image displayed in the current shooting interface, and the captured image can be used as the first image.
[0042] Optionally, in response to an image setting operation, obtaining a first image includes: in response to an image upload operation, obtaining the uploaded image as the first image.
[0043] The image upload operation can be understood as the operation of obtaining an image to be processed by uploading an image and / or video. Optionally, the image upload operation may include at least one of the following: triggering an upload control; receiving an image upload instruction; including a wake-up word associated with the image upload operation in the audio information, etc. Optionally, the first image may be an image pre-stored in the target storage space (e.g., the application software's image library, or the terminal's photo album, etc.); or it may be an image received from an external device, etc.
[0044] As another optional implementation of this embodiment, an image upload control can be pre-set in the display interface. Upon detecting a trigger operation on this image upload control, an image upload path selection page can pop up. This selection page can include at least three image upload paths, such as the local terminal album, image library, and third-party uploads. Furthermore, upon detecting a trigger operation for selecting any image upload path, an image display page corresponding to that image upload path can be displayed in the display interface. This image display page can include multiple images to be uploaded associated with that image upload path, and a selection can be made from these multiple images. Furthermore, upon detecting a trigger operation for selecting at least one image to be uploaded, the selected at least one image can be used as the uploaded image. Furthermore, upon detecting a selection completion operation (e.g., triggering a selection confirmation control or clicking a blank area on the terminal device screen), it can be determined that an image upload operation has been detected. Then, in response to this image upload operation, the uploaded image is used as the first image.
[0045] In this embodiment of the disclosure, the target object included in the first image can be determined in a variety of ways, such as user-defined settings or automatic recognition.
[0046] Optionally, if the target object is determined through user-defined settings, after acquiring the first image, the method further includes: displaying the first image; and determining the target object in the first image based on the object setting operation in response to the object setting operation for the first image.
[0047] The object setting operation can be understood as an operation that sets an object included in an image as a target object. In the embodiments of this disclosure, the object setting operation can be any operation capable of setting any object as a target object. Optionally, the object setting operation includes object selection operations and / or text input operations, etc.
[0048] As an optional implementation of this disclosure, upon obtaining the first image, it can be displayed on a display interface. The displayed first image may include at least one selectable object, and each selectable object is in a selectable state. Further, an object selection operation can be input for any selectable object, and in response to the object selection operation, the selected selectable object is determined. Then, the selected selectable object can be used as the target object in the first image. The object selection operation can be a click operation on any selectable object; or, the object selection operation can be an object bounding box drawing operation on any selectable object; or, the object selection operation can be any other operation capable of selecting the selectable object.
[0049] As another optional implementation of this disclosure, when the first image is displayed on the display interface, an object input box can be displayed based on the display interface. Furthermore, an object editing operation can be performed on the object input box, and the edited object can be displayed in the object input box. Then, upon detecting that the editing operation is complete, the object displayed in the object input box can be used as the target object.
[0050] It should be noted that the advantages of allowing users to customize target objects are that it enhances the flexibility of the object setting process, improves the intelligence of the application software in object setting functions, and enhances the interactivity between users and the application software.
[0051] Optionally, if the target object is determined by automatic identification, after acquiring the first image, the method further includes: performing object detection on the first image according to a preset object detection algorithm, and taking the detected object as the target object in the first image; wherein the preset object detection algorithm includes at least one object feature of a preset object.
[0052] Among them, object detection algorithms can be understood as algorithms that detect objects included in an image in order to determine the target object.
[0053] As an optional implementation of this disclosure, upon obtaining the first image, the first image can be processed according to a preset object detection algorithm to detect objects included in the first image. Furthermore, if a preset object is detected among the objects included in the first image based on object features of at least one preset object, the preset object present in the first image can be used as the target object.
[0054] S120. In response to the image processing request, the target object in the first image is redrawn in the target style to obtain the second image, and the second image is displayed.
[0055] The image processing request can be understood as an instruction to process an image. Image processing requests can be generated in various ways. Optionally, an image processing request can be generated when a user trigger operation on a preset image processing control is detected; or, an image processing request can be generated when the received audio information includes a trigger word associated with image processing; or, an image processing request can be generated when the received instruction includes an image processing instruction, etc. The target style can be understood as an effect that can change the style of an image; that is, after processing the image based on the target style, the resulting image can present an image style different from the original image style but consistent with the target style. The target style can be any stylized effect capable of changing the image style. Optionally, the target style can be a comic book style, ink painting style, cyberpunk style, or sketch style, etc. The second image can be an image where the target object is redrawn based on the target style, and other image areas besides the target object retain their original image style. At least a portion of the target object in the second image is presented at a second size, and the target object presents the target style. At least a portion of the region can be the entire image region corresponding to the target object; or, it can be a portion of the image region corresponding to the target object. The second size can be the size determined after performing a size transformation on at least a portion of the target object. It should be noted that the second size can be different from the first size; optionally, the second size can be smaller than the first size or larger than the first size. For example, assuming that at least a portion of the target object is reduced in size, and thus the size of at least a portion of the target object in the obtained second image is smaller than the first size, it can be determined that the second size is smaller than the first size; or, assuming that at least a portion of the target object is enlarged in size, and thus the size of at least a portion of the target object in the obtained second image is larger than the first size, it can be determined that the second size is larger than the first size. Of course, the second size can also be the same as the first size, and this disclosure does not specifically limit this.
[0056] In this embodiment of the disclosure, upon receiving an image processing request, in response to the image processing request, at least a portion of the target object is subjected to size transformation processing, and the size-transformed target object is redrawn in the target style to obtain a redrawn image. The redrawn image can be used as a second image, in which at least a portion of the target object is presented at a second size and the target object presents the target style.
[0057] It should be noted that the first image includes not only the target object but also other image regions. To accurately resize the target object and redraw it according to the target style, the target object can first be separated from the first image to obtain a mask image corresponding to the target object. Then, at least a portion of the mask image corresponding to the target object can be resized to obtain a transformed mask image. Further, the target object can be redrawn based on the first image, the transformed mask image, and the target style to obtain a target foreground image corresponding to the target object. A target background image is then determined based on the mask image corresponding to the target object and the first image. This target background image can be an image that retains all image regions in the first image except for the target object. Finally, the target foreground image and the target background image can be fused, and the fused image can be used as the second image.
[0058] The technical solution of this disclosure, in response to an image setting operation, acquires a first image, wherein the first image includes a target object and the target object is presented at a first size. This achieves the effect of supporting multiple methods to acquire the original image to be processed, enabling users to customize the original image settings, and enhancing the flexibility of image acquisition methods. Furthermore, in response to an image processing request, the target object in the first image is redrawn according to a target style to obtain a second image, which is then displayed. In the second image, at least a portion of the target object is presented at a second size and the target object exhibits the target style. This solves the problem of unstable special effects processing in related technologies, ensuring that when different first images are processed based on the same target style, the generated second images corresponding to the first images can all present the specific effect of target object size transformation. Furthermore, it achieves the effect of performing targeted local stylization redrawing of the target object in the image while simultaneously transforming the target object's size, enhancing the flexibility of the special effects processing process, enriching the special effects processing methods, and enriching the special effects display effects of the special effects images.
[0059] Figure 2This is a schematic flowchart illustrating another image processing method provided in this embodiment. Based on the above embodiments, the technical solution of this embodiment processes a first image to obtain at least one first mask image corresponding to a target object; performs size transformation processing on the target region corresponding to the target object in the at least one first mask image to obtain a second mask image; redraws the target object based on the first image and the second mask image to obtain a target foreground image; and determines a target background image corresponding to the target object based on the first mask image; performs image fusion between the target foreground image and the target background image to obtain a second image corresponding to the first image, and displays the second image. For detailed implementation, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 2 As shown, the method in this embodiment may specifically include:
[0060] S210. In response to an image setting operation, a first image is acquired, wherein the first image includes a target object and the target object is presented at a first size.
[0061] S220. In response to the image processing request, the first image is processed to obtain at least one first mask image corresponding to the target object, and the target region corresponding to the target object in the at least one first mask image is subjected to size transformation processing to obtain a second mask image.
[0062] The first mask image can be an image representing the contour of a region of the first object. Generally, a mask image can be a binary image composed of two different pixel values. It can be understood that when determining the first mask image corresponding to the target object, the pixel values of the image regions where the target object is located in the first image can be adjusted to a first preset pixel value, and the pixel values of the image regions other than the target object in the first image can be adjusted to a second preset pixel value, thus obtaining the first mask image corresponding to the target object. It should be noted that the first preset pixel value and the second preset pixel value are two different pixel values so that the target object can be distinguished from the image. The second mask image can be a mask image obtained by resizing the target region of the target object in the first mask image.
[0063] It should be noted that the number of first mask images can be one or more. When there is only one first mask image, the first mask image corresponding to the target object can be a mask image representing the overall region contour of the first object. When there are multiple first mask images, the first mask image corresponding to the target object can include a mask image representing the region contour of the target area of the first object and mask images representing the region contours of other areas of the first object besides the target area.
[0064] As an optional implementation of this disclosure, upon receiving an image processing request, in response to the request, the image region containing the target object in the first image is determined. The pixel values of the pixels in this image region are then adjusted to different pixel values from the pixel values of pixels in other image regions of the first image, thereby distinguishing the target object from the image. Furthermore, the image obtained after pixel value adjustment can be used as a first mask image corresponding to the target object.
[0065] As another optional implementation of this disclosure, upon receiving an image processing request, the target region in the target object to be resized can be determined based on the image processing request. Then, in response to the image processing request, the image region containing the target region corresponding to the target object in the first image is determined, and the pixel values of the pixels in this image region are adjusted to different pixel values from the pixel values of the pixels in other image regions in the first image, thereby distinguishing the target region corresponding to the target object from the image. Furthermore, the image obtained after adjusting the pixel values can be used as the first mask image corresponding to the target object. Further, the first image can be masked again, and other image regions in the image region containing the target object in the first image, excluding the target region, can be determined. The pixel values of the pixels in these other image regions are adjusted to different pixel values from the pixel values of the pixels in the image regions in the first image, excluding the target region, thereby distinguishing the other regions of the target object from the image. Furthermore, the image obtained after adjusting the pixel values can also be used as the first mask image corresponding to the target object.
[0066] In this embodiment of the disclosure, given at least one first mask image, a mask image to be processed can be determined from the at least one first mask image, and a target region corresponding to the target object in the mask image to be processed can be determined. Further, at least a portion of the target region can be resized to obtain a second mask image.
[0067] It should be noted that the method for determining the corresponding second mask image when there is only one first mask image is different from the method for determining the corresponding second mask image when there are multiple first mask images. These two methods will be explained below.
[0068] Optionally, a second mask image is obtained by performing a size transformation on the target region corresponding to the target object in at least one first mask image, including: if there is only one first mask image, performing image segmentation on the first mask image to obtain a first number of local mask images; obtaining a second number of mask images to be processed from the local mask images; performing a size transformation on the target region corresponding to the target object in the mask images to be processed to obtain a target mask image; and stitching the target mask image and the local mask images other than the mask images to be processed to obtain the second mask image.
[0069] In this embodiment of the disclosure, when the number of first mask images is one, it can be indicated that the first mask image can be a mask image representing the overall region contour of the target object. Correspondingly, the local mask image can be a mask image representing the local region contour of the target object. The first number can be any number greater than 1. Optionally, the first number can be 2, 3, or 4, etc. The second number can be any number less than the first number. That is, the second number can be determined based on the first number. Optionally, when the first number is 2, the second number can be 1; when the first number is 3, the second number can be 1 or 2; when the first number is 4, the second number can be 1, 2, or 3. The mask image to be processed can be a mask image to be subjected to size transformation processing; or, the mask image to be processed can also be a mask image including the target region corresponding to the target object. The target mask image can be a mask image obtained after size transformation of the target region of the target object included in the mask image to be processed. Alternatively, the target mask image can also be a mask image that meets the size transformation requirements for the target region of the target object.
[0070] As an optional implementation of this disclosure, when there is only one first mask image, the first mask image can be segmented according to the target region to be resized, dividing the first mask image into a mask image including the target region and a mask image excluding the target region, and the segmented mask image is used as a local mask image. Further, local mask images including the target region can be selected from the obtained local mask images, and the selected local mask images are used as the mask images to be processed. Further, the target region corresponding to the target object in the mask image to be processed can be resized, and the mask image obtained after resizing is used as the target mask image. Further, the target mask image and the local mask images other than the mask image to be processed can be stitched together, and the stitched image is used as the second mask image.
[0071] For example, suppose the target object is a person, and the target area to be resized is the person's face. If there is only one first mask image, the resulting first mask image can be a mask image representing the overall outline of the person. Further, by performing image segmentation on the first mask image based on the target area, the first mask image can be segmented into a local mask image that only includes the person's face mask and a local mask image that includes masks of other parts besides the person's face; that is, two local mask images are obtained. Further, the local mask image including the person's face mask can be selected from the two local mask images as the mask image to be processed; that is, one mask image to be processed is obtained. Further, the person's face in the mask image to be processed can be enlarged, and the enlarged mask image can be used as the target mask image. Further, the target mask image can be stitched together with the local mask image including masks of other parts besides the person's face to obtain the second mask image.
[0072] As another optional implementation of this disclosure, when the number of first mask images is one, the first mask image can be segmented according to a preset default segmentation rule to obtain a first number of local mask images. Further, local mask images containing the target region mask can be selected from the local mask images, and the selected local mask images are used as mask images to be processed. Further, the target region corresponding to the target object in the mask image to be processed is resized, and the resulting mask image is used as the target mask image. Further, the target mask image and the local mask images other than the mask image to be processed can be stitched together, and the stitched image is used as the second mask image. The default segmentation rule may include random segmentation and dividing the image equally into a preset number of local images, etc.
[0073] Optionally, the target region corresponding to the target object in at least one first mask image is subjected to size transformation processing to obtain a second mask image, including: when there are multiple first mask images, the target regions corresponding to the target object in a portion of the multiple first mask images are subjected to size transformation processing to obtain a second mask image.
[0074] In this embodiment of the disclosure, when there are multiple first mask images, each of the multiple first mask images may be an image in which each first mask image includes a region mask of the target region; or, the multiple first mask images may also include an image in which the target region is included and an image in which the target region is not included. Correspondingly, the "partial number" may be the total number of first mask images, or it may be at least one of the first mask images.
[0075] As an optional implementation of this embodiment, when there are multiple first mask images, a first mask image including the target region can be selected from the multiple first mask images, and the selected first mask image is used as a portion of the first mask images. Further, the target region corresponding to the target object in the portion of the first mask images can be resized, and the resulting mask image is used as the second mask image. The advantage of this approach is that it simplifies the process of determining the second mask image, thereby improving the efficiency of determining the second mask image.
[0076] It should be noted that when the number of first mask images is one or more, the advantage of determining the second mask image according to the corresponding method is that it enriches the methods for determining the second mask image and enhances the flexibility of determining the second mask image, so as to select the appropriate determination method in different application scenarios.
[0077] For example, Figure 3 This is a schematic flowchart illustrating the process of generating a second mask image using an image processing method according to an embodiment of this disclosure. Assuming... Figure 3 Image a in the diagram is the first image, and the person included in this image is the target object. Furthermore, the target area within the target object to be resized is the person's head region. We can first perform a masking process on the person's head region in the first image to obtain a mask image corresponding to the person's head region, such as... Figure 3 As shown in Figure b. Further, we can... Figure 3 The head region of the person in image b is enlarged to obtain an enlarged mask image, as shown below. Figure 3 As shown in Figure c. Further, a masking process can be performed on the person in the first image to obtain a mask image corresponding to the person, such as... Figure 3 As shown in Figure d. Then, the enlarged mask image can be merged into the mask image corresponding to the person to obtain the second mask image, as shown in Figure d. Figure 3 As shown in Figure e. It should be noted that the advantage of determining the second mask image based on the above method is that, when the boundary between the target area and the non-target area is not obvious, by performing local masking and size transformation on the target area, the target area can be accurately identified, and the size transformation of the target area can be accurately performed so that the target area after size transformation can present a clearer size transformation effect.
[0078] S230. The target object is redrawn based on the first image and the second mask image to obtain the target foreground image, and the target background image corresponding to the target object is determined based on the first mask image.
[0079] The target foreground image can consist only of the pixel region of the target object, where the target object is redrawn and its corresponding target region is resized. The target background image can be an image that includes pixel regions other than the target object. Generally, when the first image includes the target object, the pixel region of the target object can be used as the foreground image, and the pixel regions other than the target object can be used as the background image.
[0080] In this embodiment of the disclosure, in order to obtain the pixel information of the target object after obtaining the second mask image, the target object can be redrawn based on the first image and the second mask image to obtain the target foreground image.
[0081] In practical applications, when redrawing a target object in an image, there may be a situation where the target object in the redrawn image has a low degree of matching with the target object in the original image (e.g., mismatch in action posture, mismatch in facial expression, etc.), which in turn affects the display effect of the image.
[0082] In response to the above situation, in this embodiment of the disclosure, the target object is redrawn based on the first image and the second mask image to obtain the target foreground image, including: redrawing the target object based on the first image, the second mask image and the target pose information of the target object in the first image to obtain the target foreground image.
[0083] The target pose information can be used to characterize the position and posture of the target object in the image. Optionally, the target pose information may include information such as the target object's spatial location, orientation, action, and / or facial expression.
[0084] In this embodiment of the disclosure, the method of obtaining the target pose information of the target object in the first image may include a variety of methods, such as determining the target pose information based on a key point detection algorithm; or determining the target pose information based on a pre-trained pose detection model, etc.
[0085] As an optional implementation of this embodiment, the first image can be processed according to a key point detection algorithm to obtain the pose key point information of the target object in the first image, and this pose key point information can be used as the target pose information. Further, the target object can be redrawn based on the first image, the second mask image, and the target pose information to obtain the target foreground image. The advantage of this setup is that it improves the correlation between the target foreground image and the first image, achieving effective control over the generation process of the target foreground image based on the target pose information, thereby improving the image display effect of the target foreground image.
[0086] It should be noted that there are at least two ways to determine the target foreground image, and these two methods will be explained below.
[0087] One approach is to redraw the target object based on the first image, the second mask image, and the target pose information of the target object in the first image to obtain a target foreground image, including: determining the target pose information of the target object in the first image, and obtaining a first model corresponding to the target style; inputting the first image, the second mask image, and the target pose information into the first model to redraw the target object to obtain a target foreground image.
[0088] The first model can be understood as a neural network model that stylizes the target object in the image based on the image, the corresponding mask image, and the object's pose information. The first model can be a neural network model with any model structure. It should be noted that the first model corresponding to the target style can be a neural network model that can only render the target object in the image with a stylized effect consistent with the target style; or, the first model can also be a neural network model that can render the target object in the image with stylized effects corresponding to multiple styles, including the target style. Whether the first model includes one target style or multiple styles depends on the expected effect image used when training the first model. If the training samples used to train the first model only include expected effect images corresponding to the target style, the trained first model is a neural network model that can only render the target object in the image with a stylized effect consistent with the target style; if the training samples used to train the first model include expected effect images corresponding to multiple styles, and the multiple styles include the target style, the trained first model can also be a neural network model that can render the target object in the image with stylized effects corresponding to multiple styles, including the target style.
[0089] In this embodiment, the first model is trained using sample images, corresponding sample mask images, and sample pose information. The sample images can be images captured by a camera; or images reconstructed by an image reconstruction model; or images pre-stored in storage space. The sample images may include a target object. The sample mask image is an image obtained by masking the sample images. The sample pose information can be information characterizing the pose of the target object in the sample images.
[0090] It should be noted that before applying the first model provided in the embodiments of this disclosure, a pre-built deep learning model can be trained first. Before training the model, multiple training samples can be constructed to train the model based on the training samples. To improve the accuracy of the first model, as many and rich training samples as possible can be constructed. Optionally, the training process of the first model can be as follows: acquiring multiple training samples; wherein, the training samples include sample images, sample mask images corresponding to the sample images, sample pose information, and expected effect images corresponding to the target style; for each training sample, inputting the sample image, sample mask image, and sample pose information in the training sample into the pre-built deep learning model, and obtaining the actual output image; determining the loss value based on the actual output image and the expected effect image in the training samples; correcting the model parameters in the deep learning model based on the loss value, and taking the convergence of the loss function in the deep learning model as the training objective, so as to use the trained deep learning model as the first model corresponding to the target style.
[0091] As an optional implementation of this embodiment, multiple styles to be selected can be predetermined, and a first model corresponding to each style can be trained. Further, a style identifier corresponding to each style to be selected can be obtained, and the style identifier can be associated with and stored in the first model. Further, when a target style is determined, the first model matching the target style can be retrieved based on the style identifier of the target style, and the target pose information of the target object in the first image can be determined based on a preset keypoint detection algorithm. Further, the first image, the second mask image, and the target pose information can be input into the first model corresponding to the target style to redraw the target object based on the first model, and the image output by the model can be used as the target foreground image. The advantage of this setup is that determining the target foreground image based on the first model improves the accuracy and efficiency of target foreground image generation.
[0092] Another approach is to redraw the target object based on the first image, the second mask image, and the target pose information of the target object in the first image to obtain a target foreground image, including: determining target cue information; inputting the target cue information, the first image, and the second mask image into the second model to redraw the target object to obtain a target foreground image.
[0093] The target cue information can be used to guide the model to perform specific tasks or generate corresponding outputs. Target cue information helps the model more accurately understand the processing intent and generate results that better meet expectations. Target cue information includes pose cue information and style cue information. Pose cue information can include the pose information of the target object expected to be presented in the output image. Pose cue information is used to indicate the target pose information of the target object in the second image. In other words, pose cue information can be used to guide the model to output a second image that presents the target pose information. Style cue information can include the style information of the target object expected to be presented in the output image. Style cue information is used to indicate the target style of the target object in the second image. Style cue information can be used to guide the model to output a second image that presents the target style. The second model can be a neural network model that stylizes the target object in the image based on the target cue information, the image, and the corresponding mask image. It should be noted that the second model can be a neural network model with any model structure. For example, the second model can be a diffusion model. In this embodiment, the second model is trained on the diffusion model using sample images, corresponding sample mask images, and sample cue information. The diffusion model includes a pose control module. The pose control module (ControlNet) can be understood as a neural network model that controls the model to generate a new image based on the pose cues in the target prompt information.
[0094] It should be noted that before applying the second model provided in the embodiments of this disclosure, the diffusion model can be trained first. Before training the model, multiple training samples can be constructed to train the model based on the training samples. To improve the accuracy of the second model, as many and rich training samples as possible can be constructed. Optionally, the training process of the second model can be as follows: acquiring multiple training samples; wherein, the training samples include sample images, sample mask images corresponding to the sample images, sample prompt information, and expected effect images corresponding to the target style; for each training sample, inputting the sample image, sample mask image, and sample prompt information in the training sample into the pre-constructed diffusion model, and obtaining the actual output image; determining the loss value based on the actual output image and the expected effect image in the training samples; correcting the model parameters in the diffusion model based on the loss value, and taking the convergence of the loss function in diffusion as the training objective, so as to use the trained diffusion model as the second model.
[0095] As an optional implementation of this embodiment, the target pose information of the target object in the second image can be determined based on the target pose information of the target object in the first image, and pose cue information can be determined based on the target pose information. Furthermore, style cue information can be determined based on the determined target style. Further, target cue information can be constructed based on the pose cue information and style cue information. Then, the target cue information, the first image, and the second mask image can be input into the second model to redraw the target object in the first image based on the target cue information, the second mask image, and the first image, and the image output by the model is used as the target foreground image. The advantage of this setup is that generating the target foreground image based on the target cue information and the second model improves the generation quality and efficiency of the target foreground image. Consequently, the final target foreground image can meet the user's stylization processing requirements.
[0096] In practical applications, when at least a portion of a target object is resized, the second size corresponding to that portion of the target object after the size transformation is smaller than its corresponding first size, and the pixel area corresponding to the target object after the size transformation is smaller than the pixel area of the target object in the first image. This may result in a lack of pixel information in the target background image determined based on the first mask image corresponding to the first image.
[0097] To address the above situation, in this embodiment of the disclosure, determining the target background image corresponding to the target object based on the first mask image includes: determining an initial background image corresponding to the target object based on the first mask image, and performing image restoration on the initial background image to obtain the target background image. The initial background image may be an image including pixel regions other than the target object in the first image.
[0098] As an optional implementation of this embodiment, the target object in the first image can be segmented based on the first mask image to obtain an initial background image corresponding to the target object. Further, the initial background image can be repaired according to a preset image repair method, and the repaired image can be used as the target background image. The image repair method can be any method capable of image repair. Optionally, the image repair method can be based on a large mask repair model. The advantage of this setting is that it enables image repair of the background image even when the size is reduced or the display position changes, making the repaired background image fit the corresponding foreground image better, thereby improving the special effects display effect of the final generated second image.
[0099] It should be noted that, in the embodiments of this disclosure, the background image restoration process is particularly applicable when the second size corresponding to at least a portion of the target object after size transformation is smaller than its corresponding first size. Also applicable when the displayed position of the processed target object in the second image is different from its position in the first image.
[0100] In practical applications, since the second mask image is obtained by resizing the first mask image, and the first mask image is obtained by masking the first image, there may be inaccurate mask segmentation when masking the first image. Consequently, when redrawing the target object based on the second mask image and the first image, there may be unclear edge contours of the target object in the redrawn image.
[0101] In view of the above situation, in this embodiment of the present disclosure, before redrawing the target object according to the first image and the second mask image, the method further includes: determining the edge information corresponding to the target object in the second mask image, adjusting the transparency of at least some pixels in the edge information, and updating the second mask image according to the adjusted transparency.
[0102] Edge information can be information used to characterize the edge contours of the target object. Edge information can include pixel information corresponding to multiple edge pixels. At least some pixels can be all edge pixels or only some edge pixels. It is understood that transparency is usually represented by an alpha channel. With an alpha channel value ranging from 0 to 255, 0 represents complete transparency, and 255 represents complete opacity.
[0103] As an optional implementation of this embodiment, after obtaining the second mask image, the second mask image can be processed according to a preset edge detection algorithm to obtain edge information corresponding to the target object. Furthermore, the transparency of at least some pixels in the edge information can be adjusted, and then the second mask image can be updated according to the adjusted transparency. The advantage of this setup is that it achieves the effect of repairing imperfections in the edge contour of the target object, and accurately redrawing the target object in the first image based on the repaired second mask image, thereby improving the special effects display effect of the second image.
[0104] S240. The target foreground image and the target background image are fused to obtain a second image corresponding to the first image, and the second image is displayed.
[0105] In this embodiment of the disclosure, when a target foreground image and a target background image are obtained, the target foreground image and the target background image can be image fused. Then, the image obtained after fusion can be used as a second image corresponding to the first image and the second image can be displayed.
[0106] The technical solution of this embodiment processes a first image to obtain at least one first mask image corresponding to the target object. It then performs a size transformation on the target region corresponding to the target object in the at least one first mask image to obtain a second mask image. Further, it redraws the target object based on the first and second mask images to obtain a target foreground image, and determines a target background image corresponding to the target object based on the first mask image. Finally, it fuses the target foreground image and the target background image to obtain a second image corresponding to the first image, and displays the second image. This achieves the effect of size transformation and local stylized redrawing of the target object based on its mask image, thereby achieving accurate segmentation and redrawing of the target object in the image, and ultimately improving the special effects display effect of the second image.
[0107] Figure 4 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure, as shown below. Figure 4 As shown, the device includes: an image acquisition module 310 and an image processing module 320. The image acquisition module 310 is configured to acquire a first image in response to an image setting operation, wherein the first image includes a target object and the target object is presented at a first size; the image processing module 320 is configured to redraw the target object in the first image according to a target style in response to an image processing request, to obtain a second image, and to display the second image; wherein at least a portion of the target object in the second image is presented at a second size and the target object presents the target style.
[0108] The technical solution of this disclosure embodiment, through the image acquisition module 310 responding to an image setting operation, acquires a first image, wherein the first image includes a target object and the target object is presented at a first size, thereby achieving the effect of supporting multiple methods to acquire the original image to be processed, realizing the effect of user customization of the original image, and enhancing the flexibility of the image acquisition method. Furthermore, through the image acquisition module 310 responding to an image processing request, the target object in the first image is redrawn according to the target style to obtain a second image, and the second image is displayed, wherein at least a portion of the target object in the second image is presented at a second size and the target object presents the target style, solving the problem of unstable special effects processing in related technologies, and ensuring that the generated second images corresponding to the first images can all present the target object size transformation effect when processing different first images based on the same target style. Moreover, it achieves the specific effect of performing targeted local stylization redrawing of the target object in the image while simultaneously performing size transformation on the target object, enhancing the flexibility of the special effects processing process, enriching the special effects processing methods, and enriching the special effects display effect of the special effects image.
[0109] Optionally, based on any of the above-mentioned optional technical solutions, the device further includes: an image display module and an object determination module. The image display module is used to display the first image after acquiring the first image; the object setting module is used to determine a target object in the first image according to an object setting operation for the first image.
[0110] Based on any of the above optional technical solutions, optionally, the image processing module 320 includes: a mask image determination submodule, an object redrawing submodule, and an image fusion submodule. The mask image determination submodule is used to process the first image to obtain at least one first mask image corresponding to the target object, and to perform size transformation processing on the target region corresponding to the target object in the at least one first mask image to obtain a second mask image; the object redrawing submodule is used to redraw the target object based on the first image and the second mask image to obtain a target foreground image, and to determine a target background image corresponding to the target object based on the first mask image; the image fusion submodule is used to perform image fusion of the target foreground image and the target background image to obtain a second image corresponding to the first image, and to display the second image.
[0111] Based on any of the above optional technical solutions, optionally, the mask image determination submodule includes: a local mask image determination unit, a region size transformation unit, and an image stitching unit. The local mask image determination unit is used to segment the first mask image to obtain a first number of local mask images when the number of first mask images is one; the region size transformation unit is used to obtain a second number of mask images to be processed from the local mask images, and to perform size transformation processing on the target region corresponding to the target object in the mask images to be processed to obtain a target mask image, wherein the second number is less than the first number; the image stitching unit is used to stitch the target mask image and the local mask images other than the mask images to be processed to obtain a second mask image.
[0112] Based on any of the above optional technical solutions, optionally, the mask image determination submodule is specifically used to perform size transformation processing on a portion of the first mask images corresponding to the target object when there are multiple first mask images, to obtain a second mask image.
[0113] Based on any of the above optional technical solutions, optionally, the object redrawing submodule is used to redraw the target object according to the first image, the second mask image, the target style, and the target pose information of the target object in the first image, so as to obtain the target foreground image.
[0114] Based on any of the above optional technical solutions, the optional object redrawing submodule includes: a pose information determination unit and an object redrawing unit. The pose information determination unit is used to determine the target pose information of the target object in the first image, and to obtain a first model corresponding to the target style, wherein the first model is obtained by training a deep learning model using sample images, sample mask images corresponding to the sample images, and sample pose information; the object redrawing unit is used to input the first image, the second mask image, and the target pose information into the first model to redraw the target object, thereby obtaining a target foreground image.
[0115] Based on any of the above optional technical solutions, the optional object redrawing submodule includes: a prompt information determination unit and a model processing unit. The prompt information determination unit is used to determine target prompt information, wherein the target prompt information includes pose prompt information and style prompt information. The pose prompt information is used to indicate the target pose information of the target object in the second image, and the style prompt information is used to indicate the target style of the target object in the second image. The model processing unit is used to input the target prompt information, the first image, and the second mask image into a second model to redraw the target object and obtain a target foreground image. The second model is obtained by training a diffusion model using sample images, sample mask images corresponding to the sample images, and sample prompt information. The diffusion model includes a pose control module.
[0116] Based on any of the above optional technical solutions, optionally, the object redrawing submodule is used to determine the initial background image corresponding to the target object according to the first mask image, and to perform image repair on the initial background image to obtain the target background image.
[0117] Optionally, based on any of the above-mentioned optional technical solutions, the device further includes: an edge information determination module. The edge information determination module is configured to, before redrawing the target object based on the first image and the second mask image, determine the edge information corresponding to the target object in the second mask image, adjust the transparency of at least a portion of the pixels in the edge information, and update the second mask image based on the adjusted transparency.
[0118] The image processing apparatus provided in this disclosure can execute the image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0119] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0120] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 5 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0121] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0122] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0123] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0124] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0125] The electronic device provided in this embodiment and the image processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0126] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the image processing method provided in the above embodiments.
[0127] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0128] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0129] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0130] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a first image in response to an image setting operation, wherein the first image includes a target object and the target object is presented at a first size; redraw the target object in the first image in a target style in response to an image processing request to obtain a second image, and display the second image; wherein at least a portion of the target object in the second image is presented at a second size and the target object presents a target style.
[0131] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0133] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the units are not necessarily limiting in certain circumstances; for example, an image acquisition module can also be described as a "module that responds to image setting operations".
[0134] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0135] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0136] According to one or more embodiments of this disclosure, [Example 1] provides an image processing method, including: in response to an image setting operation, acquiring a first image, wherein the first image includes a target object and the target object is presented at a first size; in response to an image processing request, redrawing the target object in the first image with a target style to obtain a second image, and displaying the second image; wherein at least a portion of the target object in the second image is presented at a second size and the target object presents a target style.
[0137] According to one or more embodiments of this disclosure, [Example 2] provides the method of Example 1, which further includes: optionally, after acquiring the first image, further including: displaying the first image; and in response to an object setting operation for the first image, determining a target object in the first image according to the object setting operation.
[0138] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 1, which further includes: Optionally, redrawing the target object in the first image with a target style to obtain a second image includes: processing the first image to obtain at least one first mask image corresponding to the target object; performing size transformation processing on the target region corresponding to the target object in the at least one first mask image to obtain a second mask image; redrawing the target object according to the first image and the second mask image to obtain a target foreground image, and determining a target background image corresponding to the target object according to the first mask image; performing image fusion of the target foreground image and the target background image to obtain a second image corresponding to the first image, and displaying the second image.
[0139] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 3, which further includes: Optionally, performing a size transformation process on the target region corresponding to the target object in at least one first mask image to obtain a second mask image includes: when the number of first mask images is one, performing image segmentation on the first mask image to obtain a first number of local mask images; obtaining a second number of mask images to be processed from the local mask images; performing a size transformation process on the target region corresponding to the target object in the mask images to be processed to obtain a target mask image, wherein the second number is less than the first number; and stitching the target mask image and the local mask images other than the mask images to be processed together to obtain the second mask image.
[0140] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 3, which further includes: Optionally, performing size transformation processing on the target region corresponding to the target object in at least one first mask image to obtain a second mask image includes: when there are multiple first mask images, performing size transformation processing on a portion of the first mask images corresponding to the target object to obtain a second mask image.
[0141] According to one or more embodiments of this disclosure, Example Six provides the method of Example Three, which further includes: Optionally, redrawing the target object based on the first image and the second mask image to obtain a target foreground image includes: redrawing the target object based on the first image, the second mask image, the target style, and the target pose information of the target object in the first image to obtain a target foreground image.
[0142] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 6, which further includes: optionally, redrawing the target object based on the first image, the second mask image, the target style, and the target pose information of the target object in the first image to obtain a target foreground image, including: determining the target pose information of the target object in the first image, and obtaining a first model corresponding to the target style, wherein the first model is obtained by training a deep learning model through sample images, sample mask images corresponding to the sample images, and sample pose information; inputting the first image, the second mask image, and the target pose information into the first model to redraw the target object to obtain a target foreground image.
[0143] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 6, which further includes: optionally, redrawing the target object based on the first image, the second mask image, the target style, and the target pose information of the target object in the first image to obtain a target foreground image, including: determining target cue information, wherein the target cue information includes pose cue information and style cue information, the pose cue information being used to indicate the target pose information of the target object in the second image, and the style cue information being used to indicate the target style of the target object in the second image; inputting the target cue information, the first image, and the second mask image into a second model to redraw the target object to obtain a target foreground image, wherein the second model is trained on a diffusion model using sample images, sample mask images corresponding to the sample images, and sample cue information, and the diffusion model includes a pose control module.
[0144] According to one or more embodiments of this disclosure, [Example Nine] provides the method of Example Three, which further includes: optionally, determining the target background image corresponding to the target object based on the first mask image includes: determining the initial background image corresponding to the target object based on the first mask image, and performing image restoration on the initial background image to obtain the target background image.
[0145] According to one or more embodiments of this disclosure, Example 10 provides the method of Example 3, which further includes: optionally, before redrawing the target object based on the first image and the second mask image, the method further includes: determining edge information in the second mask image corresponding to the target object, adjusting the transparency of at least some pixels in the edge information, and updating the second mask image based on the adjusted transparency.
[0146] According to one or more embodiments of this disclosure, [Example 11] provides an image processing apparatus, comprising: an image acquisition module, configured to acquire a first image in response to an image setting operation, wherein the first image includes a target object and the target object is presented at a first size; and an image processing module, configured to redraw the target object in the first image in a target style in response to an image processing request to obtain a second image, and display the second image; wherein at least a portion of the target object in the second image is presented at a second size and the target object presents the target style.
[0147] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0148] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0149] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image processing method, characterized by, The method comprises: in response to an image setting operation, obtaining a first image, wherein the first image comprises a target object and the target object is presented in a first size; in response to an image processing request, redrawing the target object in the first image in a target style to obtain a second image, and displaying the second image; wherein at least part of the target object in the second image is presented in a second size and the target object is presented in the target style.
2. The image processing method of claim 1, wherein, After obtaining the first image, the method further comprises: displaying the first image; in response to an object setting operation on the first image, determining a target object in the first image according to the object setting operation.
3. The image processing method of claim 1, wherein, The redrawing of the target object in the first image in the target style to obtain the second image comprises: processing the first image to obtain at least one first mask image corresponding to the target object, performing size transformation processing on a target region corresponding to the target object in at least one first mask image to obtain a second mask image; redrawing the target object according to the first image and the second mask image to obtain a target foreground image, and determining a target background image corresponding to the target object according to the first mask image; performing image fusion on the target foreground image and the target background image to obtain a second image corresponding to the first image.
4. The image processing method of claim 3, wherein, The size transformation processing on the target region corresponding to the target object in at least one first mask image to obtain a second mask image comprises: in the case where the number of first mask images is one, performing image segmentation on the first mask image to obtain a first number of local mask images; obtaining a second number of to-be-processed mask images from the local mask images, performing size transformation processing on a target region corresponding to the target object in the to-be-processed mask images to obtain a target mask image, wherein the second number is less than the first number; splicing the target mask image and the local mask images except the to-be-processed mask images to obtain a second mask image.
5. The image processing method of claim 3, wherein, The size transformation processing on the target region corresponding to the target object in at least one first mask image to obtain a second mask image comprises: in the case where the number of first mask images is multiple, performing size transformation processing on a target region corresponding to the target object in a part of the first mask images in the multiple first mask images to obtain a second mask image.
6. The image processing method of claim 3, wherein, The redrawing of the target object according to the first image and the second mask image to obtain a target foreground image comprises: redrawing the target object according to the first image, the second mask image, a target style, and target pose information of the target object in the first image to obtain a target foreground image.
7. The image processing method of claim 6, wherein, The re-drawing of the target object according to the first image, the second mask image, a target style, and target pose information of the target object in the first image to obtain a target foreground image includes: determining target pose information of the target object in the first image, and obtaining a first model corresponding to the target style, wherein the first model is obtained by training a deep learning model based on a sample image, a sample mask image corresponding to the sample image, and sample pose information; inputting the first image, the second mask image, and the target pose information into the first model to re-draw the target object and obtain a target foreground image.
8. The image processing method of claim 6, wherein, The re-drawing of the target object according to the first image, the second mask image, a target style, and target pose information of the target object in the first image to obtain a target foreground image includes: determining target prompt information, wherein the target prompt information includes pose prompt information and style prompt information, the pose prompt information is used to prompt the target pose information of the target object presented in the second image, and the style prompt information is used to prompt the target style of the target object presented in the second image; inputting the target prompt information, the first image, and the second mask image into a second model to re-draw the target object and obtain a target foreground image, wherein the second model is obtained by training a diffusion model based on a sample image, a sample mask image corresponding to the sample image, and sample prompt information, and the diffusion model includes a pose control module.
9. The image processing method of claim 3, wherein, The determination of a target background image corresponding to the target object according to the first mask image includes: determining an initial background image corresponding to the target object according to the first mask image, and performing image inpainting on the initial background image to obtain a target background image.
10. The image processing method of claim 3, wherein, Before the re-drawing of the target object according to the first image and the second mask image, the method further includes: determining edge information corresponding to the target object in the second mask image, adjusting the transparency of at least some of the pixel points in the edge information, and updating the second mask image based on the adjusted transparency.
11. An image processing apparatus characterized by comprising: The method includes: an image acquisition module configured to acquire a first image in response to an image setting operation, wherein the first image includes a target object and the target object is presented in a first size; an image processing module configured to re-draw the target object in the first image in a target style in response to an image processing request to obtain a second image, and display the second image, wherein at least part of the target object in the second image is presented in a second size and the target object is presented in the target style.
12. An electronic device, comprising: The electronic device includes: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method of any one of claims 1-10.
13. A storage medium containing computer-executable instructions, wherein: The computer executable instructions, when executed by a computer processor, are for performing the image processing method as claimed in any one of claims 1-10.
14. A computer program product, characterised in that, The computer program product comprises a computer program which, when executed by a processor, implements the image processing method according to any one of claims 1-10.