Image processing method, device and storage medium

By detecting and adjusting salient areas, the screenshots with the highest aesthetic scores are generated, solving the problems of unprominent and incomplete photo objects, and realizing the automated generation of high-quality collage images and photo movies.

CN114529495BActive Publication Date: 2026-01-02BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011240356.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-09
Publication Date
2026-01-02
Estimated Expiration
2040-11-09

AI Technical Summary

Technical Problem

In photo editing, existing technologies struggle to automatically adjust prominent areas to generate striking and complete mosaic images or photo movies, resulting in incomplete and poor-quality photo objects that require manual adjustment by the user.

Method used

By detecting salient regions in the original image, adjusting them to meet the target shape parameters and performing aesthetic scoring, the region with the highest aesthetic score is selected for cropping to generate the target image.

Benefits of technology

It automatically detects and adjusts prominent areas to generate better-looking mosaic images or photo movies that do not require manual adjustments by the user, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529495B_ABST
    Figure CN114529495B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image processing method, device and storage medium; wherein the image processing method comprises: detecting a first salient region of an original image; wherein the first salient region is an image region satisfying a saliency condition; adjusting the first salient region to obtain a second salient region different from the first salient region and satisfying a target shape parameter; performing an aesthetic score on the second salient region, and determining a second salient region with the highest aesthetic score; and performing a screenshot on the original image according to the second salient region with the highest aesthetic score to generate a target image. In this way, the processing of the original image retains the salient region of the image, and the region with the highest aesthetic score is selected, so that the display effect of the displayed image is better.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of display, and in particular, to an image processing method, device and storage medium. BACKGROUND

[0002] When sharing a group of photos, the group of photos is often displayed in the form of a collage image or a photo movie, so that multiple photos can be displayed in one display unit. However, in the current photo editing process, the photos are directly selected according to the layout scheme selected by the user to make a collage or a photo movie, which easily leads to the objects in the photos becoming unobtrusive and incomplete, and the quality of the photo area being poor. In order to achieve a better presentation effect, the user needs to manually adjust, and the final result may not be satisfactory. SUMMARY

[0003] The present disclosure provides an image processing method, device and storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, an image processing method is provided, comprising:

[0005] detecting a first salient region of an original image; wherein the first salient region is an image region satisfying a saliency condition;

[0006] adjusting the first salient region to obtain a different second salient region satisfying a target shape parameter and containing the first salient region;

[0007] performing an aesthetic score on the different second salient region, and determining a second salient region with the highest aesthetic score;

[0008] performing a screenshot on the original image according to the second salient region with the highest aesthetic score to generate a target image.

[0009] Optionally, the image region satisfying the saliency condition comprises one of:

[0010] an image region in which a target object in an object contained in the original image is imaged;

[0011] an image region in which the most image features are contained in the original image;

[0012] an image region in which the highest image feature clarity is contained in the original image.

[0013] Optionally, the detecting the first salient region of the original image comprises:

[0014] detecting the original image according to a preset mask satisfying the saliency condition to obtain the first salient region of the original image.

[0015] Optionally, the first salient region in different second salient regions has different positions.

[0016] and / or;

[0017] The shape parameter of the first salient region in different second salient regions is unchanged.

[0018] Optionally, the adjusting the first salient region to obtain a second salient region satisfying a target shape parameter and containing the first salient region comprises:

[0019] determining the target shape parameter according to the received image processing instruction;

[0020] adjusting the first salient region according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

[0021] Optionally, the adjusting the first salient region according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region comprises:

[0022] enlarging the first salient region of the original image to the outside according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

[0023] Optionally, the image processing instruction comprises a first instruction of making a jigsaw image.

[0024] The determining the target shape parameter according to the received image processing instruction comprises:

[0025] determining layout information of a jigsaw image to be made and shape parameters of images in each position in the layout information according to the first instruction;

[0026] determining the target shape parameter corresponding to the original image according to the shape parameters of the images in each position.

[0027] Optionally, the image processing instruction comprises a second instruction of making a photo movie.

[0028] The determining the target shape parameter according to the received image processing instruction comprises:

[0029] determining an appearance order of each picture in a photo movie to be made and shape parameters of images corresponding to each appearance order according to the second instruction;

[0030] determining the target shape parameter corresponding to the original image according to the shape parameters of the images corresponding to each appearance order.

[0031] Optionally, the different second salient regions are aesthetically scored, including:

[0032] The different second salient regions are processed based on a preset aesthetic evaluation model to obtain aesthetic scores corresponding to the different second salient regions; the aesthetic evaluation model is obtained by training a target neural network model based on a composition layout and a corresponding aesthetic score as sample data.

[0033] According to a second aspect of the embodiments of the present disclosure, an image processing apparatus is provided, including:

[0034] A detection module is configured to detect a first salient region of an original image; the first salient region is an image region satisfying a saliency condition.

[0035] An adjustment module is configured to adjust the first salient region to obtain a different second salient region satisfying a target shape parameter and containing the first salient region.

[0036] A scoring module is configured to aesthetically score the different second salient regions and determine a second salient region with the highest aesthetic score.

[0037] A generation module is configured to take a screenshot of the original image according to the second salient region with the highest aesthetic score to generate a target image.

[0038] Optionally, the image region satisfying the saliency condition includes one of:

[0039] An image region in which a target object of an object contained in the original image is imaged;

[0040] An image region in which the most image features are contained in the original image;

[0041] An image region in which the highest image feature clarity is contained in the original image.

[0042] Optionally, the detection module includes:

[0043] The original image is detected according to a preset mask satisfying the saliency condition, and the first salient region of the original image is detected.

[0044] Optionally, the position of the first salient region in the different second salient regions is different.

[0045] And / or;

[0046] The shape parameter of the first salient region in the different second salient regions is unchanged.

[0047] Optionally, the adjustment module includes:

[0048] determining a target shape parameter according to the received image processing instruction;

[0049] adjusting the first salient region according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

[0050] Optionally, the adjusting sub-module comprises:

[0051] expanding the first salient region of the original image according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

[0052] Optionally, the image processing instruction comprises a first instruction of making a jigsaw image.

[0053] The determining module is further configured to:

[0054] determine layout information of the jigsaw image to be made and shape parameters of images at respective positions in the layout information according to the first instruction.

[0055] determine the target shape parameter corresponding to the original image according to the shape parameters of the images at the respective positions.

[0056] Optionally, the image processing instruction comprises a second instruction of making a photo movie.

[0057] The determining module is further configured to:

[0058] determine an occurrence order of respective pictures in the photo movie to be made and shape parameters of images corresponding to the respective occurrence orders according to the second instruction.

[0059] determine the target shape parameter corresponding to the original image according to the shape parameters of the images corresponding to the respective occurrence orders.

[0060] Optionally, the scoring module comprises:

[0061] a scoring sub-module configured to process different second salient regions based on a preset aesthetic evaluation model to obtain aesthetic scores corresponding to the different second salient regions, wherein the aesthetic evaluation model is obtained by training a target neural network model based on a composition layout and a corresponding aesthetic score as sample data.

[0062] According to a third aspect of the embodiments of the present disclosure, an image processing apparatus is provided, comprising:

[0063] a processor;

[0064] a memory for storing processor-executable instructions;

[0065] The processor is configured to implement the method of any one of the first aspect when executing the processor-executable instructions stored in the memory.

[0066] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided, and the computer-readable storage medium stores computer-executable instructions. The computer-executable instructions are executed by a processor to implement the steps in the method provided in any one of the first aspect.

[0067] The technical solutions provided by the embodiments of the present disclosure can have the following beneficial effects:

[0068] The image processing method provided by the embodiments of the present disclosure can detect the first salient region of the original image, adjust the first salient region to obtain a second salient region meeting the condition, and select the second salient region with the highest aesthetic score to generate a target image by taking a screenshot of the original image. In this way, since the salient region meeting the saliency condition in the original image is detected first, the important part contained in the original image is retained in the processing of the original image, which can effectively improve the incomplete picture caused by selecting the middle region of the picture in the current picture processing. In addition, since there are various ways to obtain the second salient region by adjusting the first salient region, the aesthetic score of different second salient regions is calculated, and the target image is generated by taking a screenshot of the second salient region with the highest aesthetic score, so that the display effect of the target image is better. In addition, since the salient region of the image is detected automatically, manual adjustment by the user is not required, and the user experience is better.

[0069] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0070] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0071] Figure 1 is a flowchart of an image processing method according to an exemplary embodiment Figure One .

[0072] Figure 2 is a schematic diagram of the first salient region and the second salient region.

[0073] Figure 3 is a schematic diagram of three different second salient regions obtained by enlarging.

[0074] Figure 4 is a flow of an image processing method according to an example embodiment Figure Two .

[0075] Figure 5 is a flow of an image processing method according to an example embodiment Figure Three .

[0076] Figure 6 is a structural schematic diagram of an image processing apparatus according to an example embodiment.

[0077] Figure 7 is a block diagram of an image processing apparatus according to an example embodiment. DETAILED DESCRIPTION

[0078] The example embodiments will be described in detail herein with reference to the attached drawings. When the following description refers to arrangements in the drawings, identical numbers on different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following example embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0079] The embodiments of the present disclosure provide an image processing method, Figure 1 is a flow of an image processing method according to an example embodiment Figure One As shown in Figure 1 , the image processing method comprises the following steps:

[0080] Step 101, detecting a first salient region of an original image; wherein the first salient region is an image region satisfying a saliency condition;

[0081] Step 102, adjusting the first salient region to obtain a second salient region different from the first salient region and satisfying a target shape parameter;

[0082] Step 103, performing an aesthetic score on the second salient region, and determining a second salient region with the highest aesthetic score;

[0083] Step 104, according to the second salient region with the highest aesthetic score, performing a screenshot on the original image to generate a target image.

[0084] It should be noted that the image processing method can be applied to any electronic device; for example, a smart phone, a tablet computer, a desktop computer or a digital camera, etc.

[0085] In the embodiments of the present disclosure, the first salient region refers to an image region in the original image that satisfies a saliency condition. The image region that satisfies the saliency condition is a region that can reflect the characteristics of the original image.

[0086] For example, assuming that the original image is a cat photo, the first salient region of the original image refers to an image region in which the cat is imaged. For another example, the original image is a landscape photo containing a large tree, and if the image features of the large tree are the most, the first salient region of the original image refers to an image region in which the large tree is located. For another example, if the original image is a landscape image, the clarity of the scene in the distance is low, and at this time, the image region with the highest image feature clarity can be determined as the first salient region.

[0087] The first salient region can be a connected image region in the original image, or can be composed of multiple sub-regions that are not connected in the original image, and at this time, each part of a single sub-region is connected.

[0088] In order to detect the first salient region of the original image, in an embodiment, a preset neural network model can be used to process the original image to determine the first salient region.

[0089] In some embodiments, the preset neural network model can be a convolutional neural network (CNN) or a residual neural network (ResNet).

[0090] Taking the CNN as an example, the detection of the first salient region of the original image can be: determining a mask corresponding to the region of interest to be extracted; multiplying a two-dimensional matrix array corresponding to the mask with an image matrix of the original image to determine the region of interest. The specific implementation manner is described later. The present disclosure does not limit the manner of detecting the first salient region.

[0091] After the first salient region is determined, since the first salient region is a region in the original image that can represent the characteristics of the image, after the original image is cropped by the first salient region to generate a target image, the generated target image can retain important parts contained in the image. Therefore, the incomplete image caused by selecting the middle region of the image in the current picture sharing process can be effectively improved, and the display effect of the target image obtained by cropping can be better based on the detected salient region.

[0092] Here, the region shape of the first salient region can be a regular shape such as a rectangle, a circle, or an ellipse. It can also be a contour shape of the target object; for example, the contour shape of a kitten in a photo of a kitten. Or it can also be a shape that is cropped according to the generation requirements of the target image.

[0093] In the embodiments of the present disclosure, in order to make the generated target image have better display effect and meet the actual processing needs, after the first salient region of the original image is detected, the first salient region can be adjusted to obtain a second salient region that satisfies a target shape parameter and contains the first salient region.

[0094] The target shape parameter refers to a shape parameter of a region to be cropped on the original image in a target application scenario.

[0095] The shape parameter is used to indicate the shape of the region.

[0096] In some embodiments, the target shape parameter can include a target aspect ratio, a target diameter, or a target major and minor axis length, etc. For example, assuming that the first salient region is a rectangle, the shape parameter of the first salient region can be represented by the aspect ratio of the rectangle. For another example, assuming that the first salient region is a circle, the shape parameter of the first salient region can be represented by the radius or diameter parameter of the circle.

[0097] The target application scenario here at least includes a scenario of making a jigsaw image or a scenario of making a photo movie.

[0098] In the scenario of making a jigsaw image or making a photo movie based on the original image, since the shape parameter of the first salient region detected in the original image can be different from the shape parameter of the position corresponding to the original image in the layout information selected by the user, when the original image is used to make a jigsaw image or a photo movie, the first salient region of the original image needs to be adjusted so that the target image finally displayed on the jigsaw image or the photo movie can meet the shape requirement in the layout information.

[0099] For example, taking the three original images A, B and C as an example, assuming that the selected W layout mode, A is located in the upper left corner of the jigsaw image, and the corresponding display shape is a rectangle with a length-width ratio of 4:3, B and C are located in the lower half of the image area, wherein the corresponding display shape of B and C is also a rectangle, the length-width ratio of B is 1:1, and the length-width ratio of C is 1:3. Assuming that the first salient region detected in the original image A is a rectangle with a length-width ratio of 1:1, and since the corresponding length-width ratio of A in the jigsaw image to be made is 4:3. In order to make the effect of the jigsaw image made by the original image better, it is necessary to adjust the first salient region (adjust the first salient region from 1:1 to the target length-width ratio 4:3), so that the second salient region with a length-width ratio of 4:3 and containing the complete first salient region can be obtained after adjustment.

[0100] In this way, since the first salient region is an image region in the original image that meets the saliency condition, it can reflect the characteristics of the original image. Therefore, adjusting the first salient region of the original image to make the adjusted second salient region meet the target shape parameter and contain the first salient region can achieve the processing of the original image while retaining important parts, and the display effect meets the requirements of the target shape parameter.

[0101] Further, in the embodiments of the present disclosure, there are many ways to adjust the first salient region, and these many adjustment methods can all obtain a second salient region that meets the target shape parameter and contains the first salient region, that is, the second salient region circled by different expansion methods may be different. Regardless of the second salient region, it contains the first salient region. For example, assuming that the original image is a cat image, if the first salient region is expanded to the left to obtain a second salient region containing the first salient region, then the cat in the second salient region is located on the right. If the first salient region is expanded upwards to obtain a second salient region containing the first salient region, then the cat in the second salient region is located on the lower side. Then, the cat is in different second salient regions, and different visual effects will be presented due to the different positions.

[0102] In order to select the best visual effect area to generate a target image, the embodiments of the present disclosure give different second display regions, and then give an aesthetic score to these second display regions to determine the second display region with the highest aesthetic score. The second display region with the highest aesthetic score is used to take a screenshot of the original image to generate a target image.

[0103] In this way, the target image can be obtained according to the second salient region with the highest aesthetic score, so that the display effect is better.

[0104] Here, the target image can be generated by taking a screenshot of the original image according to the second salient region with the highest aesthetic score.

[0105] Here, the shape of the target image can be a regular shape such as a rectangle or a circle. It can also be the contour shape of the target object, for example, the contour shape of a kitten in a kitten photo. Or it can also be a shape according to the generation requirements of the target image. For example, when making a photo movie based on the original image, since the image frames in the movie are rectangular, the target image is also rectangular. When making a jigsaw image based on the original image, the shapes of each position in the jigsaw image to be generated can be different, and the shapes of each position can be any shape, so the target image is also set to the corresponding shape.

[0106] In some embodiments, the image region that meets the saliency condition includes one of the following:

[0107] An image region where the target object in the object contained in the original image is imaged;

[0108] An image region where the image features contained in the original image are the most;

[0109] An image region where the image features contained in the original image have the highest clarity.

[0110] Here, if the original image is an image generated by capturing the target object, the first salient region of the original image is the image region where the target object is imaged. For example, if the original image is a kitten photo, the first salient region of the original image refers to the image region where the kitten is imaged.

[0111] If the original image contains many objects, the region where the object with the most image features is located can be determined as the image region. For example, if the original image is a landscape photo containing a large tree, and the image features of the large tree are the most, the first salient region of the original image is the image region where the large tree is located.

[0112] If the original image is a landscape image, the clarity of the distant scenery is low, and at this time the image region with the highest image feature clarity can be determined as the first salient region.

[0113] Since the first salient region is a region in the original image that can represent the image features, taking a screenshot of the original image through the first display region can retain the region in the original image that can reflect the original image features, so that the first salient region of the original image can be applied completely in subsequent processing.

[0114] In some embodiments, the first salient region of the original image is detected by:

[0115] The original image is detected according to a preset mask satisfying a saliency condition, and a first salient region of the original image is detected.

[0116] In the embodiments of the present disclosure, the preset mask satisfying the saliency condition includes a mask corresponding to an image feature of a region of interest. For example, in the above cat image, the cat is the region of interest, and the image feature of the cat can be used as the preset mask.

[0117] In some embodiments, the mask can be at least represented by a two-dimensional matrix array.

[0118] Based on the two-dimensional matrix array, the original image is detected according to the preset mask satisfying the saliency condition, and a first salient region of the original image is detected, including:

[0119] The two-dimensional matrix array corresponding to the preset mask is multiplied with an image matrix of the original image;

[0120] According to the values in the matrix array obtained by the multiplication, the first salient region of the original image is determined.

[0121] Here, in the matrix array obtained by the multiplication, the region with a value of 1 is the first salient region, and the region with a value of 0 is a non-salient region. In this way, the first salient region can be determined based on the values in the matrix array.

[0122] In some embodiments, the position of the first salient region in different second salient regions is different.

[0123] And / or;

[0124] The shape parameter of the first salient region in different second salient regions is unchanged.

[0125] Due to different adjustment methods, different second salient regions can be obtained, and therefore, the position of the first salient region in different second salient regions will be different according to different adjustment methods.

[0126] In addition, for example, if the original image is a cat photo, the aspect ratio of the first salient region corresponding to the cat image is 1:1, and the target shape parameter is represented by a target aspect ratio, and the target aspect ratio is 4:3. If the length is compressed to change 1:1 to 4:3, the object in the first salient region will be distorted.

[0127] In order to completely retain the first salient region and avoid the situation that the display shape of the first salient region is greatly changed and distorted due to adjustment to the target shape parameter, the shape parameter of the first salient region in the second salient region obtained after adjustment needs to be kept unchanged, so that the shape distortion in adjustment can be avoided and the display effect can be ensured.

[0128] In some embodiments, the adjusting the first salient region to obtain a second salient region satisfying the target shape parameter and containing the first salient region comprises:

[0129] According to the received image processing instruction, determining the target shape parameter;

[0130] According to the target shape parameter, adjusting the first salient region to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

[0131] Here, in the embodiments of the present disclosure, according to the received image processing instruction, the shape parameter of the region to be cropped on the original image in the corresponding application scenario is determined.

[0132] The application scenarios here at least include a scenario of making a jigsaw image or a scenario of making a photo movie.

[0133] In the scenario of making a jigsaw image or making a photo movie based on the original image, the corresponding image processing instruction is triggered based on the layout information selected by the user, and according to the image processing instruction, the target shape parameter of the position corresponding to the original image in the jigsaw image or the photo movie to be made can be known.

[0134] In some embodiments, the image processing instruction comprises a first instruction of making a jigsaw image.

[0135] According to the received image processing instruction, determining the target shape parameter comprises:

[0136] According to the first instruction, determining the layout information of the jigsaw image to be made and the shape parameter of the image at each position in the layout information;

[0137] According to the shape parameter of the image at each position, determining the target shape parameter corresponding to the original image.

[0138] In other embodiments, the image processing instruction comprises a second instruction of making a photo movie.

[0139] According to the received image processing instruction, determining the target shape parameter comprises:

[0140] According to the second instruction, the appearance order of each picture in the photo movie to be made and the shape parameter of the image corresponding to each appearance order are determined.

[0141] According to the shape parameter of the image corresponding to each appearance order, the target shape parameter corresponding to the original image is determined.

[0142] Here, the first instruction or the second instruction carries layout information obtained based on a detection layout selection operation. The target shape parameter is the shape parameter of the image at each position shown in the layout information.

[0143] The layout mode of the jigsaw image includes the positional relationship of multiple images and the shape parameter corresponding to each position.

[0144] The layout mode of the photo movie includes the appearance order of multiple images and the shape parameter corresponding to each order.

[0145] Here, the jigsaw image is an image presented by placing multiple images on the same picture according to a preset layout mode; the photo movie is a movie formed by playing multiple images according to a preset order. The multiple images in the jigsaw image or the photo movie have different presentation shapes and different positions in different layouts.

[0146] For example, in the scenario of making a jigsaw image, a first instruction for making a jigsaw image is received, and a jigsaw processing is performed on multiple original images according to the first instruction. In the scenario of making a photo movie, a second instruction for making a photo movie is received, and a processing for making a photo movie is performed on multiple original images according to the second instruction.

[0147] In this way, according to the different received image processing instructions, after the shape parameter of the image at the corresponding position is determined, the target shape parameter corresponding to the original image is determined based on the corresponding position.

[0148] In some embodiments, the adjusting the first salient region according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region includes:

[0149] According to the target shape parameter, the first salient region of the original image is expanded to the surrounding to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

[0150] Since the adjustment of the first salient region of the original image needs to keep the first salient region, and the second salient region contains the first salient region, the display area of the second salient region is larger than the display area of the first salient region. Then the first salient region of the original image can be expanded to the four directions to obtain the second salient region which satisfies the image proportion and contains the first salient region.

[0151] For example, assuming that the first salient region is a rectangle, the aspect ratio of the first salient region is 1:1, the application scenario is to make a jigsaw image, and the position corresponding to the original image in the jigsaw image is also a rectangle, the target shape parameter is represented by a target aspect ratio, and assuming that the target aspect ratio is 4:3, the first salient region of the original image can be expanded to the four directions to enclose the surrounding region, so that the aspect ratio changes from 1:1 to 4:3.

[0152] As a specific example, assuming that the original image is a cat photo, i.e., an image containing a cat; assuming that the aspect ratio of the first salient region corresponding to the cat imaging is detected to be 1:1, and according to the received image processing instruction of making a jigsaw image, it is determined that the display aspect ratio of the position corresponding to the cat photo in the jigsaw image to be made is 4:3. In order to satisfy a better display effect and not distort the original image after adjustment, the first salient region is adjusted to obtain a second salient region with an aspect ratio of 4:3 and containing the complete first salient region. After the cat imaging region in the cat photo is enclosed as the first salient region by a surrounding frame, the region around the cat imaging region can also be included in the surrounding frame, and at this time, the surrounding frame is the surrounding frame containing the cat imaging region and the background region beside the cat.

[0153] As shown in Figure 2 , the left side of FIG. 1 is a first salient region containing a smiley face, and the right side of FIG. 1 is a second salient region containing the smiley face and the background region beside the smiley face. Figure 2 As shown in , the left side of FIG. 2 is a first salient region containing a smiley face, and the right side of FIG. 2 is a second salient region containing the smiley face and the background region beside the smiley face.

[0154] Figure 2 Figure 2 As shown in , the left side of FIG. 3 is a first salient region containing a smiley face, and the right side of FIG. 3 is a second salient region containing the smiley face and the background region beside the smiley face.

[0155] In this way, after the expansion processing, the original image characteristics can be kept, and the display layout requirement can be met.

[0156] Here, since the first salient region of the original image needs to be expanded to the four directions to obtain the second salient region which satisfies the target shape parameter and contains the first salient region. But the expansion can be expanded in different directions, and different expansion methods obtain different second salient regions.

[0157] Figure 3 As shown in Figure 3are schematic diagrams of the second salient region obtained by three different expansion manners, wherein the first diagram shows that the first salient region is expanded upward to achieve 1:1 to 4:3, the second diagram shows that the first salient region is expanded leftward to achieve 1:1 to 4:3, and the third diagram shows that the first salient region is expanded rightward to achieve 1:1 to 4:3.

[0158] However, in the original image, the second salient region obtained by different expansion manners may be different.

[0159] For example, the original image is a cat photo, the first salient region corresponding to the cat imaging is a rectangle, and the aspect ratio is 1:1. The first salient region is expanded to obtain a second salient region containing the first salient region, and the aspect ratio of the second salient region is 4:3. Since the first salient region is just the region surrounding the cat imaging, if the first salient region is expanded rightward to obtain a second salient region containing the first salient region, the cat in the second salient region is located on the left side.

[0160] For another example, if the first salient region is expanded leftward to obtain a second salient region containing the first salient region, the cat in the second salient region is located on the right side.

[0161] For still another example, if the first salient region is expanded upward to obtain a second salient region containing the first salient region, the cat in the second salient region is located on the lower side.

[0162] Here, the different content includes different layouts or different image features.

[0163] For example, assuming that the original image is a cat photo, the background in the cat picture is the sky, and the cat is on the grass, when the first salient region is expanded upward to obtain a second salient region containing the first salient region, the sky may exist in the second salient region. When the first salient region is expanded downward to obtain a second salient region containing the first salient region, the grass may exist in the second salient region.

[0164] In this way, different second salient regions containing different content can be obtained according to different expansion manners.

[0165] In some embodiments, the aesthetic score of the different second display regions includes:

[0166] The different second salient regions are processed based on a preset aesthetic evaluation model to obtain the aesthetic score of the different second salient regions, wherein the aesthetic evaluation model is obtained by training a target neural network model based on a composition layout and a corresponding aesthetic score as sample data.

[0167] In this embodiment, the electronic device pre-stores an aesthetic evaluation model. After obtaining different second salient regions, the aesthetic evaluation model processes the second salient region and outputs the corresponding aesthetic score.

[0168] The pre-stored aesthetic evaluation model is obtained by training the target neural network model based on the composition layout and the corresponding aesthetic scores as sample data. In this way, the aesthetic scores corresponding to different second salient regions can be determined conveniently and quickly.

[0169] The aesthetic evaluation model can be a convolutional neural network model, etc.

[0170] In some embodiments, there can be multiple aesthetic evaluation models, each with different aesthetic evaluation criteria. The corresponding aesthetic evaluation model can be selected based on the detected user input to assign an aesthetic score to the second salient region.

[0171] Aesthetic evaluation criteria may include: criteria based on composition and layout or criteria based on color matching.

[0172] Different aesthetic evaluation models are applied to different image styles. These image styles include: natural landscape style, anime style, or portrait style.

[0173] To select a more suitable aesthetic evaluation model for aesthetic scoring in the second salient region and obtain a more reasonable aesthetic score, the image processing method in this embodiment further includes the following steps before performing aesthetic scoring in the second salient region based on the aesthetic evaluation model:

[0174] Identify the image style of the original image;

[0175] Based on the image style, an aesthetic evaluation model corresponding to the image style is determined.

[0176] Thus, by using a more suitable aesthetic evaluation model to score the second salient region, the resulting aesthetic score is more consistent with the actual image characteristics and is also more intelligent.

[0177] Figure 4 This is a flowchart of an image processing method according to an exemplary embodiment. Figure Two ,like Figure 4 As shown, the method includes:

[0178] Step 201: Detect the first salient region of the original image;

[0179] Step 202: Determine the target shape parameters according to the received image processing instructions;

[0180] Step 203, adjusting the first salient region of the original image to obtain a second salient region satisfying the target shape parameter and containing the first salient region;

[0181] Step 204, performing an aesthetic score on different second salient regions and determining a second salient region with the highest aesthetic score;

[0182] Step 205, taking a screenshot of the original image according to the second salient region with the highest aesthetic score to generate a target image.

[0183] Here, the first salient region is an image region in the original image satisfying a saliency condition, for example, a cat imaging region in a cat photo. If the cat imaging region in the cat photo is surrounded by a rectangular bounding box, the aspect ratio of the bounding box is the aspect ratio of the cat imaging region, which is also the aspect ratio of the first salient region, such as 1:1.

[0184] The second salient region is a region after adjusting the first salient region on the original image. Since the second salient region contains the first salient region, the display area of the second salient region is larger than that of the first salient region. And the aspect ratio of the second salient region is the target aspect ratio of the target position in the layout information selected by the user.

[0185] In this way, the original image is taken a screenshot based on the second salient region with the highest aesthetic score, so that the target image generated can both satisfy the requirement of retaining the characteristics of the original image and meet the display layout requirement.

[0186] Thus, the present disclosure solves the problem that the selected photo region subject is not prominent and incomplete and has poor quality, which needs to be manually adjusted in the related application of making a jigsaw image or a photo movie. The present disclosure can automatically detect the salient region in the original image, and automatically locate the region with the highest aesthetic score for making a jigsaw image or a photo movie under the condition of ensuring the completeness and prominence of the salient region. In this way, the generation quality of the jigsaw image or the photo movie is improved, the manual adjustment process of the user is removed, and the user experience is improved.

[0187] The present disclosure also provides the following embodiments:

[0188] When sharing a group of photos, it is often to make a collage image or a photo movie. Suppose a kitten is photographed, usually deliberately to find the angle to shoot the most cute state of the kitten, and when making a collage image or a photo movie, it is also hoped that the collage image or photo movie application can automatically and completely retain the kitten part in the photo, and require the kitten in the image to be beautiful and lively. The collage image or photo movie application selects a part of the area in the photo according to the layout scheme selected by the user. This direct selection of the middle area of the photo can easily lead to the kitten being not prominent and not complete, and the quality of the photo area being poor. In order to achieve better presentation effect, manual adjustment is required, and even the final adjustment cannot be adjusted to a satisfactory state.

[0189] Based on this, an image processing method is provided, Figure 5 Fig. 1 shows a flow of an image processing method according to an example embodiment Figure Three As shown in Figure 5 The method comprises the following steps:

[0190] Step 501, detecting a first salient region in an original image.

[0191] The first salient region can be realized by a target detection algorithm, which determines the first salient region by processing the original image through a preset neural network model.

[0192] Specifically, an original image is input, and a two-dimensional mask image (Mask) is output. The region with a value of 1 in the Mask is the salient region of the original image, and the region with a value of 0 is the non-salient region.

[0193] Step 502, framing the first salient region by a bounding box, and adjusting the first salient region based on a target shape parameter to obtain a second salient region.

[0194] After obtaining the first salient region of the photo, the first bounding box surrounding the first salient region can be obtained.

[0195] When the user selects a certain layout scheme to make a collage image or a photo movie, the first bounding box at this time may not meet the proportion required by the layout, so the shape parameter of the bounding box needs to be adjusted appropriately to obtain a second bounding box. The region surrounded by the second bounding box at this time is the second salient region.

[0196] In order to find the best photo area, under the condition of ensuring the completeness and prominence of the salient region, a plurality of second bounding boxes with different sizes and positions are generated by scaling or sliding the second bounding box, that is, different second salient regions are obtained.

[0197] Step 503, performing aesthetic scoring on different second salient regions, and determining the second salient region with the highest aesthetic score.

[0198] Based on the implementation of the aesthetic evaluation model, input an image, output the aesthetic quality score of the image, get the score of each region, and retain the highest score of the photo region.

[0199] Step 504, based on the target image determined by the second salient region, a jigsaw image or a photo movie is made.

[0200] Here, based on the second salient region, the original image is captured to obtain a screenshot area, and then a target image is generated. The target image obtained at this time is the highest score of the photo region, which can be applied to the jigsaw and photo movie to obtain a jigsaw image or photo movie with better display effect.

[0201] Thus, the image processing method provided by the embodiment of the disclosure detects the first salient region of the original image, adjusts the first salient region to obtain a second salient region meeting the condition, and selects the second salient region with the highest aesthetic score to capture the original image to generate a target image. In this way, since the salient region meeting the saliency condition in the original image is detected first, the important part contained in the original image is retained during the processing of the original image, which can effectively improve the incomplete picture caused by selecting the middle region of the picture in the current picture processing. Moreover, since there are various ways to adjust the first salient region to obtain the second salient region, the second salient region is scored, and the target image is captured through the second salient region with the highest aesthetic score, which can make the display effect of the captured target image better. In addition, since the salient region of the image is automatically detected, manual adjustment by the user is not required, and the user experience is better.

[0202] The disclosure also provides an image processing device, Figure 6 is a structural schematic diagram of an image processing device according to an example embodiment, as Figure 6 As shown in the figure, the image processing device 600 comprises:

[0203] The detection module 601 is configured to detect a first salient region of an original image; wherein the first salient region is an image region meeting a saliency condition;

[0204] The adjustment module 602 is configured to adjust the first salient region to obtain a second salient region meeting a target shape parameter and containing the first salient region;

[0205] The scoring module 603 is configured to score the different second salient regions in terms of aesthetics, and determine the second salient region with the highest aesthetic score;

[0206] The generating module 604 is configured to generate a target image by taking a screenshot of the original image according to the second salient region with the highest aesthetic score.

[0207] In some embodiments, the image region satisfying the saliency condition comprises one of:

[0208] An image region in which a target object included in an object in the original image is imaged;

[0209] An image region in which the original image contains the most image features;

[0210] An image region in which the original image contains the highest image feature clarity.

[0211] In some embodiments, the detecting module comprises:

[0212] The original image is detected according to a preset mask satisfying the saliency condition, and a first salient region of the original image is detected.

[0213] In some embodiments, the position of the first salient region in different second salient regions is different;

[0214] and / or;

[0215] The shape parameter of the first salient region in different second salient regions is unchanged.

[0216] In some embodiments, the adjusting module comprises:

[0217] The determining module is configured to determine a target shape parameter according to the received image processing instruction;

[0218] The adjusting sub-module is configured to adjust the first salient region according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

[0219] In some embodiments, the adjusting sub-module comprises:

[0220] The expanding processing module is configured to expand the first salient region of the original image to the surrounding according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

[0221] In some embodiments, the image processing instruction comprises a first instruction for making a jigsaw image.

[0222] The determining module is further configured to:

[0223] According to the first instruction, determine layout information of the jigsaw image to be made and shape parameters of images at each position in the layout information.

[0224] determine target shape parameters corresponding to the original images according to shape parameters of the images at the respective positions.

[0225] In some embodiments, the image processing instructions include second instructions for making a photo movie.

[0226] The determining module is further configured to:

[0227] determine, according to the second instructions, an occurrence order of each frame in the photo movie to be made and shape parameters of images corresponding to each occurrence order.

[0228] determine target shape parameters corresponding to the original images according to shape parameters of the images corresponding to each occurrence order.

[0229] In some embodiments, the scoring module includes:

[0230] The scoring sub-module is configured to process different second salient regions based on a preset aesthetic evaluation model to obtain aesthetic scores corresponding to the different second salient regions, wherein the aesthetic evaluation model is obtained by training a target neural network model based on a composition layout and a corresponding aesthetic score as sample data.

[0231] As to the apparatus in the above-mentioned embodiments, specific manners in which various modules perform operations have been described in details in the embodiments of the method, and thus will not be described in details here.

[0232] Figure 7 is a block diagram of an image processing apparatus 1800 according to an example embodiment. The apparatus 1800 can be a mobile phone, computer, digital broadcast terminal, message communicator, game console, tablet device, medical device, fitness device, personal digital assistant, or the like, for example.

[0233] Referring to Figure 7 The apparatus 1800 can include one or more of the following components: a processing component 1802, a memory 1804, a power supply component 1806, a multimedia component 1808, an audio component 1810, an input / output (I / O) interface 1812, a sensor component 1814, and a communication component 1816.

[0234] The processing component 1802 generally controls the overall operations of the device 1800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 1802 can include one or more processors 1820 to execute instructions and to complete all or part of steps of the above-described methods. In addition, the processing component 1802 can include one or more modules to facilitate interaction between the processing component 1802 and other components. For example, the processing component 1802 can include a multimedia module to facilitate the interaction between the multimedia component 1808 and the processing component 1802.

[0235] The memory 1804 is configured to store various types of data to support operations of the device 1800. Examples of these data include instructions for any application or methods operating on the device 1800, contact data, phonebook data, messages, images, videos, and the like. The memory 1804 can be implemented by any type of volatile or nonvolatile memory devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0236] The power component 1806 provides power to various components of the device 1800. The power component 1806 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 1800.

[0237] The multimedia component 1808 includes a screen to provide an output interface between the device 1800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensors can not only sense a boundary of a touching or sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 1808 includes a front camera and / or a rear camera. The front camera and / or the rear camera can receive external multimedia data when the device 1800 is in an operation mode, such as a shooting mode or a video mode. Each of the front camera and / or the rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0238] The audio component 1810 is configured to output and / or input audio signals. For example, the audio component 1810 includes a microphone (MIC) that is configured to receive an external audio signal when the device 1800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 1804 or transmitted via the communication component 1816. In some embodiments, the audio component 1810 also includes a speaker for outputting audio signals.

[0239] The I / O interface 1812 provides an interface between the processing component 1802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0240] The sensor component 1814 includes one or more sensors for providing status assessments for various aspects of the device 1800. For example, the sensor component 1814 can detect an open / closed position of the device 1800, relative positioning of components, such as a display and a keypad of the device 1800, a change of position of the device 1800 or a component of the device 1800, presence or absence of user contact with the device 1800, the orientation or acceleration / deceleration / g-force and a temperature change of the device 1800. The sensor component 1814 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 1814 can further include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0241] The communication component 1816 is configured to facilitate wired or wireless communication between the device 1800 and other devices. The device 1800 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 1816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 1816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, or other technology.

[0242] In exemplary embodiments, the apparatus 1800 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronics units to perform the above methods.

[0243] In exemplary embodiments, a non-transitory computer readable storage medium including instructions, such as the memory 1804 including instructions, is also provided, which can be executed by the processor 1820 of the apparatus 1800 to complete the above methods. For example, the non-transitory computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.

[0244] A non-transitory computer readable storage medium, when instructions in the storage medium are executed by a processor, enables the above methods to be performed.

[0245] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure following the general principles thereof and including such departures from the present disclosure that come within known

[0246] It should be understood that the present disclosure is not limited to the precise structures herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the claims that follow.

Claims

1. An image processing method, characterized by, The method comprises: detecting a first salient region of an original image; wherein the first salient region is an image region satisfying a saliency condition; the detecting of the first salient region of the original image comprises: detecting the original image according to a preset mask satisfying the saliency condition, to obtain the first salient region of the original image; the preset mask satisfying the saliency condition comprises: a mask corresponding to an image feature of a region of interest; adjusting the first salient region to obtain a second salient region different from the first salient region and satisfying a target shape parameter; the method further comprises: determining a shape parameter of a region to be cropped on the original image in a corresponding application scenario according to a received image processing instruction; after determining the shape parameter of the image at the corresponding position, determining the target shape parameter of the original image based on the corresponding position; wherein the application scenario at least comprises: a scenario of making a jigsaw image and a scenario of making a photo movie; in the scenario of making a jigsaw image, a first instruction of making a jigsaw image is received, and a jigsaw processing is performed on a plurality of original images according to the first instruction; in the scenario of making a photo movie, a second instruction of making a photo movie is received, and a processing of making a photo movie is performed on a plurality of original images according to the second instruction; performing aesthetic scoring on the different second salient regions and determining a second salient region with the highest aesthetic score; the performing of the aesthetic scoring on the different second salient regions comprises: identifying an image style of the original image; determining a target aesthetic evaluation model from a plurality of aesthetic evaluation models according to the image style; and processing the different second salient regions based on the target aesthetic evaluation model to obtain the aesthetic scores corresponding to the different second salient regions; cropping the original image according to the second salient region with the highest aesthetic score to generate a target image.

2. The method of claim 1, wherein, The image region satisfying the saliency condition comprises one of: an image region in which a target object included in an object of the original image is imaged; an image region in which the original image contains the most image features; an image region in which the original image contains the highest image feature clarity.

3. The method of claim 1, wherein: the positions of the first salient regions in the different second salient regions are different; and / or; the shape parameters of the first salient regions in the different second salient regions are unchanged. The adjusting of the first salient region to obtain the second salient region satisfying the target shape parameter and containing the first salient region comprises:

4. The method of claim 1, wherein, determining the target shape parameter according to the received image processing instruction; adjusting the first salient region according to the target shape parameter to obtain the second salient region satisfying the target shape parameter and containing the first salient region. The adjusting of the first salient region according to the target shape parameter to obtain the second salient region satisfying the target shape parameter and containing the first salient region comprises:

5. The method of claim 4, wherein, ​ According to the target shape parameter, the first salient region of the original image is enlarged to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

6. The method of claim 4, wherein, The determining of the target shape parameter according to the received image processing instruction comprises: According to the first instruction, the layout information of the to-be-made jigsaw image and the shape parameter of the image at each position in the layout information are determined; According to the shape parameter of the image at each position, the target shape parameter corresponding to the original image is determined.

7. The method of claim 4, wherein, The determining of the target shape parameter according to the received image processing instruction comprises: According to the second instruction, the appearance order of each frame in the to-be-made photo movie and the shape parameter of the image corresponding to each appearance order are determined; According to the shape parameter of the image corresponding to each appearance order, the target shape parameter corresponding to the original image is determined.

8. An image processing apparatus characterized by comprising: Comprise: The detection module is used for detecting a first salient region of an original image; wherein the first salient region is an image region satisfying a saliency condition; the detection module is specifically used for detecting the original image according to a preset mask satisfying the saliency condition to obtain the first salient region of the original image; the preset mask satisfying the saliency condition comprises an image feature corresponding to a region of interest; The adjusting module is used for adjusting the first salient region to obtain a different second salient region satisfying a target shape parameter and containing the first salient region; and is further used for determining, according to a received image processing instruction, a shape parameter of a region to be cropped on the original image in a corresponding application scenario, and determining, after determining the shape parameter of the image at the corresponding position, the target shape parameter corresponding to the original image based on the corresponding position; wherein the application scenario at least comprises a jigsaw image making scenario and a photo movie making scenario; in the jigsaw image making scenario, a first instruction for making a jigsaw image is received, and jigsaw processing is performed on multiple original images according to the first instruction; in the photo movie making scenario, a second instruction for making a photo movie is received, and photo movie making processing is performed on multiple original images according to the second instruction; The scoring module is used for performing aesthetic scoring on the different second salient regions and determining a second salient region with the highest aesthetic score; the scoring module is specifically used for identifying an image style of the original image; determining a target aesthetic evaluation model from multiple aesthetic evaluation models according to the image style; and processing the different second salient regions based on the target aesthetic evaluation model to obtain the aesthetic score corresponding to the different second salient regions; The generation module is used for cropping the original image according to the second salient region with the highest aesthetic score to generate a target image.

9. The apparatus of claim 8, wherein, The image region satisfying the saliency condition comprises one of: An image region in which a target object included in the original image is imaged; An image region in which the original image contains the most image features; An image region in which the original image contains the highest image feature clarity.

10. The apparatus of claim 8, wherein, the position of the first salient region in different second salient regions is different; and / or; a shape parameter of the first salient region in different second salient regions is unchanged.

11. The apparatus of claim 8, wherein, The adjusting module comprises: a determining module configured to determine a target shape parameter according to the received image processing instruction; an adjusting sub-module configured to adjust the first salient region according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

12. The apparatus of claim 11, wherein, The adjusting sub-module comprises: an expanding processing module configured to expand the first salient region of the original image to the periphery according to the target shape parameter to obtain a second salient region satisfying the target shape parameter and containing the first salient region.

13. The apparatus of claim 11, wherein, The determining module is further configured to: determine layout information of a to-be-made jigsaw image and shape parameters of images at respective positions in the layout information according to the first instruction; and determine a target shape parameter corresponding to the original image according to the shape parameters of the images at the respective positions.

14. The apparatus of claim 11, wherein, The determining module is further configured to: determine an appearance order of each picture in a to-be-made photo movie and a shape parameter of an image corresponding to each appearance order according to the second instruction; and determine a target shape parameter corresponding to the original image according to the shape parameters of the images corresponding to the respective appearance orders.

15. An image processing apparatus characterized by comprising: comprise: a processor and a memory for storing executable instructions capable of running on the processor, wherein: when the processor is used to run the executable instructions, the executable instructions perform steps in the method provided in any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium, comprising: The computer readable storage medium stores computer executable instructions, and the computer executable instructions are executed by the processor to implement steps in the method provided in any one of claims 1 to 7.

Citation Information

Patent Citations

  • High-quality thumbnail for multi-target image

    CN110909724A