Image processing method and electronic equipment
By performing convolution processing on the input image before affine transformation and using a suitable target convolution kernel for smoothing, the jagged edges of the image are solved, achieving high-quality image processing results.
Patent Information
- Application Number
- CN202410917667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2026-01-16
AI Technical Summary
During affine transformation of an image, jagged edges appear at the image edges, resulting in poor image processing quality and making it impossible to simultaneously guarantee image smoothness and quality.
By obtaining a target convolution kernel that is adapted to the input image and the predicted output image, the input image is first convolved, and then an affine transformation is performed to resolve the contradiction between image smoothness and quality, thus achieving an anti-aliasing effect.
While achieving anti-aliasing of the image after affine transformation, it also ensures excellent image quality and solves the problems of uneven and discontinuous image edges.
Smart Images

Figure CN121353059A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of image processing, and in particular to an image processing method and an electronic device. BACKGROUND
[0002] In the process of image processing (such as image rectification, image stitching, image registration, image enhancement, image post-processing, etc.), affine transformation processing of the image is often involved. Affine transformation refers to geometric transformation of the image, and common operations of affine transformation include rotation, scaling, translation, tilting, shearing, etc.
[0003] However, in the process of affine transformation of the image, the pixel positions in the image change, and discretization errors are generated when the pixels are redistributed. These discretization errors cause the image edges to appear non-smooth and discontinuous, resulting in the problem of jagged edges in the image processed by affine transformation, and the effect of image processing is poor. SUMMARY
[0004] Embodiments of the present application provide an image processing method and an electronic device, which are used to solve the problem of jagged edges generated in the process of affine transformation of the image, and solve the contradiction between the anti-jagged effect and the anti-jagged efficiency.
[0005] To achieve the above object, embodiments of the present application adopt the following technical solutions:
[0006] In a first aspect, an image processing method is provided, which comprises:
[0007] The electronic device obtains an input image. The electronic device performs convolution on the input image based on a target convolution kernel to obtain an intermediate image; the target convolution kernel is a convolution kernel adapted to the input image and / or a predicted output image; the predicted output image is an output image obtained after affine transformation of the input image. The electronic device performs affine transformation on the intermediate image to obtain an output image.
[0008] In embodiments of the present application, the convolution operation is performed on the input image before the affine transformation processing of the input image, which can achieve the effect of smoothing the input image. The target convolution kernel is a convolution kernel adapted to the input image and the predicted output image (output image), and the convolution of the input image based on the target convolution kernel can effectively solve the contradiction between image smoothing and image quality, while achieving the anti-jagged effect of the output image of the affine transformation operation, the image quality is also guaranteed to be good.
[0009] In a possible implementation manner of the first aspect, the method further comprises:
[0010] The electronic device obtains a reference range of the target convolution kernel by a preset size numerical box in a predicted output image; the reference range is used to indicate a range of a convolution kernel adapted to the input image.
[0011] The electronic device obtains a weight matrix of the target convolution kernel by a preset size position box in the input image; the weight matrix is used to indicate a weight of a convolution kernel adapted to the predicted output image.
[0012] The electronic device obtains the target convolution kernel by the reference range of the target convolution kernel and the weight matrix of the target convolution kernel.
[0013] In the embodiments of the present application, the reference range of the target convolution kernel adapted to the input image is obtained by performing reverse affine transformation based on the first numerical box of the predicted output image, and the weight matrix of the target convolution kernel adapted to the predicted output image is obtained by performing affine transformation based on the first position box of the input image. Thus, the target convolution kernel is obtained based on the reference range and the weight matrix of the target convolution kernel. The target convolution kernel is adapted to the input image and the predicted output image (output image), and thus the obtained target convolution kernel is adapted to the input image subjected to affine transformation and is also adapted to the output effect of the output image subjected to affine transformation, which can effectively solve the contradiction between image smoothing and image quality, and can realize image anti-aliasing while ensuring good image quality.
[0014] In a possible implementation form of the first aspect, the electronic device obtains the target convolution kernel in the following manner:
[0015] The electronic device obtains a weight matrix of the target convolution kernel by a preset size position box in the input image; the weight matrix is used to indicate a range of a convolution kernel adapted to the input image.
[0016] The electronic device obtains the target convolution kernel by the reference range of the target convolution kernel and the weight matrix of the target convolution kernel.
[0017] The target convolution kernel obtained in the embodiments of the present application is adapted to the predicted output image (output image), and the input image is convolved based on the target convolution kernel to obtain an intermediate image, which can also solve the contradiction between smoothing processing and image quality; and the intermediate image is subjected to affine transformation to obtain a final output image, which can also solve the problem of jaggies in the final output image.
[0018] In a possible implementation form of the first aspect, the electronic device obtains the target convolution kernel in the following manner:
[0019] The electronic device obtains a reference range of the target convolution kernel by a preset size numerical box in a predicted output image; the reference range is used to indicate a range of a convolution kernel adapted to the input image.
[0020] The electronic device obtains the target convolution kernel by a reference range of the target convolution kernel and a preset weight matrix of the target convolution kernel.
[0021] In the embodiments of the present application, the obtained target convolution kernel is adapted to the input image, the input image is convolved based on the target convolution kernel to obtain an intermediate image, and the intermediate image is also; and the final output image is obtained by performing affine transformation on the intermediate image, which can also solve the sawtooth problem in the final output image.
[0022] In a possible implementation of the first aspect, the electronic device obtains the reference range of the target convolution kernel by a preset size of a numerical box in the predicted output image, and the method comprises the following steps.
[0023] The electronic device selects a first numerical box of a preset size in the predicted output image, and sets a plurality of first pixel points included in the first numerical box to a first value.
[0024] The electronic device obtains a second numerical box in the input image by performing reverse affine transformation on the first numerical box; the second numerical box comprises a plurality of second pixel points, and the values corresponding to the second pixel points included in the second numerical box are the first value.
[0025] The electronic device obtains the minimum circumscribed rectangle of the second numerical box, and sets the values corresponding to the pixel points not belonging to the second numerical box in the minimum circumscribed rectangle to a second value.
[0026] The pixel points contained in the minimum circumscribed rectangle and the values of the pixel points form the reference range of the target convolution kernel.
[0027] In the embodiments of the present application, the reference range of the target convolution kernel adapted to the input image is obtained by performing reverse affine transformation on the first numerical box of the predicted output image, so that the obtained target convolution kernel is adapted to the input image, and the image quality of the intermediate image is ensured.
[0028] In a possible implementation of the first aspect, in the case that the number of the second pixel points included in the second numerical box is less than a preset number threshold, the minimum circumscribed rectangle of the second numerical box is obtained, and the method comprises the following steps.
[0029] The electronic device performs region correction on the region corresponding to the second numerical box to obtain the minimum circumscribed rectangle of the corrected region.
[0030] In the embodiments of the present application, the reverse affine transformation operation based on the first numerical box of the predicted output image may be a downsampling operation, and the region where the second numerical box is located may have the problem of being too small and containing too few pixel points. The region correction on the second numerical box can make the reference size of the finally obtained target convolution kernel more accurate and reliable.
[0031] In a possible implementation manner of the first aspect, the electronic device performs region correction on the region corresponding to the second numerical value box, including:
[0032] The electronic device selects a correction region of a preset size in the input image.
[0033] The electronic device corrects the region corresponding to the second numerical value box through the correction region; the corrected region is a region formed by the pixel points included in the correction region and the pixel points included in the second numerical value box.
[0034] In the embodiment, the reference size of the target convolution kernel obtained finally is more accurate and reliable, and the target convolution kernel obtained is more adaptive to the input image.
[0035] In a possible implementation manner of the first aspect, the electronic device obtains the weight matrix corresponding to the target convolution kernel through a position box of a preset size in the input image, including:
[0036] The electronic device selects a first position box of a preset size in the input image; the first position box includes a plurality of third pixel points.
[0037] The electronic device obtains a second position box mapped in the predicted output image by performing affine transformation on the first position box; the second position box includes a plurality of fourth pixel points.
[0038] The electronic device obtains distances between the fourth pixel points and a preset mapping point based on the coordinates of the fourth pixel points and the coordinates of the preset mapping point.
[0039] The weight of each fourth pixel point is determined based on the distance between the fourth pixel point and the preset mapping point; the weights of all the fourth pixel points form the weight matrix.
[0040] In the embodiment, the weight matrix of the target convolution kernel adaptive to the predicted output image is obtained by performing affine transformation on the first position box of the input image, so that the target convolution kernel obtained finally is adaptive to the predicted output image, and the image quality of the intermediate image is ensured.
[0041] In a second aspect, an image processing method is provided, including:
[0042] The electronic device displays a first interface in response to a first operation of a user; the first interface includes a first image.
[0043] The electronic device displays an image editing interface corresponding to the first image in response to a user editing operation on the first image; the image editing interface includes an image processing control;
[0044] The electronic device performs image processing on the first image to obtain a second image in response to a user operation on the image processing control; the image processing at least includes affine transformation processing on the first image by a target convolution kernel; the target convolution kernel is a convolution kernel adapted to the first image and / or a predicted output image; the predicted output image is an output image obtained after affine transformation on the first image;
[0045] The electronic device displays the second image in the image editing interface.
[0046] In the embodiments of the present application, the input image is subjected to a convolution operation before being subjected to affine transformation processing, which achieves the effect of smoothing the input image. The target convolution kernel is a convolution kernel adapted to the input image and a predicted output image (output image), and convolution of the input image based on the target convolution kernel effectively solves the contradiction between image smoothing and image quality, so that the second image output by the image processing has high image quality while solving the jaggy problem caused by affine transformation.
[0047] In a possible implementation of the second aspect, the affine transformation processing on the first image by the target convolution kernel to obtain the second image includes:
[0048] The electronic device convolves the first image based on the target convolution kernel to obtain an intermediate image after convolution;
[0049] The electronic device performs affine transformation on the intermediate image to obtain the second image.
[0050] In the embodiments of the present application, the input image is subjected to a convolution operation before being subjected to affine transformation processing, which achieves the effect of smoothing the input image. The target convolution kernel is a convolution kernel adapted to the input image and a predicted output image (output image), and convolution of the input image based on the target convolution kernel effectively solves the contradiction between image smoothing and image quality, so that the second image output by the image processing has high image quality while solving the jaggy problem caused by affine transformation.
[0051] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the method of any one of the first aspect.
[0052] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores instructions. The computer program / instructions are executed by a processor to implement the steps of the method according to any one of the first aspect.
[0053] In a fifth aspect, a computer program product is provided, and the computer program product includes instructions. The computer program / instructions are executed by a processor to implement the steps of the method according to any one of the first aspect.
[0054] In a sixth aspect, a chip is provided, and the chip includes a processor. The processor is configured to invoke a computer program in a memory to execute the method according to any one of the first aspect.
[0055] It can be understood that the electronic device according to the third aspect, the computer readable storage medium according to the fourth aspect, the computer program product according to the fifth aspect, and the chip according to the sixth aspect can achieve the beneficial effects of the first aspect and any one of the possible design manners of the first aspect, the second aspect and any one of the possible design manners of the second aspect, which will not be described herein. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 A schematic diagram of a sawtooth problem after portrait enhancement is provided for the embodiments of the present application;
[0057] Figure 2 Another schematic diagram of a sawtooth problem after portrait enhancement is provided for the embodiments of the present application;
[0058] Figure 3 A flowchart of an image processing method is provided for the embodiments of the present application;
[0059] Figure 4 A flowchart of smoothing processing of an input image is provided for the embodiments of the present application;
[0060] Figure 5 A schematic diagram of a numerical block is provided for the embodiments of the present application;
[0061] Figure 6 A schematic diagram of correcting a region corresponding to a second numerical block is provided for the embodiments of the present application;
[0062] Figure 7 A flowchart of obtaining a target convolution kernel is provided for the embodiments of the present application;
[0063] Figure 8 Another flowchart of obtaining a target convolution kernel is provided for the embodiments of the present application;
[0064] Figure 9A result diagram of normal down-sampling kernel provided by an embodiment of the present application;
[0065] Figure 10 A result diagram of down-sampling kernel provided by an embodiment of the present application;
[0066] Figure 11 A comparison diagram of image with sawtooth problem and image without sawtooth problem provided by an embodiment of the present application;
[0067] Figure 12 A structure diagram of an electronic device provided by an embodiment of the present application;
[0068] Figure 13 A scene diagram of an image processing method provided by an embodiment of the present application;
[0069] Figure 14 A software structure block diagram of an electronic device provided by an embodiment of the present application;
[0070] Figure 15 A timing diagram of each module of an electronic device provided by an embodiment of the present application;
[0071] Figure 16 A structure diagram of a chip system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0072] In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments of the present application, and are not intended to be limiting to the present application. As used in the specification and the appended claims of the present application, the singular forms "a", "an" and "the" are intended to include the plural forms, such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" refer to one or more than two (including two). The term "and / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships; for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it.
[0073] Reference within the specification to "one embodiment" or "an embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places within specified
[0074] In the present application, the word "exemplary" or "for example" is used to mean serving as an example, instance, or illustration. Any implementation described as "exemplary" or "for example" in the present application is not necessarily to be construed as preferred or advantageous over other implementations. The
[0075] Before introducing the image processing method provided by the embodiments of the present application, some basic concepts involved in the embodiments of the present application are briefly introduced.
[0076] Affine transformation: refers to geometric transformation of an image. Common operations of affine transformation include rotation, scaling, translation, skewing, shearing, etc. Affine transformation operations of an image are involved in various image processing scenarios. For example, image correction, image stitching, image registration, image enhancement, image beautification, etc. may involve one or more affine transformation operations of an image such as rotation, scaling, translation, skewing, shearing, etc.
[0077] In the process of affine transformation of an image, the problem of sawtooth of the processed image often occurs.
[0078] Sawtooth problem: in affine transformation (such as rotation, scaling, skewing, translation, shearing, etc.), the positions of each pixel in the image change, and the pixels in the image will produce discretization errors when redistributed, and these discretization errors will cause the edges of part of the image region to look not smooth and discontinuous.
[0079] Exemplarily, Figure 1A schematic diagram of a jaggies problem after portrait enhancement is given. When performing portrait enhancement on image 1, a typical portrait enhancement algorithm has a strong portrait position prior. The portrait position prior refers to the fact that if the portrait region image in image 1 is not in the target alignment position, the portrait region image needs to be rotated to the target alignment position before subsequent portrait enhancement operations. At the same time, a general end-side portrait enhancement model (such as open neural network exchange (ONNX)) adopts a static inference model, which can only process images of a fixed size. Therefore, after the portrait position prior (i.e., after rotating the portrait region image to the target alignment position), if the size of the portrait region image does not match the fixed size of the static inference model, the portrait region image needs to be up-sampled or down-sampled (or referred to as scaling processing) to obtain a portrait region image that matches the fixed size of the static inference model. Then, the portrait region image is enhanced. After the enhancement of the portrait region image, the enhanced portrait region image needs to be down-sampled or up-sampled (or referred to as inverse scaling processing) and inverse-rotated to match the position, angle, and size of the portrait region image in the original image 1, and then pasted back to image 1, thereby obtaining an image 2 after portrait enhancement.
[0080] During the entire portrait enhancement process, affine transformation processing such as rotation and scaling of the portrait region image is involved. These affine transformation processing can cause aliasing of the portrait region image, thereby causing a jaggies problem of the edge image of the portrait region image. For example, the jaggies problem of the image can be manifested as an increase in the edge jaggies of a line image. As shown in Figure 1 , the lines of the eyebrows (i.e., an image including multiple lines) of the face in the enhanced image 2 have severe jaggies. For example, the jaggies problem of the image can also be manifested as the disappearance of some edge pixels of the image. Referring to Figure 2 , Figure 2 Another schematic diagram of a jaggies problem after portrait enhancement is given. The processing process of the image is similar to that of the example shown in Figure 1 , and the difference is that the input image (image 1) includes a person wearing glasses. After the portrait enhancement of image 1, the affine transformation operation in the portrait enhancement process causes a jaggies problem in the enhanced image 2, and the glasses frame (edge of the glasses image) of the person disappears.
[0081] In other image processing scenarios, such as image beautification (one-click beautification), image rectification, image registration, and the like, as long as affine transformation is involved in the processing of the image or part of the image, a jaggies problem can occur in the processed image.
[0082] For the sawtooth problem of the image in the affine transformation, the embodiment of the application provides an image processing method. The electronic device first performs smoothing processing on the input image to obtain an intermediate image. Thus, the affine transformation is performed based on the intermediate image, which can effectively solve the sawtooth problem that may occur in the output image. The process of performing smoothing processing includes that the electronic device performs convolution on the input image by using a target convolution kernel. The acquisition of the target convolution kernel includes that the first value box of the predicted output image is subjected to inverse affine transformation to obtain a reference range of the target convolution kernel adapted to the input image, and the first position box of the input image is subjected to affine transformation to obtain a weight matrix of the target convolution kernel adapted to the predicted output image. Thus, the target convolution kernel is obtained based on the reference range and the weight matrix of the target convolution kernel. The target convolution kernel is adapted to the input image and the predicted output image (output image), and thus the obtained target convolution kernel is adapted to the input image subjected to the affine transformation and is also adapted to the output effect of the output image after the affine transformation, which can effectively solve the contradiction between image smoothing and image quality, and can guarantee good image quality while realizing image anti-sawtooth.
[0083] The image processing method provided by the application can be applied to various image processing scenes. Before introducing the image processing method provided by the embodiment of the application in combination with the scene, the steps of the image processing method are described.
[0084] Reference Figure 3 , Figure 3 The flowchart of the steps of the image processing method in the embodiment is given. The execution subject of the image processing method in the embodiment of the application can be a server or an electronic device. Taking the electronic device as an example, the steps of the image processing method include:
[0085] S101, the electronic device acquires an input image.
[0086] The input image can be any type of image. For example, the input image can be a medical image, a portrait image, a landscape image, a pet image, a food image, etc. The input image refers to an image to be subjected to affine transformation processing. The input image can be an image to be subjected to affine transformation in any image processing scene. For example, the input image can be an image to be subjected to affine transformation in the image enhancement process; the input image can also be an image to be subjected to affine transformation in the image registration process; and the input image can also be an image to be subjected to affine transformation in the image stitching process.
[0087] In the embodiment, the input image can be an image received by the electronic device from other devices, the input image can also be an image read by the electronic device from a database, and the input image can also be an image determined in response to a selection operation of the user on the images in the gallery of the electronic device.
[0088] In S102, the electronic device performs smoothing processing on the input image to obtain an intermediate image.
[0089] In this embodiment, the electronic device first performs smoothing processing on the input image, which can weaken the problem of pixel jumping caused in the affine transformation process. The output image obtained based on the intermediate image after the smoothing processing has a relatively low probability of appearing the jaggy problem, which can effectively solve the jaggy problem in the output image after the affine transformation, and achieve the effect of anti-jaggy.
[0090] In one possible manner, the electronic device performs smoothing processing on the input image can be a convolution operation on the input image.
[0091] The electronic device performs the convolution operation on the input image, and needs to determine a convolution kernel used for the convolution operation.
[0092] In this embodiment, the electronic device can determine a target convolution kernel suitable for the affine transformation of the input image based on the input image and the predicted output image corresponding to the input image.
[0093] In some embodiments, the affine transformation matrix can be determined by the electronic device based on the input image and the predicted output image. The electronic device can obtain an image after the affine transformation based on the input image and the affine transformation matrix, and can also obtain an image before the affine transformation based on the image after the affine transformation and the inverse matrix of the affine transformation matrix.
[0094] In this embodiment, the electronic device performs direct affine transformation (without smoothing processing) on the input image based on the affine transformation matrix to obtain the predicted output image. It should be noted that, in this embodiment, the final output image (referred to as output image for short) refers to the image without the jaggy problem after the smoothing processing and the affine transformation on the input image, and the predicted output image refers to the image with the jaggy problem after the direct affine transformation on the input image without the smoothing processing. The output image and the predicted output image have the same size.
[0095] After the predicted output image is determined, the electronic device can determine the affine transformation matrix T between the input image and the predicted output image according to the coordinates of at least three first pixel points in the input image and the coordinates of the second pixel points corresponding to each first pixel point in the predicted output image.
[0096] Exemplarily, the coordinates of the first pixel point of the input image (Input) can be represented as [v, w, 1], the coordinates of the second pixel point of the predicted output image (Output_pre) can be represented as [x, y, 1], and the affine transformation matrix is T. Wherein, 1 in the coordinates is used to represent a pixel point, and the value is 0, which is used to represent a vector. The relationship between the coordinates of the first pixel point, the coordinates of the second pixel point, and the affine transformation matrix can be represented as:
[0097] Output_pre[x, y, 1] = Input[v, w, 1] * T.
[0098] That is, each pixel point in the image before and after the affine transformation has a one-to-one correspondence, and the pixel point (the first pixel point) in the image before the affine transformation (the input image) corresponds to the pixel point (the second pixel point) in the image after the affine transformation (the predicted output image) one by one. Knowing any two parameters in the above formula, the other unknown parameter can be calculated and solved.
[0099] In other embodiments, the affine transformation matrix can also be preset. The electronic device can calculate the coordinates of each corresponding pixel point in the predicted output image based on the preset affine transformation matrix and the coordinates of the pixel point in the input image to obtain the predicted output image.
[0100] After obtaining the input image, the predicted output image, and the affine transformation matrix, the electronic device can obtain the target convolution kernel adapted to the input image based on these known information. Wherein, the process of the electronic device performing smoothing processing on the input image can refer to Figure 4 , the smoothing processing on the input image includes the steps of obtaining the target convolution kernel and convolving the input image based on the target convolution kernel, as shown in Figure 4 , including:
[0101] S1021, the electronic device obtains a first numerical block of the predicted output image.
[0102] The purpose of this embodiment is to solve the problem of jaggies in the output image. Therefore, it is necessary to calculate the convolution kernel adapted to the predicted output image based on the predicted output image.
[0103] Wherein, the numerical block refers to a region space with the same size as the target convolution kernel, and the values corresponding to the pixel points contained in the region space are all set to 1. The electronic device can select a region space of the same size at any position of the predicted output image. For all pixel points contained in the region space, the numerical value corresponding to the pixel point is set to 1. Wherein, the size of the target convolution kernel can be a preset fixed value. For example, the size of the target convolution kernel is M*N, and the size of the first numerical block is also M*N.
[0104] For example, such as Figure 5 As shown, Figure 5 A schematic diagram of a numerical bounding box is provided. For example, when the target convolutional kernel size is 2*2, the first numerical bounding box in the predicted output image can be represented as follows: Figure 5 As shown in 401. For example, the first numerical box includes the first pixel points (pixel 1, pixel 2, pixel 3, pixel 4) of the predicted output image, and the value corresponding to each pixel point is a first value, where the first value can be 1. In addition, in some other feasible embodiments, the size of the target convolution kernel can also be 3*3, 4*4, etc., and the size of the target convolution kernel can be set empirically.
[0105] S1022, The electronic device performs an inverse affine transformation on the first numerical box to obtain the second numerical box.
[0106] In this context, forward affine transformation refers to performing a dot product between each pixel in the input image and the affine transformation matrix to obtain the predicted output image after the affine transformation. Similarly, backward affine transformation refers to performing a dot product between each pixel in the predicted output image and the inverse of the affine transformation matrix to obtain the input image after the backward affine transformation.
[0107] In this embodiment, as Figure 5 As shown, after acquiring the first numerical box of the predicted output image, the electronic device performs an inverse affine transformation on the first numerical box based on the inverse of the affine transformation matrix, thereby obtaining the second numerical box 402 in the input image corresponding to the first numerical box. Figure 5 As shown, the electronic device acquires a second numerical bounding box. The pixel values contained within the second numerical bounding box are first values, where the first value is 1. The second numerical bounding box is obtained by inverse affine transformation of the first numerical bounding box, and the shapes of the second and first numerical bounding boxes may differ. The second numerical bounding box includes the second pixels (pixels 1', 2', 3', and 4') in the input image that correspond to the first pixels (pixels 1, 2, 3, and 4) in the predicted output image, with each pixel having a value of 1. The region containing the second numerical bounding box can be represented as region md.
[0108] In this embodiment, the electronic device acquires the second numerical bounding box to obtain a reference range for the convolution kernel corresponding to the input image. That is, the matrix formed by the arrangement of pixels contained in the second numerical bounding box can characterize the range matrix of the convolution kernel corresponding to the input image.
[0109] Optionally, in some embodiments, the reverse affine transformation operation based on the first numerical box of the predicted output image can be a down-sampling operation, and the second numerical box obtained after the reverse affine transformation can be too small in area and contain too few pixel points. In the case where the area of the second numerical box is smaller than a preset area threshold, or the number of pixel points contained in the second numerical box is smaller than a preset number, the area where the second numerical box is located can be corrected in the embodiments of the present application.
[0110] For this possible case, the area where the second numerical box is located can be corrected in the embodiments of the present application. Exemplarily, reference is made to Figure 6 , Figure 6 A schematic diagram of correcting the area corresponding to the second numerical box is given.
[0111] Specifically, reference is made to Figure 6 , and it is assumed that the area corresponding to the first numerical box is a 4*4 area 501. After the reverse affine transformation, the area where the second numerical box is located is an irregular area 502. The electronic device can select a correction area ms (such as area 503 in Figure 6 ) of the same size as the target convolution kernel at any position of the input image. It should be noted that the sizes of the area 501 and the area 503 are both 4*4, but the pixel points included in the area 501 are different from the pixel points included in the area 503.
[0112] The pixel points of the correction area ms and the pixel points of the area md corresponding to the second numerical box are taken as a union (ms∪md), and the irregular area 504 formed by the pixel points contained in the union is obtained.
[0113] Among the irregular area 504, all the pixel points included in the area 502 and the area 503 are included. The values corresponding to each pixel point in the irregular area 504 are all 1. In this way, the area where the second numerical box is located after correction (the irregular area 504) is obtained, and the values of all the pixel points in the area are 1.
[0114] In some embodiments, the area 505 corresponding to the minimum circumscribed rectangle of the irregular area 504 can be taken as the complete area corresponding to the second numerical box. The values corresponding to the pixel points in the area 505 that do not belong to the area 504 are set to 0. Thus, a regular area 505 including pixel points with values of 0 or 1 is obtained. The regular area 505 can form a matrix, and the matrix can be taken as a target convolution kernel. The range of the matrix is the reference range of the target convolution kernel (reference Figure 6 is given a schematic diagram of the target convolution kernel).
[0115] In the embodiment, the irregular region 502 is a region in the input image corresponding to the first numerical block in the prediction output image after inverse affine transformation; the region 503 is a region in the original input image with the same size as the convolution kernel. The intersection of the pixels in the irregular region 502 and the region 503 forms an irregular region 504. The formation of the irregular region 504 takes into account the first numerical block in the prediction output image and the second numerical block in the input image. The irregular region 504 forms a corresponding relationship with the prediction output image and the input image, and can represent the image features of the prediction output image and the input image. The reference size of the target convolution kernel is obtained based on the irregular region 504 (or the regular region 505), which is more accurate and reliable.
[0116] In some other possible embodiments, the electronic device can also correct the region where the second numerical block is located, and obtain an N times region of the region where the second numerical block is located as the corrected region, taking the geometric center of the region where the second numerical block is located as the midpoint. For example, N times can be 2 times. For example, the region where the second numerical block is located is m, and a 2m region (for example, a rectangular region) is obtained as the corrected region where the second numerical block is located, taking the geometric center of the region where the second numerical block is located. For example, the region where the second numerical block is located contains k pixel points, and a region formed by 2k pixel points is obtained as the corrected region where the second numerical block is located, taking the geometric center of the region where the second numerical block is located.
[0117] S1023, the electronic device obtains a first position block of the input image.
[0118] The position block refers to a region covered by the target convolution kernel when the target convolution kernel slides on the input image in the convolution operation or the affine transformation process. The size of the position block is consistent with the size of the numerical block and the size of the target convolution kernel. When the size of the target convolution kernel is 2*2, the size of the first position block is also 2*2.
[0119] In the embodiment, the first position block includes third pixel points in the input image, and each third pixel point corresponds to a first coordinate. For example, the first position block includes the first coordinates of the third pixel points (pixel point 5, pixel point 6, pixel point 7, and pixel point 8) in the input image in the position block. For example, the coordinates of the pixel point 5 are [x5, y5], the coordinates of the pixel point 6 are [x6, y6], the coordinates of the pixel point 7 are [x7, y7], and the coordinates of the pixel point 8 are [x8, y8].
[0120] S1024, the electronic device performs affine transformation on the first position box to obtain second coordinates of the fourth pixel points corresponding to each third pixel point in the first position box.
[0121] In this embodiment, after obtaining the first position box of the input image, the electronic device calculates the coordinates of the fourth pixel points corresponding to each third pixel point in the first position box based on the affine transformation matrix. The fourth pixel points can form a second position box. The second position box includes the fourth pixel points (pixel point 5', pixel point 6', pixel point 7', and pixel point 8') corresponding to the third pixel points (pixel point 5, pixel point 6, pixel point 7, and pixel point 8) of the input image, and each fourth pixel point corresponds to a second coordinate.
[0122] For example, the coordinate of pixel point 5 is [x5, y5], and after affine transformation, the corresponding pixel point 5' is obtained, and its coordinate is [x5', y5']; the coordinate of pixel point 6 is [x6, y6], and after affine transformation, the corresponding pixel point 6' is obtained, and its coordinate is [x6', y6']; the coordinate of pixel point 7 is [x7, y7], and after affine transformation, the corresponding pixel point 7' is obtained, and its coordinate is [x7', y7']; the coordinate of pixel point 8 is [x8, y8], and after affine transformation, the corresponding pixel point 8' is obtained, and its coordinate is [x8', y8'].
[0123] S1025, the electronic device obtains the weight corresponding to each fourth pixel point to obtain a weight matrix.
[0124] Specifically, in this embodiment, the electronic device can use a traditional interpolation weight calculation algorithm to calculate the distance between each fourth pixel point and the preset mapping point according to the coordinates of each fourth pixel point and the coordinates of the preset mapping point, and determine the weight of each fourth pixel point based on the distance between each fourth pixel point and the preset mapping point. The closer the distance between the fourth pixel point and the preset mapping point, the greater the weight corresponding to the fourth pixel point; the farther the distance between the fourth pixel point and the preset mapping point, the smaller the weight. For example, the interpolation weight calculation algorithm can be bicubic interpolation algorithm, Mitchell algorithm, Lanczos algorithm, etc. This embodiment uses the traditional interpolation weight calculation method, and the algorithm principle is not described here.
[0125] After obtaining the weight corresponding to each fourth pixel point, the weights of all fourth pixel points form a weight matrix corresponding to the target convolution kernel.
[0126] In this embodiment, the weight of the fourth pixel point is determined based on the second position box obtained by affine transformation of the first position box of the input image. The second position box is a position box in the predicted output image. Therefore, the weight of the target convolution kernel obtained based on the second position box can be adapted to the predicted output image (output image).
[0127] It can be understood that the execution order of the above steps S1021-S1025 is not limited.
[0128] In a possible embodiment, after the electronic device obtains the second position box of the predicted output image, the position of the first numerical box in the predicted output image can be determined according to the coordinate range of the pixel points in the second position box. For example, the coordinate range is determined according to the coordinate position of the first pixel point and the coordinate position of the last pixel point in the second position box. All the pixel points in the coordinate range are selected to form the first numerical box.
[0129] That is, the pixel points contained in the first numerical box and the second position box can be the same, except that each pixel point in the first numerical box corresponds to a first value (for example, 1), and each pixel point in the second position box corresponds to a coordinate.
[0130] S1026, the electronic device point-multiplies the reference size of the target convolution kernel indicated by the second numerical box and the weight matrix of the target convolution kernel to obtain the target convolution kernel.
[0131] Wherein, the second numerical box indicates the reference size of the target convolution kernel, and the numerical value of each pixel point in the second numerical box includes 1 (first value) or 0 (second value), which can form a matrix representing the reference size of the target convolution kernel. The numerical value of each second pixel point (input image) in the matrix of the reference size is point-multiplied with the weight of each fourth pixel point (predicted output image) in the weight matrix, so that the target convolution kernel adapted to the input image and the predicted output image can be obtained.
[0132] S1027, the electronic device convolves the input image according to the target convolution kernel to obtain an intermediate image.
[0133] In this embodiment, the electronic device convolves the input image based on the target convolution kernel to obtain a smoothed intermediate image.
[0134] In some other possible embodiments, the reference Figure 7 , Figure 7 A flowchart for obtaining a target convolution kernel is given, which includes:
[0135] S301, the electronic device obtains a first numerical box of a predicted output image.
[0136] Reference the above step S1021, which will not be repeated here.
[0137] S302, the electronic device performs reverse affine transformation on the first value box to obtain a second value box.
[0138] Reference the above step S1022, which will not be repeated here.
[0139] S303, the electronic device obtains the target convolution kernel through the reference range of the target convolution kernel and the preset weight of the target convolution kernel.
[0140] The preset weight can be a weight of the convolution kernel set according to experience.
[0141] The target convolution kernel obtained through S303 is adapted to the predicted output image (output image), and based on the target convolution kernel, the input image is convolved to obtain an intermediate image, which can also solve the contradiction between smoothing processing and image quality; and the intermediate image is subjected to affine transformation to obtain a final output image, which can also solve the problem of jaggies in the final output image.
[0142] In some other feasible embodiments, the reference Figure 8 , Figure 8 Another flowchart for obtaining a target convolution kernel is given, including:
[0143] S401, the electronic device obtains a first position box of an input image.
[0144] Reference the above step S1023, which will not be repeated here.
[0145] S402, the electronic device performs affine transformation on the first position box to obtain the second coordinates of the fourth pixel points corresponding to each third pixel point in the first position box.
[0146] Reference the above step S1024, which will not be repeated here.
[0147] S403, the electronic device obtains the weights corresponding to each fourth pixel point to obtain a weight matrix.
[0148] Reference the above step S1025, which will not be repeated here.
[0149] S404, the electronic device obtains the target convolution kernel through the weight matrix of the target convolution kernel and the preset range of the target convolution kernel.
[0150] The preset range can be a range of the convolution kernel set according to experience.
[0151] The target convolution kernel obtained through S404 is adapted to the input image, and the input image is convolved based on the target convolution kernel to obtain an intermediate image. The intermediate image can also solve the contradiction between smoothing processing and image quality. Furthermore, the intermediate image is subjected to affine transformation to obtain a final output image, which can also solve the sawtooth problem in the final output image.
[0152] In the embodiment, first, the reference range of the target convolution kernel adapted to the input image is obtained based on the inverse affine transformation of the first numerical box of the predicted output image. The weight matrix of the target convolution kernel adapted to the predicted output image is obtained based on the affine transformation of the first position box of the input image. The intermediate image obtained by convolving the input image based on the target convolution kernel adapted to the input image and the output image has higher image quality than the intermediate image obtained based on the target convolution kernel adapted to the input image or the output image. The contradiction between smoothing operation and image quality can be solved by smoothing the input image through the target convolution kernel. In addition, the above embodiment can complete the convolution operation on the discrete image (input image) through one convolution kernel, without the need to construct multiple different convolution kernels for all pixel points in the discrete image. The convolution processing through one convolution kernel involves less calculation data, realizes fast convolution operation, and improves the efficiency of the entire affine transformation.
[0153] In some other possible embodiments, the electronic device can also perform smoothing processing on the input image through other methods to obtain an intermediate image.
[0154] Considering that the sawtooth problem of the image subjected to affine transformation may be caused by the discontinuity of the pixels in the image in the downsampling process, which leads to the jump problem between the pixels of the output image. Therefore, the probability of the sawtooth problem of the output image can be reduced by expanding the downsampling kernel so that the pixel result does not have obvious jump during each downsampling.
[0155] wherein, Figure 9 A result diagram of downsampling by using a normal downsampling kernel is given. As shown in Figure 9 , the left image is an input image. Taking downsampling processing by using a downsampling kernel with a size of 4*4 as an example, the input image is subjected to downsampling processing based on the downsampling kernel, and the output result after downsampling is the right image. As shown in Figure 9 , it can be seen that the pixels in the output result after downsampling have jump in pixel value, which causes the sawtooth problem of the output image.
[0156] In some embodiments, as shown in Figure 10 , the reference Figure 10A schematic diagram of the result of downsampling after expanding the downsampling kernel is given. By expanding the downsampling kernel (from 4*4 to 12*12), the pixel value change of each pixel point is relatively smooth at each downsampling, and the pixel value of the pixel point in the output result after downsampling will not jump obviously.
[0157] Specifically, in this embodiment, by expanding the downsampling kernel, the problem of pixel value jumping of the pixel point in the output image can be avoided by using the expanded downsampling kernel for downsampling processing in the affine transformation of the input image, and to some extent, the problem of jaggies in the output image can also be avoided.
[0158] In some other possible embodiments, the intermediate image can also be obtained by performing upsampling processing on the input image based on a preset upsampling kernel in the affine transformation, and then performing downsampling processing on the intermediate image in the non-affine transformation. The upsampling processing in the affine transformation can effectively solve the problem of jaggies in the output image caused by the pixel value jumping of the pixel point.
[0159] It can be understood that the above specific algorithms of using the downsampling kernel for downsampling processing and using the upsampling kernel for upsampling processing can use traditional algorithms. This embodiment does not repeat the algorithms, and the focus is on the anti-jaggies effect formed by the smoothing processing before the affine transformation of the input image.
[0160] Compared with expanding the downsampling kernel and performing upsampling processing in the affine transformation, the method of obtaining the intermediate image by performing convolution on the input image based on the target convolution kernel provided in steps S1021-S1027, and the method of obtaining the intermediate image by performing convolution on the input image based on the target convolution kernel obtained through steps S301-S303 and steps S401-S404, can all achieve a smaller calculation amount, can solve the contradiction between image smoothing processing and calculation amount, can shorten the time for obtaining the intermediate image, and can improve the efficiency of obtaining the intermediate image.
[0161] S103, the electronic device performs affine transformation on the intermediate image according to the affine transformation matrix to obtain an output image.
[0162] In this embodiment, after the electronic device obtains the intermediate image, the electronic device performs affine transformation on the intermediate image based on the affine transformation matrix to obtain a final output image. Since the intermediate image is a smoothed image, performing affine transformation on the smoothed image can avoid the problem of pixel value jumping in the output image, that is, solve the problem of jaggies in the output image, and achieve the anti-jaggies effect of the output image.
[0163] Exemplarily, reference is made to Figure 11 , Figure 11A comparison diagram of the image with the sawtooth problem and the image without the sawtooth problem is given. In the image 1 obtained after the traditional affine transformation operation, the eyebrows (line edges) of the person have a serious sawtooth problem, which affects the quality of the image. After the image processing method provided in the embodiment is used to perform smoothing processing on the input image, and further affine transformation operation is performed on the intermediate image after the smoothing processing, the line at the eyebrows of the person in the image 2 obtained after the affine transformation operation is smooth and has no sawtooth problem.
[0164] It can be understood that the image processing method provided in the embodiment is for the affine transformation operation. In the affine transformation operation, the input image is smoothed to obtain an intermediate image, and then the intermediate image is subjected to affine transformation to obtain an output image after the affine transformation. In any scenario involving affine transformation, the image processing method provided in the embodiment can be used to perform affine transformation processing to obtain an affine transformation result with high image quality and no sawtooth problem.
[0165] Exemplarily, the scenario involving affine transformation includes a scenario with only affine transformation operation. For example, a scenario in which at least one of scaling, rotation, translation, and clipping is performed on an input image. More specifically, for example, when image 1 is to be clipped and pasted on the layer where image 2 is located, rotation or clipping operation of image 1 is involved. At this time, the image processing method provided in the application can be called to perform rotation, clipping, and other processing on image 1, so as to avoid the sawtooth problem when image 1 is pasted on image 2.
[0166] Exemplarily, the scenario involving affine transformation includes a scenario with affine transformation operation and other image processing operations. For example, a scenario in which portrait enhancement is performed on an input image. When image 1 is subjected to portrait enhancement, if the portrait position prior operation is involved, the portrait image region is subjected to affine transformation operation such as rotation, clipping, and scaling. At this time, the image processing method provided in the application can be called to perform rotation, clipping, scaling, and other processing on image 1. Thus, image 1 after the processing is subjected to portrait enhancement, so as to avoid the sawtooth problem of the image after the portrait enhancement. The application scenario of the image processing method is not limited in the embodiment.
[0167] In this embodiment, the electronic device performs convolution processing on the input image by using the target convolution kernel calculated from the input image and the predicted output image. In this embodiment, the first numerical box of the predicted output image is subjected to reverse affine transformation to obtain a reference size of the target convolution kernel adapted to the input image, and the first position box of the input image is subjected to affine transformation to obtain a weight matrix of the target convolution kernel adapted to the predicted output image. Thus, the target convolution kernel is obtained based on the reference size and the weight matrix of the target convolution kernel. The target convolution kernel is adapted to the input image and the predicted output image (output image). The image convolution processing based on the target convolution kernel can effectively solve the contradiction between image smoothing and image anti-aliasing, so that the anti-aliasing effect is better. In addition, the convolution based on the convolution kernel plays a role in smoothing the input image, and compared with other smoothing processing such as up-sampling of the input image, the calculation amount can be saved, the efficiency of the smoothing processing is improved, and the processing efficiency of the entire affine transformation is improved.
[0168] The image processing method provided in the embodiments of the present application can be applied to an electronic device. The electronic device can be a portable computer (such as a mobile phone), a tablet computer, a notebook computer, a personal computer (PC), a wearable electronic device (such as a smart watch), an augmented reality (AR) \ virtual reality (VR) device, a vehicle-mounted computer, and the like. The specific form of the electronic device is not specially limited in the following embodiments.
[0169] Figure 12 A structural schematic diagram of an electronic device 100 is shown.
[0170] The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charge management module 140, a power management module 141, a battery 142, a camera 193, a display screen 194, and the like.
[0171] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0172] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.
[0173] Among them, the controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of instruction fetching and instruction execution.
[0174] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. Avoiding repeated access, reducing the waiting time of the processor 110, thus improving the efficiency of the system.
[0175] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0176] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a limitation on the structure of the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection modes or a combination of multiple interface connection modes in the above embodiments.
[0177] The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input through a wireless charging coil of the electronic device 100. The charging management module 140 can charge the battery 142 while also supplying power to the electronic device through the power management module 141.
[0178] The power management module 141 is configured to connect the battery 142 and the charging management module 140 to the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 can also be configured to monitor parameters such as the battery capacity, the battery cycle count, the battery health status (leakage, impedance), etc. In some other embodiments, the power management module 141 can also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.
[0179] The electronic device 100 can implement the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is configured to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.
[0180] The display screen 194 is configured to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Mini led, Micro Led, Micro-oLed, quantum dot light-emitting diodes (QLED), etc. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1.
[0181] In this embodiment, the display screen 194 can display the output image after image processing. In specific scenarios, the display screen 194 can also display pictures in a gallery, an editing interface for image processing, etc.
[0182] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0183] ISP is used to process the data feedback by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and is converted into a visible image. ISP can also optimize the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, ISP can be provided in the camera 193.
[0184] The camera 193 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, or other format image signal. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0185] In this embodiment, the user can take pictures through the camera 193, and after taking pictures, the user can perform image processing (image enhancement, image cropping, image stitching) on the captured images through an image processing application.
[0186] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize data storage functions. For example, music, video, and other files can be saved in the external memory card.
[0187] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0188] The image processing methods provided in the above embodiments can be applied to different image processing application scenarios. The following describes the application of the image processing methods through specific applicable scenarios.
[0189] Taking a mobile phone as an example, the electronic device includes an application for image processing. For instance, the application could be a gallery, which could invoke image processing methods to perform image processing. For example, refer to... Figure 13 , Figure 13 A scene diagram of an image processing method is given.
[0190] The mobile phone can respond to the user's image selection and editing operations in the gallery interface, displaying an image editing interface 901. The image editing interface 901 includes multiple image processing controls, each corresponding to an image processing algorithm. In this embodiment, the "Perfect Face" button (…) Figure 13 Button 902 shown corresponds to the image processing method provided in this embodiment. "Perfect Face" is used to perform portrait enhancement processing on faces in an image (the portrait enhancement processing includes face position priors and involves affine transformation processing). During the affine transformation of the image, applying the image processing method provided in this embodiment can avoid jagged edges in the output image of the "Perfect Face" function.
[0191] The mobile phone enters a detailed editing interface 903 of the "perfect face" in response to a selection operation of the user on the "perfect face" button 902. The detailed editing interface 903 can include a face mark 903 detected on the selected picture. The mobile phone calls a portrait enhancement algorithm and an image processing algorithm to process the picture when the mobile phone receives an operation of the user on the "confirm" control 904. The mobile phone displays an output image 905 in the detailed editing interface 903 after completing the image processing of the picture. The mobile phone saves the output image 905 to a gallery or other designated storage space in response to a selection operation of the user on the "save" button.
[0192] The image processing method is introduced by taking an electronic device with an Android system as an example and in combination with various modules included in a software architecture of the electronic device.
[0193] Figure 14 is a software structure block diagram of the electronic device 100 of the embodiment of the present application.
[0194] The layered architecture divides the software into several layers, each of which has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, an application layer, an application framework layer, and a kernel layer.
[0195] The application layer can include a series of application packages.
[0196] As shown in Figure 14 , the application packages can include a camera application, a gallery application, and other applications. In addition, the applications can also include third-party applications for image processing and the like.
[0197] In the embodiment, the electronic device can display an image browsing interface of the gallery application on the display interface in response to an operation of the user entering the gallery application from the camera application, and display a picture editing interface in response to a selection operation of the user on a picture in the gallery application. Alternatively, the electronic device can display an image browsing interface of the gallery in response to an operation of the user opening the gallery application, and display a picture editing interface in response to a selection operation of the user on a picture in the gallery application.
[0198] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. As shown in Figure 14 , the application framework layer includes a preset interface through which the camera application / gallery application can communicate data / instructions with the image processing module.
[0199] For example, the gallery application can send an image processing request to the image processing module in the kernel layer through the preset interface in response to the selection operation of the user on the image processing control in the picture editing interface, and perform image processing; after the image processing module obtains the output image, the output image is returned to the gallery application through the preset interface, and the gallery application displays the output image on the display interface.
[0200] In addition, the application framework layer can further include a window manager, a content provider, a view system, a resource manager, a notification manager, etc.
[0201] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, and intercept the screen, etc. The content provider is used to store and obtain data, and enable the data to be accessed by the application program. The data can include videos, image albums, etc. The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The resource manager provides various resources for the application program, such as localized strings, icons, pictures, layout files, video files, etc. The notification manager enables the application program to display notification information in the status bar, which can be used to convey notification type messages, which can automatically disappear after a short stay without user interaction.
[0202] The image processing module and the media library are deployed in the kernel layer. The media library can support playback and recording of various commonly used audio, video formats, and static image files, etc.
[0203] The image processing module is used to execute the image processing method provided in the embodiments of the present application. When receiving the image processing request conveyed by the upper layer application through the preset interface, the corresponding input image is obtained for image processing; after obtaining the output image, the output image is fed back to the upper layer application for display through the preset interface.
[0204] The application of the image processing method in the embodiments of the present application is introduced in combination with the execution timing of each module of the electronic device. Exemplarily, the execution timing diagram of each module of the electronic device given in Figure 15
[0205] Exemplarily, the camera application collects an image in response to a user shooting operation, and stores the collected image into the gallery application. The camera application invokes the gallery application in response to a user selection of the gallery, and the electronic device displays an image browsing interface of the gallery application. The gallery application displays an editing interface of the picture in response to a user selection of the picture. Meanwhile, the gallery application sends an image processing request to the image processing module, so that the image processing module acquires the user-selected picture for image processing. After the image processing module completes the image processing of the user-selected picture to obtain an output picture, the image processing module can store the output picture into the media library. The media library returns the output picture to the upper application (here, the gallery application) after receiving the output picture. The gallery application displays the output picture after receiving the output picture.
[0206] Exemplarily, if the gallery application receives a user selection of the "perfect face" button, the gallery application sends an image processing request (for example, a portrait enhancement request) to the image processing module through a preset interface. The image processing module acquires the user-selected picture as an input picture in response to the image processing request, and performs portrait enhancement processing on the input picture. In the process of performing the portrait enhancement processing, a position prior processing of a face image region in the input picture is involved, and the position prior processing involves an affine transformation of the input picture. The image processing module can call the image processing method provided in the embodiments of the present application to perform affine transformation processing on the input picture. Specifically, a target convolution kernel adaptive to the input picture and a predicted output picture (output picture) is acquired, the input picture is convolved based on the target convolution kernel to obtain an intermediate picture. The intermediate picture is subjected to affine transformation to obtain an image after affine transformation. Thus, the portrait in the image after affine transformation is enhanced based on a traditional portrait enhancement algorithm, and finally an output picture after portrait enhancement is obtained. The image processing module stores the output picture after portrait enhancement into the media library. The media library can return the output picture after portrait enhancement to the gallery application for display.
[0207] It should be noted that the personal information (for example, the images in the gallery) used in the technical solutions of the present application is limited to information that has obtained individual consent, including but not limited to informing and reminding the user to read the relevant user agreement (notification) before the user uses the function, and signing the agreement (authorization) including authorization of relevant user information.
[0208] In the technical solutions disclosed in the present application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations, and do not violate public order and good customs.
[0209] The embodiments of the present application also provide a chip system (for example, a system on a chip (SoC)), such asFigure 16 As shown in FIG. 7, the chip system includes at least one processor 701 and at least one interface circuit 702. The processor 701 and the interface circuit 702 can be interconnected by a line. For example, the interface circuit 702 can be used to receive a signal from another device (e.g., a memory of an electronic device). For another example, the interface circuit 702 can be used to send a signal to another device (e.g., a processor 701 or a camera of an electronic device). Illustratively, the interface circuit 702 can read an instruction stored in a memory and send the instruction to the processor 701. When the instruction is executed by the processor 701, the electronic device can perform various steps in the above-described embodiments. Of course, the chip system can also include other discrete devices, which are not limited in the embodiments of the present application.
[0210] The embodiments of the present application also provide a computer readable storage medium, which includes computer instructions, when the computer instructions are run on the above-described electronic device, the electronic device performs various functions or steps performed by the electronic device 100 in the above-described method embodiments.
[0211] The embodiments of the present application also provide a computer program product, when the computer program product is run on a computer, the computer performs various functions or steps performed by the electronic device 100 in the above-described method embodiments. For example, the computer can be the above-described electronic device 100.
[0212] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-described division of the functional modules is taken as an example for illustration, and in actual application, the above-described functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0213] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the above-described device embodiments are only illustrative, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed ones can be through some interfaces, indirect coupling or communication connection between devices or units, which can be electrical, mechanical or other forms.
[0214] The units described as separate components may or may not be physically separate, and the components displayed as units may be a physical unit or multiple physical units, that is, may be located in one place, or also can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0215] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0216] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical scheme of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing an apparatus (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0217] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An image processing method, characterized by, The method comprises: An electronic device acquires an input image; The electronic device convolves the input image based on a target convolution kernel to obtain an intermediate image; the target convolution kernel is a convolution kernel adapted to the input image and / or a predicted output image; the predicted output image is an output image obtained after affine transformation of the input image; The electronic device performs affine transformation on the intermediate image to obtain an output image.
2. The method of claim 1, wherein, The method further comprises: The electronic device acquires a reference range of the target convolution kernel through a preset size numerical box in the predicted output image; the reference range is used to indicate the range of the convolution kernel adapted to the input image; The electronic device acquires a weight matrix of the target convolution kernel through a preset size position box in the input image; the weight matrix is used to indicate the weight of the convolution kernel adapted to the predicted output image; The electronic device acquires the target convolution kernel through the reference range of the target convolution kernel and the weight matrix of the target convolution kernel.
3. The method of claim 2, wherein, The electronic device acquires the reference range of the target convolution kernel through a preset size numerical box in the predicted output image, comprising: The electronic device selects a first numerical box of the preset size in the predicted output image, and sets a plurality of first pixel points included in the first numerical box to a first value; The electronic device obtains a second numerical box in the input image by performing inverse affine transformation on the first numerical box; the second numerical box includes a plurality of second pixel points, and the values of the second pixel points included in the second numerical box correspond to the first value; The electronic device acquires the minimum bounding rectangle of the second numerical box, and sets the values of the pixel points in the minimum bounding rectangle that do not belong to the second numerical box to a second value; Wherein, the pixel points contained in the minimum bounding rectangle and the values of each pixel point form the reference range of the target convolution kernel.
4. The method of claim 3, wherein, In the case that the number of the second pixel points included in the second numerical box is less than a preset number threshold, the acquisition of the minimum bounding rectangle of the second numerical box comprises: The electronic device performs regional correction on the region corresponding to the second numerical box to obtain the minimum bounding rectangle of the corrected region.
5. The method of claim 4, wherein, The electronic device performs regional correction on the region corresponding to the second numerical box, comprising: The electronic device selects a correction region of the preset size in the input image; The electronic device corrects the region corresponding to the second numerical box through the correction region; the corrected region is the region formed by the pixel points included in the correction region and the pixel points included in the second numerical box.
6. The method of claim 2, wherein, The electronic device acquires the weight matrix corresponding to the target convolution kernel through the preset size position box in the input image, comprising: The electronic device selects a first position box of the preset size in the input image; the first position box includes a plurality of third pixel points; The electronic device obtains a second position box mapped in the predicted output image by performing affine transformation on the first position box; the second position box includes a plurality of fourth pixel points; The electronic device obtains distances between each fourth pixel point and a preset mapping point based on coordinates of each fourth pixel point in the second position box and coordinates of the preset mapping point; The electronic device determines weights corresponding to each fourth pixel point based on the distances between each fourth pixel point and the preset mapping point; The weights of all fourth pixel points form the weight matrix.
7. An image processing method characterized by, The method comprises: The electronic device displays a first interface in response to a first operation of a user; the first interface includes a first image; The electronic device displays an image editing interface corresponding to the first image in response to an editing operation of the user on the first image; the image editing interface includes an image processing control; The electronic device performs image processing on the first image to obtain a second image in response to an operation of the user on the image processing control; the image processing at least includes affine transformation processing on the first image by a target convolution kernel; the target convolution kernel is a convolution kernel adapted to the first image and / or a predicted output image; the predicted output image is an output image obtained after affine transformation on the first image; The electronic device displays the second image in the image editing interface.
8. The method of claim 7, wherein, The affine transformation processing on the first image by the target convolution kernel to obtain the second image comprises: The electronic device performs convolution on the first image based on the target convolution kernel to obtain an intermediate image after convolution; The electronic device performs affine transformation on the intermediate image to obtain the second image.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program comprises instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1-8. The processor executes the computer program to implement the steps of the method of any one of claims 1-8.
10. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the steps of the method of any one of claims 1-8.
11. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the steps of the method of any one of claims 1-8.
Citation Information
Patent Citations
Method and device for generating sharp image based on blurred image
CN105469363A
Image processing method, electronic equipment and computer readable storage medium
CN117729445A