Image processing method, and device and storage medium

By using high-dimensional hyperplanes in the cutout technology to define the distribution range of background colors and determine their probability based on the distance between pixel points and hyperplanes, the problems of unnatural edge transition and nonlinear color distribution in the prior art are solved, and a more accurate and natural cutout effect is achieved.

WO2025111789A1PCT designated stage expired Publication Date: 2025-06-05SHENZHEN HOLLYLAND TECH CO LTD

Patent Information

Application Number
PCT/CN2023/134666
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing cutout technology is difficult to accurately pinch the foreground and background in the case of unnatural edge transitions and nonlinear color distribution, resulting in poor image effects.

Method used

By predetermining a high-dimensional hyperplane as the classification boundary of background color and non-background color, combining the classification boundary to define the distribution range of background color, and determining the probability that it is the foreground area based on the distance between the color coordinates of the pixel point and the hyperplane, the cutout process is performed.

Benefits of technology

It achieves a more accurate definition of the distribution range of background colors, improves the accuracy of the foreground and background, and the edge transition of the resulting cutout image is more natural.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023134666_05062025_PF_FP_ABST
    Figure CN2023134666_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present description are an image processing method, and a device and a storage medium. The method comprises: acquiring an image to be subjected to matting; determining the distance between color coordinates of each pixel point in said image and a predetermined hyperplane, wherein the hyperplane is a classification boundary of a background color and a non-background color of said image; and on the basis of the distance which corresponds to each pixel point, determining the probability of each pixel point being a foreground area of the said image, so as to perform matting processing on said image on the basis of the probability. In the embodiments of the present application, the distribution range of a background color accurately defined by means of a high-dimensional hyperplane can be used, so that a foreground area and a background area, which are obtained by means of matting, are also more accurate, thereby obtaining a matted image which is natural in transition.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, device and storage medium Technical Field

[0001] This specification relates to the field of image processing technology, and in particular to an image processing method, device, and storage medium. Background Art

[0002] Image cutout technology is widely used in various fields. By identifying the foreground and background areas of the image to be cutout, and then replacing the background area with the user's desired background, special effects images or images in specific scenes can be obtained without the need for additional dedicated scenes.

[0003] Among the current cutout technologies, some technologies pre-set the value range of each channel of the background color, and then determine whether the pixel value of each channel of each pixel point in the image to be cutout is within the value range of the channel corresponding to the background color. If so, the pixel point is determined as the background, otherwise, it is determined as the foreground. This cutout technology adopts a one-size-fits-all approach, resulting in unnatural edge transitions in the final cutout image.

[0004] There are also some technologies in which the user sets a reference color representing the background color in a certain color space, and then calculates the distance between the color of each pixel in the image to be cut out and the reference color. Based on the distance and a pre-set similarity threshold, the probability (i.e., the Alpha value) that each pixel in the image to be cut out is the foreground area is determined, and the Alpha image corresponding to the image to be cut out can be obtained, and then the image is cut out based on the Alpha image. The edge of the image cut out in this way is smoother and the processing speed is faster, but because the color distribution is nonlinear, the range defined by the reference color and a single similarity threshold (i.e., the circular range) cannot cover the entire color distribution of the background color. Therefore, if the similarity threshold is set too low, the background cannot be completely cut out. If the similarity threshold is set too high, some foreground areas that are close to the background color will also be cut out, affecting the effect of the final image after replacing the background.

[0005] It can be seen that it is necessary to provide a cutout solution that can not only more accurately cut out the foreground and background, but also make the edge transition of the cutout image more natural.

[0006] Summary of the Invention

[0007] Based on this, this specification provides an image processing method, device and storage medium.

[0008] According to a first aspect of the embodiments of this specification, there is provided an image processing method, the method comprising:

[0009] Get the image to be cut out;

[0010] Determining the distance between the color coordinates of each pixel in the image to be cutout and a predetermined hyperplane, wherein the hyperplane is a classification boundary between the background color and the non-background color of the image to be cutout;

[0011] The probability that each pixel point is a foreground area of ​​the image to be cut out is determined based on the distance corresponding to each pixel point, so as to perform cutout processing on the image to be cut out based on the probability.

[0012] According to a second aspect of the embodiments of this specification, an electronic device is provided, comprising a processor, a memory, and computer instructions stored in the memory, wherein the processor implements the method mentioned in the first aspect when executing the computer instructions.

[0013] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed, the method mentioned in the first aspect is implemented.

[0014] By applying the embodiment scheme of this specification, the present application can predetermine a high-dimensional hyperplane, use the hyperplane as the classification boundary of background color and non-background color, to more accurately define the distribution range of the background color, and then determine the distance between the color coordinates of the pixel points in the image to be cut out and the hyperplane, and determine the probability (i.e., Alpha value) of each pixel point being the foreground area based on the distance. Compared with using a circular or spherical area determined by a reference color and a preset single similarity threshold to define the range of the background color, the embodiment of the present application can use a high-dimensional hyperplane and a similarity threshold to define the range of the background color, so that the defined background color range is more accurate and can fully cover the distribution range of the background color, and thus the foreground area and background area cut out are also more accurate, resulting in a cutout image with a natural transition.

[0015] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.

[0017] FIG. 1 is a schematic diagram of defining a background color distribution range using a reference color and a similarity threshold according to an embodiment of the present disclosure.

[0018] FIG. 2 is a schematic diagram of defining a background color distribution range using a reference color and a similarity threshold according to another embodiment of the present disclosure.

[0019] FIG3 is a flow chart of an image processing method according to an embodiment of the present disclosure.

[0020] FIG4 is a schematic diagram of classifying positive and negative samples using a hyperplane according to an embodiment of this specification.

[0021] FIG5 is a schematic diagram of obtaining first sample pixels by interval sampling according to an embodiment of this specification.

[0022] FIG6 is a schematic diagram of the logical structure of an electronic device according to an embodiment of this specification. DETAILED DESCRIPTION

[0023] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this specification, as detailed in the appended claims.

[0024] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. As used in this specification and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0025] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information without departing from the scope of this specification. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0026] Image cutout technology is widely used in various fields. By identifying the foreground and background areas of the image to be cutout, and then replacing the background area with the user's desired background, special effects images or images in specific scenes can be obtained without the need for additional dedicated scenes.

[0027] The key to the cutout technology is to accurately identify the foreground and background areas from the image to be cutout. For any image to be cutout I, the subject part can be called the foreground F, and the rest is the background B. Then the image to be cutout I can be regarded as a weighted fusion of F and B: I = Alpha*F+(1-Alpha)*B, where Alpha is a continuous value between [0, 1], which can be understood as the probability that each pixel belongs to the foreground, or the opacity of the foreground image (that is, the higher the probability that the pixel belongs to the foreground, the more opaque it is). The task of cutout is to find the appropriate weight Alpha for each pixel, that is, to obtain an Alpha image corresponding to the image to be cutout. The pixel value of each pixel in the Alpha image represents the probability that the pixel at the corresponding pixel position of the image to be cutout is the foreground. Then, the foreground area can be cut out from the image to be cutout based on the Alpha image.

[0028] Among current image cutout technologies, some pre-set the value ranges of each background color channel, then determine whether the pixel values ​​of each channel of each pixel in the image to be cutout are within the value range of the background color channel. If so, the pixel is determined to be the background; otherwise, it is determined to be the foreground. For example, taking the background color as green, under the HSV color model, the value ranges of the three channels of green (H, S, and V) can be set to H∈[35°, 77°], [0.168, 1], and [0.180, 1], respectively. For the image to be cutout, it can first be converted to an HSV format image, and then the pixel values ​​of the H, S, and V channels of each pixel are determined to be within the above value ranges. If so, the pixel is determined to be the background. This one-size-fits-all approach, because there is no transition area, can easily lead to unnatural edge transitions in the cutout image.

[0029] In some other technologies, the user sets a reference color (i.e., a single color coordinate point) representing the background color in a certain color space, then calculates the distance between the color of each pixel in the image to be cut out and the reference color. Based on this distance and a pre-set similarity threshold, the probability of each pixel in the image to be cut out being the foreground area is determined, which can then be used to obtain the alpha image corresponding to the image to be cut out, and then the cutout is performed based on the alpha image. For example, taking the background color as green as an example, the reference color can be set to (R=0, G=255, B=0), and then the distance between the color of each pixel in the image to be cut out and the reference color can be calculated. For example, assuming that the color of a certain pixel is (R=20, G=30, B=40), the distance between the three-dimensional point (0, 255, 0) and the three-dimensional point (20, 30, 40) is determined. The probability of each pixel in the image to be cut out being the foreground (i.e., the alpha value) can then be determined based on this distance and a pre-set similarity threshold.

[0030] For example, the Alpha value of each pixel can be calculated using formula (1).

[0031] Wherein, d is the distance between the color of each pixel and the reference color, thres is the preset similarity threshold, and ratio is the preset smoothing coefficient.

[0032] The image extracted in this way has smoother edges and more natural transitions due to the presence of transition areas, and the processing speed is faster. However, since the color distribution is nonlinear, the range defined by the reference color and a single similarity threshold (i.e., the circular range) cannot cover the entire color distribution range of the background color. Therefore, if the similarity threshold is set too low, the background cannot be completely removed. If the similarity threshold is set too high, some foreground areas that are close to the background color will also be removed, affecting the effect of the final image after replacing the background.

[0033] It can be seen that it is necessary to provide a cutout solution that can not only more accurately cut out the foreground and background, but also make the edge transition of the cutout image more natural.

[0034] Based on this, an embodiment of the present application provides an image processing method, taking into account that the distribution of background color is usually irregular in shape. For example, if the color is represented in two-dimensional space, the distribution range of the background color is usually polygonal. If the color is represented in three-dimensional space, the distribution range of the background color is usually a polyhedron. If the alpha value of each pixel is determined based on the distance between the color of each pixel and the reference color and a pre-set similarity threshold, it is equivalent to defining the distribution range of the background color by the reference color (equivalent to the center of a circle or a sphere) and the similarity threshold (equivalent to the radius). This distribution range is circular or spherical, while the distribution range of the background color is usually an irregular polygon or a polygon. The distribution range defined in the above manner cannot accurately cover the distribution range of the background color. Therefore, if the similarity threshold is set too low, the background cannot be completely removed (as shown in FIG1, where the color is represented in two-dimensional space as an example), and if the similarity threshold is set too high, some foreground areas that are close to the background color will also be removed (as shown in FIG2, where the color is represented in two-dimensional space as an example). In order to more accurately define the distribution range of the background color, the present application can pre-determine a high-dimensional hyperplane, use the hyperplane as the classification boundary between the background color and the non-background color, and define the distribution range of the background color in combination with the classification boundary. Then, the distance between the color coordinates of the pixel points in the image to be cutout and the hyperplane is determined, and the probability of each pixel point being the foreground area (i.e., the Alpha value) is determined based on the distance.

[0035] Compared to defining the range of background color by using a circular or spherical area determined by a reference color and a preset single similarity threshold, the embodiment of the present application can use a high-dimensional hyperplane and a similarity threshold to define the range of background color, so that the defined background color range is more accurate and can fully cover the distribution range of the background color, and thus the foreground area and background area extracted are also more accurate, resulting in a cutout image with a natural transition.

[0036] The image processing method provided in the embodiments of the present application can be executed by various electronic devices, such as mobile phones, computers, various live broadcast devices, or cloud servers, etc. For example, the image processing method can be integrated into a certain cutout APP or cutout SDK. As long as the cutout APP or cutout SDK is installed on the device, the above image processing method can be implemented.

[0037] As shown in FIG3 , the image processing method may include the following steps:

[0038] S302, obtaining the image to be cut out;

[0039] In step S302, the image to be cut out can be obtained, wherein the image to be cut out can be an image in various formats, for example, it can be an image in various formats such as RGB format, YUV format, LAB format, HSV format, etc., and the embodiment of the present application does not limit this.

[0040] S304, determining the distance between the color coordinates of each pixel in the image to be cut out and a predetermined hyperplane, wherein the hyperplane is a classification boundary between the background color and the non-background color of the image to be cut out;

[0041] In step S304, in order to accurately define the distribution range of the background color of the image to be cut out, a hyperplane for representing the classification boundary of the background color and non-background color of the image to be cut out can be determined in advance. The hyperplane can be a high-dimensional hyperplane (for example, greater than or equal to 3 dimensions). Based on the hyperplane, nonlinearly distributed colors can be distinguished more accurately.

[0042] Among them, the hyperplane can be determined by using a large number of sample pixel points whose color is the background color and sample pixel points whose color is not the background color. For example, the hyperplane can be determined by various methods such as machine learning and plane fitting, and the embodiments of this application are not limited thereto.

[0043] Among them, for the same background color (that is, the value ranges of each channel corresponding to the background color are the same, that is, they can be considered to be the same background color), a hyperplane can be determined. In the subsequent cutout process, all images with the same background color can be cutout using this hyperplane. For example, for a green background, a hyperplane can be determined, and all subsequent images to be cutout with a green background can be cutout using this hyperplane.

[0044] After obtaining the hyperplane, the distance between the color coordinates of each pixel in the image to be cut out and the hyperplane can be determined. For example, assuming that the pixel values ​​of the three channels of a certain pixel in the RGB color model are: R = 20, G = 30, B = 40, then the color coordinates of the pixel can be expressed as (20, 30, 40), which means that the distance between the three-dimensional point represented by the color coordinates and the hyperplane can be determined.

[0045] S306 : Determine the probability that each pixel point is a foreground area of ​​the image to be cut out based on the distance corresponding to each pixel point, so as to perform cutout processing on the image to be cut out based on the probability.

[0046] In step S306, after determining the distance between the color coordinates of each pixel point in the image to be cut out and the hyperplane, the probability (i.e., the Alpha value) that each pixel point is the foreground area of ​​the image to be cut out can be determined based on the distance. For example, the probability can be determined based on the distance and a preset similarity threshold and smoothing coefficient, thereby obtaining the Alpha image corresponding to the image to be cut out, so that the image to be cut out can be cut out based on the Alpha image.

[0047] In some embodiments, when determining a hyperplane for representing the classification boundary between the background color and the non-background color of the image to be cutout, a plurality of sample pixels carrying labels (hereinafter referred to as first sample pixels for ease of distinction) can be obtained, wherein the label of each first sample pixel is used to indicate whether the color of the first sample pixel is the background color. Then, the plurality of first sample pixels can be used as training samples to train a preset support vector machine model, and the first sample pixel is classified by the support vector machine model to determine whether its color is the background color. Then, based on the difference between the predicted result output by the support vector machine model and the actual result indicated by the label, the model parameters of the support vector machine model are adjusted until the accuracy of the support vector machine for the pixel color classification result is higher than the preset accuracy (for example, 95%), and the trained support vector machine model is then used as the hyperplane. In order to ensure the accuracy of the trained hyperplane, the first sample pixel can include positive samples and negative samples to ensure sample diversity. As shown in Figure 4, the determined hyperplane can distinguish positive samples from negative samples with high accuracy.

[0048] Among them, the support vector machine model is the expression of the hyperplane (for example, Y=w*X+b), and the training process is the process of determining the parameters w and b in the expression, so that the final hyperplane can accurately classify background colors and non-background colors.

[0049] In some embodiments, the first sample pixel point can be represented by a first color model. When obtaining the first sample pixel point, sampling can be performed at intervals within the value range of the pixel value of each channel of the first color model to obtain multiple sample values ​​of the channel. The sample values ​​of each channel are then combined to obtain multiple first sample pixel points. For each first sample pixel point, it can be determined whether the color of the first sample pixel point is the background color, and the determination result is used as the label of the first sample pixel point.

[0050] For example, as shown in FIG5 , taking the RGB model as the first color model, each pixel in the RGB model needs to be represented by the pixel values ​​of the three channels R, G, and B, and the pixel values ​​of the three channels range from 0 to 255. Therefore, for the R channel, sampling can be performed at intervals between 0 and 255. For example, a sampling value is obtained every 5 samples, that is, 0, 5, 10, 15, etc. Similarly, for the G channel and the B channel, sampling is performed at intervals between 0 and 255 in the same manner to obtain their respective sampling values. Then, the sampling values ​​of the three channels R, G, and B are combined. For example, R=5, G=5, and B=5 are combined to obtain a first sample pixel, that is, (R=5, G=5, B=5). Then, it can be determined whether the color of the sample pixel is the background color to obtain the label of the first sample pixel.

[0051] Of course, if the first color model is an HSV model, a LAB model or a YUV model, a similar sampling method can also be used to obtain a plurality of first sample pixel points.

[0052] In order to ensure that the trained hyperplane has high accuracy, the interval step size can be smaller during interval sampling to cover as many pixels as possible.

[0053] Typically, the background color is a color distribution range. For example, when a certain color model (such as the RGB or HSV model) is used to represent the background color, the values ​​of each channel of the color model are generally not fixed values, but rather a range of values. Therefore, in some embodiments, when determining whether the color of a first sample pixel is the background color, it can be determined whether the pixel values ​​of each channel of the first sample pixel are all within the value range of the background color within the corresponding channel. If so, the color of the first sample pixel is determined to be the background color.

[0054] For example, assuming that the background color and the first sample pixel are both represented by the HSV model, and the background color is green, then the value ranges corresponding to its H, S, and V channels can be set to: H∈[35°, 77°], [0.168, 1], [0.180, 1], respectively. Assuming that a first sample pixel P1 is (H=40,S=0.2,V=0.2), the pixel values ​​of its H, S, and V channels are all within the value ranges of the channels corresponding to the background color, so the color of P1 is determined to be the background color. However, a first sample pixel P2 is (H=10,S=0.2,V=0.2), and the pixel value of its H channel is not within the value range of the H channel of the background color, so the color of P2 is determined to be a non-background color.

[0055] In some embodiments, the background color can be represented by a second color model, wherein the second color model can be the same as the first color model, for example, both are RGB models, or different, for example, the background color is represented by the HSV model and the first sample pixel is represented by the RGB model. For scenes where the first color model and the second color model are different, before determining whether the color of the first sample pixel is the background color, the color model representing the sample pixel and the color model representing the background color can be unified before making a determination. For example, taking the case where the background color is represented by the HSV model and the first sample pixel is represented by the RGB model, the background color can be converted from the HSV format to the RGB format, or the first sample pixel can be converted from the RGB format to the HSV format, to facilitate subsequent pixel value comparison.

[0056] Considering that colors in the HSV color space or LAB color space are more intuitive and easier to distinguish, while the distribution range corresponding to the background color usually requires manual setting by the user, in some embodiments, the background color can be set in the HSV color space or LAB color space. When sampling to obtain sample pixels, sampling can be performed in the RGB color space. Therefore, the first color model mentioned above can be an RGB model, and the second color model mentioned above can be an HSV model or a LAB model.

[0057] In some embodiments, assuming that the background color is represented by a second color model, when the user sets the value range of each channel of the background color in the second color model, the user can directly enter the value range of each channel based on the background to be cut out, or the user can directly specify the background color, such as red, green, or blue, and then the device executing the image processing method can automatically generate the value range of each channel based on the color selected by the user. Alternatively, in some scenarios, to facilitate user settings, the user can also directly select an area in the background of the image to be cut out, and then the device executing the image processing method can automatically determine the value range of each channel based on the color corresponding to the area selected by the user.

[0058] In some embodiments, when determining the distance between the color coordinates of each pixel in the image to be cutout and the hyperplane, the color coordinates of each pixel can be substituted into the expression of the hyperplane to obtain the distance between the color coordinates of the pixel and the hyperplane. This method is computationally intensive and time-consuming, making it unsuitable for scenarios such as live broadcasts where real-time cutout performance is critical.

[0059] In order to reduce the amount of calculation and improve the efficiency of the cutout, in some embodiments, a distance map can be first determined based on the hyperplane and a plurality of preset sample pixel points (hereinafter referred to as second sample pixel points), and the distance map is used to indicate the distance between the color coordinates of each of the plurality of sample pixel points and the hyperplane. When determining the distance between the color coordinates of each pixel point in the image to be cutout and the hyperplane, it is not necessary to substitute the color coordinates of each pixel point into the expression of the hyperplane for calculation. Instead, for the pixel point whose color coordinates in the image to be cutout are the same as the color coordinates of the second sample image pixel point, the distance corresponding to the pixel point can be directly found from the distance map. For the pixel point whose color coordinates in the image to be cutout are different from the color coordinates of the second sample image pixel point, a plurality of second sample pixel points whose color coordinates are relatively close to the color coordinates of the pixel point can be selected, and the distances of these second sample pixel points are interpolated to obtain the distance of the pixel point.

[0060] By constructing a distance map based on the hyperplane in advance, the distance from the color coordinates of each pixel in the image to be cut out to the hyperplane can be determined by table interpolation, which can reduce the amount of calculation, improve processing efficiency, and ensure the real-time performance of the cutout processing.

[0061] Among them, for the same background color (that is, the value ranges of each channel corresponding to the background color are the same, that is, they can be considered to be the same background color), a distance map can be determined. Later, during the cutout process, for all images with the same background color, this distance map can be used to determine the distance between the color coordinates of the pixel points in the image and the hyperplane. For example, for a green background, a distance map can be determined, and all subsequent images to be cutout with a green background can use this distance map to calculate the distance.

[0062] In some embodiments, the first sample pixels and the second image pixels may be the same batch of pixels. For example, when obtaining sample pixels for training a hyperplane, these sample pixels and the trained hyperplane may be used to generate a distance map. In some embodiments, the first sample pixels and the second image pixels may also be different pixels.

[0063] In some embodiments, in order to ensure that the second sample pixel points can be evenly distributed in the color space and facilitate subsequent interpolation, when determining the second sample pixel points, they can also be determined by an interval sampling method. Taking the representation of the multiple second sample pixel points by the first color model (for example, the RGB model) as an example, when determining the multiple second sample pixel points, for each channel of the first color model, interval sampling can be performed within the value range of the pixel value of the channel to obtain multiple sampling values ​​of the channel, and then the sampling values ​​of each channel are combined to obtain multiple second sample pixel points. Among them, the specific interval sampling method can be referred to the description in the above embodiment and will not be repeated here.

[0064] In some embodiments, when calculating the distance between the color coordinates of a pixel point of the image to be cut out and the hyperplane, the pixel values ​​of all channels of the pixel point can be directly used to form the color coordinates, and the distance between the color coordinates and the hyperplane can be calculated. For example, if the image to be cut out is an image in RGB format, assuming that the pixel values ​​of the three channels R, G, and B of a certain pixel point are respectively: R=10, G=20, B=30, then the distance from the three-dimensional point (10, 20, 30) to the hyperplane can be calculated. If the image to be cut out is an image in HSV format, assuming that the pixel values ​​of the three channels H, S, and V of a certain pixel point are respectively: H=10, S=0.2, V=0.3, then the distance from the three-dimensional point (10, 20, 10) to the hyperplane can be calculated.

[0065] Taking into account that for the image to be cut out represented by certain color models, some channels in the color model are channels that are not related to color, and when calculating the distance, the main goal is to determine the degree of proximity between the color of the pixel point and the background color. Therefore, only the pixel values ​​of the channels related to color will affect the degree of proximity, while the pixel values ​​of the channels that are not related to color will not affect the degree of proximity. Therefore, in some embodiments, in order to reduce the amount of calculation and improve processing efficiency, for the pixel points in the image to be cut out, the color coordinates of the pixel points can be the coordinates composed of the pixel values ​​of the target channel, and the distance between the pixel point and the hyperplane can be the distance between the coordinates composed of the pixel values ​​of the target channel of the pixel point and the hyperplane, wherein the target channel is a channel related to color.

[0066] For example, if the image to be cut out is in YUV format, the target channels are U and V channels. Assuming that the pixel values ​​of the Y, U, and V channels of a certain pixel point are: Y=10, U=20, V=30, then the distance from the two-dimensional point (20, 30) to the hyperplane can be calculated as the distance from the pixel point to the hyperplane.

[0067] If the image to be cut out is an image in HSV format, the target channels are the H and S channels; assuming that the pixel values ​​of the H, S, and V channels of a certain pixel point are: H = 10°, S = 0.2, V = 0.3, then the distance from the two-dimensional point (10, 0.2) to the hyperplane can be calculated as the distance from the pixel point to the hyperplane.

[0068] If the image to be cut out is an image in LAB format, the target channels are A and B channels. Assuming that the pixel values ​​of the L, A, and B channels of a certain pixel point are: L=10, A=20, B=30, then the distance from the two-dimensional point (20, 30) to the hyperplane can be calculated as the distance from the pixel point to the hyperplane.

[0069] If the image to be cut out is an RGB format image, the target channels are R, G, and B channels. Assuming that the pixel values ​​of the R, G, and B channels of a certain pixel point are: R=10, G=20, B=30, the distance from the three-dimensional point (10, 20, 10) to the hyperplane can be calculated.

[0070] By adopting the above method, for scenes where the image to be cutout is in HSV, YUV, or LAB format, when calculating the distance, it is only necessary to calculate the distance from the two-dimensional point to the hyperplane. Compared with taking all pixel values ​​into account, that is, calculating the distance from the three-dimensional point to the hyperplane, this calculation method can greatly reduce the amount of calculation without affecting the final cutout effect.

[0071] In some embodiments, after determining the distance between the color coordinates of each pixel in the image to be cutout and the hyperplane, the probability of each pixel being the foreground region of the image to be cutout can be determined based on the distance and preset cutout parameters. The cutout parameters can include a similarity threshold and a smoothing coefficient, which can be preset. If the distance is less than the preset similarity threshold, the color of the pixel is within the distribution range of the background color, and thus the probability is 0. If the difference between the distance and the similarity threshold is greater than 0 and less than or equal to 1, the color of the pixel is outside the distribution range of the background color and relatively close to the background color. Therefore, these pixels can serve as the transition region between the foreground and the background. The probability of this can be determined based on the difference between the distance and the similarity threshold. The larger the difference, the greater the deviation from the background color, and thus the greater the probability of being the foreground. If the difference between the distance and the similarity threshold is greater than 1, the color of the pixel is significantly different from the background color, and thus these pixels are the foreground. The probability of this can be determined based on the preset smoothing coefficient.

[0072] For example, the following formula (1) can be used to determine the probability that each pixel is a foreground area:

[0073] Wherein, d is the distance between the color of each pixel and the reference color, thres is the preset similarity threshold, and ratio is the preset smoothing coefficient.

[0074] In some embodiments, to accurately determine the classification boundary between background and non-background colors, the dimension of the hyperplane can be greater than or equal to 3. Of course, within a certain range, the higher the hyperplane dimension, the more accurate the classification boundary obtained. However, beyond this range, as the hyperplane dimension increases, its accuracy does not significantly improve. Therefore, in some embodiments, the dimension of the hyperplane does not exceed 10.

[0075] Currently, when setting cutout parameters such as the similarity threshold and smoothing coefficient, they are generally set manually by the user. For example, the user can continuously adjust the value of the cutout parameter, and the cutout software can cut out the image to be cutout based on the cutout parameters currently set by the user, and display a preview image so that the user can observe the cutout effect with the naked eye and use the value when the cutout effect is better as the final cutout parameter. This method is cumbersome, time-consuming and labor-intensive. Moreover, when the value of the cutout parameter is close to the optimal value, the change in the cutout effect is subtle and difficult to directly detect with the naked eye. Therefore, the cutout parameters set in this way are often not the optimal cutout parameters.

[0076] Therefore, it is necessary to provide a solution that is more convenient and quick and can accurately determine the optimal or better clipping parameters.

[0077] Based on this, the embodiment of the present application provides a solution for automatically determining the cutout parameters. In order to automatically determine the optimal or better cutout parameters, the applicant designed a target optimization function. The target optimization function uses the cutout parameters as variables and one or more error functions used to measure the cutout errors caused by the cutout parameters as optimization targets. The target optimization function is then optimized based on the image to be cutout, thereby automatically determining the optimized value of the cutout parameters for subsequent cutout processing. Among them, considering that for a better cutout parameter, it should meet the following conditions as much as possible:

[0078] (1) Based on the cutout parameters, the foreground and background areas of the image to be cutout should be distinguished as accurately as possible, that is, the number of misidentified pixels should be as small as possible;

[0079] (2) In order to avoid the transition area between the foreground and background areas being too large, which would affect the overall effect of the image, the area of ​​the transition area determined based on the cutout parameters should also be as small as possible, that is, the number of pixels belonging to the transition area should be as small as possible;

[0080] (3) In addition, in order to avoid incompleteness or missing foreground areas, the pixels in the transition area should be as close to the background area as possible, that is, the cumulative value of the alpha value of all pixels in the transition area (that is, the probability that the pixel belongs to the foreground area) should be as small as possible.

[0081] When setting the optimization target, in an embodiment of the present application, one or more error functions for measuring the cutout errors caused by the cutout parameters can be set based on the above three consideration dimensions, and then one or more optimization targets can be designed based on the error function to optimize the cutout parameters.

[0082] Through the method provided in the embodiment of the present application, better or optimal cutout parameters can be automatically determined without manual setting by the user, making the determination of the cutout parameters more convenient and quick, and the determined cutout parameters are more accurate, and the effect of obtaining the cutout image using the cutout parameters is better.

[0083] The method for determining the cutout parameters may include the following steps:

[0084] S702, obtaining the value range of each channel of the image to be cut out and the background color of the image to be cut out under a specified color model;

[0085] In step S702, the image to be cut out can be obtained, wherein the image to be cut out can be an image in various formats, for example, it can be an image in RGB format, YUV format, LAB format, HSV format, etc., and the embodiment of the present application does not limit this.

[0086] Considering that the background color of the image to be cut out usually falls within a color range, the value ranges of each channel of the background color under a specified color model can be predetermined. The specified color model can be an RGB model, a YUV model, an HSV model, a LAB model, etc. For example, assuming the specified color model is the HSV model and the background color is green, the value ranges corresponding to the three channels H, S, and V can be set to: H∈[35°, 77°], [0.168, 1], [0.180, 1].

[0087] Among them, the value range of the background color in each channel is determined. Its purpose is to verify the division result of the cutout parameters based on the value range of the background color when the pixels in the image to be cutout are divided into the foreground area and the background area based on the cutout parameters, so as to determine the cutout error caused by the cutout parameters.

[0088] S704: Optimize a pre-constructed target optimization function based on the image to be cutout and the value range to obtain an optimized value of a cutout parameter of the image to be cutout, wherein the cutout parameter is used to determine a probability that a pixel in the image to be cutout is a foreground area of ​​the image to be cutout, and the probability is determined based on the cutout parameter and a difference between a color of the pixel and a background color; the variable of the target optimization function is the cutout parameter, and the optimization objective of the target optimization function includes one or more of the following:

[0089] When dividing the pixels in the image to be cut out into the foreground area and the background area of ​​the image to be cut out based on the cutout parameters, the number of misclassified pixels is minimized;

[0090] Minimizing the number of pixel points in the image to be cutout that belong to the transition area between the foreground area and the background area determined based on the cutout parameters;

[0091] Minimizing the cumulative value of the probabilities of the pixels in the transition area determined based on the cutout parameters;

[0092] In step S704, after obtaining the value ranges of each channel of the image to be cut out and the background color under the specified color model, a pre-constructed target optimization function can be optimized based on the two to obtain optimized values ​​of the cutout parameters of the image to be cut out. The cutout parameters can be used to determine the probability that a pixel in the image to be cut out is in the foreground area of ​​the image to be cut out, combined with the difference between the color of the pixel in the image to be cut out and the background color. For example, the cutout parameters can be one or more of a similarity threshold and a smoothing coefficient.

[0093] The pre-built target optimization function can use the cutout parameter as a variable and one or more of the following as optimization objectives:

[0094] (1) When dividing the pixels in the image to be cut out into foreground and background regions based on the cutout parameters, the number of misclassified pixels is minimized. This optimization objective is used to measure the accuracy of the foreground and background determined based on the cutout parameters. By constraining the cutout parameters through this optimization objective, the foreground and background regions determined based on the cutout parameters can be made more accurate.

[0095] (2) Minimize the number of pixels in the transition area determined based on the cutout parameters. The transition area is the area between the foreground area and the background area in the image to be cutout. Usually, in order to ensure that the edge of the foreground area obtained by cutout can transition smoothly and more naturally, a transition area is generally set between the foreground area and the background area. Obviously, the area of ​​the transition area cannot be too large, otherwise it will affect the display effect of the entire image. Therefore, by setting this optimization goal to constrain the cutout parameters, the area of ​​the transition area can be prevented from being too large.

[0096] (3) The cumulative value of the probability that each pixel in the transition area belongs to the foreground area, determined based on the cutout parameters, is minimized. As previously mentioned, in order to ensure the integrity and non-missing nature of the foreground area, the transition area should be as close to the background area as possible. Therefore, by constraining the cumulative value of the above-mentioned probabilities of the pixels in the transition area, we can ensure that the foreground area obtained by cutout is relatively complete and free of serious missing parts.

[0097] Among them, when setting the optimization target of the target optimization function, one or more of the above optimization targets can be selected, and then the target optimization function is optimized based on the set optimization target, that is, better or optimal cutout parameters can be obtained.

[0098] In addition, in some embodiments, multiple optimization goals can also be set to constrain the cutout parameters from multiple dimensions, so that the cutout parameters ultimately automatically determined are more accurate, and cutout processing based on the cutout parameters can achieve better cutout effects.

[0099] In some embodiments, the cutout parameter may be one or more of a similarity threshold and a smoothing coefficient. For example, the smoothing coefficient may be set to a certain empirical value, and then only the similarity threshold is optimized, or both the similarity threshold and the smoothing coefficient may be optimized simultaneously.

[0100] Among them, when the difference between the color of the pixel point of the image to be cut out and the background color is less than the similarity threshold, the probability that the pixel point is the foreground area is 0. When the difference between the color of the pixel point of the image to be cut out and the background color and the similarity threshold is greater than 0 and less than or equal to 1, the probability that the pixel point is the foreground area is positively correlated with the difference between the difference and the similarity threshold. For example, the larger the difference, the greater the probability; when the difference between the color of the pixel point of the image to be cut out and the background color and the similarity threshold is greater than 1, the probability that the pixel point is the foreground area is the smoothing coefficient.

[0101] In some embodiments, when optimizing a pre-constructed target optimization function based on the image to be cut out and the above-mentioned value range to obtain the optimized value of the cutout parameter, the optimal cutout parameter can be determined through multiple iterations. For example, the initial value of the cutout parameter can be set, and then the difference between each pixel in the image to be cut out and the background color can be determined. Then, based on the difference and the initial value of the cutout parameter, the probability that each pixel in the image to be cut out belongs to the foreground area is determined, and based on the probability and the value range of the background color in each channel, the number of misclassified pixels is determined, the number of pixels belonging to the transition area is determined based on the probability, and the cumulative value of the above-mentioned probabilities of each pixel in the transition area is determined based on the probability, i.e., the values ​​of the above-mentioned three optimization objectives under the current cutout parameters. Then, the initial value of the cutout parameter can be updated based on the value of one or more of the above-mentioned three optimization objectives under the current cutout parameters, and the updated value is used as the current value of the cutout parameter, and the numerical values ​​corresponding to the above-mentioned three optimization objectives are determined again based on the current value.

[0102] Repeat the above iterative process, continuously update the current value of the cutout parameter, and replace the current value with the updated value until the preset condition for stopping the iteration is reached. The current value of the cutout parameter obtained at this time is used as the optimized value for subsequent cutout processing.

[0103] In some embodiments, if the target optimization function includes multiple optimization objectives, the target optimization function can be optimized by adopting a multi-objective gradient descent algorithm to find a Pareto optimal solution as the optimized value of the cutout parameter. Taking into account that the values ​​of different optimization objectives in the above three optimization objectives are quite different, if multiple optimization objectives are directly superimposed and the cutout parameters are optimized based on the superposition results, there may be some optimization objectives whose values ​​are much smaller than the values ​​of other optimization objectives, and the constraint effect of the optimization objective on the cutout parameters may be masked. Based on this, in the embodiment of the present application, a multi-objective gradient descent algorithm can be used to solve the target optimization function, so that in the optimization process, the influence of each optimization objective on the cutout parameters can be taken into account, so that the optimized value of the cutout parameter finally determined is more accurate, and the cutout effect based on the cutout parameter is also better.

[0104] In some embodiments, the misclassified pixels may include the following: one is that the pixel points that originally belong to the background are classified into the foreground area, and the other is that the pixel points that originally belong to the foreground are classified into the background area. Since the value range of the background color in each channel of the specified color model has been predetermined, it is possible to roughly determine whether each pixel point in the image to be cut out is the foreground or the background based on the value range. Therefore, the misclassified pixel points can be determined in combination with the value range of the background color in each channel. For example, the misclassified pixel points may include one or more of the following: (1) The pixel values ​​of each channel are all within the value range of the background color in the corresponding channel, and the pixel points are classified into the foreground area. If the pixel values ​​of a certain pixel point in each channel are all within the value range of the background color corresponding channel, then the pixel point can be considered to be the background. If at this time, the pixel point is determined to be the foreground based on the current cutout parameters, then the pixel point is considered to be a misclassified pixel point.

[0105] For example, assuming that the background color is green and is represented by the HSV model, the value ranges corresponding to its H, S, and V channels can be set to: H∈[35°, 77°], [0.168, 1], [0.180, 1]. Suppose a pixel point P1 is (H=40°, S=0.2, V=0.2), and the channel values ​​of its H, S, and V channels are all within the value range of the channels corresponding to the background color. Therefore, P1 is considered to be the background. If P1 is determined to be the foreground based on the cutout parameters at this time, it means that P1 is a misclassified pixel point.

[0106] (2) There are pixels whose pixel values ​​in any channel are outside the range of the background color in the corresponding channel and are classified as background. If a pixel has a pixel value in any channel outside the range of the background color in the corresponding channel, the pixel is considered to be foreground. If the pixel is determined to be background based on the current cutout parameters, the pixel is considered to be misclassified.

[0107] Similarly, assuming that the background color is green and is represented by the HSV model, the value ranges corresponding to its H, S, and V channels can be set to: H∈[35°, 77°], [0.168, 1], [0.180, 1]. A certain pixel point P2 is (H=10°, S=0.2, V=0.2). The channel value of the H channel is not within the value range of the H channel of the background color, so P2 is determined to be the foreground. If P2 is determined to be the background based on the cutout parameters at this time, it means that P2 is a misclassified pixel point.

[0108] In some embodiments, the background color can be set by the user. For example, when the user sets the value range of each channel of the background color in a specified color model, the user can directly enter the value range of each channel based on the background to be cut out, or the user can directly specify the background color, such as red, green, or blue, and then the device executing the image processing method can automatically generate the value range of each channel based on the color selected by the user. Alternatively, in some scenarios, in order to facilitate user settings, the user can also directly select an area in the background of the image to be cut out, and then the device executing the method can automatically determine the value range of each channel based on the color corresponding to the area selected by the user.

[0109] The various technical features in the above embodiments can be arbitrarily combined as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of the various technical features in the above embodiments also falls within the scope of disclosure of this specification.

[0110] Corresponding to the image processing method embodiment provided in the embodiment of the present application, the embodiment of the present application further provides an image processing device, the device comprising:

[0111] An acquisition module is used to acquire the image to be cutout;

[0112] a distance determination module, configured to determine the distance between the color coordinates of each pixel in the image to be cutout and a predetermined hyperplane, wherein the hyperplane is a classification boundary between the background color and the non-background color of the image to be cutout;

[0113] The probability determination module is used to determine the probability that each pixel point is the foreground area of ​​the image to be cut out based on the distance corresponding to each pixel point, so as to perform cutout processing on the image to be cut out based on the probability.

[0114] The specific implementation of image processing by the image processing device can refer to the description in the above method embodiment, which will not be repeated here.

[0115] The present application also provides an electronic device, as shown in FIG6 , which is a hardware structure diagram of an electronic device according to an embodiment of the present specification. In addition to the processor 62 and memory 64 shown in FIG6 , the device may also include other hardware, such as a forwarding chip responsible for processing messages. From a hardware perspective, the device may also be a distributed device, possibly including multiple interface cards to expand message processing at the hardware level. The memory 64 stores computer instructions, and when the processor 62 executes the computer instructions, it implements the image processing method described in any of the above embodiments.

[0116] Accordingly, an embodiment of the present application further provides a computer storage medium, in which a program is stored. When the program is executed by a processor, the method in any of the above embodiments is implemented.

[0117] The embodiments of the present application may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs or other data. Examples of computer storage media include but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0118] Those skilled in the art will readily recognize other implementations of the embodiments of the present invention after considering the specification and practicing the instructions disclosed herein. The embodiments of the present invention are intended to cover any variations, uses, or adaptations of the embodiments of the present invention that follow the general principles of the embodiments of the present invention and include common knowledge or customary techniques in the art not disclosed in the embodiments of the present invention. The description and examples are to be considered as exemplary only, and the true scope and spirit of the embodiments of the present invention are indicated by the following claims.

[0119] It should be understood that the embodiments of the present invention are not limited to the precise structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the embodiments of the present invention is limited only by the appended claims.

[0120] The above description is only a preferred embodiment of the embodiments of this specification and is not intended to limit the embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiments of this specification should be included in the scope of protection of the embodiments of this specification.

Claims

1. An image processing method, characterized in that, the method includes: obtaining an image to be matte; determining the distance between the color coordinates of each pixel point in the image to be matte and a pre-determined hyperplane, wherein the hyperplane is the classification boundary between the background color and the non-background color of the image to be matte; determining the probability that each pixel point is the foreground area of the image to be matte based on the distance corresponding to each pixel point, so as to perform matte processing on the image to be matte based on the probability.

2. The method according to claim 1, characterized in that, the hyperplane is determined based on the following method: obtaining a plurality of first sample pixel points with labels, wherein the label is used to indicate whether the color of the first sample pixel point is the background color; training a preset support vector machine model with the first sample pixel points, and using the trained model as the hyperplane.

3. The method according to claim 2, characterized in that, the first sample pixel points are represented by a first color model, and the first sample pixel points are obtained based on the following method: for each channel of the first color model, sampling at intervals within the value range of the pixel values of this channel to obtain a plurality of sampling values of this channel; combining the sampling values of each channel to obtain a plurality of first sample pixel points; for each first sample pixel point, determining whether the color of the first sample pixel point is the background color, and using the determination result as the label of the first sample pixel point.

4. The method according to claim 3, characterized in that, determining whether the color of the first sample pixel point is the background color includes: if the pixel values of each channel of the first sample pixel point are all within the value range of the background color in the corresponding channel, then determining that the color of the first sample pixel point is the background color.

5. The method according to claim 3, characterized in that, the background color is represented by a second color model, and before determining whether the color of the first sample pixel point is the background color, the method further includes: if the first color model and the second color model are inconsistent, then first unify the color model representing the background color and the color model representing the first sample pixel point.

6. The method according to claim 1, characterized in that, the background color is represented by a second color model, and the value range of each channel of the background color in the second color model is determined based on the following method: determined based on the numerical range input by the user; or determined based on the color selected by the user; or determined based on the color corresponding to the area boxed by the user in the image to be matte.

7. The method according to claim 1, characterized in that, determining the distance between the color coordinates of each pixel point in the image to be matte and a pre-determined hyperplane includes: obtaining a pre-determined distance map, wherein the distance map is used to indicate the distance between the color coordinates of a plurality of second sample pixel points and the hyperplane; By performing interpolation processing on the distances of the multiple second sample pixel points, the distance between the color coordinates of the pixel points in the image to be matte and the hyperplane is obtained.

8. The method according to claim 7, wherein, the multiple second sample pixel points are represented by a first color model, and the multiple second sample pixel points are obtained based on the following method: For each channel of the first color model, interval sampling is performed within the value range of the pixel values of this channel to obtain multiple sampling values of this channel; The sampling values of each channel are combined to obtain the multiple second sample pixel points.

9. The method according to claim 5, 6 or 8, wherein, the first color model is an RGB color model, and the second color model is an HSV color model or an LAB color model.

10. The method according to claim 1, wherein, For the pixel points in the image to be matte, the color coordinates of the pixel points are the coordinates formed by the pixel values of the target channels of the pixel points, where the target channels are channels related to color.

11. The method according to claim 10, wherein, if the image to be matte is an image in YUV format, the target channels are the U and V channels; if the image to be matte is an image in HSV format, the target channels are the H and S channels; if the image to be matte is an image in LAB format, the target channels are the A and B channels; if the image to be matte is an image in RGB format, the target channels are the R, G, and B channels.

12. The method according to claim 1, wherein, determining the probability that each pixel point is the foreground area of the image to be matte based on the distance corresponding to each pixel point includes: when the distance is less than a preset similarity threshold, the probability is 0; when the difference between the distance and the similarity threshold is greater than 0 and less than or equal to 1, the probability is positively correlated with the difference between the distance and the similarity threshold; when the difference between the distance and the similarity threshold is greater than 1, the probability is a preset smoothing coefficient.

13. The method according to claim 1, wherein, the dimension of the hyperplane is greater than or equal to 3.

14. An electronic device, wherein, the electronic device includes a processor, a memory, and computer instructions stored on the memory, and when the processor executes the computer instructions, the method according to any one of claims 1-13 is implemented.

15. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed, the method according to any one of claims 1-13 is implemented.

Citation Information

Patent Citations

  • SVM-based interactive region division method in digital image matting processing

    CN103714539A

  • Image processing method, image processing device and chip

    CN114881904A

Cited By

  • Method for automatically transparentizing white edges of multistage image slices of three-dimensional map

    CN120726072A

  • Method and equipment for identifying damage degree of fire hose and medium

    CN120823152A