Pixel pair evaluation function fitting method applied to mapping and updating strategy of proxy model among multi-scale images
By combining the multi-scale pyramid cutting algorithm and the Gaussian process proxy model, low-scale proxy model information is used to guide the initialization and optimization of the high-scale proxy model, and dynamic update strategy is introduced, which solves the problems of large computing resource requirements and low optimization efficiency in high-resolution image cutting, and achieves efficient image cutting performance.
Patent Information
- Application Number
- CN202510302965.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
The existing pixel-optimized cutout algorithm requires a large computing resource requirement and low optimization efficiency in high-resolution image cutout tasks, resulting in limited accuracy and timeliness of image cutouts.
A method of combining multi-scale pyramid cutting algorithm with Gaussian process agent model is adopted to construct a multi-scale image pyramid, and a low-scale agent model information is used to guide the initialization and optimization of the high-scale agent model, and a dynamic update strategy is introduced to improve the fitting accuracy.
It significantly reduces the computational complexity, improves the fitting quality of pixels to the evaluation function, and improves the overall performance of image cutouts, which is especially suitable for complex scene processing of high-resolution images.
Smart Images

Figure CN120147356A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of algorithms, and particularly relates to a method for fitting a pixel pair evaluation function for a mapping and update strategy of a proxy model between multi-scale images. Background Art
[0002] Image matting technology is the basis of digital synthesis technologies such as digital image and digital video non-linear editing, and is widely used in fields such as image editing, video special effects, augmented reality, and medical imaging. Matting technology can achieve the precise extraction of the foreground of interest in an image and provide the transparency mask necessary for image synthesis. Once the transparency mask is given, the specified foreground can be synthesized into a new background through simple linear operations. The transparency mask determines the proportion of the foreground color and the background color in the synthesized image, and its accuracy directly affects the quality of the synthesized image.
[0003] Among them, the matting algorithm based on pixel pair optimization models the matting problem as a pixel pair optimization problem, and calculates the optimal pair of foreground pixels and background pixels for each unknown pixel. According to the way of pixel pair optimization, it can be further divided into a sampling-based matting algorithm and a heuristic optimization-based matting algorithm. Among them, the sampling-based matting algorithm has the characteristics of fast solving speed and wide applicable scenarios, and the heuristic optimization-based matting algorithm has the advantage of high matting accuracy.
[0004] Matting algorithms based on heuristic optimization: To address the problem of losing optimal pixel pairs caused by sampling, researchers have utilized heuristic optimization algorithms to achieve global optimization of pixel pairs, avoiding sampling in the search space and effectively solving the problem of losing high-quality pixel pairs. Liang et al. replaced the pixel pair sampling process with pixel pair optimization using an evolutionary algorithm. This algorithm adjusts the number of iterations of the evolutionary algorithm to utilize all available computing resources and theoretically eliminates the risk of losing true samples. Lv et al. applied heuristic optimization techniques to solve this problem. This research relaxed the large-scale pixel pair combinatorial optimization problem into a large-scale pixel pair continuous optimization problem and used the particle swarm optimization algorithm with faster convergence to solve it. Cai et al. found that the particle swarm algorithm is prone to premature convergence when solving large-scale pixel pair combinatorial optimization problems and designed a cooperative differential evolution algorithm with a scrambling mechanism for this problem. Liang et al. proposed a particle swarm optimization algorithm based on an adaptive convergence speed controller using a similar idea. This method monitors the state of the particle swarm. When the similarity of the transparency masks of two individuals in the population is too high, it is considered that the population has prematurely converged, and the problem of premature convergence of the particle swarm is alleviated through a scrambling mechanism. Huang et al. designed a convergence speed controller as an additional operator for particle swarm optimization to avoid premature convergence problems during pixel pair optimization. Mohaptra et al. proposed an algorithm combining a particle swarm optimizer (PSO) and a competitive swarm optimizer (CSO) algorithm. The optimization algorithm is updated by learning from the winner, and the winner's example is directly passed to the next generation to improve the quality of the solution.
[0005] Liang et al. designed a pyramid heuristic optimization matting framework for the problem of high computational resource consumption in matting algorithms based on heuristic optimization. By scaling the image and the trimap, the large-scale pixel pair optimization problem is converted into medium- and small-scale optimization problems, a matting problem pyramid is established, and the optimal pixel pair solution obtained during the optimization of small-scale problems is used as heuristic information to be passed to the solution of larger-scale problems. However, the pyramid matting framework only passes the currently found better pixel pairs, losing potentially higher-quality pixel pairs, resulting in poor alpha value accuracy for some pixels. Gou et al. based on the problem that it is difficult for pixel pair optimization methods to provide high-quality transparency masks under low computational resources. Based on the idea of surrogate models, a Gaussian process surrogate model of the pixel pair evaluation function is constructed, which can effectively improve the quality of image matting under low computational resources, but a large amount of computation is required to construct a surrogate model for each unknown pixel. Gou et al. designed a micro-scale search strategy for high-resolution image matting based on the ideas of micro-scale search and pyramid matting, providing high-quality heuristic information for the evolutionary optimization algorithm to achieve high-quality matting results, but it still requires a large amount of computational resources.
[0006] Although the matte extraction algorithm based on pixel pair optimization has achieved certain results in the academic field, it still faces challenges in terms of accuracy and timeliness. Especially in the task of high-resolution image matte extraction, there are two major pain points: high computational resource requirements and low optimization efficiency.
[0007] Therefore, designing an efficient method that combines the multi-scale pyramid matte extraction algorithm with the Gaussian process surrogate model, which can make full use of the low-scale computational information to guide the high-scale optimization process, is the key to solving the above problems. Such a method can not only reduce the computational cost but also improve the fitting quality of the pixel pair evaluation function, thereby enhancing the overall performance of image matte extraction. Summary of the Invention
[0008] To solve the above problems, the present invention adopts the following technical solutions:
[0009] A method for fitting a pixel pair evaluation function that applies the mapping and update strategy of a surrogate model between multi-scale images, comprising the following steps:
[0010] 1) Construction of a multi-scale image pyramid:
[0011] The original image is downsampled k times until the preset minimum scale is reached, and the images obtained from each downsampling are combined to construct a multi-scale pyramid, where the bottom is the original image, and the evaluation function of pixel pairs is initially fitted on the minimum-scale image to construct a Gaussian process surrogate model;
[0012] 2) Initialization of the surrogate model for high-scale images:
[0013] The fitting result parameters of the Gaussian process surrogate model in the low-scale image are used as heuristic information and transferred to the high-scale image, and the parameter information of the surrogate model fitted at the low scale is used to initialize the high-scale surrogate model, reducing the computational resource overhead;
[0014] 3) Update strategy for the surrogate model of high-scale images:
[0015] After initializing the surrogate model of the high-scale image, selective update is performed according to the fitting situation of the surrogate model after high-scale initialization;
[0016] 4) Layer-by-layer optimization and result integration:
[0017] Starting from the surrogate model of the lowest-scale image, layer-by-layer optimization fitting is performed, and the result of each layer's surrogate model is used as the initialization of the next layer's surrogate model. Finally, an accurate surrogate model is obtained in the original image for pixel pair evaluation to obtain the optimal pixel pairs.
[0018] In step 1), the surrogate model formula (1) is established for each unknown pixel at the smallest scale as follows, and the coordinate mapping between different scales is shown in formula (2):
[0019]
[0020] where g sm represents the fitted Gaussian process surrogate model, fit 1 is the Gaussian process surrogate model fitting function, are the foreground and background sample points selected from the smallest scale image is the adaptation value of the sample point corresponding pixel to the evaluation function, and are the foreground coordinates of the high-scale image, and are the background coordinates of the high-scale image, and are the foreground coordinates of the low-scale image, and are the background coordinates of the low-scale image.
[0021] In step 2), the parameter update rule is based on the resolution scale factor k, and the surrogate model mapping relationship is as follows:
[0022]
[0023] σ hl = σ l ·k, σ hf = σ f ;
[0024] σ hn = σ n ·k, μ lm = μ m ;
[0025] where and are the surrogate models on the low-scale and high-scale images respectively, σ l is the kernel function length scale, adjusted to σ l ·k to adapt to the enlarged pixel point spacing in the high-scale image and ensure that the correlation range of the kernel function is consistent with the scale change, σ f is the kernel function variance, which remains unchanged, reflecting that the global change amplitude of the output value is not significantly affected by the resolution change, σ n is the noise parameter, adjusted to σ n ·k to adapt to the noise change caused by the increased resolution in the high-scale image, μm is the mean value and remains unchanged.
[0026] In step 3), the update strategy is as shown in the following formula:
[0027]
[0028] R 1 = x ∈ {l d (x hs ) < μ min ∩ (β < (l u (x hs ) - l d (x hs )) < γ)} (7)
[0029] R 2 = x ∈ {x = V j}}, j ∈ {1, 2, 3, 4,..., n} (8)
[0030] V j = f select (l d (x hs ), j) (9)
[0031] where l d (x hs ), l u (x hs ) represent the upper limit value under the confidence level of the high-scale image sample point x hs , μ m (x hs ) represents the mean value corresponding to the surrogate model x hs , σ(x hs ) represents the standard deviation of the surrogate model, n represents the number of samples for one fitting, R is to select the update strategy according to the fitting effect of the surrogate model at different scales, L ∈ represents the threshold for judging the goodness of the fitting degree of the surrogate model, L i represents the loss value of the current surrogate model at the j-th layer, R 1 represents that the difference between the upper and lower limits is greater than the minimum threshold β and less than the threshold γ, and the mean value at this x hs is less than the minimum mean value of the current surrogate model, then there may be an optimal solution here, and this sample point needs to be further fitted, R 2 represents finding the region with the lowest lower limit of the confidence interval on the current surrogate model for further fitting, V j is the number of sample points required to find the region with the lowest lower limit of the confidence interval to be updated according to the surrogate model at the current j-th layer, x represents the sample point of the high-scale image initialization that needs to be further fitted.
[0032] Finally, incrementally fit the high-scale image surrogate model, and its formula is as follows:
[0033]
[0034] where g h is a surrogate model well-fitted for the high-scale image, and fit 2 means further fitting based on the original surrogate model according to the new sample point set x, and fitness x represents the fitness value of the sample point set x.
[0035] Compared with the prior art, the advantages of the present invention are as follows: By combining the pyramid matting algorithm and the Gaussian process surrogate model, a multi-scale image processing framework is designed, thereby greatly reducing the computational complexity. A mapping method from the low-scale surrogate model to the high-scale image is proposed, making full use of the existing information of the low-scale image and improving the initialization efficiency of the high-scale model. A dynamic update strategy is introduced to ensure the fitting accuracy of the high-scale surrogate model and reduce invalid calculations through confidence interval analysis. The efficient utilization of computing resources is achieved, significantly improving the optimization efficiency of the pixel pair evaluation function and the overall performance of image matting. It is particularly suitable for processing complex scenes of high-resolution images and also has broad application potential in other image processing tasks that require multi-scale optimization. Brief Description of the Drawings
[0036] Figure 1 is a schematic flowchart of the present invention. Detailed Embodiments
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0038] The present invention proposes a method for fitting a pixel pair evaluation function for the mapping and update strategy of a surrogate model between multi-scale images, aiming to guide the initialization and optimization of the high-scale surrogate model by layer-by-layer transmitting the information of the low-scale surrogate model. Please refer to Figure 1 , and the specific design includes the following steps:
[0039] Construction of Multi-Scale Image Pyramid
[0040] The original image is downsampled k times until the preset minimum scale is reached. The images obtained from each downsampling are combined to construct a multi-scale pyramid, with the original image at the bottom. On the minimum-scale image, a Gaussian process surrogate model is initially constructed by fitting the evaluation function for pixel pairs. The surrogate model formula (1) for each unknown pixel on the minimum scale is as follows, and the coordinate mapping between different scales is shown in formula (2):
[0041]
[0042] where g sm represents the fitted Gaussian process surrogate model, fit 1 is the Gaussian process surrogate model fitting function, are the foreground and background sample points selected from the minimum-scale image is the adaptation value of the pixel pair evaluation function corresponding to the sample point and are the foreground coordinates of the high-scale image, and are the background coordinates of the high-scale image, and are the foreground coordinates of the low-scale image, and are the background coordinates of the low-scale image.
[0043] Initialization of the surrogate model for the high-scale image
[0044] The fitting result parameters of the Gaussian process surrogate model in the low-scale image are used as heuristic information and passed to the high-scale image. The high-scale surrogate model is initialized using the parameter information of the surrogate model fitted in the low scale, reducing the computational resource overhead. The parameter update rule is based on the resolution scale factor k, and the surrogate model mapping relationship is as follows:
[0045] σ hl = σ l ·k, σ hf = σ f ;
[0046] σ hn = σ n ·k, μ lm = μ m ;
[0047] where and are the surrogate models on the low-scale and high-scale images respectively, σ l is the kernel function length scale, adjusted to σ l·k to adapt to the expansion of the pixel point spacing in the high-scale image and ensure that the correlation range of the kernel function is consistent with the scale change. σ f is the variance of the kernel function and remains unchanged, reflecting that the global change amplitude of the output value is not significantly affected by the resolution change. σ n is the noise parameter and is adjusted to σ n ·k to adapt to the noise change caused by the increased resolution in the high-scale image. μ m is the mean value and remains unchanged.
[0048] Surrogate model update strategy for high-scale images
[0049] After initializing the surrogate model of the high-scale image, we perform selective updates according to the fitting situation of the surrogate model after high-scale initialization. According to the range of the confidence space of the surrogate model after initialization, we can know which areas on the surrogate model are unlikely to be the best results and can directly abandon the fitting to reduce unnecessary calculations. Finally, we only perform further fitting on the areas where there may be better pixel pairs. The specific update strategy is shown in the following formula:
[0050]
[0051] R 1 = x ∈ {l d (x hs ) < μ min ∩ (β < (l u (x hs ) - l d (x hs )) < γ)} (7)
[0052]
[0053] where l d (x hs ), l u (x hs ) respectively represent the upper limit values under the confidence levels of the high-scale image sample points x hs , μ m (x hs ) represents the mean value corresponding to the surrogate model x hs , σ(x hs ) represents the standard deviation of the surrogate model, n represents the number of samples for one fitting, R is to select the update strategy according to the fitting effects of the surrogate models at different scales, L ∈ represents the threshold for judging the goodness of the surrogate model fitting, L i represents the loss value of the current surrogate model at which layer, R 1 represents that the difference between the upper and lower limits is greater than the minimum threshold β and less than the threshold γ, and at this x hsIf the mean value at a certain point is less than the minimum mean value of the current surrogate model, there may be an optimal solution at this point, and further fitting of this sample point is required. R 2 It means to find the region with the lowest lower limit of the confidence interval on the current surrogate model for further fitting. V j is the number of sample points required to update the region with the lowest lower limit of the confidence interval according to the surrogate model of the current j-th layer. x represents the sample points that need to be further fitted for the initialization of the high-scale image.
[0054] Finally, incrementally fit the high-scale image surrogate model, and its formula is as follows:
[0055]
[0056] where g h is the surrogate model fitted for the high-scale image, fit 2 means further fitting based on the original surrogate model according to the new sample point set x, fitness x represents the fitness value of the sample point set x.
[0057] Layer-by-layer optimization and result integration
[0058] Start optimizing and fitting layer by layer from the surrogate model of the lowest-scale image. The results of the surrogate model at each layer are used as the initialization of the surrogate model at the next layer. Finally, an accurate surrogate model is obtained in the original image for pixel pair evaluation to obtain the optimal pixel pair. In this way, we can make full use of the resources of the already calculated surrogate model, reduce the overhead of repeated calculations, and improve the matting performance.
[0059] Summary
[0060] This patent proposes a multi-scale surrogate model mapping method and update strategy for optimizing the pixel pair evaluation function in image matting. The main innovation points of this method include: combining the pyramid matting algorithm and the Gaussian process surrogate model to design a multi-scale image processing framework, thus significantly reducing the computational complexity. A mapping method from the low-scale surrogate model to the high-scale image is proposed, making full use of the existing information of the low-scale image to improve the initialization efficiency of the high-scale model. A dynamic update strategy is introduced to ensure the fitting accuracy of the high-scale surrogate model and reduce invalid calculations through confidence interval analysis. The efficient use of computing resources is realized, significantly improving the optimization efficiency of the pixel pair evaluation function and the overall performance of image matting.
[0061] This method provides an efficient solution for the field of image matting, especially suitable for processing complex scenes of high-resolution images, and also has broad application potential in other image processing tasks that require multi-scale optimization.
[0062] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0063] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0064] In the present invention, unless otherwise clearly defined and limited, the terms such as "mounted", "connected", "coupled", "fixed", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection, an electrical connection, or a communication connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0065] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. In the description of this specification, the description with reference to terms such as "one solution", "some solutions", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the solution or example are included in at least one solution or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same solution or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more solutions or examples.
Claims
1. A pixel pair evaluation function fitting method for mapping and updating strategies of proxy models between multi-scale images, characterized by: The following steps are involved: 1) Multi-scale image pyramid construction: The original image is downsampled k times multiple times until the preset minimum scale is reached. The images obtained by each downsampling are combined to construct a multi-scale pyramid, where the bottom is the original image. The evaluation function of the pixel pairs is initially fitted on the minimum scale image to construct a Gaussian process proxy model. 2) Proxy model initialization for high-scale images: The fitting result parameters of the Gaussian process proxy model in the low-scale image are used as heuristic information and transferred to the high-scale image. The high-scale proxy model is initialized using the parameter information of the proxy model fitted at the low scale, thus reducing the cost of computing resources. 3) Proxy model update strategy for high-scale images: After initializing the proxy model of the high-scale image, selective updates are performed based on the fit of the proxy model after high-scale initialization; 4) Layer-by-layer optimization and result integration: Starting from the proxy model of the lowest-scale image, the optimization fitting is performed layer by layer. The result of each level of the proxy model is used as the initialization of the proxy model of the next level. Finally, an accurate proxy model is obtained in the original image for pixel pair evaluation to obtain the optimal pixel pair.
2. The pixel pair evaluation function fitting method for mapping and updating strategies of proxy models between multi-scale images according to claim 1, characterized in that: In step 1), the proxy model formula (1) is established for each unknown pixel at the minimum scale as follows, and the coordinate mapping between different scales is shown in formula (2): Among them, g sm It represents the fitted Gaussian process surrogate model, and fit1 is the Gaussian process surrogate model fitting function. It is the foreground and background sample points selected from the minimum scale image yes The pixel corresponding to the sample point is the fitness value of the evaluation function. and is the high-scale image foreground coordinate, and is the high-scale image background coordinate, and is the low-scale image foreground coordinate, and are the low-scale image background coordinates.
3. The pixel pair evaluation function fitting method for mapping and updating strategies of proxy models between multi-scale images according to claim 2, characterized in that: In step 2), the parameter update rule is based on the resolution scale factor k, and the proxy model mapping relationship is as follows: s hl =s l ·k,s hf =s f ; s hn =s n ·k,m lm =μ m ; in and are the proxy models on low-scale and high-scale images respectively, σ l is the kernel function length scale, adjusted to σ l k, to adapt to the increase in pixel spacing in high-scale images and ensure that the relevance range of the kernel function is consistent with the scale change, σ f is the kernel function variance, which remains unchanged, and the global change amplitude of the output value is not significantly affected by the change of resolution, σ n is the noise parameter, adjusted to σ n k, to adapt to the noise changes caused by the increase in resolution in high-scale images, μ m is the mean and remains unchanged.
4. The pixel pair evaluation function fitting method for mapping and updating strategies of proxy models between multi-scale images according to claim 3, characterized in that: In step 3), the update strategy is as follows: R1=x∈{l d (x hs )<μ min ∩(β<(l u (x hs )-l d (x hs ))<γ)} (7) R2=x∈{x=V j },j∈{1,2,3,4,...,n} (8) V j =f select (l d (x hs ),j) (9) Among them l d (x hs ), l u (x hs ) represent high-scale image sample points x hs Under the confidence level, the upper limit value, μ m (x hs ) represents the proxy model x hs The corresponding mean, σ(x hs ) represents the standard deviation of the surrogate model, n represents the number of samples fitted at one time, R is the update strategy selected according to the fitting effect of the surrogate model at different scales, and L ∈ Indicates the threshold for judging the degree of fit of the surrogate model, L i Indicates the loss value of the current proxy model. R1 indicates that the difference between the upper and lower limits is greater than the minimum threshold β and less than the threshold γ, and in this x hs If the mean of the above is less than the minimum mean of the current proxy model, then there may be an optimal solution here, and further fitting is required for this sample point. R2 means finding the area with the lowest lower limit of the confidence interval on the current proxy model for further fitting. V j It is to find the number of sample points in the lowest area of the confidence interval lower limit that need to be updated based on the current proxy model of the jth layer. x represents the sample points that need to be further fitted for high-scale image initialization. The final incremental fitting high-scale image proxy model is formulated as follows: where g h The proxy model is fitted for the high-scale image. fit2 means further fitting based on the original proxy model according to the new sample point set x. x Represents the fitness value of the sample point set x.