Method and system for multi-aspect ratio image retargeting based on key semantic preservation
Patent Information
- Application Number
- CN202410935101.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-07-12
AI Technical Summary
[0005]本发明正是针对现有图像重定向技术中语义缺失,像素不连续和生成结果与原图不一致等的问题,提供一种基于关键语义保留的多宽高比图像重定向方法及系统,至少包括初步重定向模块、重绘区域选取模块与图片引导生成模块;所述初步重定向模块基于语义缝隙雕刻算法,通过引入带空间先验的显著性图与梯度算子结合,删除其中能量低的线状区域得到初步重定向结果;所述重绘区域选取模块以初步重定向模块的输出结果作为输入,对初步重定向图中像素移位大的区域进行重新生成,并对语义缝隙雕刻未达到目标宽高比的结果进行扩图;所述图片引导生成模块:基于重绘区域选取模块输出的结果,利用原图作为控制条件引导局部重绘,输出图像
[0030](1)本发明提出了一个基于关键语义保留的多宽高比图像重定向方法,在各种宽高比要求下显著提升了图像重定向的语义保留与视觉效果。
Smart Images

Figure CN118941469B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical fields of computer vision and AIGC, and mainly relates to a method and system for retargeting multi-aspect ratio images based on key semantic preservation. Background Technology
[0002] Image retargeting technology is widely used in various digital media and display devices, including mobile phones, tablets, computer monitors, and smart TVs. With the diversification of screen ratios on these devices, how to display images of consistent quality across different devices has become a pressing issue.
[0003] Traditional image retargeting methods, such as scaling, cropping, and seam carving, while meeting certain requirements, all suffer from semantic loss. Scaling, although maintaining image integrity, leads to disproportionate content and severe distortion. Cropping preserves local sharpness but loses important image information. More complex seam carving techniques achieve retargeting by removing low-energy paths, preserving some important content. However, lacking a deep understanding of semantic information, seam carving results in the loss of important objects and image discontinuities when processing complex images. With the development of deep learning, generative adversarial networks (GANs) have alleviated semantic and pixel discontinuity issues to some extent. However, the instability during training and inconsistencies in generated results limit their application in real-world scenarios.
[0004] The Stable Diffusion Model (SDM) is a popular image generation technique that generates high-quality image content by progressively denoising the image. This method provides highly consistent and detailed image images during both generation and redrawing, ensuring visual coherence and naturalness in the final result. However, in practical applications, using text as a control condition often makes it difficult to effectively control complex scenes. Summary of the Invention
[0005] This invention addresses the problems of semantic loss, pixel discontinuity, and inconsistency between the generated image and the original image in existing image retargeting techniques. It provides a multi-aspect ratio image retargeting method and system based on key semantic preservation, comprising at least a preliminary retargeting module, a redrawing region selection module, and an image-guided generation module. The preliminary retargeting module is based on a semantic gap carving algorithm, which combines a saliency map with spatial priors and a gradient operator to delete low-energy linear regions to obtain the preliminary retargeting result. The redrawing region selection module uses the output of the preliminary retargeting module as input to regenerate regions with large pixel shifts in the preliminary retargeting image and expands the image where the semantic gap carving fails to achieve the target aspect ratio. The image-guided generation module, based on the output of the redrawing region selection module, uses the original image as a control condition to guide local redrawing and output the final image. This invention solves the problem of semantic loss in traditional retargeting methods, adaptively regenerating abrupt pixels and outputting an image with key semantics preserved and conforming to the desired aspect ratio for human aesthetics.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is: a multi-aspect ratio image retargeting system based on key semantic preservation, which includes at least a preliminary retargeting module, a redrawing region selection module, and an image guidance generation module;
[0007] The initial redirection module: based on the semantic gap carving algorithm, it introduces a saliency map with spatial prior and combines it with a gradient operator to delete low-energy linear regions to obtain the initial redirection result;
[0008] The redrawing region selection module: takes the output of the initial redirection module as input, regenerates the regions with large pixel shifts in the initial redirection image, and expands the image for semantic gap carving results that do not achieve the target aspect ratio;
[0009] The image guidance generation module: based on the output of the redrawing region selection module, uses the original image as a control condition to guide local redrawing and output the image.
[0010] As an improvement of the present invention, the image guidance generation module includes at least a ControlNet and an IP-Adapter assisted stable diffusion model, wherein the ControlNet is used to control the redrawing area and the IP-Adapter is used to control the generated content.
[0011] To achieve the above objectives, the present invention also adopts the following technical solution: a multi-aspect ratio image retargeting method based on key semantic preservation, specifically including the following steps:
[0012] S1, Determine the target image: Input the RGB image and set the aspect ratio of the target;
[0013] S2, Preliminary Redirection: Based on the semantic gap carving algorithm, a saliency map is obtained using the saliency detection algorithm. After attaching spatial priors, it is combined with the gradient operator to calculate the energy of each region of the image. Then, starting from the lines with low energy, the image is cropped. After deleting the low-energy linear regions, the preliminary redirection result O is obtained. i ;
[0014] S3, Identify areas with unnatural pixel transitions: Use a sliding window to identify areas after the initial retargeting result in step S2. i The number of pixels deleted before and after each pixel is used to identify areas with large pixel shifts and to expand the image for semantic gap carving results that do not meet the target aspect ratio. The large pixel shift refers to the number of retained pixels that are surrounded by a large number of deleted pixels, which is obtained through the deletion-retention binary image of the semantic analysis cutting algorithm.
[0015] S4, Unnatural region redrawing: The unnatural regions obtained in step S3 are regenerated using an image-guided stable diffusion model, and then the unmasked regions are pasted back onto the generated regions.
[0016] S5, Image Output: Based on the target aspect ratio, output the corresponding retargeting result to obtain the target aspect ratio image.
[0017] As an improvement of the present invention, the energy calculation formula in step S2 is as follows:
[0018]
[0019] Where I(x,y) represents the pixel value of image I at position (x,y), S(x,y) represents the saliency of its saliency map at position (x,y), x_0 represents the x-coordinate of the centroid of the saliency region, and W is the width of the image.
[0020] After obtaining the energy of each pixel in the image, use A vertical line describing the set of pixels in a graph, for any I, has |x(i)-x(i-1)|≤1; the energy of a vertical line is defined as... s to be deleted * The lines are obtained by minimizing energy:
[0021]
[0022] Using dynamic programming, the lines with the least energy are identified and deleted.
[0023] As another improvement of the present invention, in step S2, the semantic gap cutting algorithm automatically stops cropping the image after one-third of the significant region has been cropped.
[0024] As another improvement of the present invention, in step S3, if the number of deleted pixels before and after each pixel exceeds a threshold, it is determined to be a region with large pixel shift. The shift region mask M is calculated by a one-dimensional sliding window of length l to count the number of deleted pixels in the neighborhood of each pixel, and a one-dimensional convolution is performed on the deletion-retention binary image S:
[0025]
[0026] Where conv1d represents one-dimensional convolution, K is a convolution kernel with length l and value 1, and the region in M that exceeds the threshold η is the region with large pixel shift.
[0027] As another improvement of the present invention, in step S3, the size of the convolution kernel K is 25, the value at each position is 1, and the threshold η is 12.
[0028] As another improvement of the present invention, in step S4, after the mask is overlaid on the preliminary redirection result, it is input into the Controlnet model; the original image is input into the IP-Adapter model as a control condition, and the two together assist the Stablediffusion stable diffusion model to generate the image.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] (1) This invention proposes a multi-aspect-ratio image repositioning method based on key semantic preservation, which significantly improves the semantic preservation and visual effect of image repositioning under various aspect ratio requirements.
[0031] (2) The semantic gap cutting algorithm proposed in this invention further explores the energy difference inside the foreground based on the energy distinction between salient foreground and background, and distinguishes the energy inside the foreground through spatial location prior, making it more in line with human visual perception.
[0032] (3) The present invention proposes a redrawing region selection module, which can efficiently find regions with inconsistent pixels in the result in the image redirection method based on pixel shifting, providing conditions for post-processing of redirection. Attached Figure Description
[0033] Figure 1 This is a flowchart of the steps of the multi-aspect ratio image repositioning method based on key semantic preservation in this invention;
[0034] Figure 2 This is a schematic diagram illustrating the working principle of the multi-aspect ratio image redirection system based on key semantic preservation according to the present invention.
[0035] Figure 3 This is a schematic diagram of the image guidance generation module in the multi-aspect ratio image retargeting system based on key semantic preservation of the present invention;
[0036] Figure 4 This is a comparison diagram of the effects of the method of the present invention with other methods in Embodiment 2 of the present invention. Detailed Implementation
[0037] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.
[0038] Example 1
[0039] A method for reshaping multi-aspect ratio images based on key semantic preservation, such as Figure 1 As shown, the specific steps include the following:
[0040] Step S1, Input RGB image and target aspect ratio: Input an image I and the target aspect ratio r;
[0041] Step S2, semantic gap cutting algorithm for initial redirection: A saliency map is obtained using a saliency detection algorithm, and after attaching spatial priors, it is combined with the gradient operator to calculate the energy of each region of the image. Cropping begins from lines with lower energy to obtain an initial redirection result O that is equal to or close to the desired aspect ratio. i ;
[0042] Semantic gap carving involves selecting and deleting lines with lower energy based on the energy map. The energy calculation formula is as follows:
[0043]
[0044] This energy formula introduces semantic information with spatial priors into traditional image operators. Here, s(x,y) represents the saliency magnitude at position (x,y) in the saliency map, and x0 represents the x-coordinate of the centroid of the salient region. Furthermore, the semantic gap cutting algorithm automatically stops cropping the image after deleting one-third of the salient region.
[0045] After obtaining the energy of each pixel in the image, use A vertical line describing the set of pixels in a graph holds for any i, where |x(i) - x(i-1)| ≤ 1. The energy of a vertical line can be defined as... s to be deleted * The lines can be obtained by minimizing energy:
[0046]
[0047] By using dynamic programming, we can identify the few lines with the least energy to delete.
[0048] Step S3, redraw the region selection module to find areas with unnatural pixel connections: use a sliding window to find the preliminary retargeting results from step S2. i If the number of pixels deleted before and after each pixel exceeds a threshold, the pixel is considered to be over-shifted and needs to be redrawn.
[0049] When selecting the redrawing region, the algorithm not only adaptively identifies areas with severe pixel shifts, but also expands the image based on the current aspect ratio to ensure that the aspect ratio of the processed image equals the target aspect ratio. Specifically, areas with severe pixel shifts are obtained through a deletion-retention binary map generated by a semantic analysis segmentation algorithm. Pixels with a large number of surrounding deleted pixels are considered to have severe pixel shifts and need to be regenerated.
[0050] The shift region mask M is calculated by using a one-dimensional sliding window of length l to count the number of deleted pixels in the neighborhood of each pixel. Specifically, this is implemented by performing a one-dimensional convolution on the deletion-preservation binary image S.
[0051]
[0052] Here, conv1d represents a one-dimensional convolution, and K is a convolution kernel of length l with a value of 1. Regions in M exceeding the threshold η are considered areas of severe pixel shift. The convolution kernel K has a size of 25, with a value of 1 at each position, and the threshold η is set to 12.
[0053] Step S4: The image-guided generation module redraws the unnatural regions: It uses the image-guided stable diffusion model to regenerate the unnatural regions of the masked pixels, and then pastes the unmasked regions back onto the generated regions.
[0054] After overlaying the initial redirection result with a mask, it is input into the ControlNet model; the original image is input into the IP-Adapter model as a control condition, and both assist the Stable diffusion model in generation. This generation process uses the original image as a control condition, which avoids generating semantics that do not match the original image and improves the consistency between the generated region and the original image.
[0055] Step S5, Output the target aspect ratio image: Based on the target aspect ratio, output the corresponding retargeting result.
[0056] Example 2
[0057] A multi-aspect-ratio image retargeting system based on key semantic preservation includes at least a preliminary retargeting module, a redrawing region selection module, and an image-guided generation module.
[0058] The initial redirection module, based on the semantic gap carving algorithm, combines a saliency map with spatial priors and a gradient operator to delete low-energy linear regions to obtain the initial redirection result. The redrawing region selection module, using the output of the initial redirection module as input, regenerates regions with large pixel shifts in the initial redirection image and expands the image for semantic gap carving results that do not achieve the target aspect ratio. The image-guided generation module, based on the output of the redrawing region selection module, uses the original image as a control condition to guide local redrawing and outputs the image.
[0059] The working principle of the above system is as follows: Figure 2 As shown, the multi-aspect ratio image repositioning method based on key semantic preservation using the above system includes the following steps:
[0060] Step S1: Input RGB image and target aspect ratio: Input an image I and the target aspect ratio r.
[0061] This embodiment's image retargeting experiment was conducted on the public dataset RetargetMe. Research indicates that the mainstream aspect ratios for current devices are: 16:9, 4:3, 1:1, and 9:16. Figure 4 As shown, the target aspect ratio selected in this embodiment is 16:9.
[0062] Step S2: Use the semantic gap cutting algorithm.
[0063] This embodiment, as described in Example 1, uses the Visual Saliency Transformer (VST) salient object detection model as the salient energy extraction method and the Sobel operator as the manual operator energy extraction method. It organically combines these two methods based on the spatial prior principle of strong energy at the center and weak energy at the edges. A dynamic programming algorithm is used to calculate the energy of each line, and lines with lower energy are deleted. The number of lines deleted is determined by the smaller of the target aspect ratio *r* and the salientity loss tolerance *λ*. After deleting low-energy lines, the preliminary redirection result is output. The salientity loss tolerance *λ* represents the maximum salient region that can be lost during the redirection process; in this embodiment, it is set to 0.33.
[0064] Step S3: The redraw area selection module identifies areas where pixel connections are unnatural.
[0065] The redrawing region selection module proposed in this invention is applied to the preliminary retargeting result output in step S2 to identify regions where pixels are discontinuous and generate masks. Simultaneously, this module compares the aspect ratio of the preliminary retargeting result with the target aspect ratio; if they are inconsistent, masks are added to the top and bottom sides to achieve the desired aspect ratio. The masked regions generated by this module will be regenerated in step S4.
[0066] Step S4: The image guidance generation module redraws unnatural areas.
[0067] An image-guided stable diffusion model is used to regenerate unnatural regions of the masked pixels, and then the unmasked regions are pasted back onto the generated image. Figure 3 The demonstrated generative model structure is generated by a stable diffusion model assisted by ControlNet and IP-Adapter. ControlNet controls the redraw area, while IP-Adapter controls the generated content.
[0068] Step S5: Output the corresponding redirection result based on the target aspect ratio.
[0069] In this embodiment, the model uses a single RTX3090 graphics card for inference sampling. The encoder of the saliency detection model is selected as the t2t-vit model pre-trained on ImageNet. The redrawing process has 50 sampling steps, the sampling method is ddim, and the IP-Adapter version is selected as the plus version.
[0070] Figure 4 This demonstrates the retargeting effect of the method of this invention compared to other methods at a 16:9 aspect ratio. The first column shows the input RGB image, the second, third, and fourth columns show the retargeting results obtained by scaling, cropping, and gap cutting, respectively, and the fifth column shows the result of this invention. Figure 4 As can be seen, image retargeting results obtained using traditional methods often suffer from distortion or loss of important regions due to a lack of semantic information. This invention solves this problem by introducing saliency priors. Furthermore, through image expansion and redrawing, this invention addresses the issue of important objects exceeding the target resolution, making it applicable to any aspect ratio.
[0071] Test case
[0072] To verify the reliability of the retargeting results under different aspect ratios, the method of this invention is compared with the most common existing image retargeting methods in performance testing: This test case compares the performance of the method with three other commonly used image retargeting methods on the public dataset RetargetMe at four aspect ratios: 16:9, 4:3, 1:1 and 9:16.
[0073] The objective evaluation metric for this test case is the Saliency Discard Ratio (SDR), which is calculated as follows:
[0074]
[0075] Among them, W s ori W represents the maximum number of saliency pixels on the width axis of the saliency map in the original image.s out This represents the maximum number of salient pixels on the width axis of the saliency map resulting from the retargeting. This metric measures the proportion of salient regions lost before and after retargeting; a smaller metric indicates a better-performing method.
[0076] The comparison results of the proposed method with three other existing image repositioning methods using quantitative methods are shown in the table below:
[0077]
[0078] As can be seen from the table above, the image retargeting method proposed in this invention significantly reduces the loss of salient regions across all aspect ratios, fully demonstrating the efficiency and versatility of this invention in preserving key semantics.
[0079] The subjective evaluation metrics for this test case include content integrity score, deformation score, local smoothness score, and aesthetic score, with a score range of 0 to 3. The subjective evaluation results are the average of 16:9 results from 20 volunteers on the RetargetMe dataset. The results are shown in the table below:
[0080]
[0081] As can be seen from the table above, the image retargeting method proposed in this invention has achieved extremely high scores on all subjective indicators, and the average result is the best among all methods.
[0082] The multi-aspect-ratio image repositioning method based on key semantic preservation provided by this invention can widely assist in image repositioning tasks of various aspect ratios. Specifically in industry, ensuring the consistency and aesthetics of images across different devices and screen sizes is crucial. In fields such as digital media, advertising, and e-commerce, image quality often directly impacts user experience and decision-making; therefore, product images need to adapt to different screen ratios on various display devices to ensure optimal display across different platforms. This invention can achieve excellent repositioning results, reduce the need for companies to repeatedly create images, and improve operational efficiency.
[0083] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.
Claims
1. A multi-aspect ratio image retargeting method based on key semantic preservation, characterized in that, Specifically, the steps include the following: S1, Determine the target image: Input the RGB image and set the aspect ratio of the target; S2, Preliminary Redirection: Based on the semantic gap carving algorithm, a saliency map is obtained using the saliency detection algorithm. After attaching spatial priors, it is combined with the gradient operator to calculate the energy of each region of the image. Then, starting from the lines with low energy, the image is cropped. After deleting the low-energy linear regions, the preliminary redirection result is obtained. The formula for calculating energy is as follows: ; in, Representing an image exist The pixel value of the location, The significance plot is shown in The salience of the position, The x-coordinate represents the centroid of the salient region. The width of the image; After obtaining the energy of each pixel in the image, use A vertical line describing the set of pixels in the graph, for any I, has The energy of a vertical line is defined as... To be deleted The lines are obtained by minimizing energy: ; Using dynamic programming, the lines with the lowest energy are identified and deleted. S3, Identify areas with unnatural pixel transitions: Use a sliding window to identify areas resulting from the initial retargeting in step S2. The number of pixels deleted before and after each pixel is used to identify areas with large pixel shifts and to expand the image for semantic gap carving results that do not meet the target aspect ratio. The large pixel shift refers to the number of retained pixels that are surrounded by a large number of deleted pixels, which is obtained through the deletion-retention binary image of the semantic analysis cutting algorithm. S4, Unnatural region redrawing: The unnatural regions obtained in step S3 are regenerated using an image-guided stable diffusion model, and then the unmasked regions are pasted back onto the generated regions. S5, Image Output: Based on the target aspect ratio, output the corresponding retargeting result to obtain the target aspect ratio image.
2. The multi-aspect ratio image retargeting method based on key semantic preservation as described in claim 1, characterized in that: In step S2, the semantic gap cutting algorithm automatically stops cropping the image after one-third of the salient region has been cropped.
3. The multi-aspect ratio image retargeting method based on key semantic preservation as described in claim 1, characterized in that: In step S3, regions where the number of deleted pixels before and after each pixel exceeds a threshold are identified as areas with large pixel shifts, and a shift region mask is used. The number of deleted pixels in the neighborhood of each pixel is counted using a one-dimensional sliding window of length l, in the deletion-retention binary map. Perform one-dimensional convolution on top: ; in, This represents a one-dimensional convolution, where K is a convolution kernel of length l with a value of 1. Exceeding the threshold The region is the area with large pixel shift.
4. The multi-aspect ratio image retargeting method based on key semantic preservation as described in claim 3, characterized in that: The convolution kernel in step S3 The size is 25, the value at each position is 1, and the threshold is... It is 12.
5. The multi-aspect ratio image retargeting method based on key semantic preservation as described in claim 3, characterized in that: In step S4, a mask is overlaid on the initial redirection result and input into the Controlnet model; the original image is used as a control condition and input into the IP-Adapter model. Both of these methods work together to assist the Stable diffusion model in generating the image.
6. A multi-aspect ratio image retargeting system based on key semantic preservation, implementing the method as described in claim 1, characterized in that: It includes at least a preliminary redirection module, a redrawing region selection module, and an image guidance generation module; The initial redirection module: based on the semantic gap carving algorithm, it introduces a saliency map with spatial prior and combines it with a gradient operator to delete low-energy linear regions to obtain the initial redirection result; The redrawing region selection module: takes the output of the initial redirection module as input, regenerates the regions with large pixel shifts in the initial redirection image, and expands the image for semantic gap carving results that do not achieve the target aspect ratio; The image-guided generation module: based on the output of the redrawing region selection module, uses the original image as a control condition to guide local redrawing and output the image.
7. The multi-aspect ratio image retargeting system based on key semantic preservation as described in claim 6, characterized in that: The image-guided generation module includes at least a ControlNet and an IP-Adapter-assisted stable diffusion model, where ControlNet is used to control the redraw area and IP-Adapter is used to control the generated content.
Citation Information
Patent Citations
Single-scale motion blurred image frame restoration method for cab environment
CN112488946A
Monocular view depth estimation method based on neural radiation field and semantic segmentation
CN115393410A