Image-Text Fusion Layout Optimization for Salient Object Preservation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image-text fusion methods fail to effectively avoid blocking visually salient objects in images and optimize text layout to achieve a better aesthetic effect, as they lack a comprehensive approach to determining the attention-grabbing potential of image pixels and balancing feature value distribution.
Innovation Solution
The method determines feature values for each pixel in an image by considering visual saliency, face, edge, and text features, and uses these values to select layout formats that minimize blocking of salient objects, balancing the distribution of feature values and optimizing text placement using cost parameters and texture features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If text is laid out in an image using conventional methods, then text placement is simple and fast, but visually salient objects (such as human faces, flowers, or buildings) are blocked by the text
Solution Approach 1:
The patent segments the image into multiple feature maps (visual saliency map, face feature map, edge feature map, text feature map), where each map represents a different type of feature. By dividing the complex layout problem into separate feature dimensions, the system can evaluate and optimize text placement across multiple criteria simultaneously, ensuring salient objects are not blocked while maintaining layout complexity management.
Solution Approach 2:
The patent changes the parameter representation from simple binary occupancy to continuous feature values ranging from 0 to 1, where each pixel's feature value represents the probability of user attention. This parameter transformation enables gradient-based optimization and allows the system to balance multiple competing requirements (avoiding salient objects vs. maintaining aesthetic layout) through continuous parameter adjustment rather than discrete decisions.
2Reliability
If text layout avoids all salient objects, then visual salience is preserved, but the aesthetic effect and visual balance of the layout deteriorates
Solution Approach 1:
The patent applies local quality by allowing different regions of the image to have different layout characteristics based on their feature values. High-feature-value regions (salient objects) are protected from text blocking, while low-feature-value regions are more flexible for text placement. The system dynamically adjusts layout constraints locally rather than applying uniform rules across the entire image, thus preserving both salience and aesthetic balance.
Solution Approach 2:
The patent implements feedback through the cost function that evaluates layout quality based on feature value distribution. The system calculates cost parameters for different layout formats, considering both the magnitude of blocked feature values and the balance degree of feature value distribution. This feedback mechanism allows the system to iteratively optimize layout decisions, adjusting text placement to achieve both salience preservation and aesthetic balance.
3Manufacturing precision
If multiple candidate layout formats are evaluated, then layout quality improves, but computational time and complexity increase
Solution Approach 1:
The patent performs preliminary action by pre-calculating feature maps (visual saliency, face features, edges, text features) and their weighted sums before evaluating layout formats. This preprocessing step creates a foundation that can be efficiently reused when evaluating multiple candidate layouts, avoiding redundant computations and reducing the time penalty associated with evaluating multiple layout options.
Solution Approach 2:
The patent substitutes mechanical enumeration and manual evaluation of layout options with an automated cost-function-based selection system. Instead of relying on manual design or simple heuristics, the system uses a computational cost function that automatically evaluates and compares multiple layout formats based on feature value distribution and aesthetic criteria, efficiently identifying the optimal layout without exhaustive search.
Data Source
AI summary
This application relates to the field of digital image processing technologies, and discloses an image-text fusion method and apparatus, and an electronic device, to minimize blockage of a saliency feature in an image by a text when the text is laid out in the image, and obtain a higher visual balance degree after the text is laid out in the first image, thereby achieving a better layout effect. According to the method of this application, first, a plurality of candidate text templates and layout positions of a plurality of corresponding texts in an image can be determined, so that a text laid out in the image does not block a visually salient object having a greater feature value, such as a human face or a building. Then based on magnitudes of feature values of pixels blocked by the text, a balance degree of feature value distribution of pixels in each region in the image in which the text is laid out, and the like when the text is laid out in the image at corresponding layout positions in the image by using different text templates, a final text template of the text and a layout position of the text in the image are determined.


