Image Object and Text Tag Matching for Automated Editing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The burden on users to produce, process, and edit images with associated text in social networking applications is increased due to the inability of systems to accurately classify or process images without understanding user intent, as they cannot distinguish which parts of an image should be accompanied by text.
Innovation Solution
A computer-implemented method that identifies objects and text portions in an image, extracts object and text tags, determines if text tags describe the objects, and performs image and text processes such as masking or emphasizing based on these determinations, using object and text tag comparisons and contextual information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual image editing is performed by users, then image customization is achieved, but user burden and time consumption increase
Solution Approach 1:
The system performs automatic image processing by itself without requiring user intervention. The server automatically identifies objects in images, extracts tags, determines relationships between text and objects, and applies appropriate processing (masking or emphasizing) based on the determined relationships, making the system serve itself rather than requiring manual user operation
Solution Approach 2:
The patent replaces manual mechanical editing operations with automated computer vision and natural language processing systems. Object identification, tag extraction, and text-object relationship determination are performed automatically using AI algorithms, substituting the need for manual user analysis and editing decisions
2Extent of automation
If systems process images without understanding user intent, then automated processing is achieved, but classification accuracy deteriorates
Solution Approach 1:
The patent introduces text portions as an intermediary that bridges the gap between automated processing and user intent. The text provides contextual information that helps the system understand what the user wants to emphasize or mask, allowing automated processing to achieve accurate classification by using the text as a mediator between the image content and processing decisions
Solution Approach 2:
The system performs preliminary analysis by extracting tags from both image objects and text portions before determining the final processing actions. This preliminary tag extraction and comparison process prepares the necessary information in advance, enabling accurate classification and appropriate image processing without requiring real-time user input
Data Source
AI summary
A computer-implemented method includes receiving an image. The image includes one or more objects and one or more text portions. The computer-implemented method further includes identifying the one or more objects. The computer-implemented method further includes, for each of the one or more objects identified, extracting an object tag. The computer-implemented method further includes, for each of the one or more text portions, extracting a text tag. The computer-implemented method further includes, for each text tag, determining whether the text tag describes any of the one or more objects based on the object tag extracted from each object to yield a determination. The computer-implemented method further includes, responsive to the determination: performing an image process to that of the one or more objects, and performing a text process to that of the one or more text portions. A corresponding computer program product and computer system are also disclosed.


