Front-End Image ROI Extraction for Faster AI Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Processing a large number of images and complex image content can severely affect the efficiency and accuracy of image analysis by AI models.
Innovation Solution
A method and system that involves a front-end device to obtain images and user voice prompts, identify a region of interest based on gestures, generate a target image by combining relevant images into a panorama or masked image, and transmit this to a back-end device for analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a large number of images are provided to AI for analysis, then the comprehensiveness of analysis coverage is improved, but the efficiency of image analysis deteriorates
Solution Approach 1:
The system segments the large set of images by identifying and extracting only the region of interest (ROI) containing the target object. Instead of processing all images, the front-end device identifies the specific area where the target appears and extracts only that portion for transmission to the back-end AI device, thereby reducing the total number of images processed while maintaining analysis completeness.
Solution Approach 2:
The system extracts the essential information from multiple images by identifying the target object's location and generating a simplified target image or bounding box representation. This extraction process removes redundant background and non-target elements, leaving only the critical information needed for accurate AI analysis.
2Loss of information
If complex image content is provided to AI, then the detail richness of analysis is improved, but the accuracy of image analysis deteriorates
Solution Approach 1:
The system extracts only the relevant target object or region from complex images, removing distracting background elements, irrelevant objects, and complex environmental details. This extraction simplifies the input to the AI model while preserving the essential features of the target, thereby improving analysis accuracy without losing critical information.
Solution Approach 2:
The system applies different processing quality levels to different parts of the image. The region of interest containing the target object is processed with high precision and detail preservation, while surrounding areas are simplified or removed. This localized quality adjustment ensures that the AI model receives high-quality input for the target while reducing the overall complexity of the input data.
3Reliability
If all first images are transmitted to back-end device, then the completeness of data is improved, but the transmission time and processing load increase
Solution Approach 1:
The system extracts only the necessary image data by identifying the target object's location and extracting either the specific region containing the target or a simplified representation such as a bounding box. This extraction reduces the data volume significantly while maintaining the completeness of target-related information, thereby reducing transmission time and back-end processing load.
Solution Approach 2:
The front-end device performs preliminary processing of images before transmission, including target identification, region extraction, and image simplification. This preliminary action prepares the data in advance, ensuring that only the essential information is transmitted to the back-end device, thereby reducing transmission time and processing requirements.
Data Source
AI summary
The embodiments of the disclosure provide a method and system for improving image analysis, and a computer readable storage medium. The method includes: obtaining, by a front-end device, a plurality of first images and a user voice prompt; generating, by the front-end device, a target image based on the user voice prompt and the plurality of first images by performing at least one of following operations: identifying a region of interest from the plurality of first images based on a first gesture and generating the target image according to the region of interest; combining at least a part of the plurality of first images into a panorama image as the target image; and transmitting, by the front-end device, the target image and the user voice prompt to a back-end device.


