GAN-Based Image Resizing for Subject Prominence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digital compression methods struggle to efficiently resize images and videos while maintaining the prominence of the subject and adhering to user-defined file size constraints, often requiring manual editing that can distort the image.

Innovation Solution

A computer-implemented method using a Generative Adversarial Network (GAN) to automatically resize captured images based on user input, enhancing the prominence of the subject and reducing the file size without distorting the main subject.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If manual editing is used to reduce image size, then file size constraint is met, but image quality and subject prominence deteriorate

Engineering Contradiction:
Improvefile sizeVSAvoidimage quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent replaces manual mechanical editing operations with an automated neural network system. The neural network automatically identifies the main subject, determines optimal cropping regions, and resizes images while preserving quality, eliminating the need for manual intervention that degrades image quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables images to resize themselves automatically through the neural network without requiring external manual editing. The neural network autonomously performs subject identification, cropping region determination, and image resizing, allowing the image processing task to serve itself.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If manual editing is used to reduce image size, then file size constraint is met, but time consumption increases

Engineering Contradiction:
Improvefile sizeVSAvoidtime consumption
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent replaces time-consuming manual editing operations with an automated neural network system that processes images rapidly. The neural network automatically performs subject identification, cropping region determination, and resizing in a single streamlined process, dramatically reducing the time required compared to manual editing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The neural network is pre-trained on large datasets to recognize subjects and determine optimal cropping regions before actual image processing is needed. This preliminary training enables the system to quickly and accurately process new images without requiring time-consuming manual analysis for each image.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If lower resolution camera is used, then file size is reduced, but image detail and quality deteriorate

Engineering Contradiction:
Improvefile sizeVSAvoidimage detail
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent extracts and isolates the main subject from the full image through neural network-based subject identification and cropping region determination. By extracting only the essential subject portion and surrounding context, the system reduces file size while preserving all critical image details of the main subject, eliminating the need to use a lower resolution camera.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250131528A1Dynamic resizing of audiovisual data
Publication Date: 2025.04.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250131528A1 patent drawing
  • US20250131528A1 patent drawing
  • US20250131528A1 patent drawing

AI summary

Disclosed are techniques of a computer implemented method for resizing a captured image. One embodiment may comprise receiving a desired size and a subject of the captured image as input from a user, automatically resizing the captured image using a generative adversarial network (GAN) to about the desired size, where the resizing enhances a prominence of the subject of the captured as compared to the captured image, and storing the automatically resized image on a computer readable storage medium.