Wide-Angle Image Generation Using Semantic Segmentation Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image-based lighting technologies lack accuracy in estimating environments outside the angle of view, resulting in suboptimal image generation for wide-angle views.

Innovation Solution

A learning device and method that incorporates semantic segmentation and generative adversarial networks to generate images with wider angles of view by coupling input images with segmentation results, using machine learning models to estimate and complement environments beyond the original image frame.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If an image is generated by using a learned machine learning model to supplement environment outside the angle of view, then the angle of view is widened, but the accuracy of estimation of the environment outside the angle of view is insufficient

Engineering Contradiction:
Improveangle of viewVSAvoidaccuracy of estimation of environment outside angle of view
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing the image processing into two distinct stages: first performing semantic segmentation to classify pixels into semantic categories (sky, ground, buildings, etc.), then using these segmentation results as guidance for generating the extended wide-angle view. This segmentation approach ensures that the estimated environment outside the original angle of view maintains semantic consistency and structural合理性, thereby improving estimation accuracy while achieving a wider angle of view.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If dedicated equipment such as omnidirectional camera is used to photograph wide angle of view image, then the accuracy of wide angle of view image is improved, but the complexity of equipment and expertise required increases

Engineering Contradiction:
Improveaccuracy of wide angle of view imageVSAvoidcomplexity of equipment and expertise
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses copying by training a machine learning model on datasets containing wide-angle or omnidirectional images. The model learns to generate synthetic wide-angle views that replicate the quality and accuracy of images captured by dedicated omnidirectional cameras, but without requiring such specialized equipment. The model copies the essential characteristics of professional wide-angle imagery through learned patterns from training data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical system of dedicated omnidirectional camera hardware with a software-based machine learning model. Instead of physically capturing wide-angle views with specialized optical equipment, the system uses computational methods to synthesize wide-angle images from ordinary input images, substituting mechanical complexity with algorithmic processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If semantic segmentation results are coupled with input image to improve estimation accuracy, then the accuracy of environment estimation is improved, but the complexity of data processing increases

Engineering Contradiction:
Improveaccuracy of environment estimationVSAvoidcomplexity of data processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing semantic segmentation on the input image before using it to generate the wide-angle view. This pre-processing step categorizes different regions of the image (sky, ground, objects) in advance, providing structured guidance for the subsequent image generation process. This preliminary segmentation ensures that the extended environment maintains semantic consistency, improving estimation accuracy while organizing data processing in a systematic sequence.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11900258B2Learning device, image generating device, learning method, image generating method, and program
Publication Date: 2024.02.13 SONY INTERACTIVE ENTERTAINMENT LLC
  • US11900258B2 patent drawing
  • US11900258B2 patent drawing
  • US11900258B2 patent drawing

AI summary

A learning device, an image generating device, a learning method, an image generating method, and a program are provided which can improve accuracy of estimation of an environment outside an angle of view of an image input to an image generating section. A second learning data obtaining section (64) obtains an input image. A second learning section (66) obtains result data indicating a result of execution of semantic segmentation on the input image. An input data generating section (38) generates input data obtained by coupling the input image and the result data to each other. By using the input data as input, the second learning section (66) performs learning of a wide view angle image generating section (28) that generates, in response to input of data obtained by coupling an image to a result of execution of semantic segmentation on the image, an image having a wider angle of view than the image.