Image Explanatory Note Generation Using Text-Guided AI Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for generating explanatory notes of images using language models suffer from low accuracy and lack the ability to support decision-making based on the generated notes.

Innovation Solution

An information processing apparatus and method that acquires text associated with an image, uses a generation model trained through machine learning to generate an explanatory note, and integrates analysis methods to improve accuracy and support decision-making.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If a language model is used to generate an explanatory note of an image, then the generation process can be automated, but the generation accuracy is low

Engineering Contradiction:
Improveautomation of explanatory note generationVSAvoidgeneration accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent combines multiple input modalities (image data and text data) into a unified processing framework. The generation model receives both the target image and associated text as inputs, merging visual and linguistic information to produce more accurate explanatory notes than image-only approaches.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an embedding process as an intermediary step that converts both image and text inputs into a common representation space. This embedding layer acts as a mediator that enables the generation model to effectively process and integrate heterogeneous input types (visual and textual data) before generating the explanatory note.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If only image data is input to a generation model, then the process is simple, but the generation accuracy of explanatory notes is insufficient

Engineering Contradiction:
Improveinput processing complexityVSAvoidexplanatory note accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system merges image data and text data into a unified input framework for the generation model. By combining multiple data types rather than processing them separately, the system achieves better accuracy without requiring complex separate processing pipelines for each modality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The generation model is designed to handle multiple input types (images and text) through a universal processing architecture. The model accepts heterogeneous inputs and processes them through a common embedding and generation pathway, achieving multi-functionality without requiring separate specialized models for each input type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20260011163A1Information processing apparatus, analysis method, and non-transitory computer-readable recording medium
Publication Date: 2026.01.08 NEC CORP
  • US20260011163A1 patent drawing
  • US20260011163A1 patent drawing
  • US20260011163A1 patent drawing

AI summary

An information processing apparatus acquires a text associated with a target image which is an analysis target, and causes a generation model to generate an explanatory note of the target image according to content of the text. The generation model is obtained by performing machine learning to generate an explanatory note of an image. The information processing apparatus causes the generation model to generate a more detailed explanatory note of the target image by using the explanatory note generated by the generation model.