Generative Model Document Division Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document summarization technologies struggle to divide document data at appropriate positions, leading to inefficient user operations such as copy operations, as they rely on preset rules like morphological and syntax analysis, which may not align with user intentions.

Innovation Solution

An information processing apparatus and method that generates a character string set by dividing document data into different lengths, derives evaluation values for each character string using a generative model, and selects optimal character strings based on these values to improve document data division accuracy, allowing for user-supported operations like copy operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If preset rules such as morphological analysis and syntax analysis are used to divide document data, then the division process is simple and fast, but the division position may not be appropriate for user intentions

Engineering Contradiction:
Improvedivision speedVSAvoiddivision position accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

An automatic selection unit is introduced as an intermediary between the character string generation unit and the user. This unit automatically selects appropriate character strings from multiple candidates based on evaluation values, mediating between the mechanical division process and user needs, thereby resolving the contradiction between fast automated division and accurate user-intent-aligned division positions

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system calculates evaluation values for each character string candidate and uses this feedback to automatically select the most appropriate character strings. This feedback mechanism enables the system to self-correct and optimize division positions based on multiple criteria, improving division accuracy while maintaining automated efficiency

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the user manually selects the range of character strings to be copied, then the division position can be precise, but the user operation time increases

Engineering Contradiction:
Improvedivision position accuracyVSAvoiduser operation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The automatic selection unit performs self-service by automatically selecting appropriate character strings from multiple candidates based on evaluation values. This eliminates the need for users to manually select character string ranges, reducing user operation time while maintaining high division position accuracy through automated intelligent selection

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The automatic selection unit acts as an intermediary that takes over the manual selection task from the user. It automatically chooses appropriate character strings based on evaluation criteria, thereby resolving the contradiction by providing precise division positions without requiring user time investment

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240330592A1Information processing apparatus, information processing method, and information processing program
Publication Date: 2024.10.03 FUJIFILM CORP
  • US20240330592A1 patent drawing
  • US20240330592A1 patent drawing
  • US20240330592A1 patent drawing

AI summary

An information processing apparatus generates a first character string set by dividing first document data into different lengths, derives an evaluation value of each character string constituting the first character string set by using the first character string set and second document data created from the first document data in accordance with a purpose, and selects, based on the derived evaluation value, a plurality of character strings from the first character string set as correct answer data of a generative model that receives input of document data and outputs a second character string set including a plurality of character strings included in the input document data.