Distributed Representation Set Modification via Relationship Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for modifying distributed-representation sets in document data processing primarily focus on learning synonym and antonym pairs, failing to effectively utilize data pairs with various relationships, such as intermediate or distant meaning connections.

Innovation Solution

A document data processing apparatus and method that modifies distributed-representation sets by learning multiple data pairs with corresponding scores, allowing for the inclusion of various relationship types and using a loss function to minimize the difference between the relationship values and target scores, thereby improving the quality of the representation set.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If only synonym pairs and antonym pairs are used for learning in distributed-representation set modification, then the learning process is simple and focused, but data pairs with various relationships (intermediate relationships, close meaning relationships, distant meaning relationships) cannot be effectively utilized

Engineering Contradiction:
Improveability to utilize various relationship typesVSAvoidcomplexity of learning process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a relationship type parameter that categorizes data pairs into different relationship types (synonym, antonym, intermediate, close meaning, distant meaning). By changing the parameter space to include multiple relationship types rather than just two, the system can utilize diverse data pairs for learning while maintaining a structured approach through parameterized relationship classification.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If multiple data pairs with various relationships are subjected to learning, then the quality of distributed-representation set is improved, but the complexity of the learning process increases

Engineering Contradiction:
Improvequality of distributed-representation setVSAvoidcomplexity of learning process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the learning process by introducing relationship type as a segmentation parameter. Data pairs are segmented into different groups based on their relationship types (synonym, antonym, intermediate, close meaning, distant meaning). This segmentation allows the system to handle complex multi-type learning by breaking it down into manageable segments, each processed with appropriate weighting, thereby improving representation quality while controlling learning complexity through structured segmentation.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If synonym pairs and antonym pairs are subjected to learning, then the loss function can be minimized with focused data, but data pairs having various relationships are not utilized as learning targets

Engineering Contradiction:
Improveprecision of relationship representationVSAvoidvariety of learning data pairs
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent makes the learning system universal by designing it to handle multiple relationship types through a unified framework. The loss function is generalized to accommodate different relationship types (synonym, antonym, intermediate, close meaning, distant meaning) with appropriate weighting. This multi-functionality allows the same learning mechanism to process diverse data pairs, improving both the precision of relationship representation and the versatility of learning data utilization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11188598B2Document data processing apparatus and non-transitory computer readable medium
Publication Date: 2021.11.30 FUJIFILM BUSINESS INNOVATION CORP
  • US11188598B2 patent drawing
  • US11188598B2 patent drawing
  • US11188598B2 patent drawing

AI summary

A document data processing apparatus includes a memory and a processor. The memory stores a distributed-representation set including multiple distributed representations corresponding to multiple pieces of data. The processor is configured to modify the distributed-representation set on the basis of multiple data pairs and multiple scores corresponding to the data pairs. The data pairs are subjected to learning. The processor is configured to modify the distributed-representation set in such a manner that, for each of the data pairs, a value indicating a relationship in a modified distributed-representation pair corresponding to the data pair comes close to a score corresponding to the data pair.