Information Extraction Using Word Vector Reliability Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data extraction techniques face challenges in securing sufficient data, achieving high precision, and determining the correctness of extraction results, often resulting in inefficient and unreliable data extraction processes.

Innovation Solution

An information extracting device and method that acquire, generate, and classify data groups using word vectors based on date, time, location, and content, calculating distances and reliability to efficiently extract relevant information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional data extraction techniques are used, then data can be extracted from information sources, but the precision of extraction results is low and it cannot be determined whether the extraction results are correct

Engineering Contradiction:
Improveprecision of extraction resultsVSAvoidreliability of extraction results
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent implements feedback by calculating reliability scores for extracted data based on word vector distances and classification results. The system continuously evaluates the quality of extracted information and uses this feedback to improve future extraction decisions, ensuring higher precision and reliability of results

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces manual judgment of data quality with automated computational methods. Word vector-based distance calculations and classification algorithms substitute for human evaluation, enabling objective assessment of extraction result correctness and significantly improving measurement precision

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If multiple pieces of data exceed the threshold value, then more data can be extracted, but it becomes difficult to determine which data to employ as extraction results

Engineering Contradiction:
Improvequantity of extracted dataVSAvoidease of selecting extraction results
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent changes the parameter of data selection by introducing reliability scores calculated from word vector distances. Instead of manually evaluating multiple candidate data pieces, the system automatically ranks them by reliability, making it easy to select the most appropriate extraction results while maintaining high productivity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs self-service by automatically evaluating and ranking multiple candidate data pieces based on their reliability scores. The extraction process autonomously determines which data to employ without requiring manual intervention, simplifying the selection process while maintaining high data extraction volume

Inventive Principle:
Principle #25Self-service

3Productivity

If data is extracted without sufficient verification, then extraction speed is high, but the correctness of extraction results cannot be determined

Engineering Contradiction:
Improveextraction speedVSAvoidcorrectness of extraction results
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by calculating word vectors and determining their distances before final data extraction. This pre-processing step establishes a foundation for reliability assessment, enabling both high extraction speed and accurate correctness determination through automated evaluation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11995115B2Information extracting device, information extracting method, and information extracting program
Publication Date: 2024.05.28 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11995115B2 patent drawing
  • US11995115B2 patent drawing
  • US11995115B2 patent drawing

AI summary

An information extracting device includes an acquiring unit that acquires, with regard to each information source, a data group made up of data including a content relating to an object, a location where the content was recorded, and a date and time at which the content was recorded, a generating unit that generates a word vector of which the date and time, the location, content using a weight based on the date and time, and a type of the content, are each components, for each of the data of the data group, a distance calculating unit that calculates a distance among the word vectors, a classifying unit that classifies each of data of the data group on the basis of the distance among the word vectors, and an extracting unit that calculates a reliability with regard to the classification, and extracts data from the data group on the basis of the reliability.