Information Extraction Using Word Vector Reliability Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data extraction techniques face challenges in securing sufficient data, achieving high precision, and determining the correctness of extraction results, often resulting in inefficient and unreliable data extraction processes.
Innovation Solution
An information extracting device and method that acquire, generate, and classify data groups using word vectors based on date, time, location, and content, calculating distances and reliability to efficiently extract relevant information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional data extraction techniques are used, then data can be extracted from information sources, but the precision of extraction results is low and it cannot be determined whether the extraction results are correct
Solution Approach 1:
The patent implements feedback by calculating reliability scores for extracted data based on word vector distances and classification results. The system continuously evaluates the quality of extracted information and uses this feedback to improve future extraction decisions, ensuring higher precision and reliability of results
Solution Approach 2:
The patent replaces manual judgment of data quality with automated computational methods. Word vector-based distance calculations and classification algorithms substitute for human evaluation, enabling objective assessment of extraction result correctness and significantly improving measurement precision
2Productivity
If multiple pieces of data exceed the threshold value, then more data can be extracted, but it becomes difficult to determine which data to employ as extraction results
Solution Approach 1:
The patent changes the parameter of data selection by introducing reliability scores calculated from word vector distances. Instead of manually evaluating multiple candidate data pieces, the system automatically ranks them by reliability, making it easy to select the most appropriate extraction results while maintaining high productivity
Solution Approach 2:
The system performs self-service by automatically evaluating and ranking multiple candidate data pieces based on their reliability scores. The extraction process autonomously determines which data to employ without requiring manual intervention, simplifying the selection process while maintaining high data extraction volume
3Productivity
If data is extracted without sufficient verification, then extraction speed is high, but the correctness of extraction results cannot be determined
Solution Approach 1:
The patent applies preliminary action by calculating word vectors and determining their distances before final data extraction. This pre-processing step establishes a foundation for reliability assessment, enabling both high extraction speed and accurate correctness determination through automated evaluation
Data Source
AI summary
An information extracting device includes an acquiring unit that acquires, with regard to each information source, a data group made up of data including a content relating to an object, a location where the content was recorded, and a date and time at which the content was recorded, a generating unit that generates a word vector of which the date and time, the location, content using a weight based on the date and time, and a type of the content, are each components, for each of the data of the data group, a distance calculating unit that calculates a distance among the word vectors, a classifying unit that classifies each of data of the data group on the basis of the distance among the word vectors, and an extracting unit that calculates a reliability with regard to the classification, and extracts data from the data group on the basis of the reliability.


