Document Issue Time Estimation via Temporal Proximity Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for estimating the time of issue of documents, especially those in HTML format or published on the Internet, face challenges due to diverse representation formats and lack of associated metadata, leading to inaccurate and costly manual determination processes.
Innovation Solution
An information estimation device and method that extracts time representations from documents, generates candidate issue times, and calculates temporal proximities between them to accurately estimate the time of issue without operator intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual determination of time of issue is performed for documents without metadata, then accuracy of time estimation can be maintained, but cost and time consumption increase significantly
Solution Approach 1:
The patent segments the time determination process into automated extraction of candidate time representations followed by manual selection. The system divides the document into searchable components and extracts potential time indicators automatically, reducing the manual workload from complete determination to verification and selection among candidates.
Solution Approach 2:
The system performs self-service by automatically extracting time representations from documents without requiring complete manual analysis. The automated extraction of candidate times from document text allows the system to serve itself in the initial determination phase, freeing human operators to focus only on selecting the correct candidate from the generated list.
2Productivity
If automated extraction of time representations is performed using frequency-based methods, then productivity increases, but measurement precision decreases due to inaccurate time identification
Solution Approach 1:
The patent applies dynamics by making the selection criterion flexible rather than static. Instead of relying solely on fixed frequency rules, the system dynamically adapts by allowing operators to select from multiple candidate time representations extracted from the document, enabling the system to handle diverse document formats and time representation styles effectively.
Solution Approach 2:
The system changes the parameter of time representation extraction from frequency-based alone to a combination of extraction and human selection. This parameter change allows the system to maintain high productivity through automated extraction while improving precision through human judgment in selecting the correct time of issue from candidates.
3Measurement precision
If diverse representation formats in HTML documents are manually interpreted, then accurate time of issue can be identified, but device complexity and operational difficulty increase
Solution Approach 1:
The patent extracts time representations from diverse HTML document formats using automated text extraction techniques. By taking out the time-related information from various document structures and presenting it as standardized candidate representations, the system reduces the complexity of interpreting diverse formats while maintaining accuracy through human selection.
Solution Approach 2:
Instead of requiring the system to fully interpret and understand diverse document formats, the patent inverts the approach by extracting raw time representations and presenting them to human operators for selection. This inversion reduces device complexity by avoiding complex format interpretation while maintaining precision through human judgment.
4Ease of operation
If collection time is used as time of issue for all documents, then operational complexity is reduced, but reliability decreases as not all documents can be collected without delay
Solution Approach 1:
The patent introduces an intermediary step between automatic time determination and final time of issue assignment. Instead of directly using collection time or fully manual determination, the system extracts candidate time representations from the document content itself and presents them as intermediaries for operator selection, ensuring both ease of operation and reliability.
Solution Approach 2:
The system incorporates feedback by allowing operators to review and select from extracted time candidates. This feedback mechanism ensures that the final time of issue assignment is both reliable (accurate to the document content) and easy to operate (automated extraction with simple selection), resolving the contradiction between simplicity and accuracy.
Data Source
AI summary
Disclosed is an information estimation device for estimating an appropriate issue time from a time representation described in a document without intervention of any operator; wherein an information estimation device (1) which is a device for estimating an issue time of a document to be estimated, includes a candidate generation unit (11) which extracts a time representation described in the document, and on the basis of the extracted time representation, generates a plurality of possible issue time candidates having possibilities corresponding to the issue time of the document; and an issue time estimation unit (12) for obtaining a temporal proximity, for each of the plurality of issue time candidates, between the issue time candidate and other issue time candidates, and on the basis of the obtained temporal proximity, estimating the issue time of the document.


