Document Review Assistance With Sentence Vectors for Different Conclusions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document review systems struggle to efficiently exclude documents with similar research backgrounds but different conclusions, leading to increased burden due to the need to review numerous documents.
Innovation Solution
A document review assistance method involving a computer system that creates sentence vectors, clusters documents, specifies subgraphs of word networks, and adjusts sentence vectors based on these subgraphs to improve classification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If keyword similarity is used to determine document relevance, then documents with similar research backgrounds are identified, but documents with different conclusions cannot be efficiently excluded
Solution Approach 1:
The patent transforms the document representation from simple keyword vectors to sentence vectors that capture semantic meaning. By changing the parameter from keyword matching to sentence-level semantic vector comparison, the system can distinguish between documents with similar backgrounds but different conclusions, improving classification accuracy while maintaining review efficiency.
Solution Approach 2:
The patent replaces the mechanical keyword-matching system with a neural network-based semantic vector system. This substitution enables the system to understand the meaning and context of sentences, allowing it to accurately differentiate documents based on their actual content and conclusions rather than just shared keywords.
2Reliability
If all documents with similar keywords are reviewed, then important papers are not overlooked, but the review burden increases significantly
Solution Approach 1:
The system changes from keyword-based document grouping to sentence vector-based clustering. This parameter change enables more precise document classification, grouping together documents with truly similar meanings rather than just shared keywords, thereby reducing the number of documents reviewers need to examine while maintaining reliability.
Solution Approach 2:
The patent segments documents into clusters based on sentence vector similarity. By dividing the large set of documents into smaller, more homogeneous clusters, the system allows reviewers to focus on specific clusters relevant to their research question, significantly reducing the total review time while ensuring important papers are not missed.
3Adaptability or versatility
If keyword vector similarity is used for document classification, then documents are grouped by background, but classification accuracy for distinguishing different conclusions is insufficient
Solution Approach 1:
The patent substitutes the keyword vector similarity approach with sentence vector comparison using neural networks. This replacement enables the system to capture nuanced semantic differences in document conclusions while maintaining the ability to cluster documents by background, thereby improving classification accuracy without losing adaptability.
Solution Approach 2:
The system creates a composite representation by combining sentence vector features with clustering algorithms. This composite approach integrates both the semantic understanding of individual sentences and the grouping capability, achieving high classification accuracy while maintaining versatile document clustering functionality.
Data Source
AI summary
In screening of documents based on a similarity between keywords, it is difficult to exclude documents having similar background but different conclusions. In a document review assistance method executed by a computer system, a storage unit stores data on a plurality of documents, and the document review assistance method includes: a step of creating, by a control unit, a sentence vector based on a sentence included in the plurality of documents; a step of classifying, by the control unit, the plurality of documents into a plurality of clusters based on the created sentence vector; a step of specifying, by the control unit, a subgraph on a network of a word in a first document set included in at least one of the clusters; and a step of controlling, by the control unit, the creation of the sentence vector based on the specified subgraph.


