Issue Network Generation from Document Co-occurrences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying and organizing issues within a document corpus are inefficient, as they rely on citation networks that do not clearly indicate the specific issues or topics discussed, leading to incomplete investigations due to the lack of understanding of interconnected issues.
Innovation Solution
A computer-implemented method and system that generate an issue network by searching for documents discussing a starting issue, determining co-occurrences of normalized issues within these documents, and linking them based on their co-occurrences to create a structured network of interconnected issues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If citation networks are used to link documents, then documents can be connected through references, but the specific issues or topics discussed in each citation cannot be identified
Solution Approach 1:
The patent extracts issue information from document citations by analyzing the textual content surrounding citations. The system identifies and extracts specific issues or topics discussed in cited documents, separating this information from the citation structure itself. This allows the system to maintain citation networks while adding detailed issue information that was previously lost.
Solution Approach 2:
The patent introduces an intermediary layer between citations and issues. Instead of directly linking citations to issues, the system creates a mediation mechanism that analyzes citation contexts, identifies relevant issues, and establishes connections between citations and issues. This intermediary process enables systematic issue identification without requiring complex direct analysis of all document content.
2Reliability
If researchers manually sift through cited documents to understand issues, then complete investigation can be achieved, but time consumption increases significantly
Solution Approach 1:
The patent implements a self-service system that automatically performs the manual analysis process. The system independently identifies issues from citations, determines relationships between issues, and organizes information without requiring researcher intervention. This automation maintains research completeness while eliminating the time-consuming manual sorting process.
Solution Approach 2:
The patent replaces the mechanical manual process of sifting through documents with an automated computational system. Instead of researchers manually reading and analyzing each citation, the system uses algorithms to process citation data, identify issues, and build networks automatically. This substitution maintains analytical thoroughness while dramatically reducing time requirements.
3Measurement precision
If documents are searched for specific issues, then relevant documents can be found, but the interconnectedness of issues cannot be understood
Solution Approach 1:
The patent merges multiple functions into a unified system: issue identification, relationship detection, and network visualization are combined into a single integrated approach. By merging these functions, the system can accurately identify specific issues while simultaneously detecting their interconnectedness, eliminating the difficulty of measuring relationships that existed when functions were separate.
Solution Approach 2:
The patent adds a new dimension to issue analysis by creating a network representation that shows relationships between issues. Instead of analyzing issues in isolation (zero-dimensional) or listing them sequentially (one-dimensional), the system creates a multi-dimensional network structure that visualizes and measures connections between issues, making previously undetectable relationships visible and measurable.
Data Source
AI summary
Systems and methods for generating issue networks are disclosed. In one embodiment, a computer-implemented method of generating an issue network from a document corpus includes searching, using a computer, the document corpus for a set of documents discussing a starting issue, wherein the starting issue is one of a plurality of normalized issues defined by the document corpus. The method further includes determining a set of normalized issues discussed by the set of documents discussing the starting issue, wherein the set of normalized issues also includes the starting issue, and determining instances of co-occurrences of individual normalized issues of the set of normalized issues within individual cases of the set of documents. The method also includes linking individual normalized issues of the set of normalized issues based on their co-occurrences within the set of documents, wherein the linked individual normalized issues at least in part define the issue network.


