Distinct Author Identification System for Scholarly Data Disambiguation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing databases and tools for scholarly research face inaccuracies in authorship identification due to name ambiguity, incomplete information, and common/identical names, leading to incorrect linkage of paper and citation records, which diminishes the integrity and reliability of search and retrieval processes.
Innovation Solution
The Distinct Author Identification System (DAIS) uses a method to disambiguate author information by applying linking and clustering techniques, utilizing email addresses and other data elements to create unique author entities, establishing a more accurate relational link and forming an authority database of authors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If abbreviated author names are used in citations, then citation formatting consistency is improved, but author identification accuracy deteriorates due to name ambiguity
Solution Approach 1:
The patent introduces an intermediary disambiguation system that uses multiple data elements (email addresses, institutional affiliations, co-author relationships) as mediators to resolve the ambiguity between abbreviated citation names and actual author identities. This intermediary layer connects the simplified citation format with the detailed author information needed for accurate identification.
Solution Approach 2:
The patent transitions from one-dimensional author identification (using only the abbreviated name string) to multi-dimensional identification by incorporating multiple attributes such as email addresses, institutional affiliations, publication years, and co-author networks. This dimensional expansion allows the system to distinguish between authors with identical or similar abbreviated names.
2Measurement precision
If multiple data elements are collected for author identification, then authorship linking accuracy is improved, but system complexity increases
Solution Approach 1:
The patent segments the author identification process into distinct modules: data extraction from citations, data element validation, similarity computation, and disambiguation decision-making. Each module processes specific aspects of the data independently, making the overall complex system more manageable and maintainable while improving accuracy through specialized processing at each stage.
Solution Approach 2:
The patent performs preliminary actions by pre-processing and validating data elements before they are used in the disambiguation process. This includes normalizing email addresses, verifying institutional affiliations, and pre-computing similarity metrics, which reduces the computational burden during the main disambiguation task and simplifies the real-time processing requirements.
3Productivity
If existing citation databases are used, then research productivity is maintained, but data integrity deteriorates due to incorrect author linkages
Solution Approach 1:
The patent implements feedback mechanisms that allow the system to learn from and correct its own errors. By monitoring disambiguation results and incorporating user corrections, the system continuously refines its author identification algorithms, improving data integrity over time while maintaining high research productivity through automated processing.
Solution Approach 2:
The patent applies beforehand cushioning by implementing robust validation and cross-checking procedures before finalizing author linkages. The system uses multiple independent data elements and similarity thresholds to cushion against potential errors, ensuring that only high-confidence author identifications are accepted, thus protecting data integrity without significantly impacting processing speed.
Data Source
AI summary
The present invention provides a Distinct Author Identification System (“DAIS”) for disambiguating data to discern author entities and link or associate authorships with such author entities. The invention provides powerful disambiguation processes applied across one or more databases to yield a disambiguated authority database of authors. An entire database of publications may be processed by the DAIS to group/link authorships and to identify author entities. The author entities may then be matched or associated with actual authors to establish an authority database of authors. After initial evaluation, the DAIS may be used to reevaluate some or all of the database(s) and/or the authority database established by the DAIS may be used to add or update information. DAIS may use “hierarchical clustering” to link authorships and identify authors based on authorship similarity. DAIS evaluates the likelihood that authorships are from the same author.


