Usage Log Cleansing for Research Impact Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional measures for assessing research impact, such as citation counts, are slow and inadequate for newly published works, while alternative metrics (altmetrics) face challenges in distinguishing between human and robotic traffic in information retrieval systems, leading to inaccurate reflection of research interest.
Innovation Solution
A method and system for analyzing usage logs to differentiate between human and automated software robot behavior by identifying patterns and characteristics, such as investment and payoff events, to cleanse data and generate accurate metrics and indicators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional citation counts are used to measure research impact, then measurement reliability is improved, but speed of providing feedback is worsened (taking years to accrue)
Solution Approach 1:
The patent applies preliminary action by capturing and analyzing usage data immediately when research artifacts are accessed, rather than waiting for citations to accumulate. Usage logs record interactions such as downloads, views, and shares in real-time, providing early indicators of research impact before traditional citation metrics become available.
Solution Approach 2:
The patent introduces an intermediary metric system that bridges the gap between immediate usage data and long-term citation impact. By analyzing usage logs as an intermediate data source, the system provides timely feedback that complements traditional citation counts, allowing researchers to gauge interest velocity without sacrificing the reliability of established metrics.
2Speed
If usage logs are used to provide high-velocity indicators of research interest, then speed of feedback is improved, but measurement precision is worsened due to robotic traffic contamination
Solution Approach 1:
The patent applies segmentation by dividing usage log analysis into distinct components: identifying investment events (search queries, refinements), payoff events (downloads, shares), and bulk acquisition patterns. This segmentation allows the system to differentiate between genuine research interest and automated robotic traffic, thereby maintaining measurement precision while preserving the speed advantage of usage-based metrics.
Solution Approach 2:
The patent implements feedback mechanisms by continuously monitoring usage patterns and adjusting metric calculations accordingly. The system analyzes the relationship between investment events and payoff events, using this feedback to distinguish human research behavior from automated traffic, thus maintaining accurate measurements at high velocity.
3Ease of operation
If pre-established IP address identification is used to filter robotic traffic, then ease of operation is improved, but adaptability is worsened due to complexity of human-robot interactions
Solution Approach 1:
The patent applies parameter changes by shifting from static IP address-based filtering to dynamic behavioral parameter analysis. Instead of relying on fixed identification methods, the system monitors multiple parameters including search query patterns, time between interactions, and sequences of events. This allows the system to adapt to evolving human-robot interaction complexities while maintaining ease of operation through automated analysis.
Data Source
AI summary
Exemplary embodiments of the present disclosure provide for cleansing data generated by one or more servers in response to database interactions resulting from an automated software robot interacting with the one or more servers via a telecommunications network. Log entries in usage logs corresponding to events during a session can be analyzed to determine relationships between events and the usage logs can be classified based on the relationships as either corresponding to human behavior or automated software robot behavior. Usage logs corresponding to automated software robot behavior can be removed from further analysis.


