AI Research Data Classification via Human-AI Collaboration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The process of searching for research data is often time-consuming, prone to errors, and biased due to the lack of effective classification, leading to irrelevant information being retrieved, which can hinder the discovery of relevant data necessary for research topics or questions.
Innovation Solution
A collaboration platform utilizing artificial intelligence and machine learning to automate or semi-automate the classification of research data, allowing for user-defined protocols and feedback mechanisms to refine search results and improve data relevance, enabling human-computer collaboration for efficient data curation and synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual searching and classification of research data is performed, then human understanding and contextual judgment are achieved, but the process becomes time-consuming and tedious
Solution Approach 1:
The patent introduces an AI module as an intermediary between the research data and human users. The AI module performs preliminary classification of research data using machine learning algorithms, producing structured outputs that humans can efficiently review and refine. This intermediary processing layer reduces the time humans spend on manual classification while maintaining high accuracy through collaborative refinement.
Solution Approach 2:
The system performs preliminary classification actions automatically using AI before human review. The AI module pre-processes large volumes of research data, identifies relevant documents, and organizes them according to predefined taxonomies. This preliminary action filters out obviously irrelevant data, allowing humans to focus their time on reviewing and refining the AI's classifications rather than starting from scratch.
2Reliability
If comprehensive searching is conducted to ensure all relevant data is found, then completeness is achieved, but irrelevant information increases and requires further curation
Solution Approach 1:
The patent applies different classification criteria and filtering mechanisms to different portions of the search results based on their relevance indicators. The AI module identifies high-confidence matches and applies streamlined processing, while lower-confidence results receive more stringent filtering and human review. This localized quality control ensures comprehensive coverage while minimizing irrelevant information in the final curated set.
Solution Approach 2:
The system implements feedback loops where human reviewers assess the AI's classification accuracy and relevance judgments. This feedback is used to continuously refine the AI's algorithms and filtering criteria. The feedback mechanism enables the system to learn from errors and improve its ability to distinguish relevant from irrelevant information, maintaining completeness while reducing noise over time.
3Productivity
If automated classification is implemented to reduce time and effort, then processing speed increases, but accuracy and reliability may decrease
Solution Approach 1:
The patent implements a dynamic classification system where the level of automation adjusts based on the confidence level of AI predictions. For high-confidence classifications, the system operates fully automatically to maximize speed. For lower-confidence cases, the system dynamically transitions to semi-automated mode, inviting human review to maintain accuracy. This dynamic approach optimizes the balance between speed and precision across different types of classification tasks.
Solution Approach 2:
The system applies partial automation selectively to portions of the data where AI confidence is high, while reserving full human review for cases where accuracy is critical or AI uncertainty is high. This partial automation strategy achieves speed improvements on the majority of straightforward cases without sacrificing accuracy on complex or ambiguous classifications. The system performs excessive classification actions (both AI and human) on uncertain cases to ensure precision.
4Reliability
If human reviewers provide feedback to improve classification accuracy, then reliability increases, but the complexity of the process increases
Solution Approach 1:
The patent segments the classification process into distinct modular stages: AI preliminary classification, automated filtering, human review of selected cases, and final validation. Each stage handles specific tasks with clear input-output definitions. This segmentation allows human reviewers to focus on specific sub-tasks rather than the entire classification process, reducing the perceived complexity while maintaining high reliability through systematic multi-stage processing.
Data Source
AI summary
Presented herein are systems and methods for automated and/or semi-automated classification of research data via a collaboration platform, said platform designed for extraction and synthesis of research data from a collection (e.g., a disciplinary repository, such as an archive comprising research works, research articles, research data, etc. for any of a variety of research disciplines or subjects).


