AI Research Data Classification via Human-AI Collaboration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The process of searching for research data is often time-consuming, prone to errors, and biased due to the lack of effective classification, leading to irrelevant information being retrieved, which can hinder the discovery of relevant data necessary for research topics or questions.

Innovation Solution

A collaboration platform utilizing artificial intelligence and machine learning to automate or semi-automate the classification of research data, allowing for user-defined protocols and feedback mechanisms to refine search results and improve data relevance, enabling human-computer collaboration for efficient data curation and synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual searching and classification of research data is performed, then human understanding and contextual judgment are achieved, but the process becomes time-consuming and tedious

Engineering Contradiction:
Improveclassification accuracyVSAvoidsearch time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an AI module as an intermediary between the research data and human users. The AI module performs preliminary classification of research data using machine learning algorithms, producing structured outputs that humans can efficiently review and refine. This intermediary processing layer reduces the time humans spend on manual classification while maintaining high accuracy through collaborative refinement.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary classification actions automatically using AI before human review. The AI module pre-processes large volumes of research data, identifies relevant documents, and organizes them according to predefined taxonomies. This preliminary action filters out obviously irrelevant data, allowing humans to focus their time on reviewing and refining the AI's classifications rather than starting from scratch.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If comprehensive searching is conducted to ensure all relevant data is found, then completeness is achieved, but irrelevant information increases and requires further curation

Engineering Contradiction:
Improvedata completenessVSAvoidirrelevant information
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent applies different classification criteria and filtering mechanisms to different portions of the search results based on their relevance indicators. The AI module identifies high-confidence matches and applies streamlined processing, while lower-confidence results receive more stringent filtering and human review. This localized quality control ensures comprehensive coverage while minimizing irrelevant information in the final curated set.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system implements feedback loops where human reviewers assess the AI's classification accuracy and relevance judgments. This feedback is used to continuously refine the AI's algorithms and filtering criteria. The feedback mechanism enables the system to learn from errors and improve its ability to distinguish relevant from irrelevant information, maintaining completeness while reducing noise over time.

Inventive Principle:
Principle #23Feedback

3Productivity

If automated classification is implemented to reduce time and effort, then processing speed increases, but accuracy and reliability may decrease

Engineering Contradiction:
Improveclassification speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements a dynamic classification system where the level of automation adjusts based on the confidence level of AI predictions. For high-confidence classifications, the system operates fully automatically to maximize speed. For lower-confidence cases, the system dynamically transitions to semi-automated mode, inviting human review to maintain accuracy. This dynamic approach optimizes the balance between speed and precision across different types of classification tasks.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system applies partial automation selectively to portions of the data where AI confidence is high, while reserving full human review for cases where accuracy is critical or AI uncertainty is high. This partial automation strategy achieves speed improvements on the majority of straightforward cases without sacrificing accuracy on complex or ambiguous classifications. The system performs excessive classification actions (both AI and human) on uncertain cases to ensure precision.

Inventive Principle:
Principle #16Partial or excessive action

4Reliability

If human reviewers provide feedback to improve classification accuracy, then reliability increases, but the complexity of the process increases

Engineering Contradiction:
Improveclassification reliabilityVSAvoidprocess complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the classification process into distinct modular stages: AI preliminary classification, automated filtering, human review of selected cases, and final validation. Each stage handles specific tasks with clear input-output definitions. This segmentation allows human reviewers to focus on specific sub-tasks rather than the entire classification process, reducing the perceived complexity while maintaining high reliability through systematic multi-stage processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230359932A1Classification process systems and methods
Publication Date: 2023.11.09 RAYYAN SYSTEMS INC
  • US20230359932A1 patent drawing
  • US20230359932A1 patent drawing
  • US20230359932A1 patent drawing

AI summary

Presented herein are systems and methods for automated and/or semi-automated classification of research data via a collaboration platform, said platform designed for extraction and synthesis of research data from a collection (e.g., a disciplinary repository, such as an archive comprising research works, research articles, research data, etc. for any of a variety of research disciplines or subjects).