Bias Detection for Unstructured Text via Structured Knowledge Base
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods are inadequate for detecting bias in unstructured documents, as they rely on manual user input and are prone to inaccuracies, especially when dealing with natural language text, where protected attributes and favorable outcomes are unknown.
Innovation Solution
A system that converts unstructured documents into a structured knowledge base by extracting entities, relationships, and other features, allowing for the application of conventional bias detection techniques to identify bias, thereby providing an accurate and automated bias detection method.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual user input methods are used for bias detection, then the detection process can handle unstructured documents, but the accuracy is reduced and the process is prone to inaccuracies
Solution Approach 1:
The patent introduces an intermediary structured representation layer between the unstructured document and the bias detection algorithm. By extracting entities, relationships, and attributes into a structured format (JSON schema), the system mediates the gap between unstructured text and automated detection, enabling accurate bias detection without manual input while preserving the nuances of natural language
Solution Approach 2:
The patent replaces the mechanical manual review process with an automated system that uses natural language processing and structured data extraction. The automated extraction of entities, relationships, and attributes substitutes human analysts, eliminating manual input requirements while maintaining or improving detection accuracy through consistent application of detection criteria
2Extent of automation
If conventional bias detection techniques are applied directly to unstructured documents, then the process can be automated, but the detection accuracy deteriorates due to unknown protected attributes and favorable outcomes
Solution Approach 1:
The patent performs preliminary action by extracting and structuring protected attributes and favorable outcomes before applying bias detection techniques. The system pre-processes unstructured documents to identify entities, relationships, and attributes, organizing them into a structured format that reveals protected attributes and favorable outcomes that would be unknown in direct unstructured analysis, enabling accurate automated detection
Solution Approach 2:
The patent changes the parameter representation from unstructured text to structured data with defined schemas. By transforming documents into JSON format with explicit fields for entities, relationships, and attributes, the system changes how information is parameterized, making protected attributes and favorable outcomes identifiable and measurable for automated bias detection algorithms
3Loss of information
If unstructured documents are analyzed directly, then the original information is preserved, but the complexity of detecting bias increases significantly
Solution Approach 1:
The patent segments unstructured documents into distinct structured components: entities, relationships, attributes, and outcomes. This segmentation organizes the information into manageable units with clear definitions, reducing detection complexity while preserving all original information through systematic extraction and representation in a structured format
Data Source
AI summary
One embodiment provides a method, including: receiving a target unstructured document for determining whether the target unstructured document comprises biased information; identifying an objective of the target unstructured document by extracting, from the target unstructured document, (i) entities and (ii) relationships between the entities; creating a structured knowledge base, wherein the creating comprises (i) creating an entry in the structured knowledge base corresponding to the target unstructured document, (ii) identifying other unstructured documents having a similarity to the target unstructured document, and (iii) generating an entry in the structured knowledge base corresponding to each of the other unstructured documents; applying a bias detection technique on the structured knowledge base; and providing an indication of whether the target unstructured document comprises bias.


