Privacy-Preserving Text Insight Mining via Dual Generative Model Bias Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional bias minimization techniques for generative language models, such as GPT-3, view biases as flaws to be managed, rather than leveraging them to gain insights in applications like text mining or sentiment analysis.
Innovation Solution
The proposed system utilizes the biases of generative language models to highlight differences between models, which are then used to practical effect in text mining or sentiment analysis applications. This is achieved by obtaining language input data and providing it to both a first and a second generative language model, with the system indicating the differences between their responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If bias minimization techniques are applied to generative language models, then model fairness and neutrality are improved, but the ability to leverage biases for gaining insights in text mining and sentiment analysis is lost
Solution Approach 1:
The patent segments the analysis into two distinct model types: a first generative language model trained on general internet text that provides baseline responses, and a second generative language model trained on specific domain text that provides domain-specific responses. By comparing outputs from these segmented models, the system recovers insight information that would otherwise be lost in bias-minimized models, while maintaining fairness through controlled comparison rather than raw biased outputs.
2Adaptability or versatility
If generative language models are trained on large bodies of text including user-generated content, then model capability and coverage are improved, but user privacy is compromised
Solution Approach 1:
The patent introduces trained generative language models as intermediaries between raw user-generated content and analysis applications. The models are trained offline on large corpora including user-generated content, then deployed to generate responses to new inputs. This intermediary approach allows the system to leverage the adaptability and versatility of models trained on extensive data while protecting user privacy, as the trained models process information without requiring access to or storage of original user data.
Data Source
AI summary
An embodiment provides a method including obtaining language input data and providing the language input data to a first generative language model and a second generative language model. A first response from the first generative language model and a second response from a second generative language model are obtained. An indication is provided of a difference between the first response from the first generative language model and the second response from the second generative language model.


