Information Graph for Inclusive Machine Learning Outputs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models often exhibit biases and lack inclusivity due to inadequate representation of underrepresented populations in their training data, leading to exclusion and lack of contextual knowledge, which are linked to biases in their outputs.
Innovation Solution
The method involves constructing an information graph based on training data, identifying areas for increased inclusion, collecting auxiliary data from underrepresented populations, and integrating this data to generate an updated graph, which is then used to produce more inclusive outputs, with feedback loops to refine the process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If machine learning models are trained using standard training data, then the models can make predictions and decisions efficiently, but the outputs exhibit biases and lack inclusivity due to inadequate representation of underrepresented populations
Solution Approach 1:
The patent applies preliminary action by constructing an information graph and identifying underrepresented populations before the machine learning model generates outputs. The system proactively collects auxiliary data about these populations and updates the information graph in advance, so that when predictions are made, the model already has access to inclusive information about diverse populations, preventing biases in the final outputs
Solution Approach 2:
The patent introduces an information graph as an intermediary structure that mediates between the training data and the machine learning model outputs. This graph stores relationships and information about various populations, including underrepresented groups, allowing the system to access and incorporate diverse population information without fundamentally changing the core machine learning model or its training process
2Measurement precision
If auxiliary data about underrepresented populations is collected and integrated, then the inclusivity and accuracy of machine learning outputs improve, but the complexity of the system increases due to additional data collection and processing steps
Solution Approach 1:
The information graph serves multiple functions: it stores training data relationships, identifies underrepresented populations, collects auxiliary data, and provides contextual information to the machine learning model. By making this structure multi-functional, the patent avoids creating separate specialized systems for each task, thereby reducing overall system complexity while still achieving improved accuracy through auxiliary data integration
Solution Approach 2:
The system applies self-service by automatically identifying areas in the information graph where inclusion should be increased and autonomously collecting auxiliary data from auxiliary data sources. This automated approach reduces the need for manual intervention and complex configuration, allowing the system to improve its own inclusivity without requiring proportionally complex management infrastructure
Data Source
AI summary
A method includes constructing an information graph based on a set of training data provided to a machine learning algorithm, identifying an area of the information graph in which to increase an inclusion of the information graph, wherein the inclusion comprises a consideration of a population that is underrepresented in the information graph, collecting, from an auxiliary data source, auxiliary data about the population for use in increasing the inclusion of the information graph, utilizing the auxiliary data to increase the inclusion of the information graph, to generate an updated information graph, using the updated information graph to generate a test output that incorporates information from the auxiliary data, generating, when the test output satisfies an inclusion criterion, a runtime output using the updated information graph, receiving user feedback regarding the runtime output, and determining, in response to the user feedback, whether to further increase inclusion of the runtime output.


