EM Model Disambiguates Regulatory Statements

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Ambiguous regulatory compliance documents pose challenges for software engineers due to varying expert interpretations, leading to inconsistent classification of statements and increased manpower requirements for disambiguation.

Innovation Solution

A processor-implemented method using the Expectation-Maximization (EM) model processes regulation statements to determine consensus on ambiguous terms, questions, and answers through crowd-sourcing, generating reference data and calculating ambiguity scores based on expert expertise and label variations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple subject matter experts are involved in disambiguation, then interpretation accuracy improves, but knowledge variation leads to inconsistent labels and increased complexity

Engineering Contradiction:
Improveinterpretation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an Expectation-Maximization (EM) model as an intermediary that processes labels from multiple experts. The EM model iteratively estimates expert expertise levels and determines ground truth labels by weighing expert contributions according to their demonstrated accuracy, thereby resolving inconsistencies without requiring manual reconciliation of expert disagreements

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback mechanisms where expert labels are compared against determined ground truth to calculate expertise scores. These expertise scores feed back into the labeling process, allowing the system to automatically weight and prioritize labels from more reliable experts in subsequent iterations, improving consistency while maintaining accuracy

Inventive Principle:
Principle #23Feedback

2Productivity

If crowd-sourcing is used to generate reference data, then processing speed improves, but label variation increases ambiguity intensity

Engineering Contradiction:
Improveprocessing speedVSAvoidambiguity intensity
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent transforms the ambiguous labeling data into structured parameters including expertise scores, ambiguity intensity metrics, and ground truth probabilities. By changing the representation from raw labels to these quantified parameters, the system can mathematically process and resolve ambiguities while preserving the valuable information contained in varied expert interpretations

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system creates a composite label determination that combines multiple expert opinions weighted by their expertise levels. Rather than selecting a single label or treating all labels equally, the patent synthesizes a composite ground truth that incorporates information from multiple sources proportionally to their demonstrated reliability, reducing information loss from the aggregation process

Inventive Principle:
Principle #40Composite materials

3Measurement precision

If ground truth determination incorporates expertise calculation, then label accuracy improves, but computational complexity increases

Engineering Contradiction:
Improvelabel accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The EM model performs preliminary iterations to estimate expert expertise levels and initial ground truth probabilities before final label determination. This preliminary action separates the complex computational task into phases, where initial estimates are refined iteratively, reducing the complexity of the final accuracy determination step

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses dynamic iterative updates where expertise scores and ground truth probabilities are continuously refined based on emerging patterns in the data. The computational complexity is managed through iterative convergence rather than complex static calculations, allowing the system to adapt to the data structure as processing progresses

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11120228B2Method and system for generating ground truth labels for ambiguous domain specific tasks
Publication Date: 2021.09.14 TATA CONSULTANCY SERVICES LTD
  • US11120228B2 patent drawing
  • US11120228B2 patent drawing
  • US11120228B2 patent drawing

AI summary

This disclosure relates generally to data processing, and more particularly to a method and system for generating ground truth labels for ambiguous domain specific tasks. The system generates reference data corresponding to a regulation statement being processed, using a crowd sourcing mechanism and then processes the reference data using an Expectation Maximization (EM) model. The EM model determines consensus with respect to ambiguity of terms/phrases, validity of questions, and validity of answers, and then based on the determined consensus, provides questions and answers to disambiguate the regulation statement.