Lean Parsing for Tax Form Automation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional electronic document preparation systems face challenges in efficiently updating and accurately populating fields in tax forms due to changes in tax laws, requiring significant human and computing resources, leading to delays, inaccuracies, and increased costs.

Innovation Solution

A method and system employing lean parsing algorithms and machine learning to analyze natural language text, determining operators, operands, and dependencies in tax forms to generate machine-executable functions, reducing the need for manual expert intervention and improving efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional electronic document preparation systems are used to update and populate tax form fields, then the system can process tax forms, but it requires significant human and computing resources leading to delays and increased costs

Engineering Contradiction:
Improveprocessing speed of tax formsVSAvoidsoftware release delays
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system uses machine learning models to automatically analyze natural language text from tax forms, extract operators and operands, and generate machine-executable functions without requiring manual expert intervention for each update, enabling the system to adapt to tax law changes autonomously

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual expert analysis and traditional parsing methods with automated machine learning-based natural language processing systems that can rapidly interpret tax form instructions and generate corresponding computational functions

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If traditional parsing methods are used to analyze tax form text, then the system can understand form requirements, but it requires significant human expert intervention and computing resources

Engineering Contradiction:
Improveability to adapt to tax law changesVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The natural language processing system segments the tax form text into discrete operators and operands, analyzing each component separately to build comprehensive machine-executable functions that capture the full complexity of tax calculations

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces natural language processing models as an intermediary layer between the human-readable tax form text and the machine-executable functions, automatically translating instructional text into computational logic without requiring manual coding

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual expert intervention is used to update electronic tax forms, then accuracy can be maintained, but it increases costs and reduces efficiency

Engineering Contradiction:
Improveaccuracy of form populationVSAvoidupdate efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system uses training datasets containing examples of correct tax form completions to train machine learning models, providing feedback mechanisms that continuously improve the accuracy of generated functions through iterative learning and validation

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12019978B2Lean parsing: a natural language processing system and method for parsing domain-specific languages
Publication Date: 2024.06.25 INTUIT INC
  • US12019978B2 patent drawing
  • US12019978B2 patent drawing
  • US12019978B2 patent drawing

AI summary

Systems and methods for lean parsing are disclosed. An example method is performed by one or more processors of a system and includes retrieving form data including first sentence segments and second sentence segments, determining a first predicate structure for each of the sentence segments based on a set of operators within the first set of sentence segments, identifying known tokens within the second set of sentence segments, each of the known tokens appearing on a list of predetermined tokens, identifying new tokens within the second set of sentence segments, each of the new tokens not on the list, mapping each known and new token to at least one operator, determining a second predicate structure for each sentence segment based on the mapping, and generating a predicate argument structure incorporating the first and second predicate structures, the predicate argument structure ready for mapping to at least one machine executable function.