Property Valuation Data Cleaning for HPI Bias Mitigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing property valuation methods face challenges in accurately estimating house price indices due to systematic and idiosyncratic errors, particularly aggregation bias and transaction type bias, which affect the reliability of marking-to-market predictions when dealing with multiple prior transactions.

Innovation Solution

The Trunk-Branch Repeat Transaction Index (TB-RTI) and Multiple-Transaction Based Property Valuation (MTV) methods are introduced, which control for systematic biases by using trusted purchase transactions to establish a base index, adjusting non-purchase transactions, and employing data cleaning to mitigate idiosyncratic errors, while utilizing a weighted combination of multiple transactions for more accurate mark-to-market predictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If mortgage refinance transactions are used to increase data coverage for HPI estimation, then the quantity of transaction data increases, but transaction type bias is introduced that renders the estimated HPI inaccurate

Engineering Contradiction:
Improvequantity of transaction dataVSAvoidaccuracy of HPI estimation
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the transaction data into different types (purchase transactions and refinance transactions) and applies different processing methods to each segment. Purchase transactions are used to establish the base HPI while refinance transactions are adjusted using transaction type bias corrections before being incorporated, allowing both data segments to contribute appropriately to the final HPI estimation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies parameter changes by introducing transaction type bias adjustment factors that modify the transaction values based on their type. Refinance transactions undergo specific bias corrections (e.g., for loan-to-value ratio effects, purpose of loan) before being integrated into the HPI calculation, transforming the raw data into bias-corrected data suitable for aggregation.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If a large geographic area is defined as a single housing market to compensate for insufficient transaction data, then the quantity of available transactions increases, but aggregation bias is created that reduces measurement precision

Engineering Contradiction:
Improvequantity of transaction dataVSAvoidaccuracy of HPI estimation
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the large geographic area into multiple smaller housing markets or submarkets based on homogeneous characteristics (neighborhoods, school districts, etc.). Each segment is analyzed separately to capture local price dynamics, and the results are aggregated to form the overall HPI, thereby reducing aggregation bias while maintaining sufficient data coverage in each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by allowing different geographic segments to have their own localized HPI estimates that reflect specific neighborhood characteristics and price dynamics. Each local market receives customized analysis with appropriate data weighting and adjustment factors, rather than applying a uniform approach across the entire large geographic area.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If multiple data sources are used to estimate separate HPIs for smaller markets, then measurement precision for local markets improves, but device complexity increases

Engineering Contradiction:
Improveaccuracy of localized HPI estimationVSAvoidcomplexity of data processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal data processing framework that handles multiple data sources (public records, mortgage data, appraisal data) through a single integrated system. The same core algorithmic structure processes different data types by applying appropriate source-specific adjustment factors, allowing the system to estimate HPIs for multiple geographic levels without requiring separate processing systems for each data source.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces intermediary components such as standardized data validation rules, common adjustment factor calculations, and unified HPI aggregation methods that mediate between diverse data sources and the final HPI estimates. These intermediaries harmonize the different data sources and processing requirements, reducing overall system complexity while maintaining local market precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Quantity of substance

If duplicate transaction records are retained to ensure data completeness, then data coverage is maintained, but idiosyncratic errors increase that reduce reliability

Engineering Contradiction:
Improvedata coverageVSAvoidreliability of transaction data
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent treats duplicate records as copies that need to be identified and consolidated. Rather than deleting duplicates entirely, the system creates a master record from the most reliable source and marks other instances as duplicates, preserving the ability to trace data provenance while eliminating redundant information that would skew aggregate calculations.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies discarding and recovering by identifying duplicate records and discarding redundant copies while recovering and preserving the essential transaction information in a consolidated format. The system maintains data completeness by keeping one authoritative version of each transaction while removing duplicate entries that would otherwise inflate error terms in the analysis.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS8738388B1Market based data cleaning
Publication Date: 2014.05.27 FANNIE MAE
  • US8738388B1 patent drawing
  • US8738388B1 patent drawing
  • US8738388B1 patent drawing

AI summary

Market based data cleaning for mitigation of idiosyncratic errors in transaction data used for property valuation. The market based data cleaning technique helps to ensure that the most accurate record is retained for a transaction when duplicate records are found, by ensuring that the retained record is the most consistent with other transactions of the same property, the local market trend, and neighborhood market. Algorithms accommodate the adoption of a value as a representative single record for a transaction where multiple records are present for a transaction. Following duplicate removal, market based data cleaning further eliminates erroneous records by eliminating transaction outliers, also preferably based upon the local market trend, along with all of the remaining transactions for each given property.