Automated Neighborhood Detection from Geocoded Web Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Updating and populating records in geographic information systems, especially for larger areas like entire countries, is expensive and difficult due to the frequent changes in neighborhood names and boundaries, as well as the need to document new neighborhoods across various languages.

Innovation Solution

A process that extracts n-grams from web documents, associates them with geographic locations, identifies clusters using algorithms like DBSCAN, determines neighborhood boundaries and names, and adds them to geographic information systems, allowing for automatic detection and updating of neighborhood information without manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual surveyors are used to catalog neighborhoods, then accuracy of geographic information can be maintained, but cost and time requirements become prohibitively expensive

Engineering Contradiction:
Improveaccuracy of neighborhood informationVSAvoidcost and time efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical manual surveying system with an automated computational system that uses web document analysis, geocoding, and clustering algorithms to detect neighborhood boundaries and names, thereby eliminating the need for human surveyors while maintaining detection accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables geographic information systems to automatically update themselves by extracting neighborhood information from publicly available web documents, allowing the system to self-populate and self-update without external manual intervention

Inventive Principle:
Principle #25Self-service

2Reliability

If manual updating of geographic information systems is performed, then data accuracy can be ensured, but the system cannot keep pace with frequent changes in neighborhood names and boundaries

Engineering Contradiction:
Improvedata accuracyVSAvoidspeed of updates
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements continuous automated monitoring and extraction of neighborhood information from web documents, enabling the system to continuously detect and update neighborhood changes without interruption, thereby keeping pace with frequent changes while maintaining accuracy through algorithmic consistency

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system performs preliminary extraction and clustering of geographic information from web documents before formal documentation, allowing neighborhood changes to be detected and prepared for integration into geographic information systems in advance

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If human surveyors document neighborhoods in multiple languages, then comprehensive coverage can be achieved, but the complexity and cost increase significantly

Engineering Contradiction:
Improvemulti-language coverageVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal automated system that processes web documents in multiple languages through a single integrated platform, using language-agnostic geocoding and clustering techniques that work across different linguistic contexts without requiring separate manual processes for each language

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9442905B1Detecting neighborhoods from geocoded web documents
Publication Date: 2016.09.13 GOOGLE LLC
  • US9442905B1 patent drawing
  • US9442905B1 patent drawing
  • US9442905B1 patent drawing

AI summary

Provided is a process of identifying a name and boundary of a neighborhood based on web documents, the process including: extracting, via one or more processors, an n-gram appearing in a plurality of web documents; associating the n-gram with geographic locations associated with the web documents from which the n-gram was extracted; identifying a neighborhood by identifying a cluster of geographic locations associated with the n-gram; determining a boundary for the neighborhood from the distribution of geographical locations in the cluster; determining a name for the neighborhood from the n-gram; and adding the name and boundary of the neighborhood to a geographic information system.