Geo-tagged Image Establishment Identification via OCR and Mapping Constraints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques lack an effective method for computers to identify establishments, such as buildings and signs, in geo-tagged images, which are prevalent due to the use of digital cameras and positioning devices.

Innovation Solution

A computing environment with an image server and client system that utilizes geo-tagged images, optical character recognition (OCR), and establishment databases to detect and identify known establishments by recognizing text regions, comparing OCR results with nearby establishment information, and selecting representative images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional OCR engines are used to interpret street scene images, then text recognition can be performed, but accuracy is insufficient without mapping information constraints

Engineering Contradiction:
Improvetext recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by obtaining mapping information (establishment names, locations, boundaries) before OCR interpretation. This prior knowledge is used to constrain and guide the OCR engine, significantly improving text recognition accuracy without requiring complex real-time processing during image interpretation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Mapping information serves as an intermediary between the raw image data and the OCR engine. The established boundaries, names, and location data from mapping sources act as constraints that guide the OCR interpretation process, enabling accurate text recognition while maintaining system manageability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If mapping information is used as prior constraints for OCR, then text recognition accuracy improves significantly, but the system requires integration of multiple data sources

Engineering Contradiction:
Improvedigital map data accuracyVSAvoiddata integration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system achieves multi-functionality by using mapping information for multiple purposes: constraining OCR interpretation, identifying establishments in images, and providing contextual information for image corpus analysis. This unified approach to data utilization improves accuracy while managing integration complexity through a single framework

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The mapping information database serves itself by providing constraints and contextual data that enable the system to automatically identify establishments and interpret images without requiring external intervention. The established boundaries and names from mapping sources self-constrain the OCR process, improving accuracy while the system manages its own data integration needs

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP2583201B1Identifying establishments in images
Publication Date: 2023.07.12 GOOGLE LLC
  • EP2583201B1 patent drawingFigure 1
  • EP2583201B1 patent drawingFigure 2
  • EP2583201B1 patent drawingFigure 3

AI summary

Establishments are identified in geo-tagged images. According to one aspect, text regions are located in a geo-tagged image and text strings in the text regions are recognized using Optical Character Recognition (OCR) techniques. Text phrases are extracted from information associated with establishments known to be near the geographic location specified in the geo-tag of the image. The text strings recognized in the image are compared with the phrases for the establishments for approximate matches, and an establishment is selected as the establishment in the image based on the approximate matches. According to another aspect, text strings recognized in a collection of geo-tagged images are compared with phrases for establishments in the geographic area identified by the geo-tags to generate scores for image-establishment pairs. Establishments in each of the large collection of images as well as representative images showing each establishment are identified using the scores.