Food Recognition Using Visual and Speech Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems fail to accurately and automatically identify and quantify food items on a plate, especially under diverse lighting conditions, which hinders precise caloric content assessment.

Innovation Solution

A method and system that uses a combination of offline and online feature-based learning to classify and segment food items using color and texture features, followed by 3D volume estimation, allowing for accurate caloric content calculation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automatic image analysis techniques are used for food recognition, then food item identification can be automated, but the system fails to estimate food volume accurately

Engineering Contradiction:
Improvefood recognition automationVSAvoidvolume estimation accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system segments the food plate image into multiple food items using color and texture features. Each food item is individually identified and classified, enabling separate volume estimation for each item rather than treating the entire plate as a single entity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from 2D image analysis to 3D volume estimation by computing depth information and spatial relationships. This dimensional transformation allows the system to estimate actual food volumes from flat images, resolving the limitation of prior 2D-only analysis.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If intensity-based segmentation and classification is used, then food items can be segmented using color and texture features, but the system cannot estimate food volume for accurate caloric assessment

Engineering Contradiction:
Improvefood item segmentation accuracyVSAvoidvolume information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system performs preliminary segmentation and classification of food items using color and texture features before volume estimation. This preliminary action organizes the data structure and identifies food boundaries, which then enables accurate volume computation for each segmented item.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary computational layer that bridges 2D image segmentation and 3D volume estimation. This intermediary process uses the segmented regions and their spatial relationships to compute volume metrics, preserving volume information that would otherwise be lost in pure 2D analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If a large number of food classes are recognized, then the system can handle diverse diets, but state of the art object recognition methods are unable to operate effectively

Engineering Contradiction:
Improvefood type coverageVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system uses a universal classification framework based on color and texture features that can recognize multiple food classes simultaneously. This multi-functional approach allows the same feature extraction and classification mechanisms to handle diverse food types, from fruits and vegetables to proteins and carbohydrates, maintaining reliability across a broad spectrum of food classes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8439683B2Food recognition using visual analysis and speech recognition
Publication Date: 2013.05.14 SRI INTERNATIONAL
  • US8439683B2 patent drawing
  • US8439683B2 patent drawing
  • US8439683B2 patent drawing

AI summary

A method and system for analyzing at least one food item on a food plate is disclosed. A plurality of images of the food plate is received by an image capturing device. A description of the at least one food item on the food plate is received by a recognition device. The description is at least one of a voice description and a text description. At least one processor extracts a list of food items from the description; classifies and segments the at least one food item from the list using color and texture features derived from the plurality of images; and estimates the volume of the classified and segmented at least one food item. The processor is also configured to estimate the caloric content of the at least one food item.