Multimodal Food Identification Model Using Text-Image Pairing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current food identification technologies using image search have low accuracy, especially for new images not previously learned, leading to errors in calorie counting and food type classification.
Innovation Solution
A method and apparatus using a multimodal model that pre-learns text and image features for food images, allowing for accurate identification of new food images by pairing text features with image features to enhance classification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a typical image classification model is used for food identification, then the model can classify previously learned images, but the classification accuracy for new images not learned before is low
Solution Approach 1:
The patent applies preliminary action by pre-learning a multimodal model using a diverse dataset of food images and their corresponding text descriptions before actual food identification tasks. This pre-learning phase enables the model to learn general food characteristics, textures, colors, and associations between visual and textual features, allowing it to accurately identify new food images that were not present in the training data. The model is pre-trained on various food categories including main dishes, side dishes, desserts, and beverages, establishing a foundation for handling unseen food images with high accuracy.
Solution Approach 2:
The patent implements universality by developing a multimodal model that integrates both image processing and text processing capabilities. The model simultaneously analyzes visual features from food images and textual features from descriptions, enabling it to handle a wide variety of food types and presentations. This multi-functional approach allows the system to identify foods across different categories (main dishes, side dishes, desserts, beverages) and adapt to various food preparation styles, making it universally applicable to new and unseen food images rather than being limited to specific pre-learned categories.
2Measurement precision
If simple image search is used for food identification, then the process is simple, but the search accuracy is considerably low leading to errors in calorie counting
Solution Approach 1:
The patent applies merging by combining multiple types of data processing - image processing and text processing - into a unified multimodal model. The system merges visual feature extraction from food images with text feature extraction from food descriptions, then integrates these two modalities through a fusion mechanism. This combination allows the model to leverage both visual appearance and textual information to achieve high-accuracy food identification and calorie counting, overcoming the limitations of simple image search while managing complexity through integrated architecture.
Solution Approach 2:
The patent implements composite materials by creating a composite representation that combines image features and text features into a unified model structure. The multimodal model integrates different types of data (visual and textual) into a cohesive system that processes both modalities simultaneously and combines their insights for final food identification. This composite approach enables the system to achieve high measurement precision in food identification and calorie counting by synthesizing information from multiple sources rather than relying on a single simple image search mechanism.
3Measurement precision
If deep learning is applied to image identification, then classification performance improves for trained images, but the model fails to accurately classify new images not seen during training
Solution Approach 1:
The patent applies the intermediary principle by introducing text descriptions as a mediator between the visual input and the classification output. The text feature extractor serves as an intermediary that processes food descriptions and provides textual representations that complement visual features. This intermediary textual information acts as a bridge, enabling the model to better understand and generalize food concepts, thereby improving its ability to accurately classify new images that were not present in the training data while maintaining high classification performance on seen images.
Solution Approach 2:
The patent implements another dimension by adding a textual dimension to the traditional two-dimensional image-based classification approach. Instead of relying solely on visual features from images, the system incorporates text descriptions as an additional dimension of information. This multi-dimensional approach combines visual and textual data to create a more comprehensive representation of food items, enabling the model to generalize better to new images by leveraging both visual and textual characteristics simultaneously, thus improving both classification performance and adaptability.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided are a method and an apparatus for identifying food using a multimodal model. A method for identifying food using a multimodal model according to one embodiment of the present disclosure comprises setting a plurality of food images belonging to different classes as pre-learning targets; pre-learning a multimodal model for each of the plurality of food images set as the pre-learning target so that a text feature paired with an image feature is more similar to the image feature than the other text features not paired with the image feature; and identifying food to be identified from input food images using the pre-learned multimodal model.