Novel few-sample food classification method

Through the combination of image enhancement and feature fusion layer, the accuracy problem of food recognition under the condition of few samples is solved, and high-precision food classification is achieved.

CN120299032APending Publication Date: 2025-07-11ZHEJIANG NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510417384.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing food recognition technology is difficult to achieve high-precision classification under the condition of few samples, and traditional methods are not effective, while deep learning methods require large-scale data collection and labeling, resulting in high costs and difficult to implement.

Method used

The image enhancement layers CutCenter and CornerMix are used to extract local details, feature extraction is combined with ViT network, and feature representation is optimized through the global-local feature fusion layer and the adaptive loss function to generate comprehensive global and fine-grained local features.

Benefits of technology

It improves the accuracy of food classification, reduces the impact of noise, enhances the model's attention to subtle local areas, and achieves high-precision small sample classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299032A_ABST
    Figure CN120299032A_ABST
Patent Text Reader

Abstract

The invention discloses a novel few-sample food classification method, and relates to the technical field of food identification. Comprising the steps of obtaining a food image, preprocessing the food image to obtain a preprocessing data set, constructing a food classification model comprising an image enhancement layer, a feature extraction layer and a global-local feature fusion layer, inputting the preprocessing data set into the food classification model, and performing error analysis through a loss function according to an output predicted value and an actual value. And obtaining a trained food classification model, and inputting the to-be-detected image into the trained food classification model to classify the food image. According to the method, the features with comprehensive global and fine-grained local representation can be generated, so that high classification precision is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of food recognition, and in particular to a new few-shot food classification method. Background Art

[0002] Existing food recognition technologies can be divided into traditional methods and deep learning methods. Traditional object recognition methods usually rely on manually designed feature extraction and classical machine learning algorithms, such as support vector machines and decision trees. Feature extraction engineering is a key step in traditional methods, where domain experts define and select appropriate features, such as edges, textures, and color histograms. Initially, feature extraction methods (such as color histogram methods and spatial feature methods of food ingredients) could only extract single features. However, this method had poor results, with the highest accuracy not exceeding 30%. Subsequently, multi-kernel learning was proposed to integrate various image features, such as color, texture, and scale-invariant features, and these features achieved an accuracy of over 60% on multiple datasets. Subsequently, various multi-feature extraction methods were proposed, including conventional template matching of SIFT features, bag of SURF features, and two-step recognition algorithms.

[0003] Food classification methods based on deep learning usually rely on high-quality features extracted by the network to achieve satisfactory classification accuracy. More specifically, based on the characteristics of food images, various deep learning models have been optimized and applied to food recognition, such as residual network models, and those that combine transfer learning with the Inception-V3 model. These methods directly extract deep visual features through CNN, ignoring the features of food images and making it difficult to achieve the best recognition effect. There are also some methods that use additional ingredient knowledge for food recognition. The multi-scale multi-view feature aggregation method aggregates three different types of features from class-supervised networks and ingredient-supervised networks into more robust, comprehensive, and discriminative features. Therefore, this method achieves excellent recognition performance, and as the size of the dataset increases, deep learning models show significant improvements in other fields. Although these large-scale datasets help achieve high accuracy, they also mean collecting and labeling large-scale food images, which requires high time and economic costs and is prone to difficulties in technology implementation due to problems such as excessive costs and non-profitability in the short term.

[0004] Therefore, providing a new few-shot food classification method to solve the difficulties existing in the prior art is an urgent problem for those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a new few-shot food classification method, which can generate features with comprehensive global and fine-grained local representations, thereby achieving higher classification accuracy.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A novel few-shot food classification method, comprising the following steps:

[0008] Obtain food images, and preprocess the food images to obtain a preprocessed dataset;

[0009] Construct a food classification model including an image enhancement layer, a feature extraction layer, and a global-local feature fusion layer;

[0010] Input the preprocessed dataset into the food classification model, and perform error analysis through a loss function based on the predicted value and the actual value output, to obtain a trained food classification model;

[0011] Input the image to be detected into the trained food classification model to classify the food image.

[0012] Optionally, the image enhancement layer adopts CutCenter and CornerMix:

[0013] CutCenter is used to select a central rectangular region in the original image I in the preprocessed dataset and crop it into a new enhanced image I CC ;

[0014] CornerMix is used to randomly select a rectangular region in the four corner regions of the enhanced image I CC and erase it to generate an enhanced image I CM .

[0015] Optionally, the feature extraction layer adopts a ViT network as the backbone network for feature extraction,

[0016] The ViT network includes an input module, a decomposition module, an embedding module, and a Transformer Encoder module.

[0017] Optionally, the embedding module is used to embed the blocks obtained in the decomposition module into the Transformer encoder module in the Transformer Encoder module, and the expression is:

[0018]

[0019] where, V TE is the final embedding block, V P is the extended embedding block, V Pos is the position embedding block, V CLS is the class token, and V is the embedding block.

[0020] Optionally, the Transformer Encoder module consists of L stacked encoder blocks, and each block consists of layer normalization, multi-head attention, and an MLP block.

[0021] Optionally, the global-local feature fusion layer calculates the fusion features from multiple latent representation class tokens by minimizing the adaptive feature fusion loss The expression is:

[0022] β = a 1 C O + a 2 C CC + a 3 C CM ,

[0023] where a 1 is the weight of C O a 2 is the weight of C CC a 3 is the weight of C CM weight.

[0024] Optionally, the loss function L VF includes the cross-entropy loss function L CE (X, Y) and the pairwise confusion loss D EC (x i , x k ).

[0025] According to the above technical solutions, compared with the prior art, the present invention provides a new few-shot food classification method, which has the following beneficial effects: This application combines image enhancement and adaptive global-local feature fusion. The image enhancement module mainly consists of two parts, namely CutCenter and CornerMix. Among them, the CutCenter method extracts and amplifies fine local regions to enhance fine-grained feature learning, while the CornerMix method identifies and erases irrelevant regions to reduce the negative impact of noise in food images; then, for these new enhanced images, an adaptive weighting method with theoretical guarantee is proposed to learn the complementary relationship between features from the enhanced images and the original images. The learned weighting relationship guides feature fusion to generate features with comprehensive global and fine-grained local representations, thereby achieving high classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0027] Figure 1 Flow chart of a novel few-shot food classification method disclosed by the present invention;

[0028] Figure 2 Schematic diagram of a novel few-shot food classification method disclosed by the present invention;

[0029] Figure 3 Graph of the model training algorithm disclosed by the present invention. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] Referring to Figure 1 and Figure 2 as shown, the present invention discloses a novel few-shot food classification method, including the following steps:

[0032] Obtain food images, and preprocess the food images to obtain a preprocessed data set;

[0033] Construct a food classification model including an image enhancement layer, a feature extraction layer, and a global-local feature fusion layer;

[0034] Input the preprocessed data set into the food classification model, and perform error analysis through a loss function based on the predicted value and the actual value output, to obtain a trained food classification model;

[0035] Input the image to be detected into the trained food classification model to classify the food image.

[0036] Furthermore, the image enhancement layer adopts CutCenter and CornerMix, which are respectively used to extract and magnify the local areas containing rich food details and erase the irrelevant areas, so as to pay more attention to the subtle local areas and the main areas of the target food, and enhance the global-local feature representation at the image level. It includes:

[0037] CutCenter is used to select the central rectangular area in the original image I in the preprocessed data set and crop it into a new enhanced image I CC ;

[0038] CornerMix is used to randomly select rectangular areas in the four corner areas of the enhanced image I CC and erase them to generate the enhanced image ICM 。

[0039] Specifically, the enhanced image I CC has a size that is half of the original image. The enhanced image contains less background while magnifying the food details, which helps the model capture fine-grained local features. CornerMix is used to reduce the interference information (such as non-target food and irrelevant objects) in the four corner regions of the original image. The area S of the randomly selected rectangular region r ranges from 0.02 to 0.1 times the area S of the original picture. The size of each rectangle is determined by the following formula: o .

[0040] S r = random(0.02, 0.1) * S r ,

[0041] S r = random(0.02, 0.1) * S o ,

[0042] r = random(0.3, 1 / 0.3),

[0043]

[0044] where H p is the height of the selected rectangular region, and W p is the width of the selected rectangular region. The enhanced image I CM generated by CornerMix reduces the interference background and helps the model focus on and learn the main regions of the target food.

[0045] Furthermore, the feature extraction layer uses the ViT network as the backbone network for feature extraction.

[0046] The ViT network includes an input module, a decomposition module, an embedding module, and a Transformer Encoder module.

[0047] Specifically, ViT can effectively capture the relationships between any regions in the image, which makes it suitable for solving the challenges of rich image category diversity and high shape similarity in food images. The image set (I, I CC , I CM ) output by the image enhancement module is used as the input to this module. For an input image of size W×H×3, we divide it into N non-overlapping blocks where N = W×H / P 2 . These blocks are flattened and linearly projected into block embeddings class tokens It is added as a learnable parameter to collect global information and then connected with V to obtain an extended embedding In addition, to preserve spatial information, the positional embedding introduced as a trainable parameter is directly added to the extended embedding V P .

[0048] Furthermore, the embedding module is used to embed the blocks obtained in the decomposition module into the Transformer encoder module in the Transformer Encoder module, and the expression is:

[0049]

[0050] where V TE is the final embedding block, V P is the extended embedding block, V Pos is the positional embedding block, V CLS is the class token, and V is the embedding block.

[0051] Furthermore, the Transformer Encoder module consists of L stacked encoder blocks, each block consisting of layer normalization, multi-head attention, and an MLP block. Among them, the input of the j-th encoder block is expressed as where j ∈ [1, L], which is the output of the (j - 1)-th encoder block. In addition, the output dimension of each encoder block is consistent with its input. Finally, the output of the Transformer Encoder module can be expressed as where is generated by the trainable class token V CLS , and the finally generated set of class tokens is denoted as (C O , C CC , C CM ). By extracting features from the original image and the augmented image, multi-features capable of capturing global and local features are obtained.

[0052] Furthermore, the global-local feature fusion layer calculates the fusion feature from multiple latent representation class tokens by minimizing the adaptive feature fusion loss The expression is:

[0053] β = a 1 C O + a 2 C CC + a 3 C CM ,

[0054] where a 1 is the weight of C O , and a 2 is the weight of CCC Weight, a 3 is C CM Weight.

[0055] Specifically, for food image classification, some discriminative features should be given higher weights because they play a crucial role in distinguishing similar classes. To solve this problem, an adaptive weighting method with theoretical guarantees is proposed to learn the weight relationships between different features. The fused feature is calculated by minimizing the following adaptive feature fusion loss l F from multiple latent representations C 1 ,..., C L The initial expression is:

[0056]

[0057] where F is the Frobenius norm, C = [C O , C CC , C CM , a l is the weight of C l where L = 3. Considering the constraint on a l , the Lagrange multiplier method is used to solve this optimization problem. By introducing the Lagrange multiplier η, the above equation can be reformulated as:

[0058]

[0059] where p is a variable (set to 5 in this paper). For simplicity, define Given L = 3, the partial derivatives of l F (a l , η) with respect to a l and η are as follows:

[0060]

[0061] Therefore, setting the partial derivatives to zero, we can obtain:

[0062]

[0063] where a l is updated to the following form:

[0064]

[0065] Therefore, the fused feature β can be expressed as:

[0066] β = a 1 C O + a 2 C CC+a 3 C CM 。

[0067] Furthermore, the loss function L VF includes the cross-entropy loss function L CE (X, Y) and the pairwise confusion loss D EC (x i , x k ).

[0068] Specifically, the cross-entropy loss is used as the loss function of ViT because it is widely used in image recognition. For highly similar samples from different classes, the network is forced to extract features with higher confidence to minimize the cross-entropy loss during training. This leads to feature learning based on specific samples, resulting in poor generalization ability. Since Chinese food images belong to fine-grained images with high inter-class similarity, using the cross-entropy loss alone as the loss function of the model is likely to lead to overfitting. The specific calculation of the cross-entropy function is as follows:

[0069]

[0070] Among them, consider a C-class classification problem. Let S = {s1, s2, …, s n} be the dataset. Let X ∈ R d be the feature space, and Y = {1, …, C} be the label space. There exists a set where each (s i , x i , y i ) ∈ (S × X × Y). The classifier is defined as a function that maps the feature space to the class probability space where f j (x i ) represents the probability that the feature x i is classified as class j, and y ij corresponds to the j-th element of the one-hot encoded label of the sample s i .

[0071] The pairwise confusion loss is introduced into the loss function as an additional regularization term. This function encourages the model to learn more general features rather than sample-specific features, thus solving the overfitting problem caused by relying solely on the cross-entropy loss. The formula of the pairwise confusion function is as follows:

[0072]

[0073] f(x i ) represents the class probability of the sample x i , and f(x k ) represents the class probability of the sample x k , where k ≠ i. This paper proposes LVF As the final loss function, the cross-entropy loss and pairwise confusion loss are combined to better supervise multi-feature learning. Refer to Figure 3 As shown, the proposed loss algorithm is developed in three consecutive stages. First, focus on the cross-entropy loss of the fused feature C FF . Let denote the C FF space. By replacing X with X F in the previous equation, the cross-entropy loss of C FF is given by the following formula:

[0074]

[0075] Subsequently, calculate the pairwise confusion loss of the fused feature C FF . For each pair of features and its corresponding class probability C FF , the pairwise confusion loss is added to the loss of our model as follows:

[0076]

[0077] Finally, combine the feature losses of the original image and the augmented image to supervise feature learning, so as to obtain better pre-fusion features, and then improve the fused feature representation. Specifically, we calculate the cross-entropy loss and pairwise confusion loss of these class labels (C O , C CC , C CM ). Let X O ∈R d be the class label C O space, X CC ∈R d be the class label C CC space, X CM ∈R d be the class label C CM space. There exists a set where

[0078] each and represent the feature and as well as 's class probability. We define T as (O, CC, CM). The final loss function L VF of our model is given by the following formula:

[0079]

[0080] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A novel few-shot food classification method, characterized in that, It includes the following steps: Obtain a food image, and preprocess the food image to obtain a preprocessed dataset; Construct a food classification model including an image enhancement layer, a feature extraction layer, and a global-local feature fusion layer; Input the preprocessed dataset into the food classification model, and perform error analysis through a loss function based on the predicted value and the actual value output, to obtain a trained food classification model; Input the image to be detected into the trained food classification model to classify the food image.

2. A novel few-shot food classification method according to claim 1, characterized in that The image enhancement layer uses CutCenter and CornerMix: CutCenter is used to select the central rectangular region in the original image I in the preprocessed dataset and crop it into a new enhanced image I CC ; CornerMix is used to randomly select rectangular regions in the four corner regions of the enhanced image I CC and erase them to generate the enhanced image I CM .

3. A novel few-shot food classification method according to claim 1, characterized in that The feature extraction layer uses the ViT network as the backbone network of the feature extraction layer, The ViT network includes an input module, a decomposition module, an embedding module, and a Transformer Encoder module.

4. A novel few-shot food classification method according to claim 3, characterized in that The embedding module is used to embed the blocks obtained in the decomposition module into the Transformer encoder module in the Transformer Encoder module, and the expression is: Among them, V TE is the final embedding block, V P is the extended embedding block, V Pos is the positional embedding block, V CLS is the class token, and V is the embedding block.

5. A novel few-shot food classification method according to claim 3, characterized in that The Transformer Encoder module is composed of L stacked encoder blocks, and each block is composed of layer normalization, multi-head attention, and an MLP block.

6. A novel few-shot food classification method according to claim 1, characterized in that The fused features calculated by the global-local feature fusion layer from multiple potential representation class tokens by minimizing the adaptive feature fusion loss The expression is as follows: β = a 1 C O + a 2 C CC + a 3 C CM , Among them, a 1 is the weight of C O weight, a 2 is the weight of C CC weight, a 3 is the weight of C CM weight.

7. A novel few-shot food classification method according to claim 1, characterized in that Loss function L VF including cross-entropy loss function L CE (X, Y) and pairwise confusion loss D EC (x i , x k )

Citation Information

Cited By

  • Method and device for enhancing partial discharge detection data

    CN120801946A