Food identification method

Through the multi-stream parallel feature extraction architecture and adaptive weighted fusion mechanism, the complexity and real-time problems in food recognition are solved, and the classification accuracy and robustness are improved, especially in complex backgrounds and fine-grained feature recognition tasks.

CN120356205APending Publication Date: 2025-07-22KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510455606.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has difficult to meet classification complexity and diversity in food recognition, systemic challenges in data acquisition and processing, labeling inaccuracy and real-time requirements, especially in complex backgrounds and fine-grained feature recognition tasks.

Method used

A multi-flow parallel feature extraction architecture is adopted, including color flow, texture flow, shape flow and structure flow, combining dense blocks, transition layers, feature enhancement modules and attention mechanisms, and the recognition capabilities are improved through adaptive weighting.

Benefits of technology

It significantly improves the classification accuracy and robustness of food recognition, can effectively deal with category imbalance data and complex backgrounds, reduces confusion of similar food types, and improves sensitivity and recognition ability to fine-grained features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356205A_ABST
    Figure CN120356205A_ABST
Patent Text Reader

Abstract

The invention relates to a food identification method, and belongs to the field of image identification and food identification. Comprising the following steps: respectively processing color, texture, shape and structural features of basic features through four parallel feature streams; carrying out dense block processing on a feature map output by each feature flow; inputting the feature map subjected to dense block processing into a transition layer for dimensionality reduction; inputting the feature map processed by the transition layer into a feature enhancement module, and further performing convolution processing on the feature map processed by the transition layer to obtain an enhanced feature; inputting the enhanced features into an attention mechanism for processing; fusing all the feature maps processed by the attention mechanism, and adjusting the contribution of each feature processed by the attention mechanism through an adaptive weighting module; and inputting the feature map subjected to feature fusion into a classification module, and performing final classification prediction. According to the invention, complex food images can be identified, confusion between similar food types is effectively reduced, and the food identification and classification effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a food recognition method, belonging to the technical fields of image recognition and food recognition. Background Art

[0002] Food recognition, as an interdisciplinary research direction of computer vision and artificial intelligence, is a key technical path connecting perception technology, nutritional science, and big data analysis. Its research value is not only limited to technological innovation but also lies in solving practical application challenges across disciplines.

[0003] I. Key Challenges in Application Scenarios:

[0004] 1. Classification Complexity and Diversity: There is a vast variety of food types globally, especially local or home - made foods, which have significant appearance differences. It is difficult for a single model to identify all food types simultaneously, and it is easy to confuse foods that look similar but are actually different types (such as different breads, soups, pastries, etc.).

[0005] 2. Systematic Challenges in Data Acquisition and Processing: Food recognition datasets often suffer from class imbalance problems. For example, some popular foods (such as hamburgers, pizzas, etc.) have a large number of samples, while samples of other minority foods (such as some local special dishes) may be scarce.

[0006] 3. Inaccurate Annotation: The labels in the dataset may be incorrect or inconsistent, leading to the introduction of noise during the training process. This is particularly important for the training of deep - learning models because the model may learn incorrect patterns from the wrong labels.

[0007] 4. Real - time Requirements: In many practical scenarios, food recognition needs to be carried out in real - time environments, such as dish recognition in cafeterias, food delivery, etc. Real - time processing requires the model to have a high inference speed and be able to quickly process image inputs.

[0008] II. Theoretical and Practical Difficulties at the Technical Level:

[0009] 1. Image Classification Problem: Food recognition is usually an image classification task, aiming to predict which category an input image belongs to according to the image. Traditional convolutional neural networks (CNNs) can effectively handle such tasks, but as the number of food types increases, a single classification model often has difficulty covering all possible categories.

[0010] Object Detection Problem: In some applications, food recognition is not only a classification problem but also requires detecting the presence and location of each food in the image (for example, recognizing multiple foods on a plate). This type of task requires the model to be able to handle object localization and multi - object detection, which poses higher requirements on the accuracy and speed of the network.

[0011] 2. Difficulties in training deep neural networks (DNNs): Deep learning models, especially convolutional neural networks (CNNs), may encounter problems such as vanishing gradients and overfitting during the training process. Especially when dealing with small datasets or class imbalance, the performance may decline during training.

[0012] 3. Multidimensional complexity of feature extraction: The size, background, different angles, and partial occlusion of food images can affect the performance of the model. A model may need to recognize food at different scales and angles, and how to effectively handle these problems is a technical challenge. Summary of the Invention

[0013] In view of the above-mentioned problems, the present invention provides a food recognition method. The method of the present invention can flexibly adapt to the recognition of different types of food images, and demonstrates strong performance and robustness in complex backgrounds, fine-grained features, and multi-class recognition tasks. The present invention improves the effect of food recognition and classification.

[0014] The technical solution of the present invention is: A food recognition method, the method comprising:

[0015] Step1. Obtain a food image dataset and standardize the input image;

[0016] Step2. Extract basic features from the standardized input image through an initial convolutional layer;

[0017] Step3. Process color, texture, shape, and structure features of the basic features through four parallel feature streams respectively, and each feature stream adopts different convolutional kernels and processing strategies;

[0018] Step4. Process the feature maps output by each feature stream through a dense block;

[0019] Step5. Input the feature maps processed by the dense block into a transition layer for dimensionality reduction;

[0020] Step6. Input the feature maps processed by the transition layer into a feature enhancement module to further perform convolutional processing on the feature maps processed by the transition layer to obtain enhanced features;

[0021] Step7. Input the enhanced features into an attention mechanism for processing;

[0022] Step8. Adaptive weighted sum and attention mechanism fusion: fuse all the feature maps processed by the attention mechanism, and adjust the contribution of the features processed by each attention mechanism through an adaptive weighting module;

[0023] Step9. The feature map after feature fusion is input into the classification module for final classification prediction.

[0024] Further, each image in the dataset in Step1 is adjusted to 224×224 pixels and input using the RGB channels; the numerical range of the input image X is normalized to between [0, 1], where the size of the input image X is H×W×C, where C is the number of channels, and H and W are the height and width respectively;

[0025]

[0026] where, I is the input image pixel value, I norm is the standardized image pixel value, and μ and σ are the mean and standard deviation of the image respectively.

[0027] Further, Step2 includes: the standardized input image X undergoes preliminary feature extraction through a convolutional layer, then through BatchNorm and ReLU activation processing, and then through 3*3 max pooling to complete the basic feature extraction; the specific steps are:

[0028] Step2.1. The input image X undergoes preliminary feature extraction through a convolutional operation; the process of preliminary feature extraction is expressed as:

[0029]

[0030] where Conv2d7() represents the convolutional process, the convolutional kernel size is 7, the stride is 2, and the padding is 3, where H' and W' are the height and width respectively, and X' is the result after convolution;

[0031] Step2.2. Perform BatchNorm and ReLU activation processing on the extracted preliminary features; the process of BatchNorm and ReLU activation processing is expressed as:

[0032] X″ = ReLU(BatchNorm(X′));

[0033] ReLU is the activation function:

[0034] BatchNorm is batch normalization, and the batch normalization process is:

[0035] where: γ and β allow the model to restore the necessary scaling and offset after normalization to avoid information loss caused by normalization; ∈ prevents the denominator from being zero, especially when the batch variance is close to zero, and μ and σ are the mean and standard deviation respectively, and batch is the batch size set for the input pictures;

[0036] X” is the processed feature data;

[0037] Step2.3. Perform max pooling on the result X” of the BatchNorm and ReLU activation processing to obtain the basic feature X pool : Basic feature X pool It is expressed as:

[0038] X pool = MaxPool2d3(X″);

[0039] Among them, MaxPool2d3() represents max pooling processing, the kernel size of the max pooling processing is 3, the stride is 2, and the padding is 1.

[0040] Furthermore, in the said Step3, it includes: The feature stream includes a color stream, a texture stream, a shape stream, and a structure stream. The processing process of each feature stream includes:

[0041] (1). The color stream processes color information using a 1*1 convolutional layer to obtain a color feature map X color ; Color feature map X color It is expressed as:

[0042] X color = Conv2d1(X pool );

[0043] Among them, Conv2d1() is a convolutional function;

[0044] (2). The texture stream uses multi-scale spatial convolution. First, use a 3*3 convolution, and then use a convolution with a dilation rate of 2 to capture texture features to obtain a texture feature map X texture ; Texture feature map X texture It is expressed as:

[0045] X texture = Conv2d3(X pool );

[0046] Among them, Conv2d3() is a convolutional function;

[0047] (3). The shape stream uses a relatively large convolutional kernel of 5*5 to focus on the shape of the object to obtain a shape feature map X shape ; Shape feature map X shape It is expressed as:

[0048] X shape = Conv2d5(X pool );

[0049] Among them, Conv2d5() is a convolutional function;

[0050] (4) The structural stream extracts multi-scale information through spatial pyramid pooling to focus on spatial structure features, obtaining the structural feature map X structure ; The structural feature map X structure is expressed as:

[0051] X structure = SPP(X pool )

[0052] where SPP() is pyramid pooling;

[0053] The color feature map X color , the texture feature map X texture , the shape feature map X shape and the structural feature map X structure are uniformly represented as the feature maps output by each feature stream where B is the size of the batch, C flow is the number of output channels, and H1 and W1 are the feature sizes after convolution.

[0054] Further, Step 4 includes:

[0055] The dense block contains multiple convolutional layers, and each convolutional layer is densely connected to the feature maps of all previous convolutional layers; in the dense block, the output of each layer is concatenated with the outputs of all previous layers; the feature map X dense after being processed by the dense block is expressed as:

[0056] X dense = DenseBlock(X flow );

[0057] where X flow is the feature map output by each feature stream.

[0058] Further, Step 5 includes:

[0059] The feature map X dense after being processed by the dense block is input into the transition layer. The transition layer first performs BatchNorm and ReLU processing on the feature map after being processed by the dense block, then uses a 1*1 convolution to adjust the channels, and uses a 2*2 average for dimensionality reduction;

[0060]

[0061] Then, pooling is used:

[0062]

[0063] where H2 and W2 are the feature sizes after pooling, and C transition_in is Xtransition_in The number of channels.

[0064] Furthermore, in the said Step6, it includes: the enhanced features Are expressed as:

[0065] X enhanced = Conv2d3(X transition_out );

[0066] Wherein, H3 and W3 are the feature sizes after convolution.

[0067] Furthermore, in the said Step7, it includes: the attention mechanism includes color attention, texture attention, structure attention, semantic attention; the processing process of the attention mechanism includes:

[0068] (1), Color attention uses a channel convolutional self-attention mechanism to perform more detailed modeling on color information;

[0069] X color_att = Attention color (X enhanced );

[0070] (2), Texture attention uses dilated convolution combined with self-attention mechanism to perform more detailed modeling on texture information;

[0071] X texture_att = Attention texture (X enhanced );

[0072] Wherein: Attention texture Adopts a method combining dilated convolution and self-attention, focusing on capturing fine-grained texture information;

[0073] (3), Structure attention uses a spatial pyramid pooling strategy to capture spatial information at different scales;

[0074] X structure_att = Attention structure (X enhanced );

[0075] Wherein: Attention structure Adopts an attention module combined with spatial pyramid pooling SPP to enhance the global structure features of the image;

[0076] (4), Semantic attention dynamically adjusts the weights of features through a channel attention mechanism combined with category information to improve the network's understanding at the semantic level;

[0077] X semantic_att = Attentionsemantic (X enhanced );

[0078] Among them: Attention semantic Channel attention is adopted to emphasize the importance of semantic features on different channels.

[0079] Furthermore, in the said Step8, it includes: in the adaptive weighted sum attention mechanism fusion, the adaptive weighting module in the adaptive weighted sum attention mechanism fusion dynamically adjusts the attention of the adaptive weighting module to different feature streams according to the feature importance of different inputs; the adaptive weighting module calculates the weight of each feature, and represents the weighted fusion process as:

[0080] X fused = AdaptiveWeighting(X color _ att , X texture _ att , X structure _ att , X semantic _ att );

[0081] AdaptiveWeighting() represents the processing of the adaptive weighting module.

[0082] Furthermore, in the said Step9, it includes:

[0083] Step9.1. Compress the spatial information of each feature map X after the adaptive weighted sum attention mechanism fusion into a vector X fused through global average pooling; the vector X pooled is represented as: pooled

[0084] X pooled = GlobalAvgPool(X fused );

[0085] Among them, GlobalAvgPool represents global average pooling;

[0086] Step9.2. Map the vector X pooled to the label space of food classification through a fully connected layer:

[0087] Y = Linear(X pooled ) = W·X pooled + b;

[0088] Among them, W and b are weights and biases, and Linear represents the processing of the linear layer;

[0089] ​Step 9.3: Adopt the stochastic gradient descent optimization strategy and introduce the Dropout regularization technique to suppress overfitting; use the Focal Loss function for optimal design in response to class imbalance and difficult sample learning problems; the mathematical expression of the Focal Loss function is defined as:

[0090] FL(P t ) = -α t (1 - p t ) γ log(p t );

[0091] where P t represents the predicted probability of the FoodBackbone model for the target class; (1 - p t ) γ is the adjustment factor, which achieves the goal of focusing on difficult-to-classify samples by increasing the loss weight of difficult-to-classify samples; γ is the focusing hyperparameter, which is used to reduce the loss weight of easy samples and enhance the learning intensity of difficult samples; α t is the parameter for adjusting class imbalance, usually set as where N is the total number of samples, and N t is the number of samples of class t.

[0092] The beneficial effects of the present invention are:

[0093] 1. The present invention proposes an innovative multi-stream parallel feature extraction architecture, which is specifically designed for the food recognition task; this architecture combines multi-stream feature learning and an efficient feature fusion module, can comprehensively extract multi-dimensional features from food images, and significantly improve the classification accuracy and robustness through feature fusion and enhancement strategies;

[0094] 2. The present invention ingeniously utilizes the advantages of different feature streams (including four independent feature streams (color stream, texture stream, shape stream, structure stream)), and each stream focuses on capturing different features in the image; the color stream can identify the color information in the image, the texture stream captures fine-grained texture features through dilated convolution and self-attention mechanisms, the shape stream focuses on the discrimination of object shapes and the overall contour, and the structure stream processes spatial information at different scales through pyramid pooling technology to extract structured features in the image; this design allows the network to perform feature modeling on the input image from different perspectives, thereby improving the understanding ability of complex food images. By fusing these diverse features, the present invention can identify complex food images and effectively reduce the confusion between similar food types;

[0095] 3. The present invention adjusts the weights of different feature streams through an adaptive weighting method; this enables the network to automatically identify which features are the most critical according to the different characteristics of the image, thereby improving the sensitivity to detailed features and further enhancing the recognition ability for complex food categories;

[0096] 4. The datasets in food recognition tasks usually have the problem of class imbalance. In particular, the data samples of some popular food categories are abundant, while the samples of other rare local special dishes are fewer; the FoodBackbone of the present invention provides multi-stream feature extraction and adaptive weighted fusion, which provides support for dealing with class imbalance; when dealing with class imbalance, the adaptive weighting module of the FoodBackbone network of the present invention can dynamically adjust the weights of each class during the training process, so that the features of minority-class foods can receive sufficient attention, thereby balancing the differences between classes; through this method, the network can still maintain good recognition performance in the case of class imbalance;

[0097] 5. The datasets in food recognition tasks may have the problem of inaccurate annotation. Especially during the training process, incorrect labels will have a negative impact on model learning; to improve the robustness of the model, the FoodBackbone of the present invention introduces an enhanced attention mechanism;

[0098] The FoodBackbone network of the present invention adopts Texture Attention, Structure Attention, and Semantic Attention; these attention modules can help the network adaptively focus on key information regions in the feature map and suppress the interference of irrelevant regions; when there are errors in the data labels, the attention mechanism can automatically reduce the influence of the incorrect label regions, thereby enhancing the sensitivity of the model to correct features;

[0099] 6. Through multi-stream parallel processing, the FoodBackbone of the present invention can comprehensively capture the diverse features of food images, and at the same time introduces DenseBlock and feature fusion modules to efficiently integrate the outputs of each stream; the output of each feature stream depends not only on the input of the current layer, but also uses all the outputs of the previous layer, ensuring the full propagation of information in the network, thereby efficiently utilizing the features of each layer; this dense connection design effectively alleviates the problem of gradient disappearance and enhances the sensitivity of the model to local details and global information;

[0100] 7. The core innovation of the FoodBackbone of the present invention lies in the collaborative work of its multi-stream parallel feature extraction and deep feature fusion, enabling the network to understand the information in food images from multiple dimensions; by gradually refining and enhancing features, the network generates high-dimensional feature representations, thereby improving the classification ability on complex food images; in addition, the network design also incorporates a dynamic feature fusion strategy to enhance the processing ability for complex backgrounds and fine-grained features.

[0101] 8. The design of the FoodBackbone of the present invention breaks through the limitations of traditional food recognition methods and provides broad application prospects for practical application scenarios such as catering services, food quality monitoring, and nutritional analysis; in the fine-grained food classification task, the model demonstrates excellent performance, proving its effectiveness in complex food recognition tasks. Brief Description of the Drawings

[0102] Figure 1 It is the network structure diagram of the FoodBackbone of the present invention;

[0103] Figure 2 It is the structure diagram of the food recognition backbone network of the present invention;

[0104] Figure 3 It is the network structure diagram of the food attention mechanism of the present invention;

[0105] Figure 4 It is the schematic diagram of the classification confusion matrix of the present invention. Detailed Embodiments

[0106] Embodiment 1: As Figures 1-4 shown, a food recognition method, the method includes:

[0107] Step1. Obtain a food image dataset and standardize the input image;

[0108] Further, each image in the dataset in Step1 is adjusted to 224×224 pixels and input using the RGB channels; the division of the dataset follows the conventional ratio of 70% training set, 15% validation set, and 15% test set to ensure the generalization ability of the model; the numerical range of the input image X is normalized to between [0, 1], where the size of the input image X is H×W×C, where C is the number of channels, H and W are the height and width respectively, the shape of the input image X is (224, 224, 3), and the goal is the standardized input to ensure the consistency of the image size and channels;

[0109]

[0110] where, I is the input image pixel value, I normTo standardize the pixel values of the image, μ and σ are the mean and standard deviation of the image respectively.

[0111] Step2. Extract basic features from the standardized input image through an initial convolutional layer;

[0112] Furthermore, Step2 includes: the standardized input image X undergoes preliminary feature extraction through a convolutional layer, using a 7*7 convolution, a stride of 2, and a padding of 3, then through BatchNorm and ReLU activation processing, and then through 3*3 max pooling to complete the basic feature extraction; the specific steps are as follows:

[0113] Step2.1. The input image X undergoes preliminary feature extraction through a convolution operation; the process of preliminary feature extraction is expressed as:

[0114]

[0115] Among them, Conv2d7() represents the convolution process, the convolution kernel size is 7, the stride is 2, and the padding is 3. Among them, H' and W' are the height and width respectively, and X' is the result after convolution;

[0116] Step2.2. Perform BatchNorm and ReLU activation processing on the extracted preliminary features; the process of BatchNorm and ReLU activation processing is expressed as:

[0117] X″ = ReLU(BatchNorm(X′));

[0118] ReLU is the activation function:

[0119] BatchNorm is batch normalization, and the batch normalization process is:

[0120] Among them: γ and β allow the model to recover the necessary scaling and offset after normalization, avoiding information loss caused by normalization; ∈ prevents the denominator from being zero, especially when the batch variance is close to zero. μ and σ are the mean and standard deviation respectively, and batch is the batch size set for the input pictures;

[0121] X” is the processed feature data;

[0122] Step2.3. Perform max pooling processing on the result X” of BatchNorm and ReLU activation processing to obtain the basic feature X pool : The basic feature X pool Is expressed as:

[0123] X pool = MaxPool2d3(X″);

[0124] Among them, MaxPool2d3() represents max pooling processing, where the kernel size of the max pooling processing is 3, the stride is 2, and the padding is 1.

[0125] Step 3, Multi-Stream Feature Extraction: The basic features are processed through four parallel feature streams for color, texture, shape, and structure features respectively. Each feature stream uses different convolution kernels and processing strategies;

[0126] Each stream focuses on extracting different types of features (color, texture, shape, and structure). This design allows the network to model the features of the input image from different perspectives, thereby improving the ability to understand complex food images.

[0127] Furthermore, in the above-mentioned Step 3, it includes: The feature streams include a color stream, a texture stream, a shape stream, and a structure stream. The processing process of each feature stream includes:

[0128] (1), The Color Stream uses a 1*1 convolutional layer to process color information and obtains a color feature map X color ; The color feature map X color is expressed as:

[0129] X color = Conv2d1(X pool );

[0130] Among them, Conv2d1() is a convolution function;

[0131] (2), The Texture Stream uses multi-scale spatial convolution. First, it uses a 3*3 convolution, and then a convolution with a dilation rate of 2 to capture texture features and obtains a texture feature map X texture ; The texture feature map X texture is expressed as:

[0132] X texture = Conv2d3(X pool );

[0133] Among them, Conv2d3() is a convolution function;

[0134] (3), The Shape Stream uses a relatively large convolution kernel of 5*5 to focus on the shape of the object and obtains a shape feature map X shape ; The shape feature map X shape is expressed as:

[0135] X shape = Conv2d5(X pool );

[0136] Among them, Conv2d5() is a convolution function;

[0137] (4), The Structure Stream extracts multi-scale information through Spatial Pyramid Pooling (SPP) to focus on spatial structure features, obtaining the structure feature map X structure ; The structure feature map X structure is expressed as:

[0138] X structure = SPP(X pool )

[0139] Among them, SPP() is pyramid pooling;

[0140] The color feature map X color , the texture feature map X texture , the shape feature map X shape and the structure feature map X structure are uniformly represented as the feature maps output by each feature stream Among them, B is the size of the batch, C flow is the number of output channels, and H1 and W1 are the feature sizes after convolution.

[0141] The process of Spatial Pyramid Pooling (SPP) is as follows:

[0142] (1) Divide into multiple sub-regions according to different levels l, and each sub-region is S l *S l

[0143]

[0144] where N l is the number of sub-regions at each level;

[0145] (2) Max pooling for each sub-region:

[0146]

[0147] where, x l is the feature of the sub-region;

[0148] (3) Concatenate the sub-regions:

[0149] SPP(X pool ) = Concat(pool1(x1), pool2(x2), …, pool l(x l ))

[0150] where Concat is the concatenation of feature vectors.

[0151] Step 4. The feature maps output by each feature stream are processed through a DenseBlock;

[0152] Furthermore, Step 4 includes:

[0153] The DenseBlock contains multiple convolutional layers, and each convolutional layer is densely connected (dense connection) to the feature maps of all previous convolutional layers; in the DenseBlock, the output of each layer is concatenated with the outputs of all previous layers; the feature map X dense is expressed as:

[0154] X dense = DenseBlock(X flow );

[0155] where X flow is the feature map output by each feature stream, and the DenseBlock enhances the richness of the feature map through skip connections.

[0156] The detailed internal implementation process of the DenseBlock is as follows:

[0157] (1). The basic building unit of dense connection:

[0158] F l = ReLU(BatchNorm(X flow ))

[0159] where l is the layer number;

[0160] (2). Sequential processing and feature concatenation:

[0161] First layer: X1 = F1(X flow );

[0162] Second layer: X2 = F2([X flow , X1]);

[0163] Third layer: X3 = F3([X flow , X2, X1]);

[0164] ......

[0165] The l-th layer: X l = F l ([X flow , X l-1 ,..., X l );

[0166] (3) Feature Fusion Mechanism:

[0167]

[0168] where k is the growth rate, representing the number of channels in each layer;

[0169] The result is X dense = Concat(X flow , X1, X2, …, X l ).

[0170] Step5. Input the feature map processed by the dense block into the transition layer (Transition) for dimensionality reduction;

[0171] Input the feature map X processed by the dense block dense into the transition layer. The transition layer first performs BatchNorm and ReLU processing on the feature map processed by the dense block, then adjusts the channels using 1*1 convolution, and performs dimensionality reduction using 2*2 average pooling;

[0172]

[0173] Then, use pooling:

[0174]

[0175] where H2 and W2 are the feature sizes after pooling, and C transition_in is the number of channels of X transition_in .

[0176] Step6. Input the feature map processed by the transition layer into the Feature Enhancement module to further perform convolution processing on the feature map processed by the transition layer to make the features more discriminative and obtain the enhanced features;

[0177] Furthermore, in Step6, it includes: the enhanced features are represented as:

[0178] X enhanced = Conv2d3(X transition_out );

[0179] where H3 and W3 are the feature sizes after convolution.

[0180] Step7. Input the enhanced features into the Attention Mechanism for processing;

[0181] Further, in the Step 7, it includes: the attention mechanism includes Color Attention, Texture Attention, Structure Attention, and Semantic Attention; the processing process of the attention mechanism includes:

[0182] (1), Color Attention conducts more detailed modeling on color information through a channel convolutional self-attention mechanism;

[0183] X color_att = Attention color (X enhanced );

[0184] (2), Texture Attention conducts more detailed modeling on texture information through dilated convolution combined with self-attention mechanism, adapts to variable texture patterns, and enhances the network's ability to capture details;

[0185] X texture_att = Attention texture (X enhanced );

[0186] Among them: Attention texture Adopts a method combining dilated convolution and self-attention, focusing on capturing fine-grained texture information;

[0187] (3), Structure Attention uses the spatial pyramid pooling (SPP) strategy to capture spatial information at different scales, helping the network understand multi-level structure information in the image;

[0188] X structure_att = Attention structure (X enhanced );

[0189] Among them: Attention structure Adopts an attention module combined with spatial pyramid pooling SPP to enhance the global structure features of the image, improve the network's understanding at the semantic level, and help the model focus on key regions; captures spatial information at different scales through the spatial pyramid pooling (SPP) strategy to help the network understand multi-level structure information in the image;

[0190] (4), Semantic Attention dynamically adjusts the weights of features by combining channel attention mechanism with category information to improve the network's understanding at the semantic level;

[0191] X semantic_att = Attention semantic (X enhanced );

[0192] Where: Attention semantic Adopts channel attention to emphasize the importance of semantic features on different channels.

[0193] Step8, Adaptive Weighting and Attention Fusion: Fuse the feature maps processed by all attention mechanisms, and adjust the contributions of the features processed by each attention mechanism through an adaptive weighting module; The Adaptive Weighting Module network learns the weights of the input features through an adaptive weighting module. This enables the network to dynamically adjust its attention to different feature streams according to the importance of different input features, making the final feature representation more accurate and useful.

[0194] Furthermore, in Step8, it includes: The adaptive weighting module in the Adaptive Weighting and Attention Fusion dynamically adjusts the attention of the adaptive weighting module to different feature streams according to the importance of different input features; The adaptive weighting module calculates the weights of each feature, and represents the weighted fusion process as:

[0195] X fused = AdaptiveWeighting(X color _ att , X texture _ att , X structure _ att , X semantic _ att )

[0196] AdaptiveWeighting() represents the processing of the adaptive weighting module.

[0197] The detailed internal implementation process of AdaptiveWeighting is as follows:

[0198] (1), All feature extraction:

[0199]

[0200] (2), Feature splicing and weight calculation:

[0201] Feature splicing:

[0202] g concat = [gcolor , g texture , g structure , g semantic

[0203] Calculate the attention weights:

[0204] z = ReLU(W·g concat + b)

[0205] where W is the weight and b is the bias;

[0206] Normalization:

[0207] α = softmax(z) = [α color , α texture , α structure , α semantic

[0208] (3), Weighted fusion:

[0209] X fused = α color ·g color + α texture ·g texture + α structure ·g structure + α semantic ·g semantic .

[0210] This fusion method combines multiple features, enabling the network to comprehensively integrate information from multiple perspectives, thereby providing a more comprehensive representation of the food image;

[0211] Step9. The feature map after feature fusion is input into the Classification Module for final classification prediction.

[0212] Furthermore, Step9 includes:

[0213] Step9.1. Compress the spatial information of each feature map X after adaptive weighted sum and attention mechanism fusion into a vector X fused through Global Average Pooling (GAP); the vector X pooled is expressed as: pooled X pooled = GlobalAvgPool(X fused ;

[0215] ;

[0216] where GlobalAvgPool represents global average pooling; ​​

[0216] Step 9.2. Map the vector X to the label space of food classification through a fully connected layer: pooled

[0217] Y = Linear(X pooled ) = W · X pooled + b;

[0218] where W and b are weights and biases, and Linear represents the processing of the linear layer;

[0219] Step 9.3. Adopt the Stochastic Gradient Descent optimization strategy and introduce the Dropout regularization technique to suppress overfitting; use the Focal Loss function to optimize the design for the problems of class imbalance and difficult sample learning; the mathematical expression of the Focal Loss function is defined as:

[0220] FL(P t ) = -α t (1 - p t ) γ log(p t );

[0221] where P t represents the predicted probability of the FoodBackbone model for the target class; (1 - p t ) γ is the adjustment factor, which realizes the goal of focusing on difficult-to-classify samples by increasing the loss weight of difficult-to-classify samples; γ is the focusing hyperparameter, which is used to reduce the loss weight of easy samples and enhance the learning intensity of difficult samples; α t is the parameter for adjusting class imbalance, usually set as where N is the total number of samples, and N t is the number of samples of class t.

[0222] The present invention proposes an innovative multi-stream food recognition neural network architecture - FoodBackbone. This architecture combines multiple feature streams (color stream, texture stream, shape stream, and structure stream), and through unique Dense Blocks and Transition Layers, effectively extracts and fuses features, thereby significantly improving the accuracy and robustness of food image classification. The design of FoodBackbone fully considers the complementarity of different types of features (such as color, texture, shape, and structure), and through Spatial Pyramid Pooling (SPP) and multi-scale convolution techniques, enhances the recognition ability for fine-grained food images.

[0223] ​In the design of the network, first, a basic feature is extracted through an initial convolutional layer, and then the color, texture, shape, and structure features are processed through four dedicated streams respectively. Each stream uses different convolutional kernels and processing strategies to ensure a comprehensive analysis of food images from multiple perspectives. After all the streams pass through the dense blocks, dimensionality reduction is performed through the transition layer, ensuring that the feature maps maintain important information without excessive redundancy. Subsequently, the network further enhances the discriminability of the feature maps through the feature enhancement module, and finally, the attention mechanism module (including texture attention, structure attention, and semantic attention) automatically focuses on important regions, improving the performance of the model.

[0224] The FoodBackbone of the present invention has achieved a significant performance improvement in traditional food recognition tasks. Especially in complex food image classification and fine-grained recognition tasks, it shows extremely strong robustness. The network of the present invention can automatically adjust the weights of various features, making it have broad application potential in practical applications (such as automatic dietary analysis, nutritional assessment, and food safety monitoring, etc.). The present invention elaborates in detail on the architecture design principle, feature stream processing strategy, and the implementation method of the attention mechanism of FoodBackbone, and verifies its efficiency in diverse food recognition tasks through experiments. The core innovation of FoodBackbone lies in the combination of its multi-stream feature extraction and adaptive feature fusion mechanism, which greatly improves the classification performance of complex food images and promotes the development of fine-grained food recognition technology.

[0225] The design goal of the present invention, FoodBackbone, is to provide stronger feature expression ability and adaptability in food recognition tasks by combining feature extraction networks of multiple streams and an enhanced attention mechanism. Traditional food recognition methods usually rely on a single convolutional neural network (CNN) to extract image features. However, these methods may be limited when dealing with complex backgrounds, different food categories, and tasks with different levels of detail importance. To overcome these challenges, the present invention, FoodBackbone, adopts a multi-stream network structure and an enhanced attention mechanism, enabling the network to adaptively focus on different features in the image and enhancing its adaptability to complex food images. Multi-Stream Feature Extraction: FoodBackbone employs four parallel feature streams (color stream, texture stream, shape stream, and structure stream), with each stream focusing on extracting different aspects of the image. The color stream focuses on the color information of the image, the texture stream captures texture features, the shape stream focuses on the shape of the object, and the structure stream focuses on spatial structure features. These different streams can extract different food features simultaneously, enhancing the overall feature expression ability. The FoodBackbone network provides an architecture with efficient feature expression ability by combining multi-stream feature extraction, an enhanced attention mechanism, and an adaptive weighted feature fusion method. It can flexibly adapt to different types of food images, especially demonstrating strong performance and robustness in complex backgrounds, fine-grained features, and multi-category recognition tasks.

[0226] Taking the TestSet of the food2k food recognition dataset as an example, the present invention uses classification accuracy as the evaluation metric, which is defined as: for a given test dataset, the ratio of the number of samples correctly classified by the classifier to the total number of samples. The larger the value, the better the performance of the model; the lower the value, the worse the performance of the model.

[0227] Table 1 shows the effects of the same model under different evaluation metrics

[0228] Serial number Verify different evaluation metrics Algorithm classification accuracy 1 OverallAccuracy 91.72% 2 Top-3Accuracy 98.51% 3 Top-5Accuracy 99.15%

[0229] Serial number 1: The proportion of samples for which the first prediction result of the model is correct among all samples is close to 92%.

[0230] Serial number 2: It indicates that among the top 3 predictions, 98.51% of the samples contain the correct class label.

[0231] Serial number 3: Among the top 5 predictions, 99.15% of the samples contain the correct class label.

[0232] The results in Table 1 show that the FoodBackbon model has extremely strong classification ability. In most cases, the model can not only give correct predictions, but also include the correct category in the top few predictions;

[0233] Especially, the Top-3 and Top-5 accuracies are very high, indicating that the model has a strong ability to distinguish between different categories. For a complex classification task like food recognition, this result is very excellent.

[0234] Figure 4 This is a schematic diagram of the classification confusion matrix of the present invention; it can be seen that the classification effect is extremely good.

[0235] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A food recognition method, characterized in that: The method includes: Step1. Obtain a food image dataset and standardize the input image; Step2. Extract basic features from the standardized input image through an initial convolutional layer; Step3. Process color, texture, shape, and structure features respectively for the basic features through four parallel feature streams, and each feature stream adopts different convolutional kernels and processing strategies; Step4. The feature maps output by each feature stream are processed through a dense block; Step5. Input the feature maps processed through the dense block into a transition layer for dimensionality reduction; Step6. Input the feature maps processed through the transition layer into a feature enhancement module to further perform convolutional processing on the feature maps processed through the transition layer to obtain enhanced features; Step7. Input the enhanced features into an attention mechanism for processing; Step8. Adaptive weighted sum and attention mechanism fusion: fuse all the feature maps processed by the attention mechanism, and adjust the contribution of the features processed by each attention mechanism through an adaptive weighting module; Step9. The feature maps after feature fusion will be input into a classification module for final classification prediction.

2. The food recognition method according to claim 1, wherein: Each image in the dataset described in the above Step1 is adjusted to 224×224 pixels and input using the RGB channels; the numerical range of the input image X is normalized to between [0, 1], where the size of the input image X is where C is the number of channels, and H and W are the height and width respectively; where I is the input image pixel value, and I norm is the normalized image pixel value, and μ and σ are the mean and standard deviation of the image, respectively.

3. The food recognition method according to claim 1, wherein: In Step2, it includes: the standardized input image X is subjected to preliminary feature extraction through a convolutional layer, then processed through BatchNorm and ReLU activation, and then through 3*3 max pooling to complete basic feature extraction; the specific steps are: Step2.

1. The input image X is subjected to preliminary feature extraction through a convolutional operation; the process of preliminary feature extraction is expressed as: Among them, Conv2d7() represents convolutional processing, the convolutional kernel size is 7, the stride is 2, and the padding is 3. Among them, H' and W' are the height and width respectively, and X' is the result after convolution; Step2.

2. Perform BatchNorm and ReLU activation processing on the extracted preliminary features; the process of BatchNorm and ReLU activation processing is expressed as: X″ = ReLU(BatchNorm(X′)); ReLU is the activation function: BatchNorm is batch normalization, and the batch normalization process is as follows: Among them: γ and β allow the model to recover the necessary scaling and offset after normalization, avoiding information loss caused by normalization; ∈ prevents the denominator from being zero, especially when the batch variance is close to zero. μ and σ are the mean and standard deviation respectively, and batch is the batch size set for the input pictures; X” is the processed feature data; Step 2.3: Perform max pooling on the result X” of BatchNorm and ReLU activation processing to obtain the basic feature X pool : Basic feature X pool It is expressed as: X pool = MaxPool2d3(X″); Among them, MaxPool2d3() represents max pooling processing, the kernel size of max pooling processing is 3, the stride is 2, and the padding is 1.

4. A food recognition method according to claim 1, characterized in that: In Step3, it includes: the feature streams include a color stream, a texture stream, a shape stream, and a structure stream, and the processing process of each feature stream includes: (1), The color stream processes color information using a 1*1 convolutional layer to obtain the color feature map X color ; The color feature map X color is expressed as: X color = Conv2d1(X pool ); Among them, Conv2d1() is a convolutional function; (2) The texture stream uses multi-scale spatial convolution. First, it uses a 3×3 convolution, and then a convolution with a dilation rate of 2 to capture texture features, obtaining a texture feature map X texture ; The texture feature map X texture is expressed as: X texture = Conv2d3(X pool ); Among them, Conv2d3() is a convolutional function; (3) Shape flow uses a larger convolutional kernel of 5×5 to focus on the shape of the object and obtains a shape feature map X shape ; The shape feature map X shape is expressed as: X shape = Conv2d5(X pool ); Among them, Conv2d5() is a convolutional function; (4) The structural stream extracts multi-scale information through spatial pyramid pooling to focus on spatial structure features, obtaining the structural feature map X structure ; The structural feature map X structure is expressed as: X structure = SPP(X pool ) Among them, SPP() is pyramid pooling; The color feature map X color , the texture feature map X texture , the shape feature map X shape and the structure feature map X structure are uniformly represented as the feature maps output by each feature stream where B is the size of the batch, C flow is the number of output channels, and H1 and W1 are the feature sizes after convolution.

5. The food recognition method according to claim 1, wherein: In Step4, it includes: The dense block contains multiple convolutional layers, and each convolutional layer is densely connected to the feature maps of all previous convolutional layers; in the dense block, the output of each layer is concatenated with the outputs of all previous layers; the feature map X processed by the dense block dense is expressed as: X dense = DenseBlock(X flow ); Among them, X flow is the feature map output for each feature stream.

6. The food recognition method according to claim 1, wherein: In Step5, it includes: The feature map X after being processed by the dense block dense is input into the transition layer. The transition layer first performs BatchNorm and ReLU processing on the feature map after being processed by the dense block, then uses a 1×1 convolution to adjust the channels, and uses a 2×2 average pooling for dimensionality reduction; Then, use pooling: Among them, H2 and W2 are the feature sizes after pooling, and C transition_in is X transition_in 's number of channels.

7. A food recognition method according to claim 1, characterized in that: The Step 6 includes: the enhanced feature which is expressed as: X enhanced = Conv2d3(X transiton_out ); Among them, H3 and W3 are the feature sizes after convolution.

8. The food recognition method according to claim 1, wherein: Step 7 includes: the attention mechanism includes color attention, texture attention, structural attention, and semantic attention; the processing process of the attention mechanism includes: (1) Color attention performs more detailed modeling on color information through a channel convolutional self-attention mechanism; X color_att = Attention color (X enhanced ) (2) Texture attention performs more detailed modeling on texture information through dilated convolution combined with a self-attention mechanism; X texture_att = Attention texture (X enhanced ) Among them: Attention texture Adopt a method that combines dilated convolution and self-attention, focusing on capturing fine-grained texture information; (3) Structural attention uses a spatial pyramid pooling strategy to capture spatial information at different scales; X structure_att = Attention structure (X enhanced ); Among them: Attention structure An attention module combined with spatial pyramid pooling (SPP) is adopted to enhance the global structural features of the image; (4) Semantic attention dynamically adjusts the weights of features by combining channel attention mechanisms with category information to enhance the network's understanding at the semantic level; X semantic_att = Attention semantic (X enhanced ); Among them: Attention semantic Channel attention is adopted to emphasize the importance of semantic features on different channels.

9. A food recognition method according to claim 1, characterized in that: Step 8 includes: in the adaptive weighted sum attention mechanism fusion, the adaptive weighting module in the adaptive weighted sum attention mechanism fusion dynamically adjusts the attention of the adaptive weighting module to different feature streams according to the importance of different inputs; the adaptive weighting module calculates the weights of each feature, and the weighted fusion process is expressed as: X fused = AdaptiveWeighting(X color_att , X texture_att , X structure_att , X semantic_att ) AdaptiveWeighting() represents the processing of the adaptive weighting module.

10. A food recognition method according to claim 1, characterized in that: Step 9 includes: Step 9.

1. Compress the spatial information of the feature map X after fusing each adaptive weighted sum attention mechanism into a vector X through global average pooling fused ; The vector X pooled is expressed as: pooled ​ X pooled = GlobalAvgPool(X fused ); Among them, GlobalAvgPool represents global average pooling; Step9.

2. Map the vector X through a fully connected layer pooled to the label space of food classification: Y = Linear(X pooled ) = W·X pooled + b; Among them, W and b are weights and biases, and Linear represents linear layer processing; Step 9.3: Adopt a stochastic gradient descent optimization strategy and introduce Dropout regularization technology to suppress overfitting; Use the Focal Loss loss function for optimized design for class imbalance and difficult sample learning problems; the mathematical expression of the FocalLoss loss function is defined as: FL(P t ) = -α t (1 - p t ) γ log(p t ); Among them, P t represents the predicted probability of the target category by the FoodBackbone model; (1 - p t ) γ is a modulating factor that achieves the goal of focusing on difficult-to-classify samples by increasing the loss weight of difficult-to-classify samples; γ is a focusing hyperparameter used to reduce the loss weight of easy samples and enhance the learning intensity of difficult samples; α t is a parameter for adjusting the inter-class imbalance, usually set to where N is the total number of samples, and N t is the number of samples in class t.