Food image recognition and calorie detection method based on deep learning

By introducing an occlusion recognition module and fill algorithm into the food image recognition network, combining the ResNeXt and Inception_Resnet recognition modules and the attention mechanism of the yolov8 network, the feature loss and noise problems caused by occlusion in food image recognition are solved, and the accuracy and robustness of the recognition are improved.

CN119992538APending Publication Date: 2025-05-13HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510124871.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the process of food image recognition, there are uncertain factors such as lighting, food movement, focal length changes at different angles, and other objects such as tableware or fingers that obstruct food, which affects the accuracy of food recognition, especially the problem of local obstruction of food is more prominent.

Method used

The network includes an occlusion recognition module, through which the module recognizes whether there is occlusion in the image and adjusts the hyperparameters. The discriminated image is passed into the recognition module composed of ResNeXt and Inception_Resnet in parallel, and finally the recognition result is obtained through the decision fusion method. The yolov8 module is used for food object detection, first identifying through the occlusion recognition module, then filling the occlusion image through the filling algorithm, and then detecting the food in the image through the yolov8 network with the attention mechanism.

Benefits of technology

Through the occlusion recognition and filling algorithm, the accuracy and robustness of food image recognition are improved, and the feature loss and noise problems caused by local occlusion of food are effectively solved, which is improved and the completeness and accuracy of the final recognition results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992538A_ABST
    Figure CN119992538A_ABST
Patent Text Reader

Abstract

The invention discloses a food image recognition and calorie detection method based on deep learning, and belongs to the field of image recognition algorithms. According to an existing food and food recognition method, when food is shielded, feature loss is caused, the problems of noise, local aliasing and the like are included, and therefore extraction and matching of a recognition algorithm on facial features are interfered. A food image recognition and calorie detection method based on deep learning comprises the following steps: establishing a data set of simulated food images, and performing preprocessing; the preprocessing comprises image shielding processing, food target detection data set preprocessing, and creation of a data set required by a shielding detection network; determining a detection network and an image filling network for detecting the food image with the shielding condition; constructing a judgment network used for judging whether the to-be-detected food image has a shielding condition or not; and constructing a detection network for detecting the food image with the shielding condition to obtain a final result. According to the invention, the identification precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a food image recognition method, and in particular to a food image recognition and calorie detection method based on deep learning. Background Art

[0002] So far, the types and quantities of food consumed by people have increased with the improvement of living standards. Diseases caused by diet are becoming more and more common among the population. People hope to regulate the impact on the human body by understanding the properties of the food they eat. Therefore, research on deep learning based on diet and food image recognition has been carried out and has made rapid progress. Modeling the unique characteristics of food images will make important progress in the research of food preference learning, food image calorie estimation and personalized recipe recommendation. Good results have been achieved in both speed and accuracy. However, there are still many challenges to be solved in the image recognition method itself or in the process of food image recognition. For example, in the recognition process, the shadows caused by different angles of illumination, the blur caused by food movement and focal length change, and the occlusion of food by other objects such as tableware or fingers will affect the recognition of food to varying degrees. Among the many factors that affect the accuracy of food recognition, the problem of partial occlusion of food is particularly prominent. The extraction and comparison of key features of food are the key to the food recognition algorithm. The completeness of important features will greatly affect the final recognition results. When food is occluded, it will cause feature loss, noise and local aliasing, which will interfere with the recognition algorithm's extraction and matching of facial features. Summary of the invention

[0003] The purpose of the present invention is to solve the above problems. An occlusion recognition module is included in the network, through which it can be identified whether there is occlusion in the image, and the corresponding hyperparameters are adjusted. The discriminated image is passed into the recognition module composed of ResNeXt and Inception_Resnet in parallel, and the recognition result is finally obtained by the decision fusion method. The present invention also uses the yolov8 module to detect food targets, firstly identifies it through the occlusion recognition module, then fills the occluded image through the filling algorithm, and then detects the food in the image through the yolov8 network with an attention mechanism. Thereby forming a food image recognition and calorie detection method based on deep learning.

[0004] The above purpose is achieved through the following technical solutions: A food image recognition and calorie detection method based on deep learning, the method is implemented by the following steps: Step 1: Create a dataset of simulated food images and perform preprocessing; the preprocessing includes image occlusion processing, food target detection dataset preprocessing, and creating a dataset required for the occlusion detection network; Step 2: Determine a detection network and an image filling network for detecting food images with occlusion; Step 3: Build a judgment network to determine whether the food image to be detected is blocked: Determine whether the food image to be detected is occluded. If not, set the hyperparameter λ in the decision-level fusion method to 0.3 for decision fusion. If yes, select the ResNeXt network and the Inception-ResNet network to obtain image features, and then use the decision-level fusion method to perform decision fusion. The Inception-ResNet network improves network performance by introducing group convolution and group number methods; Step 4: Build a detection network for detecting food images with occlusion: The image is input into the occlusion judgment network for preprocessing. If there is no occlusion in the image, the image filling network is skipped directly, and the image is input into the modified yolov8 network for processing to obtain the final output; if there is occlusion in the image, the occlusion is first filled through the image filling network, and then the filled image is input into the yolov8 network for processing to obtain the final result.

[0005] Furthermore, the step 1 of establishing a dataset of simulated food images includes: ETHZ Food-101 dataset, containing 101 categories of Western food and 101,000 images of dishes, with each category containing 1,000 images, including 750 training images and 250 test images; FoodX-251 dataset, a dataset containing 251 fine-grained classes, with 118,000 training, 12,000 validation, and 28,000 test images; UEC Food100 dataset, including 100 categories of Japanese dishes and 12,905 dish images; UEC Food256 dataset, including Japanese dishes, the dish categories are expanded from the original 100 to 256, and the corresponding number of dish images increases to 25,088; The occlusion effect of dishes in different scenes is simulated by simulating the occlusion of dish images, including: Shadow occlusion, which simulates occlusion by adding shadow effects on the dish images; Random image block occlusion: randomly distribute multiple small image blocks or color blocks on the dish image to simulate the effect of the dish being partially covered or blocked. Rectangular or circular occlusion: add rectangular, circular or other shaped occlusion blocks to the dish image to simulate the occlusion effect of cutlery, fingers or other objects.

[0006] Furthermore, the steps of image occlusion processing are specifically: The images in the data set are cut to obtain the corresponding rectangular or semicircular image blocks, and the image blocks are used to occlude the data set. The processing methods are: (1) Cropping images to a fixed size: First, take images from the dataset and crop each image to a size of 448x448 pixels; (2) Image scaling: The 448x448 pixel image obtained by cropping in the previous step is further scaled to 224x224 pixels; (3) Randomly cutting image blocks: After the image is scaled to 224x224 pixels, a random cutting operation is performed on each image, that is, a portion of the area is randomly selected for cutting to generate an image block for occlusion; (4) Random occlusion processing: Use the image generated in the previous step or directly from the scaled image to randomly select one or more positions from the four vertices for occlusion; In this way, a dataset containing occlusion effects is obtained, which can be used to train the model to improve the robustness of the model to occlusion.

[0007] Furthermore, the method for preprocessing a food target detection dataset refers to selecting a dataset in a Yolo format, creating a dataset in a Yolo format, and preprocessing the dataset. The specific steps include: (1) Traverse all categories of the dataset: After creating the necessary files, loop through the folder numbers from 1 to 101. Each number corresponds to an image of a category. Create a dictionary for each category to store the width and height information of each image in that category. (2) Process images in each category: traverse all files in the current category folder; Check if the file is an image file; use PIL.Image to open the image, get its width and height, save them to the dictionary, and move the image to a new directory; (3) Processing bounding box information: Open and read the corresponding file content; Skip the first line of the file, and then process the remaining bounding box information line by line. For each line of bounding box information, parse out the category ID and the coordinates of the bounding box (x_min, y_min, x_max, y_max); find the width and height information of the image based on the file name, and calculate the coordinates of the center point and the aspect ratio of the bounding box; write this information to the file corresponding to the image, and the file name is based on the category ID but removes the file extension.

[0008] Furthermore, the process of creating the dataset required by the occlusion detection network is as follows: (1) Define the function for processing images: create a folder, traverse each category folder in the source folder, and process the images in each category folder; (2) Processing unsegmented and segmented datasets: copy the unsegmented images to the images / 1 directory and mark them as 1; copy the segmented images to the images / 0 directory and mark them as 0; the file names and labels of all images are written to the label.txt file; (3) Split the data set: Read all the lines in the label.txt file, shuffle the order, and then split it into a training set and a test set in a ratio of 3:1; write the lines of the training set and the test set into the train.txt and test.txt files respectively; write the labels of two categories in the classes.txt file: 0 and 1.

[0009] Furthermore, the detection network for detecting food images with occlusion in step 2 is specifically: using DenseNet to detect occlusion in the image; wherein, The DenseNet network uses a DenseBlock combined with a Transition structure. The Transition module connects two adjacent DenseBlocks and reduces the size of the feature map through Pooling. Bottleneck layer: DenseNet first uses 1*1 convolution to reduce the input data channel to 4k, and then uses 3*3 convolution to extract features; Transition layer: The layer between every two dense blocks is called the Transition layer, which completes the convolution and pooling operations; the transition layer consists of a BN layer, a 1x1 convolution layer, and a 2x2 average pooling layer: the 1×1 convolution layer is used to reduce the dimension, the average pooling layer reduces the size of the feature map, and the Transition layer generates the output feature map; The image filling network for determining the food image with occlusion in step 2 is specifically: The partial convolutional neural network (PCNN) is selected to process the missing area by using some convolutional layers, which normalize the convolution kernel according to the number of valid pixels in the filled area; PConv is a partial convolution layer, Filter Size is the specified convolution kernel size, and Stride is the stride; PConv1-8 is an encoder, and PConv9-16 is a decoder; the Batch Normalization column indicates whether there is a BatchNormalization layer after PConv; the Nonlinearity column indicates whether a nonlinear layer is used and what nonlinear layer is used; Concat represents a skip connection, which fuses the upsampling result with the PConv result of the corresponding stage of the encoder.

[0010] Furthermore, in the process of constructing the judgment network for judging whether the food image to be detected is blocked as described in step 3, The grouped convolution adopted by the ResNeXt network is to divide the feature maps into different groups, and then convolve each group of feature maps separately, and each convolution kernel only processes part of the channels; The network structure of the Inception network is divided into three parts: Steam, Inception-resnet and Reduction; among them, Stem reduces the resolution of the input image through multi-layer convolution operations; Reduction structure reduces the size of the feature map through convolution layers or pooling layers; The Inception-resnet structure introduces transposed convolution; by using convolution kernels of different sizes and pooling operations in parallel, it captures features of different scales in the image at the same time, and then fuses these features through the Concat operation.

[0011] Furthermore, the modified yolov8 network described in step 4 refers to the yolov8 network with the ECA attention mechanism added, wherein, The backbone network of the YOLOv8 network adopts a structure similar to that of CSPDarknet. The backbone part defines the basic architecture of the model, that is, the network structure for feature extraction; the neck network is located between the backbone network and the head network for feature fusion and enhancement; the head network defines the detection head of the model, which is the network structure for final target detection; the ConvModule module contains convolutional layers, BN and activation functions for feature extraction; the DarknetBottleneck module increases the network depth through residual connections while maintaining efficiency; the CSP Layer module is a variant of the CSP structure, which improves the training efficiency of the model through partial connections; The specific steps of adding ECA attention mechanism to the yolov8 network are as follows: Introduce ECA channel attention module in the backbone network Backbone of yolov8. By assigning higher weights to important channels, the ECA attention mechanism can enhance the feature expression of these channels, while reducing the weights of unimportant channels; The implementation process of the yolov8 network with the ECA attention mechanism added is as follows: (1) The input feature map is subjected to global average pooling, and the feature map is converted from a matrix of [h,w,c] to a vector of [1,1,c]; (2) Calculate the adaptive one-dimensional convolution kernel size kernel_size according to the number of channels of the feature map; (3) Use kernel_size in one-dimensional convolution to obtain the weight for each channel of the feature map; (4) Multiply the normalized weights and the original input feature map channel by channel to generate a weighted feature map; The overall network processing flow of the yolov8 network with the ECA attention mechanism added is: The image is input into the occlusion judgment network for preprocessing. If there is no occlusion in the image, the image filling network is skipped directly, and the image is input into the modified yolov8 network for processing to obtain the final output; if there is occlusion in the image, the occlusion is first filled through the image filling network, and then the filled image is input into the yolov8 network for processing to obtain the final result.

[0012] Furthermore, the process of introducing transposed convolution in the Inception-resnet structure is as follows: Use transposed convolution to obtain the optimal upsampling method through network learning; In the transposed convolution, let the input matrix be X, the input matrix be Y, and there is a new convolution kernel matrix C:

[0013] The transposed convolution is actually to perform the inverse operation of this process, that is, to obtain X through C and Y:

[0014] Weight matrix for transposed convolution It does not necessarily come from the original convolution matrix C, but its shape is the same as the transpose of the original convolution matrix C.

[0015] Furthermore, the decision-level fusion process is specifically: Based on the complementarity of the learned original features of the image and the occlusion features at the feature layer, the decision-level fusion strategy is selected for research, and the cross entropy loss function is selected. The calculation formula is as follows:

[0016] The final loss function is defined as:

[0017] Where λ is a hyperparameter that balances the two parts; The input image is preprocessed by the occlusion judgment network. If there is no occlusion in the image, the hyperparameter λ is set to 0.3. If there is occlusion in the image, the hyperparameter λ is set to 0.7. The image is then fused through the dish recognition network to obtain the final result on the ETHZ Food-101 occlusion dataset.

[0018] The beneficial effects of the present invention are: The present invention designs a food recognition network. The network includes an occlusion recognition module, through which it can be identified whether there is occlusion in the image and the corresponding hyperparameters can be adjusted. The discriminated image is passed into the recognition module composed of ResNeXt and Inception_Resnet in parallel, and finally the recognition result is obtained by the decision fusion method. This paper also uses the yolov8 module to detect food targets. First, the occlusion recognition module is used for recognition, and then the occluded image is filled by the filling algorithm. Finally, the food in the image is detected by the yolov8 network with an attention mechanism. Specifically: 1. The present invention introduces transposed convolution in the Inception-resnet structure. Usually, after multiple convolution operations are performed on an image, the size of the feature map will continue to shrink. For certain specific tasks, the image needs to be restored to its original size before operation. This size restoration operation of mapping an image from a small resolution to a large resolution is called upsampling. There are many upsampling methods, such as nearest neighbor interpolation, linear interpolation, bilinear interpolation, and bicubic interpolation. However, these upsampling methods are designed based on people's prior experience, and the effects are not ideal in many scenarios. Therefore, this article adopts transposed convolution. The upsampling method of transposed convolution is not a preset interpolation method, but like the standard convolution, it has learnable parameters, and the optimal upsampling method can be obtained through network learning.

[0019] 2. Design a decision-level fusion method. Feature-level fusion and decision-level fusion are two commonly used fusion methods. The former directly combines the feature vectors of the two branches into a joint feature vector to train a FER classifier; the latter combines the recognition results of the two branches. In the network of the present invention, the learned original features and occlusion features of the image are weakly complementary at the feature layer, so the decision-level fusion strategy is selected for research. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flow chart of the method involved in the present invention; Figure 2 It is a diagram showing the difference between the ResNeXt and ResNet modules involved in the present invention; Figure 3 This is the Inception-ResNet network structure diagram involved in the present invention; Figure 4 This is a schematic diagram of introducing a transposed convolution module into the Inception-resnet structure involved in the present invention; Figure 5 This is the yolov8 network structure diagram involved in the present invention; Figure 6 is a schematic diagram of a convolution module involved in the present invention; Figure 7 is a schematic diagram of an ECA attention module involved in the present invention; Figure 8 This is the overall network structure diagram of adding the ECA attention mechanism involved in the present invention; Fig. 9 It is the overall network structure diagram involved in the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0022] A food image recognition and calorie detection method based on deep learning in this embodiment, such as Figure 1 As shown, the method is implemented by the following steps: Step 1: Create a dataset of simulated food images and perform preprocessing; the preprocessing includes image occlusion processing, food target detection dataset preprocessing, and creating a dataset required for the occlusion detection network; Step 2: Determine a detection network and an image filling network for detecting food images with occlusion; Step 3: Build a judgment network to determine whether the food image to be detected is blocked: Determine whether the food image to be detected is occluded. If not, set the hyperparameter λ in the decision-level fusion method to 0.3 for decision fusion. If yes, select the ResNeXt network and the Inception-ResNet network to obtain image features, and then use the decision-level fusion method to perform decision fusion. The Inception-ResNet network improves network performance by introducing grouped convolutions and cardinality. Step 4: Build a detection network for detecting food images with occlusion: The image is input into the occlusion judgment network for preprocessing. If there is no occlusion in the image, the image filling network is skipped directly, and the image is input into the modified yolov8 network for processing to obtain the final output; if there is occlusion in the image, the occlusion is first filled through the image filling network, and then the filled image is input into the yolov8 network for processing to obtain the final result.

[0023] The step 1 of establishing a dataset of simulated food images includes: ETHZ Food-101 dataset, containing 101 categories of Western food and 101,000 images of dishes, with each category containing 1,000 images, including 750 training images and 250 test images; FoodX-251 dataset, a dataset containing 251 fine-grained classes, with 118,000 training, 12,000 validation, and 28,000 test images; UEC Food100 dataset, including 100 categories of Japanese dishes and 12,905 dish images; UEC Food256 dataset, including Japanese dishes, the dish categories are expanded from the original 100 to 256, and the corresponding number of dish images increases to 25,088; The occlusion effect of dishes in different scenes is simulated by simulating the occlusion of dish images, including: Shadow occlusion simulates occlusion by adding shadow effects to the dish image; this occlusion method can create a three-dimensional and spatial sense of the dish under a specific light source.

[0024] Random image block occlusion: randomly distribute multiple small image blocks or color blocks on the dish image to simulate the effect of the dish being partially covered or blocked. Rectangular or circular occlusion: add rectangular, circular or other shaped occlusion blocks to the dish image to simulate the occlusion effect of cutlery, fingers or other objects.

[0025] The steps of image occlusion processing are specifically: The images in the data set are cut to obtain the corresponding rectangular or semicircular image blocks, and the image blocks are used to occlude the data set. The processing methods are: (1) Cropping images to a fixed size: First, take images from the dataset and crop each image to a size of 448x448 pixels; ensure that all images have a uniform starting size before processing to facilitate subsequent uniform processing; (2) Image scaling: The 448x448 pixel image obtained by cropping in the previous step is further scaled (resized) to 224x224 pixels; (3) Randomly cutting image blocks: After the image is scaled to 224x224 pixels, a random cutting operation is performed on each image, that is, a portion of the area is randomly selected for cutting to generate an image block for occlusion; (4) Random occlusion processing: Finally, use the image generated in the previous step or directly from the scaled image to randomly select one or more positions from the four vertices (upper left corner, upper right corner, lower left corner, and lower right corner) for occlusion; In this way, a dataset containing occlusion effects is obtained, which can be used to train the model to improve the robustness of the model to occlusion.

[0026] The method for preprocessing a food target detection dataset refers to selecting a dataset in YOLO format, creating a dataset in YOLO format, and preprocessing it. The specific steps include: There are many dataset formats for target detection, mainly divided into VOC data format, COCO data format, and YOLO data format; VOC data format is a standard format for image annotation, which is used to store images and their related annotation information. In the VOC format, the annotation label information of each image will be saved in an XML file. The Pascal VOC dataset is one of the commonly used large-scale datasets for target detection. The format of the COCO dataset is based on JSON (JavaScript Object Notation), using a main JSON file to describe the entire dataset, and other auxiliary JSON files to store information such as images and annotations. Compared with VOC, the COCO dataset has the characteristics of more small targets, more targets in a single image, and most objects are distributed non-centrally, which is more in line with daily environments. YOLO format data generally contains image path, image width and height, object category, and object location information. Among them, the object location information is usually composed of the upper left corner coordinates (x, y) of the target box, the width and height (w, h) of the target box, collectively referred to as bounding box (BBox). The yolo dataset annotation format is mainly needed for the yolo project.

[0027] (1) Traverse all categories of the dataset: After creating the necessary files, loop through the folder numbers from 1 to 101. Each number corresponds to an image of a category. Create a dictionary for each category to store the width and height information of each image in that category. (2) Process images in each category: traverse all files in the current category folder; Check if the file is an image file; use PIL.Image to open the image, get its width and height, save them to the dictionary, and move the image to a new directory; (3) Processing bounding box information: Open and read the corresponding file content; Skip the first line of the file, and then process the remaining bounding box information line by line. For each line of bounding box information, parse out the category ID and the coordinates of the bounding box (x_min, y_min, x_max, y_max); find the width and height information of the image based on the file name, and calculate the coordinates of the center point and the aspect ratio of the bounding box; write this information to the file corresponding to the image, and the file name is based on the category ID but removes the file extension.

[0028] The process of creating the dataset required by the occlusion detection network is as follows: (1) Define a function for processing images: create a folder and traverse each category folder in the source folder to process the images in each category folder; if the number of images in a category is less than 100, select all images; otherwise, randomly select 100 images. Copy the selected images from the source folder to the target folder and record the image label and file name without extension in the output file.

[0029] (2) Processing unsegmented and segmented datasets: copy the unsegmented images to the images / 1 directory and mark them as 1; copy the segmented images to the images / 0 directory and mark them as 0; the file names and labels of all images are written to the label.txt file; (3) Split the data set: Read all the lines in the label.txt file, shuffle the order, and then split it into a training set and a test set in a ratio of 3:1; write the lines of the training set and the test set into the train.txt and test.txt files respectively; write the labels of two categories in the classes.txt file: 0 and 1.

[0030] The detection network for detecting food images with occlusion as described in step 2 is specifically: using DenseNet (densely connected convolutional network) to detect occlusion in the image; compared with ResNet, DenseNet proposes a more radical dense connection mechanism: that is, all layers are interconnected, specifically, each layer will accept all previous layers as its additional input. The core idea of ​​DenseNet is dense connection, that is, the input of each layer contains the output of all previous layers. This mechanism promotes feature reuse and information flow, allowing the network to extract features more effectively and alleviate the gradient vanishing problem. Among them, The DenseNet network uses a structure of DenseBlock combined with Transition, where DenseBlock is a module containing many layers, the feature map size of each layer is the same, and the layers are densely connected. The Transition module connects two adjacent DenseBlocks and reduces the feature map size through Pooling. As shown in the table, the network structure of DenseNet is mainly composed of DenseBlock and Transition; Bottleneck layer: Although each layer only produces k output feature maps, the following layers can get the input of all previous layers, so the number of input channels after splicing is still relatively large. Adding a 1x1 convolution before the 3x3 convolution in each dense block can reduce the number of input feature maps, reduce the amount of calculation, and fuse the features of each channel. The structure with a bottleneck layer, namely BN-ReLU-Conv(1x1)-BN-ReLU-Conv(3x3), is called DenseNet-B; DenseNet first uses 1*1 convolution to reduce the dimension of the input data channel to 4k, and then uses 3*3 convolution to extract features, thereby improving computational efficiency; Transition layer: The layer between every two dense blocks is called the Transition layer, which completes the convolution and pooling operations to reduce the number of feature maps. The transition layer consists of a BN layer, a 1x1 convolution layer, and a 2x2 average pooling layer: the 1×1 convolution is used to reduce the dimension and play the role of compressing the model, while the average pooling reduces the size of the feature map and halves the size of the feature maps. If a dense block has m feature maps, the Transition layer generates i*m output feature maps, where i represents the compression factor. The image filling network for determining the food image with occlusion in step 2 is specifically: Select Partial Convolutional Neural Network (PCNN) PCNN uses partial convolutional layers to process missing areas to maintain consistency and detail preservation of the filling results. The partial convolutional layer normalizes the convolution kernel according to the number of valid pixels in the filling area to avoid calculating the information of the occluded area; PConv is a partial convolution layer, Filter Size is the specified convolution kernel size, and Stride is the stride; PConv1-8 is an encoder, and PConv9-16 is a decoder; the Batch Normalization column indicates whether there is a BatchNormalization layer after PConv; the Nonlinearity column indicates whether a nonlinear layer is used and what nonlinear layer is used; Concat represents a skip connection, which fuses the upsampling result with the PConv result of the corresponding stage of the encoder.

[0031] In the process of constructing the judgment network for judging whether the food image to be detected is blocked as described in step 3, The group convolution adopted by the ResNeXt network is to divide the feature maps into different groups, and then convolve each group of feature maps separately, and each convolution kernel only processes part of the channels. This operation can effectively reduce the amount of calculation. The structure of the ResneXt-50 (32×4d) network is shown in the following table.

[0032] Table 3 Structure of ResneXt-50 (32×4d) network The difference between ResNeXt and ResNet modules is as follows Figure 2 As shown in Figure 2, in ResNet, the input feature with 256 channels is compressed 4 times to 64 channels by 1×1 convolution, and then a 3×3 convolution kernel is used to process the feature. The number of channels is expanded by 1×1 convolution and then connected to the original feature residual before output.

[0033] ResNeXt uses the same processing strategy, but in ResNeXt, the input features with 256 channels are divided into 32 groups, each group is compressed 64 times to 4 channels before processing. The 32 groups are added together and connected to the original feature residuals before output. Here, cardinatity refers to the number of identical branches in a block.

[0034] The original intention of the design of the Inception network is to solve the problems of surge in computation, overfitting and gradient disappearance faced by traditional convolutional neural networks when increasing depth and width. The Inception-ResNet network is a deep convolutional neural network (CNN) architecture that combines the Inception module with residual connections (Residual Connections) based on the Inception network. This network structure aims to combine the multi-scale feature extraction capability of the Inception module and the training stability advantage of the residual connection to further improve the performance of the convolutional neural network. The structure of the network is as follows Figure 3 shown.

[0035] The network structure is divided into three parts: Steam, Inception-resnet and Reduction; among them, Stem quickly reduces the resolution of the input image through multi-layer convolution operations; this design helps reduce the amount of computation in the subsequent Inception module, because lower-resolution feature maps require less computing resources to process. At the same time, reducing the resolution also helps extract high-level features in the image, which is particularly important for image recognition tasks.

[0036] The Reduction structure reduces the size of the feature map through a convolutional layer or a pooling layer, that is, reduces its width and height. Size reduction helps reduce the amount of subsequent calculations and enables the network to learn more abstract and global feature representations.

[0037] The Inception-resnet structure introduces transposed convolution; the residual connection structure of the Inception-resnet structure not only helps alleviate the gradient vanishing problem, but also significantly accelerates the network training process. By using convolution kernels of different sizes and pooling operations in parallel, the features of different scales in the image are captured at the same time, and then these features are fused through the Concat operation to form a richer and more comprehensive feature representation. This parallel processing and multi-scale feature fusion method helps to improve the model's ability to recognize complex image content. Due to the large depth of the network, in order to avoid problems such as gradient vanishing, such as Figure 4 As shown in Figure 1, a transposed convolution module is introduced into the Inception-resnet structure to improve the resolution of the feature map.

[0038] The modified yolov8 network described in step 4 refers to the yolov8 network with the ECA attention mechanism added, where (1)yolov8 network The food detection network selects yolov8 network, yolov8 network is as follows Figure 5The shown figure mainly consists of three parts.

[0039] The Backbone network is the basis of the model and is responsible for extracting features from the input image. These features are the basis for subsequent network layers to detect targets. The backbone network of the YOLOv8 network adopts a structure similar to CSPDarknet. The backbone part defines the basic architecture of the model, that is, the network structure used for feature extraction; the Neck network is located between the backbone network and the head network, and its function is to perform feature fusion and enhancement; the head network defines the detection head of the model, that is, the network structure used for final target detection; other modules, the ConvModule module contains convolutional layers, BN and activation functions (SiLU) for feature extraction; the DarknetBottleneck module increases the network depth through residual connections while maintaining efficiency; the CSP Layer module is a variant of the CSP structure, which improves the training efficiency of the model through partial connections; The yolov8 network with the ECA attention mechanism added is specifically: the images filled by the PCNN algorithm may produce various losses and noises. In order to improve the expressiveness of image features in the yolov8 network, the present invention introduces the ECA (Efficient Channel Attention) channel attention module in the backbone network Backbone of yolov8. By assigning higher weights to important channels, the ECA attention mechanism can enhance the feature expression of these channels, thereby improving the expressiveness of the overall feature map. At the same time, the weights of unimportant channels are reduced. The ECA attention mechanism helps to suppress noise and redundant information, allowing the network to focus more on features that are beneficial to the task. In order to better extract image features, the present invention adds the ECA module in front of the convolution module, such as Figure 6 shown.

[0040] The implementation process of the yolov8 network with the ECA attention mechanism added is as follows: (1) The input feature map is subjected to global average pooling, and the feature map is converted from a matrix of [h,w,c] to a vector of [1,1,c]; (2) Calculate the adaptive one-dimensional convolution kernel size kernel_size according to the number of channels of the feature map; (3) Use kernel_size in one-dimensional convolution to obtain the weight for each channel of the feature map; (4) Multiply the normalized weights and the original input feature map channel by channel to generate a weighted feature map, such as Figure 7 As shown; The overall network processing flow of the yolov8 network with the ECA attention mechanism added is: like Figure 8 As shown, the image is input into the occlusion judgment network for preprocessing. If there is no occlusion in the image, the image filling network is directly skipped, and the image is input into the modified yolov8 network for processing to obtain the final output; if there is occlusion in the image, the occlusion is first filled through the image filling network, and then the filled image is input into the yolov8 network for processing to obtain the final result.

[0041] The process of introducing transposed convolution in the Inception-resnet structure is as follows: Usually, after performing multiple convolution operations on an image, the size of the feature map will continue to shrink. For certain specific tasks, the image needs to be restored to its original size before operation. This size restoration operation of mapping an image from a small resolution to a large resolution is called upsampling. There are many upsampling methods, such as nearest neighbor interpolation, linear interpolation, bilinear interpolation, and bicubic interpolation. However, these upsampling methods are designed based on people's prior experience, and the effects are not ideal in many scenarios. Therefore, the present invention adopts transposed convolution. The upsampling method of transposed convolution is not a preset interpolation method, but like the standard convolution, it has learnable parameters, and the optimal upsampling method can be obtained through network learning; In the transposed convolution, let the input matrix be X, the input matrix be Y, and there is a new convolution kernel matrix C:

[0042] The transposed convolution is actually to perform the inverse operation of this process, that is, to obtain X through C and Y:

[0043] Weight matrix for transposed convolution It does not necessarily come from the original convolution matrix C, but its shape is the same as the transpose of the original convolution matrix C.

[0044] The process of decision-level fusion is: Feature-level fusion and decision-level fusion are two commonly used fusion methods. The former directly combines the feature vectors of the two branches into a joint feature vector to train a FER classifier; the latter combines the recognition results of the two branches. In the network of the present invention, the learned original features and occlusion features of the image are weakly complementary at the feature layer, so the decision-level fusion strategy is selected for research. The present invention selects the cross entropy loss function, and the calculation formula is as follows.

[0045]

[0046] Definition of the final loss function

[0047] where λ is a hyperparameter that balances the two parts.

[0048] like Fig. 9 As shown in the figure, the input image is preprocessed by the occlusion judgment network. If there is no occlusion in the image, the hyperparameter λ is set to 0.3. If there is occlusion in the image, the hyperparameter λ is set to 0.7. The image is then fused through the dish recognition network. The final results on the ETHZ Food-101 occlusion dataset are shown in the following table.

[0049] Table 4 Test results of ETHZ Food-101 occlusion dataset The embodiments of the present invention disclose preferred embodiments, but are not limited thereto. A person skilled in the art can easily understand the spirit of the present invention based on the above embodiments and make different extensions and changes. However, as long as they do not deviate from the spirit of the present invention, they are all within the protection scope of the present invention.

Claims

1. A food image recognition and calorie detection method based on deep learning, characterized in that: The method is implemented by the following steps: Step 1: Create a dataset of simulated food images and perform preprocessing; the preprocessing includes image occlusion processing, food target detection dataset preprocessing, and creating a dataset required for the occlusion detection network; Step 2: Determine a detection network and an image filling network for detecting food images with occlusion; Step 3: Build a judgment network to determine whether the food image to be detected is blocked: Determine whether the food image to be detected is occluded. If not, set the hyperparameter λ in the decision-level fusion method to 0.3 for decision fusion. If yes, select the ResNeXt network and the Inception-ResNet network to obtain image features, and then use the decision-level fusion method to perform decision fusion. The Inception-ResNet network improves network performance by introducing group convolution and group number methods; Step 4: Build a detection network for detecting food images with occlusion: The image is input into the occlusion judgment network for preprocessing. If there is no occlusion in the image, the image filling network is skipped directly, and the image is input into the modified yolov8 network for processing to obtain the final output; if there is occlusion in the image, the occlusion is first filled through the image filling network, and then the filled image is input into the yolov8 network for processing to obtain the final result.

2. The method for food image recognition and calorie detection based on deep learning according to claim 1, characterized in that: The step 1 of establishing a dataset of simulated food images includes: ETHZ Food-101 dataset, containing 101 categories of Western food and 101,000 images of dishes, with each category containing 1,000 images, including 750 training images and 250 test images; FoodX-251 dataset, a dataset containing 251 fine-grained classes, with 118,000 training, 12,000 validation, and 28,000 test images; UEC Food100 dataset, including 100 categories of Japanese dishes and 12,905 dish images; UEC Food256 dataset, including Japanese dishes, the dish categories are expanded from the original 100 to 256, and the corresponding number of dish images increases to 25,088; The occlusion effect of dishes in different scenes is simulated by simulating the occlusion of dish images, including: Shadow occlusion, which simulates occlusion by adding shadow effects on the dish images; Random image block occlusion: randomly distribute multiple small image blocks or color blocks on the dish image to simulate the effect of the dish being partially covered or blocked. Rectangular or circular occlusion: add rectangular, circular or other shaped occlusion blocks to the dish image to simulate the occlusion effect of cutlery, fingers or other objects.

3. A method for food image recognition and calorie detection based on deep learning according to claim 1 or 2, characterized in that: The steps of image occlusion processing are specifically: The images in the data set are cut to obtain the corresponding rectangular or semicircular image blocks, and the image blocks are used to occlude the data set. The processing methods are: (1) Cropping images to a fixed size: First, take images from the dataset and crop each image to a size of 448x448 pixels; (2) Image scaling: The 448x448 pixel image obtained by cropping in the previous step is further scaled to 224x224 pixels; (3) Randomly cutting image blocks: After the image is scaled to 224x224 pixels, a random cutting operation is performed on each image, that is, a portion of the area is randomly selected for cutting to generate an image block for occlusion; (4) Random occlusion processing: Use the image generated in the previous step or directly from the scaled image to randomly select one or more positions from the four vertices for occlusion; In this way, a dataset containing occlusion effects is obtained, which can be used to train the model to improve the robustness of the model to occlusion.

4. The method for food image recognition and calorie detection based on deep learning according to claim 3, characterized in that: The method for preprocessing a food target detection dataset refers to selecting a dataset in YOLO format, creating a dataset in YOLO format, and preprocessing it. The specific steps include: (1) Traverse all categories of the dataset: After creating the necessary files, loop through the folder numbers from 1 to 101. Each number corresponds to an image of a category. Create a dictionary for each category to store the width and height information of each image in that category. (2) Process images in each category: traverse all files in the current category folder; Check if the file is an image file; use PIL.Image to open the image, get its width and height, save them to the dictionary, and move the image to a new directory; (3) Processing bounding box information: Open and read the corresponding file content; Skip the first line of the file, and then process the remaining bounding box information line by line. For each line of bounding box information, parse out the category ID and the coordinates of the bounding box (x_min, y_min, x_max, y_max); find the width and height information of the image based on the file name, and calculate the coordinates of the center point and the aspect ratio of the bounding box; write this information to the file corresponding to the image, and the file name is based on the category ID but removes the file extension.

5. A method for food image recognition and calorie detection based on deep learning according to claim 1, 2 or 4, characterized in that: The process of creating the dataset required by the occlusion detection network is as follows: (1) Define the function for processing images: create a folder, traverse each category folder in the source folder, and process the images in each category folder; (2) Processing unsegmented and segmented datasets: Copy the unsegmented images to the images / 1 directory and mark them as 1; The segmented images are copied to the images / 0 directory and marked as 0; the file names and labels of all images are written to the label.txt file; (3) Split the data set: Read all the lines in the label.txt file, shuffle the order, and then split it into a training set and a test set in a ratio of 3:1; write the lines of the training set and the test set into the train.txt and test.txt files respectively; write the labels of two categories in the classes.txt file: 0 and 1.

6. The method for food image recognition and calorie detection based on deep learning according to claim 5, characterized in that: The detection network for detecting food images with occlusion in step 2 is specifically: using DenseNet to detect occlusion in the image; wherein, The DenseNet network uses a DenseBlock combined with a Transition structure. The Transition module connects two adjacent DenseBlocks and reduces the size of the feature map through Pooling. Bottleneck layer: DenseNet first uses 1*1 convolution to reduce the input data channel to 4k, and then uses 3*3 convolution to extract features; Transition layer: The layer between every two dense blocks is called the Transition layer, which completes the convolution and pooling operations; the transition layer consists of a BN layer, a 1x1 convolution layer, and a 2x2 average pooling layer: the 1×1 convolution layer is used to reduce the dimension, the average pooling layer reduces the size of the feature map, and the Transition layer generates the output feature map; The image filling network for determining the food image with occlusion in step 2 is specifically: The partial convolutional neural network (PCNN) is selected to process the missing area by using some convolutional layers, which normalize the convolution kernel according to the number of valid pixels in the filled area; PConv is a partial convolution layer, Filter Size is the specified convolution kernel size, and Stride is the stride; PConv1-8 is an encoder, and PConv9-16 is a decoder; the Batch Normalization column indicates whether there is a BatchNormalization layer after PConv; the Nonlinearity column indicates whether a nonlinear layer is used and what nonlinear layer is used; Concat represents a skip connection, which fuses the upsampling result with the PConv result of the corresponding stage of the encoder.

7. A method for food image recognition and calorie detection based on deep learning according to claim 1, 2, 4 or 6, characterized in that: In the process of constructing the judgment network for judging whether the food image to be detected is blocked as described in step 3, The grouped convolution adopted by the ResNeXt network is to divide the feature maps into different groups, and then convolve each group of feature maps separately, and each convolution kernel only processes part of the channels; The network structure of the Inception network is divided into three parts: Steam, Inception-resnet and Reduction; in, Stem reduces the resolution of the input image through multi-layer convolution operations; Reduction structure reduces the size of the feature map through convolution layers or pooling layers; The Inception-resnet structure introduces transposed convolution; by using convolution kernels of different sizes and pooling operations in parallel, it captures features of different scales in the image at the same time, and then fuses these features through the Concat operation.

8. The method for food image recognition and calorie detection based on deep learning according to claim 7, characterized in that: The modified yolov8 network described in step 4 refers to the yolov8 network with the ECA attention mechanism added, where The backbone network of the YOLOv8 network adopts a structure similar to that of CSPDarknet. The backbone part defines the basic architecture of the model, that is, the network structure for feature extraction; the neck network is located between the backbone network and the head network for feature fusion and enhancement; the head network defines the detection head of the model, which is the network structure for final target detection; the ConvModule module contains convolutional layers, BN and activation functions for feature extraction; the DarknetBottleneck module increases the network depth through residual connections while maintaining efficiency; the CSP Layer module is a variant of the CSP structure, which improves the training efficiency of the model through partial connections; The specific steps of adding ECA attention mechanism to the yolov8 network are as follows: Introduce ECA channel attention module in the backbone network Backbone of yolov8. By assigning higher weights to important channels, the ECA attention mechanism can enhance the feature expression of these channels, while reducing the weights of unimportant channels; The implementation process of the yolov8 network with the ECA attention mechanism added is as follows: (1) The input feature map is subjected to global average pooling, and the feature map is converted from a matrix of [h,w,c] to a vector of [1,1,c]; (2) Calculate the adaptive one-dimensional convolution kernel size kernel_size according to the number of channels of the feature map; (3) Use kernel_size in one-dimensional convolution to obtain the weight for each channel of the feature map; (4) Multiply the normalized weights and the original input feature map channel by channel to generate a weighted feature map; The overall network processing flow of the yolov8 network with the ECA attention mechanism added is: The image is input into the occlusion judgment network for preprocessing. If there is no occlusion in the image, the image filling network is skipped directly, and the image is input into the modified yolov8 network for processing to obtain the final output; if there is occlusion in the image, the occlusion is first filled through the image filling network, and then the filled image is input into the yolov8 network for processing to obtain the final result.

9. The method for food image recognition and calorie detection based on deep learning according to claim 7, characterized in that: The process of introducing transposed convolution in the Inception-resnet structure is as follows: Use transposed convolution to obtain the optimal upsampling method through network learning; In the transposed convolution, let the input matrix be X, the input matrix be Y, and there is a new convolution kernel matrix C: ; The transposed convolution is actually to perform the inverse operation of this process, that is, to obtain X through C and Y: ; Weight matrix for transposed convolution It does not necessarily come from the original convolution matrix C, but its shape is the same as the transpose of the original convolution matrix C.

10. The method for food image recognition and calorie detection based on deep learning according to claim 9, characterized in that: The decision-level fusion process is specifically: Based on the complementarity of the learned original features of the image and the occlusion features at the feature layer, the decision-level fusion strategy is selected for research, and the cross entropy loss function is selected. The calculation formula is as follows: ; The final loss function is defined as: ; Where λ is a hyperparameter that balances the two parts; The input image is preprocessed by the occlusion judgment network. If there is no occlusion in the image, the hyperparameter λ is set to 0.

3. If there is occlusion in the image, the hyperparameter λ is set to 0.

7. The image is then fused through the dish recognition network to obtain the final result on the ETHZ Food-101 occlusion dataset.