Lightweight food image recognition method based on inter-group feature map calculation

By building an IRB module and a LIT module, combining feature fusion and efficient channel attention mechanism, the problems of difficulty in deployment, weak global optimization capabilities and overload of parameters in food image recognition are solved, and efficient performance and precise control of lightweight food image recognition methods are achieved.

CN120107957AInactive Publication Date: 2025-06-06LUDONG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510239051.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

At this stage, the food image recognition method has problems such as deployment difficulties, weak global optimization capabilities and overloading of parameters, making it difficult to maintain high performance in lightweight models.

Method used

The lightweight food image recognition method based on intergroup feature map calculation is adopted, and local feature extraction and global feature extraction are realized by constructing an inverse residual module (IRB module) and a lightweight intergroup transformer module (LIT module), and local feature extraction and global feature extraction are realized, and the "marginal effect" caused by feature grouping is alleviated through feature fusion and efficient channel attention mechanism.

Benefits of technology

Under the premise of lightweighting, the efficient performance of food image recognition is achieved, the parameter volume is precisely controlled from 0.7M to 10.4M, and the accuracy rate is 88.24% on data sets such as ETHZFood-101. Grad-CAM visualization confirms the expansion of the model's receptive field, which improves the inference speed of mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107957A_ABST
    Figure CN120107957A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight food image recognition method based on inter-group feature map calculation, and belongs to the technical field of computer vision. According to the method, local feature extraction is realized through an IRB module, and global feature extraction is completed in combination with a grouped transformer structure of an FGT module in an LIT module. Wherein the FGT module adopts Unfold linear transformation to convert a feature pattern into a Patch sequence, realizes global relationship capture through a spatial self-attention mechanism and a feedforward network, and reconstructs a feature dimension by utilizing a Fold operation, so that the calculation complexity is reduced. Channel shuffling, inter-group shuffling and irregular grouping strategies are innovatively introduced, and a high-efficiency channel attention mechanism (ECA) is matched to carry out channel re-calibration on splicing features, so that the multi-scale feature fusion capability is effectively improved. The method solves the problems that a traditional food image recognition model is large in parameter quantity, too high in calculation complexity and incapable of balancing precision and speed, and is particularly suitable for mobile device food recognition scenes with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a lightweight food image recognition method based on inter-group feature map calculation, and belongs to the technical field of computer vision. Background Art

[0002] With the development of human society, people's demand for healthy diet is getting higher and higher. More and more researchers have begun to focus on identifying food images and calculating energy based on the recognition results, so as to effectively evaluate people's daily nutritional intake and ultimately realize smart diet and nutrition management. Food image recognition is the basic and core task in the whole process. The key lies in extracting discriminative visual features. Deep learning has gradually become a universal solution for food image recognition due to its powerful ability in discriminative feature learning. Food images have unique visual features, with small differences between categories and large differences within categories. In addition, due to the diversity of cooking methods and ingredients, the local features of the same food category are different, and the differences between different food categories are extremely small. How to effectively calculate its global features and fuse local features with global features is a characteristic of food image recognition tasks; Convolutional Neural Network (CNN) has advantages in extracting local features, but extracting long-distance correlation features requires stacking deep networks, which will greatly increase the number of parameters and computation, which is not conducive to lightweighting. In addition, its global modeling ability is weak, and the existing food image recognition methods based on lightweight CNNs have difficulty in achieving optimal performance. Although Vision Transformer (ViT) is good at capturing long-distance correlation information, it is difficult to lightweight due to vector dot product calculations, more training data, and the number of iterations required. The latest models based on Vision Transformer often have a large number of parameters, so they often face great difficulties in deployment. Therefore, the current methods for food recognition mainly have problems such as deployment difficulties, weak global optimization capabilities, and parameter overload. Therefore, how to use grouping operations to reduce the model size while maintaining high performance of the model is a problem worth exploring. Summary of the invention

[0003] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a lightweight food image recognition method based on inter-group feature map calculation; The technical solution provided by the present invention is as follows: a lightweight food image recognition method based on inter-group feature map calculation, characterized in that it comprises the following steps: Step 1: Build a lightweight food image recognition model, including: Construct an inverted residual module, namely the IRB module, for local feature extraction and feature map downsampling operations; Construct a lightweight inter-group transformer module, namely the LIT module, for global feature extraction and fusion of inter-channel information; Step 2: Train the model using open source datasets; Step 3: Deploy the trained model to the client device.

[0004] Furthermore, the IRB module is an inverted residual structure in the lightweight model MobileNetV2, which includes several convolution layers with convolution kernels of 3*3 and 1*1 sizes; The LIT module includes several convolutional layers, channel shuffling operations, linear transformation of feature maps, feature map grouping transformer modules, feature fusion layers and efficient channel attention mechanisms.

[0005] Furthermore, the lightweight food image recognition model architecture is layered: The first layer is composed of a convolution module and an IRB module. The convolution module has a kernel size of 3*3, a step size of 2, and an output channel number of 32α. The IRB module has a step size of 1 and an output channel number of 64α. The second layer is composed of two IRB modules; the step size of the first IRB module is 2, and the number of output channels of both IRB modules is 128α; The third layer is composed of a mixture of an IRB module and two LIT modules; the step size of the IRB module is 2, the number of output channels of the three modules is 256α, and the number of groups of the two LIT modules is 4 and 3 respectively; The fourth layer is composed of a mixture of one IRB module and four LIT modules; the step size of the IRB module is 2, the number of output channels of the five modules is 384α, and the number of groups of the four LIT modules is 4, 3, 4, and 3 respectively; The fifth layer is composed of a mixture of an IRB module and three LIT modules; the step size of the IRB module is 2, the number of output channels of the four modules is 512α, and the number of groups of the three LIT modules is 4, 2, and 1 respectively; The last layer is the global pooling fully connected layer, which consists of a global average pooling operation and a fully connected layer. The output dimension is a one-dimensional vector of cls, which represents the probability distribution of the category. in, It is the model width hyperparameter setting. By setting this parameter, the model scale can be adjusted; cls represents the number of data set categories.

[0006] Furthermore, the calculation formula of the IRB module is as follows: ; in, Represents the value of the output feature map at position (i, j) and channel k, Represents the value of the input feature map at position (i+u, j+v) and channel m, represents the weight of the depth convolution kernel at position (u,v) and channel m, Represents the value of the input feature map at position (i, j) and channel n, Represents the weight of the point-by-point convolution kernel in channel n to k.

[0007] Further, the LIT module calculation process is as follows: The input feature map is , where C is the number of input feature map channels; H and W are the length and width of the feature map respectively; First, the input feature map is sent to the convolution layer for convolution calculation, that is: ; in Represents the feature map after convolution calculation; Conv represents the convolution operation, where the convolution kernel size is 3, the number of convolution kernels is C, and the step size is 1; then the feature map after convolution calculation is; Unfold; calculated; after the Unfold calculation, the input feature map is converted into Patch for feature grouping transformer calculation; in the Unfold calculation, for each position (i, j) in the output feature map Y on channel k, the following calculation is performed: ; in, Represents the Patch after Unfold calculation. and Indicates the length and width of the Patch obtained after Unfold calculation; Indicates that the input feature map is at a specific position , the feature value on channel c; c represents the channel index of the feature map, and its value range is from 0 to C, which is used to traverse all channels of the input feature map; x represents the position index in the height direction of the patch, and its value range is from 0 to , used to traverse each position in the height direction of the Patch; y represents the position index in the width direction of the Patch, and the value range is from 0 to , used to traverse each position in the width direction of the Patch; Represents the convolution kernel weight at position (i, j) on the cth channel in the Unfold calculation; then Resize it to a size of Patch, that is, , where d represents the number of channels of the patch, and P represents the ratio of the patch size to the feature map size before calculation. The calculation formula is as follows: ; N represents the number of elements in each Patch, and its calculation formula is as follows: ; Then, the patch after Unfold calculation is sent to the feature graph grouping transformer module (FGT module) for global feature calculation. In the FGT module, the patch is divided into g groups of patches through feature grouping operation, where g is the number of groups and the dimension of each group of patches is ; Therefore, the calculation formula of the FGT module is as follows: ; in, It represents the Patch obtained after calculation by the FGT module. Indicates the first patch in group g. N eigenvalues ​​of a patch; The function represents the spatial self-attention mechanism and feed-forward neural network calculation; The function represents the feature fusion operation, that is, combining each group of patches to restore the feature map dimension to ; The function represents the inter-group shuffle operation, that is, shuffle each group of patches; Then, the Patch calculated by the FGT module is subjected to the Fold operation. The Fold operation is the inverse transformation of the Unfold operation. The calculation process is as follows: First, the Patch calculated by the FGT module is reshaped into shape, followed by a folding operation, i.e. ; In the above formula, Represents the feature map obtained after Fold calculation, Represents the convolution kernel weight of the pth convolution kernel at position (i, j) in the Fold calculation, Represents the reshaped Patch; after this calculation, the Patch is transformed back to a size of The feature map of The feature map obtained in the above steps has been subjected to global feature calculation, and the inter-channel information of the calculated feature map is fused through the following formula, namely: ; In the above formula, Represents the output feature map of the LIT module, Represents the feature map obtained after Fold calculation, Conv represents the convolution operation with a convolution kernel of 1 and a number of C. It represents the shuffling operation between channels of the feature map after the convolution operation. Concat represents the feature map concatenation operation to concatenate the input of the LIT module with the shuffled feature map along the channel dimension. ECA represents efficient channel attention calculation.

[0008] Furthermore, the global pooling fully connected layer calculation steps are as follows: The input feature map is , perform global average pooling calculation; global average pooling will average all pixel values ​​of each channel to obtain a one-dimensional feature vector with a dimension of C , the specific calculation formula is: ; Where C is the number of input feature map channels; H and W are the length and width of the feature map respectively; The feature vector obtained in the above steps is converted into recognition results through the fully connected layer and the Softmax function. The specific calculation formula is: ; in It is a one-dimensional vector with dimension cls, representing the probability distribution of the category; cls represents the number of categories in the data set; Represents the weight matrix of the fully connected layer; Represents the bias vector of the fully connected layer; output The highest probability category can be used to obtain the recognition result of the food image.

[0009] The beneficial effects of the present invention are as follows: the recognition effect can be achieved under the premise of lightweight due to the effective fusion of local and global features of food images. The local feature extraction is realized by the deep convolution and point-by-point convolution structure of the lightweight IRB module, and the feature map downsampling is completed by step size control; the global feature extraction is calculated by the FGT module in the LIT module through the grouping transformer structure, in which the Unfold linear transformation is used to convert the feature map into a Patch sequence, and the spatial self-attention mechanism and the feedforward neural network are combined to realize global relationship modeling, and the feature map dimension is restored by the Fold operation, which significantly reduces the computational complexity. The LIT module innovatively integrates channel shuffling, inter-group shuffling and irregular grouping strategies, effectively alleviates the "marginal effect" caused by feature grouping, and enhances the model's ability to capture multi-scale features. At the same time, by introducing the efficient channel attention mechanism (ECA), the channel recalibration of the spliced ​​features is carried out to strengthen the key feature response. The model uses an elastic adjustment mechanism for the width hyperparameter α to achieve precise control of the number of parameters from 0.7M to 10.4M, and achieves an accuracy of 88.24% on datasets such as ETHZFood-101. Ablation experiments show that the ECA mechanism improves the accuracy of the UEC Food-256 dataset by 0.35%, while the irregular grouping strategy improves the accuracy by 0.24% compared with the conventional grouping, verifying the effectiveness of the module design. Finally, it is confirmed through Grad-CAM visualization that the design expands the receptive field of the model, effectively improves the inference speed when deployed on mobile devices, and achieves a balance between accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 It is a schematic diagram of heat map comparison between an embodiment of the present invention and a general lightweight recognition network. DETAILED DESCRIPTION

[0011] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further explained in detail below in conjunction with embodiments; it should be noted that the specific embodiments described herein are only intended to explain the present invention, but not to limit and define the present invention; like Figure 1 As shown, a lightweight food image recognition method based on inter-group feature map calculation includes the following steps: Step 1: Build a lightweight food image recognition model on the computing platform, including: Construct an inverted residual block (IRB module) for local feature extraction and feature map downsampling operations; Construct a Lightweight Inter-Group Transformer (LIT) module for global feature extraction and fusion of inter-channel information; The IRB module is an inverted residual structure in the lightweight model MobileNetV2, which includes several convolution layers with convolution kernels of 3*3 and 1*1 sizes.

[0012] The LIT module includes several convolutional layers, channel shuffling operations, linear transformations of feature maps (such as expansion and folding operations), a feature grouping transformer module (Feature Grouping Transformer, FGT module), a feature fusion layer and an efficient channel attention mechanism (Efficient Channel Attention, ECA).

[0013] Layering the model architecture: The first layer is composed of a convolution module and an IRB module; the convolution module has a kernel size of 3*3, a step size of 2, and an output channel number of 32α; the IRB module has a step size of 1 and an output channel number of 64α.

[0014] The second layer is composed of two IRB modules; the step size of the first IRB module is 2, and the number of output channels of both IRB modules is 128α.

[0015] The third layer is composed of a mixture of one IRB module and two LIT modules; the step size of the IRB module is 2, the number of output channels of the three modules is 256α, and the number of groups of the two LIT modules is 4 and 3 respectively.

[0016] The fourth layer is composed of a mixture of one IRB module and four LIT modules; the step size of the IRB module is 2, the number of output channels of the five modules is 384α, and the number of groups of the four LIT modules is 4, 3, 4, and 3 respectively.

[0017] The fifth layer is composed of a mixture of one IRB module and three LIT modules; the step size of the IRB module is 2, the number of output channels of the four modules is 512α, and the number of groups of the three LIT modules is 4, 2, and 1 respectively.

[0018] The last layer is the global pooling fully connected layer, which consists of a global average pooling operation and a fully connected layer. The output dimension is a one-dimensional vector of cls, representing the probability distribution of the category.

[0019] in, It is the model width hyperparameter setting. By setting this parameter, the model scale can be adjusted; cls represents the number of data set categories.

[0020] The calculation formula of the IRB module in the model architecture is as follows: ; in, Represents the value of the output feature map at position (i, j) and channel k, Represents the value of the input feature map at position (i+u, j+v) and channel m, represents the weight of the depth convolution kernel at position (u,v) and channel m, Represents the value of the input feature map at position (i, j) and channel n, Represents the weight of the point-by-point convolution kernel in channel n to k.

[0021] The IRB module uses depthwise convolution and point-by-point convolution to extract local features. When the step size of IRB is 2, the size of the output feature map will be halved, achieving the purpose of feature map downsampling.

[0022] The calculation process of the LIT module in the model architecture is as follows: The input feature map is , where C is the number of input feature map channels; H and W are the length and width of the feature map respectively.

[0023] First, the input feature map is sent to the convolution layer for convolution calculation, that is: ; in Represents the feature map after convolution calculation; Conv represents the convolution operation, where the convolution kernel size is 3, the number of convolution kernels is C, and the step size is 1; then the feature map after convolution calculation is unfolded (Unfold) calculation; after the Unfold calculation, the input feature map is converted into Patch for feature grouping transformer calculation; in the Unfold calculation, for each position (i, j) in the output feature map Y on channel k, the following calculation is performed: ; in, Represents the Patch after Unfold calculation. and Indicates the length and width of the Patch obtained after Unfold calculation; in this embodiment, both are set to 2. Indicates that the input feature map is at a specific location ( , the feature value on channel c; c represents the channel index of the feature map, and its value range is from 0 to C, which is used to traverse all channels of the input feature map; x represents the position index in the height direction of the patch, and its value range is from 0 to , used to traverse each position in the height direction of the Patch; y represents the position index in the width direction of the Patch, and the value range is from 0 to , used to traverse each position in the width direction of the Patch; Represents the convolution kernel weight at position (i, j) on the cth channel in the Unfold calculation; then Resize it to a size of Patch, that is, , where d represents the number of channels of the patch, and P represents the ratio of the patch size to the feature map size before calculation. The calculation formula is as follows: ; N represents the number of elements in each Patch, and its calculation formula is as follows: ; Then, the patch after Unfold calculation is sent to the feature graph grouping transformer module (FGT module) for global feature calculation. In the FGT module, the patch is divided into g groups of patches through feature grouping operation, where g is the number of groups and the dimension of each group of patches is ; Therefore, the calculation formula of the FGT module is as follows: ; in, It represents the Patch obtained after calculation by the FGT module. Indicates the first patch in group g. N eigenvalues ​​of a patch; The function represents the spatial self-attention mechanism and feedforward neural network calculation; the Combine function represents the feature fusion operation, that is, combining each group of patches to restore the feature map dimension to ; The Shuffle function represents the inter-group shuffle operation, that is, each group of patches is shuffled.

[0024] Then, the Patch calculated by the FGT module is folded. The Fold operation is the inverse transformation of the Unfold operation. The calculation process is as follows: First, the Patch calculated by the FGT module is reshaped into shape, followed by a folding operation, i.e. ; In the above formula, Represents the feature map obtained after Fold calculation, Represents the convolution kernel weight of the pth convolution kernel at position (i, j) in the Fold calculation, Represents the reshaped Patch; after this calculation, the Patch is transformed back to a size of feature map.

[0025] The feature map obtained in the above steps has been subjected to global feature calculation, and the inter-channel information of the calculated feature map is fused through the following formula, namely: ; In the above formula, Represents the output feature map of the LIT module, Represents the feature map obtained after Fold calculation, Conv represents the convolution operation with a convolution kernel of 1 and a number of C. It represents the shuffling operation between channels of the feature map after the convolution operation. Concat represents the feature map concatenation operation to concatenate the input of the LIT module with the shuffled feature map along the channel dimension. ECA represents efficient channel attention calculation.

[0026] Specifically, the LIT module extracts global features through the FGT module and uses linear transformation operations, such as Unfold and Fold, to reduce the computational complexity of the model. In addition, the performance of the model is enhanced by introducing ECA, inter-group shuffling operations, inter-channel shuffling operations, and feature fusion operations, thereby achieving efficient global feature extraction.

[0027] The calculation steps of the global pooling fully connected layer in the above model architecture are as follows: The input feature map is , perform global average pooling calculation; global average pooling will average all pixel values ​​of each channel to obtain a one-dimensional feature vector with a dimension of C , the specific calculation formula is: ; Where C is the number of input feature map channels; H and W are the length and width of the feature map respectively.

[0028] The feature vector obtained in the above steps is converted into recognition results through the fully connected layer and the Softmax function. The specific calculation formula is: ; in It is a one-dimensional vector with dimension cls, representing the probability distribution of the category; cls represents the number of categories in the data set; Represents the weight matrix of the fully connected layer; Represents the bias vector of the fully connected layer; output The highest probability category can be used to obtain the recognition result of the food image.

[0029] Step 2: Train the model in an open source dataset. Based on the model built in step 1, you can train the model on the server platform using open source food datasets such as ETHZ Food-101, Vireo Food-172, and UEC Food-256.

[0030] The ETHZ Food-101 dataset contains 101 categories of Western food, using 75,750 images for training and 25,250 for validation.

[0031] The Vireo Food-172 dataset contains 172 categories of Chinese food, using 66,071 images for training and 44,170 images for validation.

[0032] The UEC Food-256 dataset contains 256 Asian food categories, with 22,095 images for training and 9,300 images for validation.

[0033] In order to verify the effectiveness of the LIT module of this embodiment in the food image recognition model, as well as the effective role of ECA and irregular grouping on the "marginal effect", we conducted an ablation experiment on the above dataset;

[0034]

[0035] In order to improve the convergence efficiency of the model and prevent it from falling into the local optimal situation during the calculation process, during the training process, this method adopts the cosine annealing strategy to adjust the learning rate, so that the learning rate changes according to the shape of the cosine function during the training process, so that the learning rate can change periodically within a larger range, avoiding the model from falling into the local optimal solution, and finely adjusting the parameters in the later stage. The calculation formula is as follows: ); in is the learning rate of the tth training cycle, and are the minimum and maximum values ​​of the learning rate after restart, and T is the total length of the learning rate annealing cycle.

[0036] In the training process of this method, since the gradient descent algorithm is used to update the model parameters, these model parameters may become unstable due to factors such as gradient noise and fluctuations in training data; therefore, in this method, the exponential moving average (EMA) is used to smooth the model parameters, which can reduce this instability to a certain extent.

[0037] Step 3: Deploy the trained model to the end-side device; determine the task deployment end and select a reasonable model width for deployment; since the model width is adjustable, the model sizes are: 0.7M, 1.5M, 2.6M, 4.1M, 5.9M, 8.0M, and 10.4M. As the number of model parameters increases, the performance of the model will continue to improve. Therefore, when facing resource-constrained scenarios, it is necessary to reasonably determine the model width so that the recognition method can achieve better performance in limited scenarios.

[0038] In order to verify the competitiveness of this method in the food image recognition task, this example uses the Grad-CAM method to generate a heat map, such as Figure 2 As shown, the brighter area represents the receptive field of the model; the MobileViTv2 model is a lightweight network proposed by Apple in 2022, which is highly competitive in a variety of mobile vision tasks; the first row of pictures comes from the ETHZ Food-101, ETHZ Food-172, and UEC Food-256 datasets, and the second and third rows correspond to the heat maps of the MobileViTv2 model and this method, respectively; by comparison, it can be seen that the receptive field of this method is wider; the attention of the MobileViTv2 model is relatively limited. The reason is that we use more Shuffle operations in the model, which further expands the receptive field of the model and can pay attention to more food image feature information; in addition, it can be seen that our model has obtained better food attention while being lighter.

[0039] It should be understood that the parts not elaborated in detail in this specification belong to the prior art; the above embodiments only describe the preferred implementation mode of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solution of the present invention by ordinary engineering and technical personnel in this field should fall within the protection scope determined by the claims of the present invention.

Claims

1. A lightweight food image recognition method based on inter-group feature map calculation, characterized in that: It includes the following steps: Step 1: Build a lightweight food image recognition model, including: Construct an inverted residual module, namely the IRB module, for local feature extraction and feature map downsampling operations; Construct a lightweight inter-group transformer module, namely the LIT module, for global feature extraction and fusion of inter-channel information; Step 2: Train the model using open source datasets; Step 3: Deploy the trained model to the client device.

2. According to claim 1, a lightweight food image recognition method based on inter-group feature map calculation is characterized in that: The IRB module is an inverted residual structure in the lightweight model MobileNetV2, which includes several convolution layers with convolution kernels of 3*3 and 1*1 sizes; The LIT module includes several convolutional layers, channel shuffling operations, linear transformation of feature maps, feature map grouping transformer modules, feature fusion layers and efficient channel attention mechanisms.

3. The lightweight food image recognition method based on inter-group feature map calculation according to claim 1 is characterized in that: The lightweight food image recognition model architecture is divided into layers: The first layer is composed of a convolution module and an IRB module. The convolution module has a kernel size of 3*3, a step size of 2, and an output channel number of 32α. The IRB module has a step size of 1 and an output channel number of 64α. The second layer is composed of two IRB modules; the step size of the first IRB module is 2, and the number of output channels of both IRB modules is 128α; The third layer is composed of a mixture of an IRB module and two LIT modules; the step size of the IRB module is 2, the number of output channels of the three modules is 256α, and the number of groups of the two LIT modules is 4 and 3 respectively; The fourth layer is composed of a mixture of one IRB module and four LIT modules; the step size of the IRB module is 2, the number of output channels of the five modules is 384α, and the number of groups of the four LIT modules is 4, 3, 4, and 3 respectively; The fifth layer is composed of a mixture of an IRB module and three LIT modules; the step size of the IRB module is 2, the number of output channels of the four modules is 512α, and the number of groups of the three LIT modules is 4, 2, and 1 respectively; The last layer is the global pooling fully connected layer, which consists of a global average pooling operation and a fully connected layer. The output dimension is a one-dimensional vector of cls, which represents the probability distribution of the category. in, It is the model width hyperparameter setting. By setting this parameter, the model scale can be adjusted; cls represents the number of data set categories.

4. The lightweight food image recognition method based on inter-group feature map calculation according to claim 1 is characterized in that: The IRB module calculation formula is as follows: ; in, Represents the value of the output feature map at position (i, j) and channel k, Represents the value of the input feature map at position (i+u, j+v) and channel m, represents the weight of the depth convolution kernel at position (u,v) and channel m, Represents the value of the input feature map at position (i, j) and channel n, Represents the weight of the point-by-point convolution kernel in channel n to k.

5. The lightweight food image recognition method based on inter-group feature map calculation according to claim 1 is characterized in that: The LIT module calculation process is as follows: The input feature map is , where C is the number of input feature map channels; H and W are the length and width of the feature map respectively; First, the input feature map is sent to the convolution layer for convolution calculation, that is: ; in Represents the feature map calculated by convolution; Conv represents the convolution operation, where the convolution kernel size is 3, the number of convolution kernels is C, and the step size is 1; then the feature map after convolution calculation is Unfolded; Calculation; After the Unfold calculation, the input feature map is converted into a patch for feature grouping transformer calculation; in the Unfold calculation, for each position (i, j) in the output feature map Y on channel k, the following calculation is performed: ; in, Represents the Patch after Unfold calculation. and Indicates the length and width of the Patch obtained after Unfold calculation; Indicates that the input feature map is at a specific position , the feature value on channel c; c represents the channel index of the feature map, and its value range is from 0 to C, which is used to traverse all channels of the input feature map; x represents the position index in the height direction of the patch, and its value range is from 0 to , used to traverse each position in the height direction of the Patch; y represents the position index in the width direction of the Patch, and the value range is from 0 to , used to traverse each position in the width direction of the Patch; Represents the convolution kernel weight at position (i, j) on the cth channel in the Unfold calculation; then Resize it to a size of Patch, that is, , where d represents the number of channels of the patch, and P represents the ratio of the patch size to the feature map size before calculation. The calculation formula is as follows: ; N represents the number of elements in each Patch, and its calculation formula is as follows: ; Then, the patch after Unfold calculation is sent to the feature graph grouping transformer module (FGT module) for global feature calculation. In the FGT module, the patch is divided into g groups of patches through feature grouping operation, where g is the number of groups and the dimension of each group of patches is ; Therefore, the calculation formula of the FGT module is as follows: ; in, It represents the Patch obtained after calculation by the FGT module. Indicates the first patch in group g. N eigenvalues ​​of a patch; The function represents the spatial self-attention mechanism and feed-forward neural network calculation; The function represents the feature fusion operation, that is, combining each group of patches to restore the feature map dimension to ; The function represents the inter-group shuffle operation, that is, shuffle each group of patches; Then, the Patch calculated by the FGT module is subjected to the Fold operation. The Fold operation is the inverse transformation of the Unfold operation. The calculation process is as follows: First, the Patch calculated by the FGT module is reshaped into shape, followed by a folding operation, i.e. ; In the above formula, Represents the feature map obtained after Fold calculation, Represents the convolution kernel weight of the pth convolution kernel at position (i, j) in the Fold calculation, Represents the reshaped Patch; after this calculation, the Patch is transformed back to a size of The feature map of The feature map obtained in the above steps has been subjected to global feature calculation, and the inter-channel information of the calculated feature map is fused through the following formula, namely: ; In the above formula, Represents the output feature map of the LIT module, Represents the feature map obtained after Fold calculation, Conv represents the convolution operation with a convolution kernel of 1 and a number of C. It represents the shuffling operation between channels of the feature map after the convolution operation. Concat represents the feature map concatenation operation to concatenate the input of the LIT module with the shuffled feature map along the channel dimension. ECA represents efficient channel attention calculation.

6. The lightweight food image recognition method based on inter-group feature map calculation according to claim 1 is characterized in that: The calculation steps of the global pooling fully connected layer are as follows: The input feature map is , perform global average pooling calculation; global average pooling averages all pixel values ​​of each channel to obtain a one-dimensional feature vector with a dimension of C , the specific calculation formula is: ; Where C is the number of input feature map channels; H and W are the length and width of the feature map respectively; The feature vector obtained in the above steps is converted into recognition results through the fully connected layer and the Softmax function. The specific calculation formula is: ; in It is a one-dimensional vector with dimension cls, representing the probability distribution of the category; cls represents the number of categories in the data set; Represents the weight matrix of the fully connected layer; Represents the bias vector of the fully connected layer; output The highest probability category can be used to obtain the recognition result of the food image.