A food recognition method based on multi-stage feature fusion
Patent Information
- Application Number
- CN202211385292.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-11-07
AI Technical Summary
[0004]针对现有技术的上述缺陷,本发明提供一种基于多阶段特征融合的食物识别方法,旨在解决现有技术中从图像中识别目标的精度不高的问题
[0030]与现有技术相比,本发明提供了一种基于多阶段特征融合的食物识别方法,通过神经网络提取图像特征的同时,保留特征提取过程中的低层次特征并进行处理后与高层次特征进行融合,实现了利用图像特征提取过程中多阶段的融合特征,能够有效提升图像识别的精度。
Smart Images

Figure CN117373015B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a food recognition method based on multi-stage feature fusion. Background Technology
[0002] Currently, there are methods that use neural networks to extract image features and then identify targets in images. However, existing image feature extraction mainly involves stacking convolutional layers to extract features, and each layer transforms the previous layer, making the extracted features progressively more abstract from low to high levels. This makes it difficult to obtain more comprehensive and globally relevant features, resulting in a deficiency of local perception. Some low-level features are missing in high-level features, leading to low recognition accuracy.
[0003] Therefore, existing technologies still need to be improved and enhanced. Summary of the Invention
[0004] To address the aforementioned shortcomings of existing technologies, this invention provides a food recognition method based on multi-stage feature fusion, aiming to solve the problem of low accuracy in identifying targets from images in existing technologies.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0006] A first aspect of the present invention provides a food recognition method based on multi-stage feature fusion, the method comprising:
[0007] The image to be processed is obtained and input into the feature extraction layer to obtain at least one intermediate feature map and a result feature map;
[0008] The intermediate feature maps and the result feature maps are structurally transformed to obtain the feature matrices of each node. The adjacency matrices corresponding to the feature matrices of each node are constructed. The adjacency matrices are then input into the graph neural network to obtain the features of each stage.
[0009] The features of each stage and the result feature map are fused to obtain the target image features, and food recognition is performed based on the target image features.
[0010] The food recognition method based on multi-stage feature fusion is described above, wherein the feature extraction layer includes multiple extraction layers, the output of each extraction layer is the intermediate feature map, the input of each extraction layer is the output of the previous extraction layer, and the output of the last extraction layer is the target feature map.
[0011] The food recognition method based on multi-stage feature fusion includes a hierarchical extraction layer comprising multiple residual modules.
[0012] The food recognition method based on multi-stage feature fusion, wherein the number of residual modules in each extraction layer is less than the number of residual modules in the previous extraction layer.
[0013] The food recognition method based on multi-stage feature fusion, wherein the structural transformation processing of each intermediate feature map and the result feature map to obtain the feature matrix of each node includes:
[0014] Each of the intermediate feature maps is input into the corresponding pooling layer to obtain a pooling feature map;
[0015] For the target feature map, the following steps are performed to perform structural transformation processing on the target feature map to obtain the node feature matrix corresponding to the target feature map:
[0016] Each local feature in the target feature map is treated as a node feature, forming a node feature matrix corresponding to the target feature map;
[0017] The target feature map is either the pooling feature map or the result feature map, and the size of the local feature in the target feature map of size A×B×C is 1×1×C.
[0018] The food recognition method based on multi-stage feature fusion, wherein constructing the adjacency matrix corresponding to each node feature matrix includes:
[0019] Calculate the association similarity matrix corresponding to the node feature matrix;
[0020] The adjacency matrix is obtained by normalizing the correlation similarity matrix and then binarizing it.
[0021] The food recognition method based on multi-stage feature fusion, wherein the food recognition based on the target image features includes:
[0022] The target image features are input into the recognition module to obtain the food recognition result output by the recognition module;
[0023] The parameters of the recognition module, the feature extraction layer, the pooling layer, and the graph neural network are obtained after training based on multiple sets of training data.
[0024] A second aspect of the present invention provides a food recognition device based on multi-stage feature fusion, comprising:
[0025] The feature extraction module is used to acquire the image to be processed, input the image to be processed into the feature extraction layer, and obtain at least one intermediate feature map and a result feature map;
[0026] The feature transformation module is used to perform structural transformation processing on each of the intermediate feature maps and the result feature map to obtain the feature matrix of each node, construct the adjacency matrix corresponding to each of the node feature matrices, and input each of the adjacency matrices into the graph neural network to obtain the features of each stage.
[0027] The feature fusion module is used to fuse the features of each stage and the result feature map to obtain the target image features, and to perform food recognition based on the target image features.
[0028] A third aspect of the present invention provides a terminal, the terminal including a processor and a computer-readable storage medium communicatively connected to the processor, the computer-readable storage medium being adapted to store a plurality of instructions, the processor being adapted to invoke the instructions in the computer-readable storage medium to perform the steps of implementing the food recognition method based on multi-stage feature fusion as described in any of the preceding claims.
[0029] In a fourth aspect, the present invention provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the food recognition method based on multi-stage feature fusion as described in any of the preceding claims.
[0030] Compared with existing technologies, this invention provides a food recognition method based on multi-stage feature fusion. While extracting image features through a neural network, it retains and processes low-level features in the feature extraction process and then fuses them with high-level features. This achieves the utilization of multi-stage fusion features in the image feature extraction process, which can effectively improve the accuracy of image recognition. Attached Figure Description
[0031] Figure 1 A flowchart illustrating an embodiment of the food recognition method based on multi-stage feature fusion provided by the present invention;
[0032] Figure 2 A flowchart illustrating the use of datasets in one embodiment of the food recognition method based on multi-stage feature fusion provided by the present invention;
[0033] Figure 3 This is a flowchart illustrating the implementation process of the food recognition method based on multi-stage feature fusion provided by the present invention.
[0034] Figure 4 A schematic diagram of the multi-stage feature extraction structure in an embodiment of the food recognition method based on multi-stage feature fusion provided by the present invention;
[0035] Figure 5A framework diagram of a food recognition model in an embodiment of the food recognition method based on multi-stage feature fusion provided by the present invention;
[0036] Figure 6 The diagram shows the effect verification of an embodiment of the food recognition method based on multi-stage feature fusion provided by the present invention.
[0037] Figure 7 This is a schematic diagram of the structural principle of an embodiment of the food recognition device based on multi-stage feature fusion provided by the present invention;
[0038] Figure 8 A schematic diagram illustrating the principle of an embodiment of the terminal provided by the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0040] The food recognition method based on multi-stage feature fusion provided by this invention can be applied to terminals with computing capabilities. The terminal can execute the food recognition method based on multi-stage feature fusion provided by this invention to perform food recognition. The terminal can be, but is not limited to, various computers, mobile terminals, smart home appliances, wearable devices, etc.
[0041] Example 1
[0042] like Figure 1 As shown, one embodiment of the food recognition method based on multi-stage feature fusion includes the following steps:
[0043] S100. Obtain the image to be processed, input the image to be processed into the feature extraction layer, and obtain at least one intermediate feature map and a result feature map.
[0044] The food recognition method based on multi-stage feature fusion provided in this embodiment can be implemented based on a food recognition model, which includes the feature extraction layer. The food recognition model can be trained based on multiple sets of training data, each set of training data including sample images and food recognition result labels in the sample images.
[0045] The image to be processed can be an image input into the food recognition model for training during the training process of the food recognition model, i.e., an image in the training set; it can also be an image used to verify the food recognition model after the food recognition model has been trained, i.e., an image in the verification set; or it can be a specified image from which food needs to be identified.
[0046] like Figure 2As shown, before training the food recognition model, data acquisition and processing are required to obtain a dataset and divide it into training and validation sets. Currently, there are many existing food recognition datasets, such as Food-101 and Vireo Food-172. These public datasets can be collected using mainstream search engines and databases. The food image datasets can be downloaded and saved, and the downloaded data can be analyzed to determine the components of these datasets. Furthermore, the images should be cropped to a specific size (e.g., 224x224), and the entire dataset should be divided into training and validation sets in a 3:1 ratio.
[0047] Since convolution calculations are performed in regions corresponding to the kernel size each time, the input and output are locally connected. Furthermore, convolutional networks are typically constructed by stacking convolutional layers, with each layer transforming the previous one, resulting in increasingly abstract features from low to high levels. In the method provided in this embodiment, features from different levels at each stage are extracted and fused to compensate for missing local features, thereby improving the accuracy of image recognition.
[0048] Specifically, for the image to be processed, multi-stage feature extraction is first performed, which can be achieved through the feature extraction layer. The feature extraction layer can include multiple hierarchical extraction layers, each outputting features for one stage, and each hierarchical extraction layer sequentially extracts features at higher levels. Each hierarchical extraction layer includes multiple residual modules. In this embodiment, the residual module is a Bottleneck module, which is a residual module with a special structure, including a 1×1 convolutional kernel, a 3×3 convolutional kernel, and finally a 1×1 convolutional kernel. The number of residual modules included in each hierarchical extraction layer decreases sequentially.
[0049] Taking a three-layer extraction layer as an example, the first layer extracts low-level features, the second layer extracts mid-level features, and the third layer extracts high-level features. The following examples illustrate the process of each layer extracting the intermediate feature map.
[0050] The first layer of extraction can include 7 bottleneck modules and can be divided into two layers. The first layer is mainly composed of 3 bottleneck modules. After inputting the image to be processed, a first feature map can be obtained. The size of the feature map is 56×56×256. The second layer is mainly composed of 4 bottleneck modules. After inputting the first feature map, further features are extracted to obtain a second feature map. The size of the second feature map is 28×28×512. This feature map is the first intermediate feature map, which is the extracted low-level feature.
[0051] The second layer of extraction may include 6 bottleneck modules. The second feature map is input into the second layer of extraction to further extract features and obtain a fifth feature map. The fifth feature map is the second intermediate feature map, which is the extracted intermediate-level feature. The size of the fifth feature map is 14×14×1024.
[0052] The third layer of extraction may include three bottleneck modules. The fifth feature map is input into the third layer of extraction to further extract features and obtain an eighth feature map. The eighth feature map is the result feature map, which is the extracted high-level feature. The size of the eighth feature map is 7×7×2048.
[0053] After obtaining the intermediate feature map and the result feature map, in order to fuse the features extracted at each stage, this embodiment further processes the intermediate feature map and the result feature map. Specifically, the steps include:
[0054] S200. Perform structural transformation processing on each of the intermediate feature maps and the result feature maps to obtain the feature matrices of each node, construct the adjacency matrix corresponding to each of the node feature matrices, and input each of the adjacency matrices into the graph neural network to obtain the features of each stage.
[0055] Specifically, the structural transformation processing of each intermediate feature map and the result feature map to obtain the feature matrix of each node includes:
[0056] Each of the intermediate feature maps is input into the corresponding pooling layer to obtain a pooling feature map;
[0057] For the target feature map, the following steps are performed to perform structural transformation processing on the target feature map to obtain the node feature matrix corresponding to the target feature map:
[0058] Each local feature in the target feature map is treated as a node feature, forming a node feature matrix corresponding to the target feature map;
[0059] The target feature map is either the pooling feature map or the result feature map, and the size of the local feature in the target feature map of size A×B×C is 1×1×C.
[0060] The construction of the adjacency matrix corresponding to each of the node feature matrices includes:
[0061] Calculate the association similarity matrix corresponding to the node feature matrix;
[0062] The adjacency matrix is obtained by normalizing the correlation similarity matrix and then binarizing it.
[0063] The following example, using two intermediate feature maps, illustrates the processing of the intermediate feature maps and the result feature maps.
[0064] For the first intermediate feature map, i.e., the second feature map mentioned above, it first passes through a pooling layer using Max Pooling, with a filter size of 3×3, a stride of 2, and padding of 1. After inputting the second feature map, features are further extracted to obtain a third feature map, which has a size of 14×14×512. This third feature map then passes through another pooling layer using Max Pooling, with a filter size of 3×3, a stride of 2, and padding of 1. After inputting the third feature map, features are further extracted to obtain a fourth feature map, which has a size of 7×7×512. This fourth feature map is the pooled feature map corresponding to the second feature map. After obtaining the fourth feature map, it is input into a graph representation module. The graph representation module is built using a graph neural network and mainly consists of a graph result transformation module, a graph neural layer, and a ReLU layer. The input feature dimension is 512, and the output feature dimension is 256. The graph structure transformation module extracts the fourth feature map mentioned above, treats each smaller local feature as a node, and flattens it to obtain a 49×512 feature vector, resulting in a 49-node 512-dimensional feature matrix F. Then, the association similarity matrix R is calculated, and the matrix R is normalized using the softmax function. Finally, a suitable threshold is set, and the matrix is binarized using a threshold function to obtain a sparse matrix, which is the constructed adjacency matrix A. This process can be expressed by the formula:
[0065] R = F × F T
[0066]
[0067] A = f(softmax(R))
[0068] Here, F is the node feature matrix, and R is the association similarity matrix. The association degree is calculated to determine the connection between nodes, which is used to mine the spatial relationship between local features. f(x) is a threshold function, and δ is the threshold value. The threshold is obtained through parameter tuning and can be set to 0.6. The sparse matrix can be obtained through binarization, and the resulting matrix A is the constructed adjacency matrix. The adjacency matrix A is input into the graph convolutional network and expanded to obtain the fifth feature map. The fifth feature map has a size of 7×7×256, which is the extracted first-stage feature, denoted as F1.
[0069] For the second intermediate feature map, namely the fifth feature map mentioned above, it first passes through a pooling layer using Max Pooling with a filter size of 3×3, a stride of 2, and padding of 1. After inputting the fifth feature map, features are further extracted to obtain a sixth feature map, which has a size of 7×7×1024. This sixth feature map is the pooled feature map corresponding to the fifth feature map. Then, the sixth feature map is input into a graph representation module, which, like the graph representation module mentioned above, consists of a graph structure transformation module, a graph convolutional layer, and a ReLU layer. The input feature dimension is 1024, and the output feature dimension is 512. Through the graph structure transformation module, the sixth feature map is extracted, and each smaller local feature is treated as a node. This is flattened to obtain a 49×1024 feature vector, resulting in 49 1024-dimensional node feature matrices. Then, the association similarity matrix is calculated, and the matrix is normalized using the softmax function. Finally, a suitable threshold is set, and the matrix is binarized using the threshold function to obtain a sparse matrix, which is the constructed adjacency matrix. The adjacency matrix is input into the graph convolutional network and expanded to obtain the seventh feature map, which has a size of 7×7×512. This feature map is the extracted second-stage feature, denoted as F2.
[0070] No pooling is performed on the resulting feature map. The resulting feature map is directly input into the graph representation module, which mainly consists of a graph structure transformation module, a graph convolutional layer, and a ReLU layer. The input feature dimension is 2048, and the output feature dimension is 1024. Through the graph structure transformation module, the sixth feature map is extracted, and each smaller local feature is treated as a node. The vector is flattened to obtain a 49×2048 feature vector, resulting in 49 2048-dimensional node feature matrices. Then, the association similarity matrix is calculated, and the matrix is normalized using the softmax function. Finally, a suitable threshold is set, and the matrix is binarized using a threshold function to obtain a sparse matrix, which is the constructed adjacency matrix. The adjacency matrix is input into the graph convolutional network to obtain the eighth feature map, which has a size of 7×7×1024. This feature map is the extracted second-stage feature, denoted as F3.
[0071] Please refer to it again. Figure 1 The method provided in this embodiment further includes the following steps:
[0072] S300: The features of each stage and the result feature map are fused to obtain the target image features, and food is identified based on the target image features.
[0073] like Figure 4 As shown, after obtaining the features of each stage, the features of each stage and the result feature map are fused to obtain the target image features. Specifically, the features of each stage and the result feature map are concatenated and fused using the concat function, and then dimensionality reduction is performed. For example, a convolutional layer with a kernel of 1×1 and a stride of 1 is used for dimensionality reduction. After dimensionality reduction, a feature with more global characteristics than the original result feature map is obtained.
[0074] The food identification based on the target image features includes:
[0075] The target image features are input into the recognition module to obtain the food recognition result output by the recognition module;
[0076] The parameters of the recognition module, the feature extraction layer, and the graph neural network are obtained after training based on multiple sets of training data.
[0077] The obtained target image features are merely more discriminative features, such as... Figure 5 As shown, the obtained target image features need to be used as input features and input into the recognition module. The recognition module, the feature extraction layer, the pooling layer, and the graph representation module mentioned above constitute a food recognition model. That is, the food recognition model implements the entire processing of the image to be processed and outputs the food recognition result. The constructed food recognition model needs to be trained to determine suitable parameters, such as... Figure 3 As shown, the images in the training set are input into the food recognition model. The parameters are tuned and optimized by the difference between the food recognition results output by the food recognition model and the image labels. The prediction accuracy of the model is verified by the validation set data.
[0078] The inventors have verified the effectiveness of the method provided in this embodiment through experiments, such as... Figure 6 As shown, various existing models were trained and validated on the VireoFood172 dataset. Using the method provided in this embodiment (taking the Resnent50 residual module as the residual module and the GCN graph convolutional network as the graph neural network as an example), the classification accuracy reached 87.22%, which is significantly better than other existing models.
[0079] In summary, this embodiment provides a food recognition method based on multi-stage feature fusion. While extracting image features through a neural network, it retains and processes low-level features during the feature extraction process and then fuses them with high-level features. This achieves the utilization of multi-stage fusion features in the image feature extraction process, which can effectively improve the accuracy of image recognition.
[0080] It should be understood that although the steps in the flowcharts shown in the accompanying drawings are displayed sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0081] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0082] Example 2
[0083] Based on the above embodiments, the present invention also provides a food recognition device based on multi-stage feature fusion, such as... Figure 7 As shown, the food recognition device based on multi-stage feature fusion includes:
[0084] The feature extraction module is used to acquire the image to be processed, input the image to be processed into the feature extraction layer, and obtain at least one intermediate feature map and a result feature map, as described in Embodiment 1.
[0085] The feature transformation module is used to perform structural transformation processing on each of the intermediate feature maps and the result feature map to obtain each node feature matrix, construct the adjacency matrix corresponding to each node feature matrix, and input each adjacency matrix into the graph neural network to obtain the features of each stage, as specifically described in Embodiment 1;
[0086] The feature fusion module is used to fuse the features of each stage and the result feature map to obtain the target image features, and to perform food recognition based on the target image features, as described in Embodiment 1.
[0087] Example 3
[0088] Based on the above embodiments, the present invention also provides a terminal, such as... Figure 8 As shown, the terminal includes a processor 10 and a memory 20. Figure 8 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0089] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a food recognition program 30 based on multi-stage feature fusion, which can be executed by the processor 10 to implement the food recognition method based on multi-stage feature fusion described in this application.
[0090] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other chip, used to run program code stored in the memory 20 or process data, such as executing the food recognition method based on multi-stage feature fusion.
[0091] In one embodiment, when the processor 10 executes the food recognition program 30 based on multi-stage feature fusion in the memory 20, the following steps are performed:
[0092] The image to be processed is obtained and input into the feature extraction layer to obtain at least one intermediate feature map and a result feature map;
[0093] The intermediate feature maps and the result feature maps are structurally transformed to obtain the feature matrices of each node. The adjacency matrices corresponding to the feature matrices of each node are constructed. The adjacency matrices are then input into the graph neural network to obtain the features of each stage.
[0094] The features of each stage and the result feature map are fused to obtain the target image features, and food recognition is performed based on the target image features.
[0095] The feature extraction layer includes multiple extraction layers, the output of each extraction layer is the intermediate feature map, the input of each extraction layer is the output of the previous extraction layer, and the output of the last extraction layer is the target feature map.
[0096] The hierarchical extraction layer includes multiple residual modules.
[0097] The number of residual modules in each extraction layer is less than the number of residual modules in the previous extraction layer.
[0098] The step of performing structural transformation processing on each of the intermediate feature maps and the result feature maps to obtain the feature matrices of each node includes:
[0099] Each of the intermediate feature maps is input into the corresponding pooling layer to obtain a pooling feature map;
[0100] For the target feature map, the following steps are performed to perform structural transformation processing on the target feature map to obtain the node feature matrix corresponding to the target feature map:
[0101] Each local feature in the target feature map is treated as a node feature, forming a node feature matrix corresponding to the target feature map;
[0102] The target feature map is either the pooling feature map or the result feature map, and the size of the local feature in the target feature map of size A×B×C is 1×1×C.
[0103] The construction of the adjacency matrix corresponding to each of the node feature matrices includes:
[0104] Calculate the association similarity matrix corresponding to the node feature matrix;
[0105] The adjacency matrix is obtained by normalizing the correlation similarity matrix and then binarizing it.
[0106] The food identification based on the target image features includes:
[0107] The target image features are input into the recognition module to obtain the food recognition result output by the recognition module;
[0108] The parameters of the recognition module, the feature extraction layer, the pooling layer, and the graph neural network are obtained after training based on multiple sets of training data.
[0109] Example 4
[0110] The present invention also provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the food recognition method based on multi-stage feature fusion as described above.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A food recognition method based on multi-stage feature fusion, characterized in that, The method includes: The image to be processed is obtained and input into the feature extraction layer to obtain at least one intermediate feature map and a result feature map; The intermediate feature maps and the result feature maps are subjected to structural transformation processing to obtain the feature matrices of each node. The adjacency matrices corresponding to the feature matrices of each node are constructed. The adjacency matrices are input into the graph neural network to obtain the stage features corresponding to each intermediate feature map and the result feature map respectively. The features of each stage and the result feature map are fused to obtain the target image features, and food is identified based on the target image features; The feature extraction layer includes multiple extraction layers. Each extraction layer outputs features for a stage, and each extraction layer extracts features for a higher level in sequence.
2. The food recognition method based on multi-stage feature fusion according to claim 1, characterized in that, The feature extraction layer includes multiple extraction layers. The output of each extraction layer is the intermediate feature map, the input of each extraction layer is the output of the previous extraction layer, and the output of the last extraction layer is the result feature map.
3. The food recognition method based on multi-stage feature fusion according to claim 2, characterized in that, The hierarchical extraction layer includes multiple residual modules.
4. The food recognition method based on multi-stage feature fusion according to claim 3, characterized in that, The number of residual modules in each extraction layer is less than the number of residual modules in the previous extraction layer.
5. The food recognition method based on multi-stage feature fusion according to claim 1, characterized in that, The structural transformation process performed on each of the intermediate feature maps and the result feature map to obtain the feature matrix of each node includes: Each of the intermediate feature maps is input into the corresponding pooling layer to obtain a pooling feature map; For the target feature map, the following steps are performed to perform structural transformation processing on the target feature map to obtain the node feature matrix corresponding to the target feature map: Each local feature in the target feature map is treated as a node feature, forming a node feature matrix corresponding to the target feature map; The target feature map is either the pooling feature map or the result feature map, and the size of the local feature in the target feature map of size A×B×C is 1×1×C.
6. The food recognition method based on multi-stage feature fusion according to claim 5, characterized in that, The construction of the adjacency matrix corresponding to each of the node feature matrices includes: Calculate the association similarity matrix corresponding to the node feature matrix; The adjacency matrix is obtained by normalizing the correlation similarity matrix and then binarizing it.
7. The food recognition method based on multi-stage feature fusion according to claim 5, characterized in that, The food identification based on the target image features includes: The target image features are input into the recognition module to obtain the food recognition result output by the recognition module; The parameters of the recognition module, the feature extraction layer, the pooling layer, and the graph neural network are obtained after training based on multiple sets of training data.
8. A food recognition device based on multi-stage feature fusion, characterized in that, The food recognition device based on multi-stage feature fusion is used to implement the food recognition method based on multi-stage feature fusion as described in any one of claims 1-7, including: The feature extraction module is used to acquire the image to be processed, input the image to be processed into the feature extraction layer, and obtain at least one intermediate feature map and a result feature map; The feature transformation module is used to perform structural transformation processing on each of the intermediate feature maps and the result feature map to obtain each node feature matrix, construct the adjacency matrix corresponding to each node feature matrix, and input each adjacency matrix into the graph neural network to obtain the features of each stage. The feature fusion module is used to fuse the features of each stage and the result feature map to obtain the target image features, and to perform food recognition based on the target image features; The feature extraction layer includes multiple extraction layers. Each extraction layer outputs features for a stage, and each extraction layer extracts features for a higher level in sequence.
9. A terminal, characterized in that, The terminal includes: a processor and a computer-readable storage medium communicatively connected to the processor. The computer-readable storage medium is adapted to store multiple instructions, and the processor is adapted to call the instructions in the computer-readable storage medium to execute the steps of implementing the food recognition method based on multi-stage feature fusion as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the food recognition method based on multi-stage feature fusion as described in any one of claims 1-7.
Citation Information
Patent Citations
Graph structure representation high-order correlation discovery fine-grained image recognition method and device
CN113222041A
Target identification method and apparatus, and electronic device
CN114359796A