Food classification component identification method and system based on multi-scale feature attention synergy
By employing a multi-scale feature attention collaborative method, the problem of the separation between food identification and component analysis tasks is solved, achieving high accuracy and generalization ability in food classification and component identification, which is suitable for applications such as intelligent nutrition analysis.
Patent Information
- Application Number
- CN202511108736.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies suffer from problems such as the fragmentation of food recognition and component analysis tasks, insufficient local feature representation, and inadequate cross-task information complementarity, making it difficult to achieve semantically guided component localization and key region focusing in complex scenarios.
A multi-scale feature attention collaborative method is adopted, which combines a global feature extraction module, a layer-by-layer progressive local feature extraction module, and a bidirectional cross-task attention module with the KL divergence loss function to achieve feature interoperability and collaborative optimization for food classification and component identification.
It significantly improves the recognition effect of complex food images, achieves high-precision food classification and component identification, and has a certain generalization and zero-shot inference capability, making it suitable for applications such as intelligent nutrition analysis.
Smart Images

Figure CN120997823A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of computer vision and food computing, specifically to a method and system for food classification and component recognition based on multi-scale feature attention collaboration. Background Technology
[0002] In recent years, food computation, as an emerging interdisciplinary research field, has received widespread attention due to its significant potential in assisting users with food selection, healthy diet planning, and nutrition management. As a fundamental task in food computation, food recognition not only directly serves the basic needs of daily human diet but also plays a crucial role in numerous practical applications such as nutritional analysis, dietary monitoring, and health management. Furthermore, from a computer vision perspective, food image recognition itself belongs to a special and challenging sub-problem of fine-grained visual classification tasks, possessing profound theoretical research value and attracting increasing attention from computer vision and artificial intelligence researchers.
[0003] Despite the abundance of recipe information online, accurately identifying the corresponding dish and matching it to a recipe from a food image remains a significant challenge. The core difficulty lies in precisely identifying the food category from an image and then analyzing its composition. Since the breakthroughs in Convolutional Neural Networks (CNNs), numerous studies have attempted to identify food categories and infer their components or nutritional value from images using visual analysis methods. However, these methods generally face the challenge of insufficient dataset coverage. According to Wikipedia, there are over 8,000 types of dishes worldwide, but in reality, most of these dishes are formed by different combinations of a limited number of basic ingredients. Therefore, analyzing from a component perspective can not only greatly simplify the identification task but also effectively handle unfamiliar food types, thereby improving the system's generalization performance and practical application value.
[0004] The existing invention patent application document CN119380059A, entitled "Self-Supervised Clustering Method for Hyperspectral Images Based on Local-Global Dual-Branch Networks," describes a method that includes: constructing 3D pixel blocks; obtaining shallow depth features through a ResNet module; injecting the shallow features into a dual-path network module; calculating the similarity between local features and cluster centers to obtain a local semantic probability distribution, and then obtaining the target distribution; feeding global features into a feedforward neural network to obtain the probability of each pixel and obtain a global semantic probability distribution; and constructing a network loss function through a dual self-supervised mechanism to guide the update of the entire network model. However, the aforementioned prior art has the following drawbacks:
[0005] 1. Limitations of task objectives and application scenarios:
[0006] The aforementioned existing technologies focus on unsupervised clustering of hyperspectral images, such as land cover classification. Their technical solutions extract features through local-global dual-branch networks and optimize the cluster distribution using KL divergence loss. However, these existing methods only solve a single task (image clustering) and do not involve multi-task collaboration, such as the joint optimization of food classification and component identification. Furthermore, their application scenarios are limited to hyperspectral remote sensing images and cannot be transferred to the field of food computation.
[0007] Root cause analysis:
[0008] Its dual-branch network (CNN+Transformer) design only serves single-task clustering objectives and lacks cross-task feature interaction mechanisms, which cannot meet the needs of bidirectional information complementarity for food classification and component identification in this application.
[0009] 2. Lack of feature interaction mechanism:
[0010] The aforementioned existing technologies extract local (CNN) and global (Transformer) features through parallel dual-branch extraction, but there is no active interaction between the branches. They only constrain feature distribution differences through KL divergence, failing to establish a dynamic correlation channel between semantic and component information, resulting in insufficient cross-modal information complementarity.
[0011] as a result of:
[0012] In complex scenarios (such as strong correlations between ingredients and categories in food images), it is difficult to achieve semantically guided ingredient localization or focus on key regions for ingredient-assisted classification.
[0013] 3. Insufficient representation of multi-scale features:
[0014] While the aforementioned existing technologies extract multi-scale features, they do not explicitly enhance the discriminative power between scales. Their KL divergence is only used to optimize the cluster distribution, failing to address the problem of multi-scale features tending to be homogeneous, thus limiting fine-grained recognition capabilities.
[0015] Specific manifestations:
[0016] Small components in food images, such as spices and side dishes, are easily overwhelmed by global features, lacking a mechanism to constrain the diversity of local features.
[0017] In summary, existing technologies suffer from technical problems such as the fragmentation of food identification and component analysis tasks, insufficient representation of local features, and inadequate cross-task information complementarity. Summary of the Invention
[0018] The technical problem to be solved by this invention is: how to solve the technical problems of food identification and component analysis tasks being fragmented, insufficient expression of local features, and insufficient cross-task information complementarity in the prior art.
[0019] This invention solves the above-mentioned technical problems by employing the following technical solution: A food classification and component recognition method based on multi-scale feature attention collaboration includes:
[0020] S1. Using the overall feature extraction module, high-level semantic information is extracted from the input food image to obtain the overall features;
[0021] S2. Using a progressive local feature extraction module, multi-scale detailed features are extracted from the input food image step by step, and the KL divergence loss function is used to enhance the discriminative power between features of different scales.
[0022] S3. Design a bidirectional cross-task attention module and adopt a bidirectional cross-attention mechanism to perform feature exchange and collaborative optimization operations for food classification tasks and component recognition tasks.
[0023] S4. Integrate overall features and multi-scale detailed features to generate a unified feature representation for food category prediction and multi-label component recognition.
[0024] This invention improves the detail representation capability of multi-scale features through progressive feature fusion, significantly enhancing the recognition effect of complex food images;
[0025] This invention has the advantage of system integration, with a single model simultaneously outputting food classification results and component analysis, making it suitable for application scenarios that require real-time feedback, such as intelligent nutrition analysis apps;
[0026] The trained model can simultaneously achieve high-precision food classification and component identification, and possesses a certain degree of generalization and zero-shot inference capabilities. It is suitable for applications such as mobile health management and intelligent nutrition analysis.
[0027] In a more specific technical solution, the bidirectional cross-attention mechanism in S3 uses food image features as queries and ingredient label features as keys to dynamically filter visually relevant ingredient information along the visual-to-ingredient path.
[0028] In a more specific technical solution, the bidirectional cross-attention mechanism in S3 uses component features as the query to reverse locate key regions in the image along the component-to-visual path.
[0029] In a more specific technical solution, S3 generates cross-modal consistent feature representations through bidirectional attention interaction:
[0030]
[0031] f′ i =A f→i V i
[0032]
[0033] f i =A i→f V f
[0034] In the formula, Qf,Kf,Vf and Qi,Ki,Vi are the learnable weight matrices of food features and component features, respectively.
[0035] The cross-attention collaboration module of the present invention establishes a bidirectional attention path, wherein food category features dynamically guide the selection of component features, and the component features are used in reverse to locate key visual regions in the image, thereby achieving semantic information alignment and cross-task feature supplementation.
[0036] This invention utilizes the dual-task synergy of a bidirectional attention-guided feature interaction mechanism. In the vision-to-component path, food category features serve as the query vector (Q_f), dynamically filtering key information (V_i) from component features, significantly improving the relevance of component recognition;
[0037] In the component-to-visual path, component features are used as query vectors (Q_i) to reverse locate key regions of the image, solving the problem of semantic separation between vision and components in traditional methods;
[0038] Cross-modal feature alignment is achieved through bidirectional cross-attention, establishing semantic associations between food categories and ingredients.
[0039] In a more specific technical solution, S3 employs a multi-task joint optimization strategy;
[0040] Calculate the KL divergence loss;
[0041] The total loss function is obtained by weighting and combining the single-label cross-entropy loss function and the KL divergence loss.
[0042] In a more specific technical solution, a single-label cross-entropy loss function is used to perform the food recognition task.
[0043] In a more specific technical solution, the single-label cross-entropy loss function is expressed using the following logic:
[0044]
[0045] In a more specific technical solution, the following logic is used to perform a weighted combination to obtain the total loss function:
[0046] L=αL food +βL ing +γL KL
[0047] In the formula, α, β, and γ are equilibrium parameters.
[0048] This invention employs an end-to-end joint optimization strategy, using a multi-task loss function to balance the food category prediction loss, component multi-label prediction loss, and feature distribution difference loss, thereby achieving collaborative updating of model parameters.
[0049] This invention employs a ternary loss collaborative optimization approach, integrating food classification loss (L_food), ingredient identification loss (L_ing), and KL divergence loss (L_KL). It dynamically balances the optimization directions of the two tasks through weight parameters (α, β, γ). Compared to the binary loss approach of a single-task scheme (Document 1 only contains L_ing and L_KL), this achieves collaborative error correction.
[0050] In a more specific technical solution, S4 utilizes a fully connected network structure to concatenate and fuse multi-scale detailed features and overall features acquired at different training stages to obtain a unified feature representation:
[0051] f food =concat(f Glo ,f′ U - ′S+1 ,...,f′ U )
[0052] f ing =concat(f Glo ,f U-S+1 ,...,f U )
[0053] In the formula, U is the total number of network stages, and S is the number of training steps.
[0054] In more specific technical solutions, the food classification and component recognition system based on multi-scale feature attention collaborative methods includes:
[0055] The overall feature extraction module is used to extract high-level semantic information from the input food image to obtain overall features;
[0056] A progressive local feature extraction module is used to extract multi-scale detailed features from the input food image step by step, and the KL divergence loss function is used to enhance the discriminative power between features of different scales.
[0057] A bidirectional cross-task attention module is designed to perform feature exchange and collaborative optimization operations for food classification and component recognition tasks using a bidirectional cross-attention mechanism.
[0058] The feature fusion module is used to fuse overall features and multi-scale detailed features to generate a unified feature representation for food category prediction and multi-label component recognition. The feature fusion module is connected to the overall feature extraction module, the progressive local feature extraction module, and a bidirectional cross-task attention module.
[0059] The present invention has the following advantages over the prior art:
[0060] By introducing a cross-task bidirectional attention mechanism, the bottleneck of traditional task separation is broken, enabling collaborative interaction between food semantics and component information;
[0061] By employing KL divergence to constrain feature diversity, the problem of multi-scale feature convergence is solved, and it performs excellently on various datasets.
[0062] The cross-attention collaboration module of the present invention establishes a bidirectional attention path, wherein food category features dynamically guide the selection of component features, and the component features are used in reverse to locate key visual regions in the image, thereby achieving semantic information alignment and cross-task feature supplementation.
[0063] This invention employs an end-to-end joint optimization strategy, using a multi-task loss function to balance the food category prediction loss, component multi-label prediction loss, and feature distribution difference loss, thereby achieving collaborative updating of model parameters.
[0064] This invention solves the technical problems of food identification and component analysis tasks being fragmented, insufficient local feature representation, and inadequate cross-task information complementarity in the prior art. Attached Figure Description
[0065] Figure 1 This is a schematic diagram of the basic steps of the food classification and component identification method based on multi-scale feature attention collaboration in Embodiment 1 of the present invention;
[0066] Figure 2 This is a schematic diagram of the food recognition model in Embodiment 2 of the present invention. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] Example 1
[0069] like Figure 1As shown, the food classification and component recognition method based on multi-scale feature attention collaboration includes the following basic steps:
[0070] S1. Use the overall feature extraction module to extract high-level semantic information from the input food image;
[0071] In this embodiment, the overall feature extraction module is based on the Vision Transformer architecture. It captures the overall semantic information of the image through a self-attention mechanism and performs global average pooling on the output of the last layer to obtain the global feature vector.
[0072] In this embodiment, the overall feature extraction module is based on a self-supervised training (Vision Transformer, ViT) architecture. It aggregates the final layer features of the Transformer through a global average pooling (GAP) operation to form a global feature representation.
[0073] f Glo =GAP(f out )
[0074] Here, fout is the output feature of the last layer of ViT.
[0075] S2. Through a progressive local feature extraction module, multi-scale detailed features are extracted step by step, and the KL divergence loss function is used to enhance the discriminative power between features of different scales.
[0076] In this embodiment, the progressive local feature learning module adopts a phased activation mechanism, gradually unlocking deep network layers during forward propagation, extracting multi-scale local features by combining global max pooling, and constraining the distribution differences of features at different stages by KL divergence regularization.
[0077] In this embodiment, the progressive local feature module adopts a stepwise activation method, unlocking deeper network structures sequentially during the forward propagation of the network, obtaining multi-scale local features through global max pooling, and using KL divergence to constrain the distribution of features at different stages.
[0078] In this embodiment, deeper Transformer modules are activated progressively in multiple training phases to capture fine-grained feature information in a gradual manner.
[0079] In each training phase, local feature vectors are extracted using global max pooling (GMP) operations:
[0080]
[0081] By leveraging KL divergence to maximize the differences in feature distribution at different stages, the model is encouraged to focus on diverse visual regions.
[0082] Specifically, the implementation of KL divergence optimization includes mapping local features at different stages into probability distributions and effectively suppressing the convergence of features at different stages by increasing the KL divergence gap between different stages.
[0083] S3. Design a bidirectional cross-task attention module to promote feature interoperability and collaborative optimization between food classification tasks and component recognition tasks;
[0084] In the bidirectional cross-attention mechanism of this embodiment, in the vision-to-component path, food image features are used as queries and component label features are used as keys to dynamically filter visually relevant component information.
[0085] By employing an end-to-end joint optimization strategy, a multi-dimensional loss function is used to balance the component identification loss for multi-label loss and the feature difference constraint loss, thereby achieving overall updating of model parameters.
[0086] In the component-to-visual path, key regions in the image are located in reverse by using component features as queries;
[0087] Cross-modal consistent feature representations are generated through bidirectional attention interaction, and the formula is as follows:
[0088]
[0089] f′ i =A f→i V i
[0090]
[0091] f i =A i→f V f
[0092] Where Qf,Kf,Vf and Qi,Ki,Vi are the learnable weight matrices for food features and component features, respectively.
[0093] In this embodiment, a multi-task joint optimization strategy is adopted; specifically, the food recognition task uses the single-label cross-entropy loss function:
[0094]
[0095] The component identification task uses an improved multi-label cross-entropy loss function to address the label imbalance problem.
[0096]
[0097] Calculate the KL divergence loss:
[0098]
[0099] The total loss function is a weighted combination:
[0100] L=αL food +βL ing +γL KL
[0101] Where α, β, and γ are equilibrium parameters.
[0102] S4. Integrate overall and local features to generate a unified and efficient feature representation, which is used to complete food category prediction and multi-label identification of ingredients.
[0103] A fully connected network structure is used to concatenate and fuse local features acquired at different training stages with global features to obtain the final feature representation. The formula is:
[0104] f food =concat(f Glo ,f′ U-S+1 ,...,f′ U )
[0105] f ing =concat(f Glo ,f U-S +1,...,f U )
[0106] Where U is the total number of network stages and S is the number of training steps.
[0107] Example 2
[0108] In this embodiment, a multi-task progressive feature aggregation network model is constructed. Specifically, in the progressive local feature learning module, multi-scale local features are extracted at layers 6, 9, and 12 of the Transformer encoder, and a progressive training strategy is adopted to optimize the model in three stages.
[0109] Phase 1: Freeze the last 6 layers, then train the first 6 Transformer layers to extract the underlying texture features;
[0110] Phase 2: Unfreeze layers 7-9 and constrain feature diversity using KL divergence loss;
[0111] Phase 3: Jointly optimize all 12 layers, integrating global and local features;
[0112] In the cross-task attention interaction module of this embodiment, a bidirectional cross-attention mechanism is adopted; specifically, the multi-task classifier includes, but is not limited to:
[0113] Food classification branch: The fully connected layer outputs 101 / 172-dimensional class probabilities;
[0114] Component identification branch: The multi-label classifier outputs the probability of the existence of N-dimensional components (N≥100), and the LogSumExp loss function is used to handle label imbalance.
[0115] In this embodiment, model training and optimization are performed; specifically, data preprocessing is performed, wherein the input image is uniformly scaled to 256×256 pixels, a 224×224 region is randomly cropped, and horizontal flipping, ±15° rotation and HSV color gamut enhancement are applied.
[0116] The component label is encoded as an N-dimensional binary vector, with 1 indicating the presence of a component and 0 otherwise.
[0117] The training strategy used in this embodiment includes:
[0118] Initialization: Fine-tuning based on ImageNet-21K pre-trained weights;
[0119] Optimizer: SGD is used, with an initial learning rate of 1e-3, momentum of 0.9, and weight decay of 1e-4.
[0120] Progressive training: divided into 3 stages (10 epochs per stage), with a total of 30 epochs of training, and the learning rate decays by 0.8 times in each stage;
[0121] Determine the multi-task loss function:
[0122] L = 0.6L food +0.3L ing +0.1L KL .
[0123] In summary, by introducing a cross-task bidirectional attention mechanism, the bottleneck of traditional task separation is broken, enabling collaborative interaction between food semantics and component information;
[0124] By employing KL divergence to constrain feature diversity, the problem of multi-scale feature convergence is solved, and it performs excellently on various datasets.
[0125] The cross-attention collaboration module of the present invention establishes a bidirectional attention path, wherein food category features dynamically guide the selection of component features, and the component features are used in reverse to locate key visual regions in the image, thereby achieving semantic information alignment and cross-task feature supplementation.
[0126] This invention employs an end-to-end joint optimization strategy, using a multi-task loss function to balance the food category prediction loss, component multi-label prediction loss, and feature distribution difference loss, thereby achieving collaborative updating of model parameters.
[0127] This invention solves the technical problems of food identification and component analysis tasks being fragmented, insufficient local feature representation, and inadequate cross-task information complementarity in the prior art.
[0128] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A food classification and component recognition method based on multi-scale feature attention collaboration, characterized in that, The method includes: S1. Using the overall feature extraction module, high-level semantic information is extracted from the input food image to obtain the overall features; S2. Using a progressive local feature extraction module, multi-scale detail features are extracted from the input food image step by step, and the KL divergence loss function is used to enhance the discriminability between features of different scales. S3. Design a bidirectional cross-task attention module and adopt a bidirectional cross-attention mechanism to perform feature exchange and collaborative optimization operations for food classification tasks and component recognition tasks. S4. Integrate the overall features and the multi-scale detailed features to generate a unified feature representation for food category prediction and multi-label component recognition.
2. The food classification and component recognition method based on multi-scale feature attention collaborative method according to claim 1, characterized in that, In the bidirectional cross-attention mechanism described in S3, on the visual-to-component path, food image features are used as queries and component label features are used as keys to dynamically filter visually relevant component information.
3. The food classification and component recognition method based on multi-scale feature attention collaborative method according to claim 1, characterized in that, In the bidirectional cross-attention mechanism described in S3, key regions in the image are located in reverse order using component features as the query in the component-to-visual path.
4. The food classification and component identification method based on multi-scale feature attention collaborative method according to claim 1, characterized in that, In step S3, cross-modal consistent feature representations are generated through bidirectional attention interaction: f′ i =A f→i V i f i =A i→f V f In the formula, Qf,Kf,Vf and Qi,Ki,Vi are the learnable weight matrices of food features and component features, respectively.
5. The food classification and component identification method based on multi-scale feature attention collaborative method according to claim 1, characterized in that, In S3, a multi-task joint optimization strategy is adopted; Calculate the KL divergence loss; The total loss function is obtained by weighting and combining the single-label cross-entropy loss function and the KL divergence loss.
6. The food classification and component identification method based on multi-scale feature attention collaborative method according to claim 5, characterized in that, The food recognition task is performed using the single-label cross-entropy loss function.
7. The food classification and component identification method based on multi-scale feature attention collaborative method according to claim 6, characterized in that, The single-label cross-entropy loss function can be expressed using the following logic:
8. The food classification and component identification method based on multi-scale feature attention collaborative method according to claim 5, characterized in that, The total loss function is obtained by weighting and combining the results using the following logic: L=αL food +βL ing +γL KL In the formula, α, β, and γ are equilibrium parameters.
9. The food classification and component identification method based on multi-scale feature attention collaborative method according to claim 1, characterized in that, In step S4, a fully connected network structure is used to concatenate and fuse the multi-scale detail features and the overall features obtained at different training stages to obtain the unified feature representation: f food =concat(f Glo ,f′ U-S+1 ,..,f′ U ) f ing =concat(f Glo ,f U-S+1 ,...,f U ) In the formula, U is the total number of network stages, and S is the number of training steps.
10. A food classification and component recognition system based on multi-scale feature attention collaboration, characterized in that, The system includes: The overall feature extraction module is used to extract high-level semantic information from the input food image to obtain overall features; A progressive local feature extraction module is used to extract multi-scale detailed features from the input food image step by step, and the KL divergence loss function is used to enhance the discriminative power between features of different scales. A bidirectional cross-task attention module is designed to perform feature exchange and collaborative optimization operations for food classification and component identification tasks using a bidirectional cross-attention mechanism. The feature fusion module is used to fuse the overall features and the multi-scale detailed features to generate a unified feature representation for food category prediction and multi-label component recognition. The feature fusion module is connected to the overall feature extraction module, the layer-by-layer progressive local feature extraction module, and the designed bidirectional cross-task attention module.
Citation Information
Patent Citations
Hyperspectral image self-supervised clustering method based on local-global double-branch network
CN119380059A
Cited By
Dairy product quality grade determination method based on multi-source information fusion
CN121352637A