Deep hash food image retrieval method based on context-aware proxy interaction and fusion
By designing a deep hash food image retrieval method based on context-aware agent interaction and fusion, using the Transformer encoder and AIP module to extract global and local features, and optimizing hash code generation through a cross-fusion module and a new loss function, the problem of insufficient fine-grained feature extraction in food image retrieval is solved, and efficient and accurate food image retrieval is achieved.
Patent Information
- Application Number
- CN202510607146.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-12
AI Technical Summary
Existing image retrieval models find it difficult to effectively capture the fine-grained features of food images when processing them, resulting in degraded retrieval performance. In addition, hashing methods may be affected by the fine-grained features of food images during the hash code generation process, resulting in suboptimal hash codes.
A deep hashing food image retrieval method based on context-aware agent interaction and fusion is designed. Global and local features are extracted through the Transformer encoder and AIP module, and hash code generation is optimized through a cross-fusion module and a new loss function to achieve more efficient retrieval.
The accuracy and efficiency of food image retrieval are improved, and it can better adapt to the fine-grained visual features and key area feature extraction of food images, generate accurate hash codes, and improve retrieval performance.
Smart Images

Figure CN120632140A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology and relates to food image retrieval technology, deep learning, and hash learning. Specifically, it provides a deep hash food image retrieval method based on context-aware agent interaction and fusion. Background Art
[0002] Food plays a vital role in human life. Food images, with their intuitive and easy-to-understand nature, have become a medium for conveying food information and a primary source of information. With the continuous advancement of image data acquisition and analysis technologies, the accumulation of large-scale food image data has rapidly increased, leading to a doubling of research demand for accurate and efficient food image processing. Food image retrieval, a major branch of image retrieval and food computing, has become a research hotspot.
[0003] However, compared to general image retrieval tasks, food images present unique challenges due to their more complex feature distribution. First, the visual features of food images are complex and diverse. Different types of food exhibit diverse shapes and distinct colors and textures. For example, the same ingredient can take on completely different appearances when cooked differently. Furthermore, foods of similar color and shape can be visually difficult to distinguish. This fine-grained variation further complicates feature extraction and classification. Furthermore, food images often contain a significant amount of irrelevant background (such as tableware and tabletops), which interferes with image feature extraction and makes it difficult for the model to focus on key food areas. In contrast, the target objects in general image retrieval tasks are mostly rigid objects with clear geometric structures (such as buildings and vehicles) or entities with distinct features (such as animals and birds), whose visual features are relatively stable and easy to extract. This difference in feature distribution makes it difficult for existing image retrieval models to effectively capture fine-grained features in food images, resulting in reduced retrieval performance. Given these challenges, food image retrieval places higher demands on the model's feature representation capabilities and ability to capture subtle differences.
[0004] Therefore, it is necessary to design a specialized model structure and algorithm that can adapt to the complex fine-grained visual features of food images and focus on the extraction of key food area features. In the field of image retrieval, many existing methods rely on content-based image retrieval technology, which uses the visual features of images for search. However, using content-based image retrieval alone usually incurs high computational costs and is inefficient. Hash image retrieval not only has significant advantages in storage, but also can greatly improve retrieval efficiency through hash codes. For complex food images, existing hashing methods may be affected by the fine-grained features of food images during the hash code generation process, resulting in suboptimal hash codes. Therefore, in the process of generating hash codes, it is necessary to design a loss function to guide the generation of hash codes to maintain the consistency of the hash space and the feature space and generate accurate and representative hash codes.
[0005] In order to learn accurate hash codes based on the uniqueness of food images to achieve higher retrieval accuracy, this application proposes a deep hashing food image retrieval method based on context-aware agent interaction and fusion. Summary of the Invention
[0006] This invention provides a deep hashing food image retrieval method based on context-aware agent interaction and fusion. This method fully utilizes the semantic information between features at different scales, aims to better adapt to the extraction of fine-grained visual features and key food regions in food images, and achieve efficient deep hashing retrieval. To achieve the above objectives, the invention is implemented through the following technical solutions: A deep hashing food image retrieval method based on context-aware agent interaction and fusion includes the following steps: Step 1: Input image and database; given dataset Include An image, whose expression is defined as ,in Indicates the first images. Given an image in the dataset With corresponding label data , labeled dataset By label data Composition, whose expression is defined as ,in Indicates the number of image categories. The sample belongs to a certain category When the label vector Bit ,otherwise .
[0007] Step 2: Feature extraction process; given an input food image , first divide it into several non-overlapping and fixed-size blocks. Then, linear mapping is performed on each block embedding to convert it into a feature vector. The feature vector is passed through the Transformer encoder module and the Aggregation-Interaction-Propagation (AIP) module for feature extraction to obtain the corresponding output representation and , respectively representing the global features and local detail features of the image.
[0008] Step 3: Feature cross fusion process; First, the output representation of the Transformer encoder module Extract class tokens on the first dimension , as the representative feature of the global image, and it is represented by the output of the AIP module Perform splicing operation to obtain the spliced features . Further, the spliced features and the output representation of the Transformer encoder Perform linear transformations respectively and perform two cross attention calculations to obtain the outputs of two cross attentions respectively. and .right and Fusion to obtain interactive features , the formula is as follows:
[0009] in is a learnable parameter used to dynamically adjust the importance of the interaction between two features.
[0010] Step 4: Optimize the network using two loss functions; using polarization loss To minimize the learned hash code Its corresponding target vector Using enhanced cross entropy loss To constrain the semantic consistency of the hash code. The goal of this function is to reduce the intra-class distance while increasing the inter-class distance, and to optimize the network parameters in the hash learning process, so that the network can learn more accurate hash codes.
[0011] Combining the above two losses, set is a weight hyperparameter that balances the two losses and retrieves the new deep hashing loss of the framework. Defined as:
[0012] Step 5: Perform model training and performance evaluation based on the above steps; the batch size of the training algorithm is set to 128, and the optimizer uses adaptive moment estimation, i.e. Adam, to optimize the network model, with a learning rate of 1×10 -5 .
[0013] Image retrieval quality evaluation uses two indicators: average retrieval precision, i.e. mAP, and precision-recall curve, i.e. PR curve.
[0014] Step 6: Output the results; use the trained model to process the input query image, extract the visual features of the food image, and compare it with the images in the food image database. Through similarity calculation, find a series of images that are most similar to the query image and output them as retrieval results.
[0015] In step 3, the spliced features and the output representation of the Transformer encoder Perform linear transformations respectively and obtain the following characteristic matrices: , generate the query matrix , key matrix , value matrix ;for , generate the query matrix , key matrix , value matrix . Subsequently, two cross-attention calculations are performed on the above feature matrix.
[0016] In step 4, polarization loss and enhanced cross entropy loss The implementation is as follows: Hash code of each sample After being learned, according to the label , for each Select the corresponding target vector The polarization loss function aims to minimize the learned hash code Its corresponding target vector The difference between the polarization loss The formula is defined as:
[0017] in is a hyperparameter that controls the polarization threshold. Indicates the batch size.
[0018] Through the classification head Mapping to probability distribution of categories:
[0019] in Indicates the The samples belong to The hash code value of each category.
[0020] For the samples, and its loss for:
[0021] Finally, define the enhanced cross entropy loss for:
[0022] In step 6, the Hamming distance similarity metric is used to calculate the distance metric, using the difference between hash codes to measure the similarity between images. This method can better improve the retrieval accuracy and make the matching more accurate.
[0023] Compared with the prior art, the beneficial effects of the present invention are: This paper proposes a deep hashing food image retrieval model based on context-aware agent interaction and fusion. This model is an end-to-end deep learning framework that combines convolutional and Transformer networks. It has the ability to capture higher-dimensional global semantic information and more detailed local feature representations. Compared to traditional image retrieval network architectures, this network effectively addresses the problem of insufficient extraction of fine-grained features and key region features in food images.
[0024] The design of the AIP module in this invention improves the model's ability to express the semantics of food images, enabling it to better distinguish different food images. Taking into account the complex semantic information of food images, a Cross-Fusion Module (CFM) is proposed. CFM uses the context of the AIP module to enhance the global features extracted by the Transformer encoder to achieve the fusion of relevant information at different scales, thereby improving the model's ability to discriminate between similar images. In order to maintain the consistency of the semantics of the hash code space, a new deep hashing loss function is designed. This function aims to optimize both intra-class similarity and inter-class differences, while simultaneously optimizing the network parameters in the hash learning process, thereby improving overall retrieval performance.
[0025] This method was evaluated on three different food datasets: ETH Food-101, VireoFood-172, and UEC Food-256. Experimental results show that the method achieves good retrieval accuracy on all of these datasets, a significant improvement in performance for food image retrieval tasks.
[0026] As described above, the design and optimization of this invention have achieved significant progress in the field of deep hashing food image retrieval. Its innovations lie in the design of the AIP module, CFM, and a new loss function. These improvements enable the model to perform well in food image retrieval tasks. These innovations provide powerful methods and technologies for applying deep learning to food image retrieval, with significant practical implications for various dietary and health management scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 This is a schematic diagram of the overall network framework provided by the present invention; Figure 2 This is a schematic diagram of the specific architecture of the AIP model of the present invention; Figure 3 These are the first eight result images retrieved by the retrieval model of the present invention on the ETH Food-101 dataset. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solutions and advantages of the present invention clearer and more specific, the present invention will be further explained in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the specific embodiments described herein are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the present application.
[0030] Example 1: Combine Figure 1 As shown, this embodiment provides a deep hashing food image retrieval network based on context-aware agent interaction and fusion. The network includes the following three designs: the Aggregation-Interaction-Propagation (AIP) module, the Cross-Fusion Module (CFM), and a new deep hashing loss function. The network includes the following main steps and optimization objectives: Step 1: Input image and database; given dataset Include An image, whose expression is defined as ,in Indicates the first images. Given an image in the dataset With corresponding label data , labeled dataset By label data Composition, whose expression is defined as ,in Indicates the number of image categories. The sample belongs to a certain category When the label vector Bit ,otherwise .
[0031] Step 2: Feature extraction process; given an input food image First, the input image is processed using two-dimensional convolution to divide it into several non-overlapping and fixed-size blocks. Then, each block embedding is linearly mapped and converted into a feature vector to obtain the initial representation of the image block. . At the same time, a class token is introduced ( ), as learnable parameters initialized with zero. Initial representation of class tokens and image patches After concatenation, the expanded embedding representation is generated To capture spatial location information, a learnable location embedding matrix is added , retaining the position embedding information. The initial feature representation of the position embedding is integrated The definition is as follows:
[0032] The Transformer encoder module utilizes multi-head attention operations and adopts layer normalization and residual connection operations to enhance the stability and transferability of feature representation. The Transformer encoder module contains A stack of repeatedly stacked Transformer blocks. Each block contains a multi-head self-attention ( ) layers and a multilayer perceptron ( )layer.
[0033] Specifically, for each Transformer block, the feature representation of the current block input is first Perform layer normalization ( ) and sent to The layer calculates the updated feature representation:
[0034] Then, the updated feature representation Perform layer normalization again and send it to The layer performs nonlinear feature transformation and finally obtains the output representation of the Transformer encoder module :
[0035] in , It is through The intermediate representation after the layer is updated, and the residual connection ensures the stability and transferability of the feature representation.
[0036] Combine Figure 2 As shown in Figure 2, the AIP module enables more efficient information communication within each individual layer, thereby enhancing the ability of feature interaction across image patches. This module contains three key mechanisms, which are as follows: Aggregation: In the aggregation stage, the features of each image block are aggregated through deep convolution operations to generate rich feature representations This operation is performed by Aggregating local information within the window can effectively mine the detailed features inside the block. The formula is as follows:
[0037]
[0038]
[0039] in is the convolution operation, is the depthwise convolution operation, For batch normalization, and It is the intermediate representation after convolution and residual connection; Interaction: In the interaction phase, the features generated in the aggregation phase are represented As a proxy block token, we sample proxy block tokens from the entire feature map space and perform multi-head self-attention operations on the selected proxy block tokens. In this process, all proxy block tokens in the space are used as queries to participate in the self-attention calculation. The formula is as follows:
[0040]
[0041] in It is the input Perform linear mapping to obtain query vectors , key vector Sum vector . For each head's dimensions, is the transpose operation. for Function, which normalizes all spatial blocks. Function to get the attention output of the interaction phase Through the above operations, the interactive stage can achieve effective modeling of global features under the premise of controllable computational complexity; Propagation: In the propagation stage, the transposed convolution operation is used to obtain the propagation characteristics , Output the attention of the interaction phase Propagate to their respective adjacent block tokens. Specifically, through transposed convolution, the module is able to communicate comprehensive information between image blocks, capture the global features of the image, and improve the coherence and consistency of feature representation. The specific formula is as follows:
[0042] in is the average pooling operation, is the transposed convolution operation. Then the features are propagated After the mapping operation, the output features of the AIP module are obtained ; Step 3: Feature cross fusion process; First, the output representation of the Transformer encoder module Remove the class token on the first dimension , as the representative features of the global image, and compare it with the output features of the AIP module Perform the splicing operation. Features after splicing Defined as:
[0043] in For splicing operation.
[0044] Furthermore, the spliced features and the output representation of the Transformer encoder module Perform linear transformations respectively and obtain the following characteristic matrices: , generate the query matrix , key matrix , value matrix ;for , generate the query matrix , key matrix , value matrix .
[0045] Then, two cross-attention calculations are performed on the above feature matrix respectively. The specific operations are as follows: 1) The query matrix and The bond matrix Sum Matrix As input, get the first output :
[0046] 2) The query matrix and The bond matrix Sum Matrix As input, get the second output :
[0047] in is a scaling factor used to stabilize the gradient computation.
[0048] Finally, the output of the two cross attention and Fusion to obtain interactive features , the formula is as follows:
[0049] in is a learnable parameter used to dynamically adjust the importance of the interaction between two features.
[0050] Step 4: Use two loss functions to optimize the network; given categories and hash bit lengths , the target vector matrix When the hash code After being learned, its goal is to be as close as possible to the target vector matrix . According to the label , for each Select the corresponding target vector .
[0051] The polarization loss function aims to minimize the learned hash code Its corresponding target vector The difference between the polarization loss The formula is defined as:
[0052] in For the The hash code of each sample. is a hyperparameter that controls the polarization threshold. Indicates the batch size.
[0053] The enhanced cross entropy loss function is designed to constrain the semantic consistency of the hash code. Its goal is to reduce the intra-class distance while expanding the inter-class distance, and optimize the network parameters in the hash learning process, so that the network can learn more accurate hash codes. First, the classification head is used to Mapping to probability distribution of categories:
[0054] in Indicates the The samples belong to The hash code value of each category.
[0055] For the samples, and its loss for:
[0056] Finally, define the enhanced cross entropy loss for:
[0057] Combining the above two losses, set is a weight hyperparameter that balances the two losses and retrieves the new deep hashing loss of the framework. Defined as:
[0058] Step 5: Perform model training and performance evaluation based on the above steps; the batch size of the training algorithm is uniformly set to 128, and the optimizer uses Adam to optimize the network model with a learning rate of 1×10 -5 .
[0059] Image retrieval quality evaluation uses two indicators: average search precision (mAP) and precision-recall curve (PR curve).
[0060] Step 6: Output the results. Use the trained model to process the query image, extract visual features from the food image, and compare it with images in the food image database. A distance metric is calculated using the Hamming distance similarity metric, measuring the similarity between images using the difference between hash codes. The set of images most similar to the query image is found and output as the retrieval results.
[0061] As attached Figure 1As shown in the figure, the present invention designs a deep hashing food image retrieval framework based on context-aware agent interaction and fusion. For the input food image data, the designed network architecture first extracts image features, and then uses a new deep hashing loss function for network optimization. The model is then trained and performance evaluated. Finally, the retrieval results are output. While extracting global features, the present invention uses the AIP module to interact with global and local detail features, and efficiently fuses local and global features through CFM.
[0062] In step 5, model training requires a data set, and model performance evaluation requires evaluation indicators. The following is an introduction to the data set and evaluation indicators in this invention.
[0063] Datasets: This paper is evaluated on three different food datasets: ETH Food-101, VireoFood-172, and UEC Food-256. (1) The ETH Food-101 dataset contains image datasets from 101 food categories, with a total of 101,000 food images, 250 test images for each category, and 750 training images. In this experiment, 120 images are randomly selected from each category as the training set and 30 images are used as the test set. (2) The Vireo Food-172 dataset divides food into 172 major categories, with 200-1000 food images in each category, for a total of 110,241 food images. In this experiment, 120 images are randomly selected from each category as the training set and 30 images are used as the test set. (3) The UEC Food-256 dataset contains 31,395 images of 256 kinds of food. The number of categories in this dataset is inconsistent, which requires the processing to be able to understand complex visual features and a wide range of category recognition capabilities to distinguish category semantic similarities. In this experiment, 80 images are randomly selected from each category as the training set and 20 images as the test set.
[0064] Evaluation Metrics: The food image retrieval quality assessment in this paper uses two metrics: mean average precision (mAP) and precision-recall (PR) curve. For the three food datasets, the mAP@1000 metric was used for evaluation.
[0065] Mean Average Precision (mAP): mAP is an overall performance metric calculated based on the Average Precision (AP) of each query. For a single query, AP reflects the average precision of the model under all possible recall rates and is typically calculated by observing the model's performance in different rankings of search results. mAP, on the other hand, averages the AP across all queries to measure the model's overall performance on the entire test set. The high or low mAP directly reflects the quality of the search. A higher mAP value means the model is able to find relevant items with greater accuracy.
[0066] Precision-Recall (PR) Curve: The PR curve visually reflects model performance by showing the relationship between precision (Precision) and recall (Recall) at different thresholds. Precision measures the proportion of results predicted as relevant that are truly relevant, while recall indicates the proportion of all truly relevant results that are successfully retrieved. The PR curve plots recall on the horizontal axis and precision on the vertical axis, with the shape of the curve illustrating the trade-off between the two. Generally speaking, the closer the curve is to the upper right corner, the better the model performance.
[0067] The present invention fully verifies the performance of the retrieval model through the above evaluation indicators. Food image retrieval has a wide range of application value and can be applied to the fields of diet and health management.
[0068] Example 2: As attached Figure 3 As shown in the figure, this paper proposes a deep hashing food image retrieval method based on dual-attention cross-fusion. The backbone network of this model combines the concept of convolutional interaction with the Transformer, which can not only extract local fine-grained semantic information but also achieve local and global information interaction, thereby generating more accurate food image hash code representations. To address the optimization issues of the food image hashing algorithm, a new deep hashing loss function is designed to maintain spatial consistency and enhance inter-class separation and intra-class aggregation capabilities.
[0069] Figure 3The visualization results presented in the figure are examples of image retrieval on the Food-101 dataset of the present invention. These results provide an intuitive understanding of the model performance. The present invention randomly selected 8 images on the Food-101 dataset for retrieval and returned the top 8 retrieval results. The first column in the figure is the query image, and the second to ninth columns represent the returned retrieval results. The correct retrieval results are marked with blue boxes, and the incorrect retrieval results are marked with red boxes. The results show that most of the retrieval results shown in the figure are correct, and very few incorrect images are included. This shows that the present invention has achieved a high retrieval accuracy on the food dataset and can accurately find images similar to the query image.
Claims
1. A deep hashing food image retrieval method based on context-aware agent interaction and fusion, comprising the following steps: Step 1: Input image and database; Given a dataset Include An image, whose expression is defined as ,in Indicates the first images; images in a given dataset With corresponding label data , labeled dataset By label data Composition, whose expression is defined as ,in Indicates the number of image categories; when The sample belongs to a certain category When the label vector Bit ,otherwise ; Step 2: Feature extraction process; Given an input food image First, it is divided into several non-overlapping and fixed-size blocks; then, each block embedding is linearly mapped and converted into a feature vector; the feature vector is extracted by the Transformer encoder module and the AIP module respectively to obtain the corresponding output representation and , respectively represent the global features and local detail features of the image; Step 3: Feature cross fusion process; First, the output representation of the Transformer encoder module Extract class tokens on the first dimension , as the representative feature of the global image, and it is represented by the output of the AIP module Perform splicing operation to obtain the spliced features ; Further, the features after splicing and the output representation of the Transformer encoder Perform linear transformations respectively and perform two cross attention calculations to obtain the outputs of two cross attentions respectively. and ;right and Fusion to obtain interactive features , the formula is as follows:
2. Among them It is a learnable parameter used to dynamically adjust the importance of the interaction between two features; Step 4: Optimize the network using two loss functions; using polarization loss To minimize the learned hash code Its corresponding target vector The difference between ; using enhanced cross entropy loss To constrain the semantic consistency of the hash code; the goal of this function is to reduce the intra-class distance while expanding the inter-class distance, and to optimize the network parameters in the hash learning process, so that the network can learn more accurate hash codes; Combining the above two losses, set is a weight hyperparameter that balances the two losses and retrieves the new deep hashing loss of the framework. Defined as:
3. Step 5: Perform model training and performance evaluation based on the above steps; the batch size of the training algorithm is set to 128, and the optimizer uses adaptive moment estimation, i.e. Adam, to optimize the network model with a learning rate of 1×10 -5 ; Image retrieval quality evaluation uses two indicators: average retrieval precision, or mAP, and precision-recall curve, or PR curve; Step 6: Output the results; use the trained model to process the input query image, extract the visual features of the food image, and compare it with the images in the food image database. Through similarity calculation, find a series of images that are most similar to the query image and output them as retrieval results.
4. The deep hashing food image retrieval method based on context-aware agent interaction and fusion according to claim 1 is characterized by: The pair of outputs represents and The linear transformation is implemented as follows: Features after splicing and the output representation of the Transformer encoder Perform linear transformations respectively and obtain the following characteristic matrices: , generate the query matrix , key matrix , value matrix ;for , generate the query matrix , key matrix , value matrix . Subsequently, two cross-attention calculations are performed on the above feature matrix.
5. The deep hashing food image retrieval method based on context-aware agent interaction and fusion according to claim 1 is characterized by: The polarization loss and enhanced cross entropy loss The implementation is as follows, When Hash code of each sample After being learned, according to the label , for each Select the corresponding target vector ; The polarization loss function aims to minimize the learned hash code Its corresponding target vector The difference between the polarization loss The formula is defined as:
6. Among them is a hyperparameter that controls the polarization threshold; Indicates the batch size; Through the classification head Mapping to probability distribution of categories:
7. Among them Indicates the The samples belong to Hash code value of each category; For the samples, and its loss for:
8. Finally, define the enhanced cross entropy loss for:
9. The deep hashing food image retrieval method based on context-aware agent interaction and fusion according to claim 1 is characterized by: The similarity calculation adopts the similarity measurement method of Hamming distance, which uses the difference between hash codes to measure the similarity between images; this method can better improve the retrieval accuracy and make the matching more accurate.