A Global and Component Fine-Grained Recipe Retrieval Method Using Segmentation and Negative Matching
Through the segmentation and negative matching method, combined with SAM and cascading Transformer models, a dynamic negative matching-triple joint loss function is constructed, which solves the problem of global feature extraction of images in the prior art that is susceptible to noise interference and information loss of recipe components, and achieves a more accurate and robust cross-modal recipe retrieval effect.
Patent Information
- Application Number
- CN202510404739.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-02
AI Technical Summary
Existing cross-modal recipe retrieval methods are easily disturbed by background noise when extracting global features of the image, resulting in dilution of food features; at the same time, forcibly merging recipe components leads to semantic confusion and loss of component information, and are too sensitive to redundant information, introducing semantic bias.
The global and component fine-grained recipe search was performed using the method of segmentation and negative matching, and the food images were segmented through the SAM model to remove background interference, and the recipe text features were extracted using the cascading Transformer model. Construct dynamic negative matching-triple joint loss function, dynamically expand the mismatched sample spacing through Monte Carlo sampling, and optimize the semantic association of component-level alignment and segmentation features.
Effectively reduce noise interference, avoid semantic confusion and component information loss, significantly improve the accuracy and robustness of cross-modal retrieval, and has significant effects compared with traditional global alignment methods.
Smart Images

Figure SMS_426 
Figure SMS_427 
Figure QLYQS_35
Abstract
Description
Technical Field
[0001] The present invention relates to a global and component fine-grained recipe retrieval method using segmentation and negative matching, belonging to the technical field of food cross-modal retrieval in food computing. Background Art
[0002] With the development of vision-language models and the increasing emphasis on diet by people, cross-modal recipe retrieval has attracted extensive attention in the industrial and academic fields. However, due to the fact that recipes usually contain multi-level structures and rich context information, cross-modal recipe retrieval faces great challenges. Cross-modal recipe retrieval is a task that combines image and text information to identify and match recipes and their related images. Through cross-modal retrieval, a series of applications can be promoted, including food exploration, precise search, food and beverage marketing, personalized recipe creation, etc. The key to this task lies in accurately identifying information such as ingredients and cooking methods in food images and understanding the corresponding descriptions in recipe texts. At the same time, a semantic association model between images and texts needs to be constructed so that the two can perform effective similarity calculations in the same semantic space to achieve precise retrieval.
[0003] Most current cross-modal recipe retrieval methods adopt the triple global alignment method, which has significant disadvantages. First, the extracted global image features are mixed with a large amount of background noise, resulting in the dilution of ingredient features. Second, the three components of the recipe are forcibly combined into an overall vector, causing semantic confusion between recipes and the loss of unique information of each component. Finally, when performing triple global alignment, it is too sensitive to redundant information, and the forced alignment method introduces semantic biases. Due to these reasons, the global feature alignment method has unsatisfactory effects. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the above-mentioned existing technologies and provide a global and component fine-grained recipe retrieval method using segmentation and negative matching.
[0005] The technical solution provided by the present invention is as follows: A global and component fine-grained recipe retrieval method using segmentation and negative matching, characterized in that it includes the following steps:
[0006] Step S1, using the Recipe1M dataset, establish datasets of two modalities of food images and recipes, where the food images include two parts: food and background, and the recipes contain three components: title, ingredients, and instructions. Divide the datasets of food images and recipes into training sets, validation sets, and test sets;
[0007] Step S2, respectively use the ViT vision model and the cascaded Transformer model to extract features from the food images and recipes in the training set, validation set, and test set, specifically including:
[0008] Step S21, extract the global features of the food image;
[0009] Step S22, extract the image features of the main food after image segmentation;
[0010] Step S23, extract the text features of the recipe;
[0011] Step S3, construct the loss function on the training set, specifically including:
[0012] Step S31, calculate the matching probabilities from the food image to the recipe and from the recipe to the food image, and use the calculated matching probabilities to construct an overall negative matching loss function to increase the distance between unmatched samples;
[0013] Step S32, calculate the bidirectional matching probabilities between each component of the three recipes and the food image, and use the calculated matching probabilities to calculate three component-level negative matching loss functions to further increase the distance between unmatched samples;
[0014] Step S33, use the segmentation model to segment out the main food, calculate the similarity matrices of the segmented main food and the ingredient components in the recipe respectively, and use the triplet loss to reduce the matching error between the main food and the recipe ingredient components to enhance the robustness of cross-modal alignment;
[0015] Step S34, construct the final dynamic negative matching - triplet joint loss function;
[0016] Step S4, perform cross-modal recipe retrieval, and use the recall rate as the evaluation index.
[0017] Furthermore, in the above-mentioned Step S2:
[0018] Step S21, extract the global features of the food image:
[0019] Input the food image set into the ViT vision model to extract the global feature vector of the food image , denotes using the ViT vision model, denotes the th food image, is the visual modality identifier, denotes the number of food images;
[0020] Step S22, extract the image features of the main food after image segmentation:
[0021] Use the SAM segmentation model to segment food images. First, generate a large number of masks in the everything mode to cover each region in the food image; then, filter out too small or meaningless regions through an area threshold, and use the CLIP large model for semantic matching to retain the masks related to food; finally, select the top-N regions with the highest semantic scores as enhanced data, extract the main food in the food image, and remove the interference of the background; the food image segments formed after segmentation are sent into the ViT vision model to obtain the average food image segment embedding vector by performing average pooling on the embeddings of all food image segments , denotes the number of food image segments, denotes the th food image segment;
[0022] Step S23, extract the text features of the recipe:
[0023] Adopt a cascaded Transformer model to process at the word level and sentence level respectively. Given a recipe set , where denotes the title component, denotes the ingredient component, denotes the instruction component, is the text identifier, denotes the number of recipes; send the title component , the ingredient component and the instruction component into the Transformer model of the first layer respectively. The Transformer model of the first layer receives each word in the sentences of the three components and encodes it to capture the relationship between each word. Finally, through average pooling, the embeddings of all words in each sentence are averaged to obtain the sentence-level embeddings of the title component, ingredient component, and instruction component = , denotes the Transformer model of the first layer; send the obtained sentence-level embeddings of the title component, ingredient component, and instruction component into the Transformer of the second layer to capture the relationship between each sentence in the title component, ingredient component, and instruction component. Through average pooling, the sentence-level embeddings of the title component, ingredient component, and instruction component are averaged to obtain the final overall fusion feature vectors of the title component, ingredient component, and instruction component = , where It is the second-layer Transformer model; the overall fusion feature vectors of the title component, ingredient component, and instruction component are concatenated to obtain the global feature vector of the final recipe. , represents a fully connected layer.
[0024] Furthermore, in the step S3:
[0025] Step S31, the global feature vector of the food image , the global feature vector of the recipe , using the Monte Carlo method to calculate the global cross-modal matching probability from the food image to the recipe , and the global cross-modal matching probability from the recipe to the food image The formula is as follows:
[0026] ;
[0027] ;
[0028] Among them, represents the similarity between the global feature vector of the food image and the global feature vector of the recipe , represents the similarity between the global feature vector of the recipe and the global feature vector of the food image , is the temperature parameter, used to control the smoothness of the probability distribution, is a subset randomly sampled from the training set, representing the th randomly sampled recipe sample and food image sample respectively, represents the normalization factor of the recipe sample, represents the normalization factor of the food image sample, used to adjust the deviation of random sampling, and N is the batch size;
[0029] Using the formula for calculating the global cross-modal matching probability from the recipe to the food image to construct the matching probability set of the negative sample from the recipe to the food image = N is the batch size, is the matching probability from the th recipe to the th food image; then construct the overall negative matching loss function from the recipe to the food image , the formula is as follows:
[0030] ;
[0031] Construct the overall negative matching loss function from food images to recipes , and the construction method is the same as that of the loss function. Use the formula for calculating the global cross-modal matching probability from food images to recipes to construct the set of matching probabilities of food images to negative samples of recipes = where N is the batch size, represents the matching probability of the th food image to the th recipe; then construct the overall negative matching loss function from food images to recipes , and the formula is as follows:
[0032] ;
[0033] And add the two to get the formula for the overall negative matching loss function as follows:
[0034] ;
[0035] Step S32, calculate the overall fusion feature vector of the title component to the global feature vector of the food image matching probability , the overall fusion feature vector of the ingredient component to the global feature vector of the food image matching probability and the overall fusion feature vector of the instruction component to the global feature vector of the food image matching probability ; The formulas are as follows respectively:
[0036] ;
[0037] ;
[0038] ;
[0039] Among them, represents the similarity between the overall fusion feature vector of the title component and the global feature vector of the food image , represents the similarity between the overall fusion feature vector of the ingredient component and the global feature vector of the food image , represents the similarity between the overall fusion feature vector of the ingredient component and the global feature vector of the food image ; is a temperature parameter used to control the smoothness of the probability distribution. is a subset randomly sampled from the training set, representing the K-th randomly sampled food image sample. represents the normalization factor for the food image sample, used to adjust the bias of random sampling, where N is the batch size.
[0040] Calculate the global feature vector of the food image to the overall fusion feature vector of the title component matching probability , the global feature vector of the food image to the overall fusion feature vector of the ingredient component matching probability , the global feature vector of the food image to the overall fusion feature vector of the title component matching probability The formula is as follows:
[0041] ;
[0042] ;
[0043] ;
[0044] where represents the similarity between the global feature vector of the image to the overall fusion feature vector of the title component , represents the similarity between the global feature vector of the image to the overall fusion feature vector of the ingredient component , represents the similarity between the global feature vector of the image to the overall fusion feature vector of the instruction component , is a temperature parameter used to control the smoothness of the probability distribution. respectively represent a subset randomly sampled from the training set, representing the -th randomly sampled overall fusion feature vector sample of the title component, overall fusion feature vector sample of the ingredient component, overall fusion feature vector sample of the instruction component. is the normalization factor for the recipe sample, used to adjust the bias of random sampling, where N is the batch size.
[0045] Construct the title negative matching loss function , the ingredient negative matching loss function and the instruction negative matching loss function , respectively using the overall fusion feature vector of the title component to the global feature vector of the food image matching probability , the overall fusion feature vector of the ingredient component to the global feature vector of the food image matching probability and the overall fusion feature vector of the instruction component to the global feature vector of the food image matching probability Construct the matching probability set of the negative samples of the title component to the food image , the matching probability set of the negative samples of the ingredient component samples to the food image , the matching probability set of the negative samples of the instruction component samples to the food image , N is the batch size, represents the th title component to the th negative sample matching probability of the food image, represents the th ingredient component to the th negative sample matching probability of the food image, represents the th instruction component to the th negative sample matching probability of the food image; then construct the negative matching loss function of the title component to the food image , the negative matching loss function of the ingredient component to the food image , the negative matching loss function of the ingredient component to the food image The formula is as follows:
[0046] ;
[0047] ;
[0048] ;
[0049] Construct the negative matching loss function of the food image to the title component , the negative matching loss function of the food image to the ingredient component , the negative matching loss function of the food image to the instruction component , respectively using the global feature vector of the food image to the overall fusion feature vector of the title component matching probability , the global feature vector of the food image to the overall fusion feature vector of the ingredient component matching probability and the global feature vector of the food image to the overall fusion feature vector of the instruction component matching probability Construct a negative sample matching probability set of the food image to the title component , a negative sample matching probability set of the food image to the ingredient component , a negative sample matching probability set of the food image to the instruction component , N is the batch size, denotes the negative sample matching probability of the food image to the th title component, denotes the negative sample matching probability of the food image to the th ingredient component, denotes the negative sample matching probability of the food image to the th instruction component recipe image negative sample; then construct the negative matching loss function of the title component to the food image , the negative matching loss function of the ingredient component to the food image , the negative matching loss function of the ingredient component to the food image The formula is as follows:
[0050] ;
[0051] ;
[0052] ;
[0053] Add the two-direction loss functions to obtain the negative matching function of the title component , the negative matching function of the ingredient component , the negative matching function of the ingredient component The formula is as follows:
[0054] ;
[0055] ;
[0056] ;
[0057] Add the negative matching function of the title component , the negative matching function of the ingredient component , the negative matching function of the ingredient component to obtain the negative matching loss function of the component The formula is as follows:
[0058] ;
[0059] Step S33, embed each average food image segment after image segmentation into a vector to form a vector set and the overall fusion feature vector of each ingredient component to form a vector set , and calculate the similarity matrices of and respectively. The specific calculation formula is as follows: ; ;
[0060] ;
[0061] ;
[0062] Use the triplet loss to minimize the difference between the matrices. The calculation formula is:
[0063] ;
[0064] where is the cosine similarity vector, is the triplet loss margin parameter, is the anchor point, is the positive sample, is the negative sample; is the transpose matrix, is the transpose matrix;
[0065] Step S34, add the three loss functions to obtain the final dynamic negative matching - triplet joint loss function :
[0066] + + .
[0067] The beneficial effects of the present invention are as follows: Firstly, the present invention uses a cascaded Transformer to hierarchically extract recipe text features. During the feature extraction process, both a global feature vector needs to be constructed and the vectors of the three components need to be saved respectively, overcoming the semantic confusion problem caused by the loss of component information in traditional methods. Then, the SAM model is introduced to segment the food image to select the main food area, and the ViT model is used to extract visual features with interference removed, effectively reducing noise interference. Finally, a dynamic negative matching - triple joint loss function is constructed. The global negative matching loss dynamically expands the distance of mismatched samples through Monte Carlo sampling, the component - level negative matching loss optimizes the fine - grained alignment of the title, ingredients, and instructions respectively, and the local triple loss strengthens the semantic association between the segmented main food features and the text ingredient items.
[0068] The present invention innovatively integrates component fine - grained feature extraction and cross - modal contrast learning mechanisms. It uses the SAM model to accurately locate the main food area in food image samples, removing the influence of background noise, and maintains the semantic independence of components in recipe samples. Through the negative matching - triple joint optimization strategy, it avoids the text redundancy problem caused by forced alignment and the semantic confusion problem caused by the loss of component information, significantly improving the accuracy and robustness of cross - modal retrieval. Compared with traditional global alignment methods, the present invention has significant effects. Detailed implementation manners
[0069] The following makes a detailed description of the specific implementation manners of the present invention:
[0070] A global and component - level fine - grained recipe retrieval method using segmentation and negative matching, which includes the following steps:
[0071] Step S1, using the Recipe1M dataset, establish datasets of two modalities, namely food images and recipes. The food images include food and background, and the recipes contain three components: title, ingredients, and instructions. Divide the datasets of food images and recipes into training sets, validation sets, and test sets.
[0072] Step S2, respectively use the ViT visual model and the cascaded Transformer model to extract features from the food images and recipes in the training set, validation set, and test set. Specifically, it includes the following steps:
[0073] Step S21, extract the global features of food images:
[0074] Input the food image set into the ViT visual model to extract the global feature vector of the food image , denotes using the ViT visual model, denotes the th food image, is a visual modality identifier, indicating the number of food images.
[0075] Step S22, extract the image features of the main food after image segmentation:
[0076] Use the SAM segmentation model to segment the food images. First, generate a large number of masks in the everything mode to cover each area in the food images. Then, filter out too small or meaningless areas through an area threshold, and use the CLIP large model for semantic matching to retain the masks related to food. Finally, select the Top-N regions with the highest semantic scores as the enhanced data, extract the main food in the food images, and remove the interference of the background. It is found through verification that the Top-4 effect is the best. Send the food image segments into the ViT vision model and perform average pooling on the embeddings of all food image segments to obtain the average food image segment embedding vector , indicating the number of food image segments, indicating the th food image segment.
[0077] Step S23, extract the text features of the recipe:
[0078] Adopt a cascaded Transformer model to process at the word level and sentence level respectively. Given a recipe set , where represents the title component, represents the ingredient component, represents the instruction component, is a text identifier, indicating the number of recipes; send the title component , the ingredient component and the instruction component into the Transformer model of the first layer respectively. The Transformer model of the first layer receives each word in the sentences of the three components and encodes them to capture the relationships between each word. Finally, through average pooling, the embedding representations of all words in each sentence are averaged to obtain the sentence-level embeddings of the title component, ingredient component, and instruction component = , Denote the first - layer Transformer model; send the sentence - level embeddings of the obtained title component, ingredient component, and instruction component into the second - layer Transformer to capture the relationships between each sentence in the title component, ingredient component, and instruction component, and obtain the overall fusion feature vectors of the final title component, ingredient component, and instruction component by averaging the sentence - level embeddings of the title component, ingredient component, and instruction component through average pooling. = , where is the second - layer Transformer model; concatenate the overall fusion feature vectors of the title component, ingredient component, and instruction component to obtain the global feature vector of the final recipe. , denotes the fully - connected layer.
[0079] Step S3, construct the loss function on the training set, which specifically includes the following steps:
[0080] The core of constructing the negative - matching loss is to find a common feature subspace to maximize the distance between negative - sample image - recipe pairs.
[0081] Take the global feature vector of the food image , the global feature vector of the recipe , and use the Monte Carlo method to calculate the global cross - modal matching probability from the food image to the recipe , and the global cross - modal matching probability from the recipe to the food image The formulas are as follows:
[0082] ;
[0083] ;
[0084] where denotes the similarity between the global feature vector of the food image and the global feature vector of the recipe , denotes the similarity between the global feature vector of the recipe and the global feature vector of the food image , is the temperature parameter, which is used to control the smoothness of the probability distribution, is a subset randomly sampled from the training set, representing the th randomly sampled recipe sample and food - image sample respectively, denotes the normalization factor of the recipe sample, denotes the normalization factor of the food - image sample, which is used to adjust the deviation of random sampling, and N is the batch size.
[0085] Using the formula for calculating the global cross-modal matching probability from the computational recipe to the food image Construct a set of matching probabilities for negative samples from the recipe to the food image = N is the batch size, is the th recipe to the th food image matching probability; then construct the overall negative matching loss function from the recipe to the food image , the formula is as follows:
[0086] ;
[0087] Since we have two tasks: retrieving recipes from food images and retrieving food images from recipes, we need to construct the overall negative matching loss function from the food image to the recipe , and the construction method is the same as that of the loss function. Using the formula for calculating the global cross-modal matching probability from the food image to the recipe Construct a set of matching probabilities for negative samples from the food image to the recipe = N is the batch size, represents the th food image to the th recipe matching probability. Then construct the overall negative matching loss function from the food image to the recipe , the formula is as follows:
[0088] ;
[0089] And add the two to get the overall negative matching loss function formula as follows:
[0090] ;
[0091] Through this overall negative matching loss function The problem that images are difficult to fully represent the content described by the text, that is, the problem of text information redundancy, can be effectively solved.
[0092] Step S32, each component in the recipe contains its own more detailed information. By performing negative matching on each component, more detailed mismatched information can be captured to further widen the mismatched distance, and calculate the overall fusion feature vector of the title component to the global feature vector of the food image matching probability , the overall fusion feature vector of the ingredient component to the global feature vector of the food image matching probability The overall fusion feature vector of the instruction component to the global feature vector of the food image matching probability . The formulas are as follows:
[0093] ;
[0094] ;
[0095] ;
[0096] wherein, represents the similarity between the overall fusion feature vector of the title component to the global feature vector of the food image , represents the similarity between the overall fusion feature vector of the ingredient component to the global feature vector of the food image , represents the similarity between the overall fusion feature vector of the ingredient component to the global feature vector of the food image , is the temperature parameter, used to control the smoothness of the probability distribution, is a subset randomly sampled from the training set, representing the Kth randomly sampled food image sample, represents the normalization factor for the food image sample, used to adjust the deviation of random sampling, and N is the batch size.
[0097] Since it is a two-way retrieval, it is necessary to calculate the matching probability of the global feature vector of the food image to the overall fusion feature vector of the title component , the matching probability of the global feature vector of the food image to the overall fusion feature vector of the ingredient component , the matching probability of the global feature vector of the food image to the overall fusion feature vector of the title component The formulas are as follows:
[0098]
[0099] ;
[0100] ;
[0100] ;
[0101] wherein, represents the global feature vector of the image The similarity of the overall fusion feature vector to the title component , where the global feature vector of the image is represented The similarity of the overall fusion feature vector to the ingredient component , where the global feature vector of the image is represented The similarity of the overall fusion feature vector to the instruction component , is the temperature parameter, used to control the smoothness of the probability distribution are respectively represented as a subset randomly sampled from the training set, representing the th overall fusion feature vector sample of the randomly sampled title component, the overall fusion feature vector sample of the ingredient component, and the overall fusion feature vector sample of the instruction component is the normalization factor of the recipe sample, used to adjust the deviation of the random sampling, and N is the batch size
[0102] Next, construct the negative matching loss function for each component, namely the title negative matching loss function , the ingredient negative matching loss function and the instruction negative matching loss function , respectively using the matching probability of the overall fusion feature vector of the title component to the global feature vector of the food image , the matching probability of the overall fusion feature vector of the ingredient component to the global feature vector of the food image , and the matching probability of the overall fusion feature vector of the instruction component to the global feature vector of the food image , construct the set of matching probabilities of the negative samples of the title component to the food image , the set of matching probabilities of the negative samples of the ingredient component samples to the food image , the set of matching probabilities of the negative samples of the instruction component samples to the food image , where N is the batch size , represents the matching probability of the th title component to the th negative sample of the food image represents the matching probability of the th ingredient component to the th negative sample of the food image represents the matching probability of the th instruction component to the The negative sample matching probability of a food image. Then construct the negative matching loss function from the title component to the food image and the negative matching loss function from the ingredient component to the food image and the negative matching loss function from the ingredient component to the food image The formula is as follows:
[0103] ;
[0104] ;
[0105] ;
[0106] Since it is a two-way retrieval, it is necessary to construct the negative matching loss function from the food image to the title component and the negative matching loss function from the food image to the ingredient component and the negative matching loss function from the food image to the instruction component , respectively using the global feature vector of the food image to the overall fusion feature vector of the title component matching probability , the global feature vector of the food image to the overall fusion feature vector of the ingredient component matching probability and the global feature vector of the food image to the overall fusion feature vector of the instruction component matching probability Construct the negative sample matching probability set from the food image to the title component , the negative sample matching probability set from the food image to the ingredient component , the negative sample matching probability set from the food image to the instruction component , N is the batch size, represents the negative sample matching probability of the food image to the th title component, represents the negative sample matching probability of the food image to the th ingredient component, represents the negative sample matching probability of the food image to the th instruction component recipe image negative sample. Then construct the negative matching loss function from the title component to the food image and the negative matching loss function from the ingredient component to the food image and the negative matching loss function from the ingredient component to the food image The formula is as follows:
[0107] ;
[0108] ;
[0109] ;
[0110] Add the loss functions in two directions to obtain the negative matching function of the title component , the negative matching function of the ingredient component , the negative matching function of the ingredient component The formula is as follows:
[0111] ;
[0112] ;
[0113] ;
[0114] Add the negative matching function of the title component , the negative matching function of the ingredient component , the negative matching function of the ingredient component to obtain the negative matching loss function of the component The formula is as follows:
[0115] ;
[0116] Step S33, embed each average food image segment after image segmentation into a vector to form a vector set and the overall fusion feature vector of each ingredient component to form a vector set , and calculate the similarity matrices of and and respectively. The specific calculation formula is as follows:
[0117] ;
[0118] ;
[0119] Use the triplet loss to minimize the difference between matrices. The calculation formula is:
[0120] ;
[0121] where is the cosine similarity vector, is the triplet loss margin parameter, is the anchor point is a positive sample, is a negative sample, is the transposed matrix, is the transposed matrix;
[0122] Step S34, add the three loss functions to obtain the final dynamic negative matching - triplet combined loss function :
[0123] + + .
[0124] Step S4, perform cross - modal recipe retrieval, and use the recall rate as the evaluation metric.
[0125] This example is mainly trained and validated on the Recipe1M dataset, which contains pairs of images and recipes. The officially divided training set, validation set, and test set contain 238,999, 51,119, and 51,303 pairs of samples respectively. To evaluate the retrieval performance of the model, we use the median rank (medR) and recall rates (Recall@1, Recall@5, Recall@10, abbreviated as R1, R5, R10 respectively) as metrics. These metrics are calculated on ranking lists of size N = {1,000, 10,000}, and the average results are reported on 10 randomly selected samples. And this case is compared with the classic model H - T, and the experimental results are shown in Table 1 and Table 2.
[0126] Table 1 N = 1000
[0127]
[0128] Table 2 N = 10000
[0129]
[0130] It should be understood that the parts not elaborated in detail in this specification belong to the prior art. The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary engineering and technical personnel in the art to the technical solution of the present invention should fall within the protection scope determined by the claims of the present invention.
Claims
1. A global and component fine-grained recipe retrieval method using segmentation and negative matching, characterized in that It includes the following steps: Step S1, using the Recipe1M dataset, establish a dataset of two modalities: food images and recipes, where the food image includes two parts: food and background, and the recipe includes three components: title, ingredients, and instructions. The dataset of food images and recipes is divided into a training set, a validation set, and a test set; Step S2, respectively using the ViT visual model and the cascaded Transformer model to extract features from the food images and recipes in the training set, validation set, and test set, specifically including: Step S21, extracting global features of food images; Step S22, extracting the image features of the main food after image segmentation; Step S23, extracting text features of the recipe; Step S3, constructing a loss function on the training set, specifically including: Step S31, calculating the matching probability between food images and recipes and between recipes and food images, and using the calculated matching probabilities to construct an overall negative matching loss function to increase the distance of unmatched samples; Step S32, calculating the bidirectional matching probability between each component of the three recipes and the food image, and using the calculated matching probability to calculate the negative matching loss function at the three component levels, further increasing the distance between unmatched samples; Step S33, using the segmentation model to segment the main food, calculating the similarity matrix of the segmented main food and the ingredient components in the recipe, and using triplet loss to reduce the matching error between the main food and the ingredient components of the recipe, thereby enhancing the robustness of cross-modal alignment; Step S34, constructing a final dynamic negative match-triplet joint loss function; Step S4, performing cross-modal recipe retrieval, using the recall rate as an evaluation indicator.
2. A global and component fine-grained recipe retrieval method using segmentation and negative matching as claimed in claim 1, characterized in that In the step S2: Step S21, extracting global features of food images: Food image collection Input to the ViT visual model to extract the global feature vector of food images , Indicates the use of ViT visual model, Indicates Food images, is the visual modality identifier, Indicates the number of food images; Step S22, extracting the image features of the main food after image segmentation: The SAM segmentation model is used to segment food images. First, a large number of masks are generated through the "everything" mode to cover various areas in the food image. Then, the area threshold is used to filter out small or meaningless areas, and the CLIP large model is used for semantic matching to retain masks related to food. Finally, the Top-N areas with the highest semantic scores are selected as enhanced data to extract the main food in the food image and remove background interference. The food image fragments formed after segmentation are Input to ViT visual model , average pooling is performed on the embeddings of all food image segments to obtain the average food image segment embedding vector , represents the number of food image segments, Indicates Food image clips; Step S23, extracting text features of the recipe: A cascaded Transformer model is used to process at the word level and sentence level. Given a recipe collection ,in, Represents the title component, Represents an ingredient component, Represents a command component, is a text identifier, Indicates the number of recipes; the title component , ingredient components and the directive component The first-layer Transformer model receives each word in the sentences of the three components and encodes them to capture the relationship between each word. Finally, the embedding representations of all words in each sentence are averaged through average pooling to obtain the sentence-level embeddings of the title component, ingredient component, and instruction component. = , Represents the first-layer Transformer model; the obtained sentence-level embeddings of the title component, ingredient component, and instruction component are sent to the second-layer Transformer to capture the relationship between each sentence in the title component, ingredient component, and instruction component. The sentence-level embeddings of the title component, ingredient component, and instruction component are averaged through average pooling to obtain the final overall fusion feature vector of the title component, ingredient component, and instruction component. = ,in It is the second-layer Transformer model; the overall fusion feature vectors of the title component, ingredient component, and instruction component are concatenated to obtain the final global feature vector of the recipe. , represents a fully connected layer.
3. A global and component fine-grained recipe retrieval method using segmentation and negative matching as claimed in claim 1, characterized in that In the step S3: Step S31: The global feature vector of the food image , the global feature vector of the recipe , using Monte Carlo methods to calculate the global cross-modal matching probability of food images to recipes , and the global cross-modal matching probability of recipes to food images The formula is as follows: ; ; in, Representing the global feature vector of food images to the global feature vector of the recipe The similarity between A global feature vector representing a recipe to the global feature vector of the food image The similarity between is the temperature parameter, which is used to control the smoothness of the probability distribution. is a subset randomly sampled from the training set, representing the randomly sampled recipe samples and food image samples, represents the normalization factor of the recipe sample, represents the normalization factor of the food image sample, which is used to adjust the bias of random sampling, and N is the batch size; Utilizes the global cross-modal matching probability of calculating recipes to food images The formula constructs the matching probability set of negative samples from recipes to food images = N is the batch size, For the Recipes to The matching probability of food images; then construct the overall negative matching loss function from recipe to food image , the formula is as follows: ; Constructing an overall negative matching loss function from food images to recipes , construction method and The loss function is consistent, using the global cross-modal matching probability of calculating food images to recipes The formula constructs the matching probability set of food images to recipe negative samples = N is the batch size, Indicates Food images to The matching probability of recipes; then construct the overall negative matching loss function from food images to recipes , the formula is as follows: ; And add the two together to get the overall negative matching loss function formula as follows: ; Step S32, calculating the overall fusion feature vector of the title component to the global feature vector of the food image The matching probability , the overall fusion feature vector of the batching components to the global feature vector of the food image The matching probability The overall fusion feature vector of the instruction component to the global feature vector of the food image The matching probability ; The formulas are as follows: ; ; ; in, Represents the overall fused feature vector of the title components to the global feature vector of the food image The similarity between Represents the overall fusion feature vector of the batching components to the global feature vector of the food image The similarity between Represents the overall fusion feature vector of the batching components to the global feature vector of the food image The similarity between is the temperature parameter, which is used to control the smoothness of the probability distribution. is a subset randomly sampled from the training set, representing the Kth randomly sampled food image sample, It is represented as the normalization factor of the food image sample, which is used to adjust the bias of random sampling, and N is the batch size; Compute global eigenvectors of food images to the overall fused feature vector of the title component The matching probability , the global feature vector of food images The overall fusion feature vector of the batching component The matching probability , the global feature vector of food images to the overall fused feature vector of the title component The matching probability The formula is as follows: ; ; ; in, Represents the global feature vector of the image to the overall fused feature vector of the title component The similarity of Represents the global feature vector of the image The overall fusion feature vector of the batching component The similarity of Represents the global feature vector of the image To the overall fused feature vector of the instruction component The similarity of is the temperature parameter, which is used to control the smoothness of the probability distribution. They are respectively represented as a subset randomly sampled from the training set, representing the The overall fusion feature vector samples of the randomly sampled title component, the overall fusion feature vector samples of the ingredient component, and the overall fusion feature vector samples of the instruction component. is the normalization factor of the recipe sample, which is used to adjust the bias of random sampling, and N is the batch size; Construct title negative matching loss function respectively , negative matching loss function And instruction negative matching loss function , respectively using the overall fusion feature vector of the title component to the global feature vector of the food image The matching probability , the overall fusion feature vector of the batching components to the global feature vector of the food image The matching probability The overall fusion feature vector of the instruction component to the global feature vector of the food image The matching probability Construct a set of matching probabilities from the title component to the negative samples of the food image , the set of negative sample matching probabilities from ingredient component samples to food images , the set of negative sample matching probabilities from instruction component samples to food images , N is the batch size, Indicates Title components to The negative sample matching probability of food images, Indicates Ingredients to The negative sample matching probability of food images, Indicates The instruction component to The negative sample matching probability of food images; then construct the negative matching loss function from the title component to the food image , negative matching loss function from ingredient components to food images , negative matching loss function from ingredient components to food images The formula is as follows: ; ; ; Constructing negative matching loss function for food image to title component , negative matching loss function from food image to ingredient component , food image to instruction component negative matching loss function , respectively using the global feature vector of the food image to the overall fused feature vector of the title component The matching probability , the global feature vector of food images The overall fusion feature vector of the batching component The matching probability and the global feature vector of the food image To the overall fused feature vector of the instruction component The matching probability Construct a set of negative sample matching probabilities from food images to title components , the set of negative sample matching probabilities from food images to ingredient components , the set of negative sample matching probabilities of food images to instruction components , N is the batch size, Indicates Food images to The negative sample matching probability of title components, Indicates Food images to The negative sample matching probability of the ingredient components, Indicates Food images to The negative sample matching probability of the instruction component recipe image; then construct the negative matching loss function from the title component to the food image , negative matching loss function from ingredient components to food images , negative matching loss function from ingredient components to food images The formula is as follows: ; ; ; Add the loss functions in both directions to get the negative matching function of the title component , ingredient component negative matching function , ingredient component negative matching function The formula is as follows: ; ; ; Negatively match the title component to the function , ingredient component negative matching function , ingredient component negative matching function Add together to get the component negative matching loss function The formula is as follows: ; Step S33, embed each average food image segment after image segmentation into a vector Composition vector set and the overall fusion feature vector of each ingredient component Composition vector set , respectively calculated The similarity matrix and The similarity matrix , the specific calculation formula is as follows: ; ; Using triplet loss To minimize The difference between matrices is calculated as: ; in is the cosine similarity vector, is the triplet loss boundary parameter, To draw a point, is a positive sample, is a negative sample; for Transpose the matrix, for Transpose a matrix; Step S34, add the three loss functions to obtain the final dynamic negative match-triplet joint loss function : + + 。
Citation Information
Patent Citations
Cross-modal pedestrian re-identification method based on inter-modal common semantic learning
CN118711217A
Recipe retrieval method based on modal interaction
CN119128052A