Menu recommendation method and device
By combining facial recognition technology, using open-set object detectors and multimodal image-text matchers to analyze user facial features, and combining image search frameworks and mapping databases, personalized and dynamic recipe recommendations are achieved, solving the problem of inaccurate recommendations in existing systems and improving user experience.
Patent Information
- Application Number
- CN202511149588.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-09-30
AI Technical Summary
Existing recipe recommendation systems lack personalization and dynamism, and are unable to make personalized adjustments based on the user's real-time status, resulting in inaccurate and limited relevance of recommendation results.
Combining facial recognition technology, the user's face and facial features are detected through an open-set object detector, false detections are eliminated using a multimodal image-text matcher, recipe recommendations are determined based on an image search framework and mapping database, and a multimodal large model is used for personalized and dynamic recommendations.
It realizes personalized recipe recommendations based on the user's real-time status, improves the accuracy and relevance of recommendations, and enhances the user experience.
Smart Images

Figure CN120723979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart home technology, and in particular to a recipe recommendation method and device. Background Art
[0002] Facial recognition technology uses computer vision and pattern recognition techniques to automatically identify or verify faces from images or videos. In recent years, the development of deep learning and neural networks has significantly improved the accuracy and efficiency of facial recognition technology. Facial recognition technology is widely used in security monitoring, identity verification, social media, human-computer interaction, and other fields.
[0003] A recipe recommendation system recommends suitable recipes based on user preferences, dietary habits, nutritional needs, and other information. Traditional recipe recommendation systems typically rely on user historical data, ratings, and tags. These systems are widely used in the catering industry, healthcare management, smart home, and other fields.
[0004] Currently, recipe recommendations based on facial recognition analysis are a way for users to improve their health. However, existing technologies have certain limitations, such as: 1. Lack of personalization: Traditional recipe recommendation systems typically rely on users' historical data and preferences, lacking real-time analysis of their current status and needs, resulting in insufficiently personalized recommendations. 2. Inability to dynamically adjust: Existing recipe recommendation systems struggle to dynamically adjust to users' real-time status (e.g., mood, health status, etc.), resulting in limited accuracy and relevance of recommendations. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a recipe recommendation method and device, so as to achieve more personalized and dynamic recipe recommendations by combining face recognition technology and a recipe recommendation system.
[0006] In a first aspect, an embodiment of the present invention provides a recipe recommendation method, which includes: inputting a user's facial image into an open set target detector to output detection results of the face and facial features; inputting the detection results into a multimodal image-text comparator to output verification results of the user's face and facial features; wherein the multimodal image-text comparator is used to remove false detections in the detection results; performing an image search on the verification results based on an image search framework to determine the attributes corresponding to the verification results; wherein the image search framework is equipped with a base library of targets of interest, and the attributes are used to characterize features of different dimensions of the face and facial features; determining at least one recipe corresponding to the attribute based on a pre-established mapping database of attributes and recipes; inputting the facial image and at least one recipe into a multimodal large model to output a recipe recommended to the user.
[0007] In an optional embodiment of the present application, the above attributes are used to characterize whether the user has blackheads, dark circles, bags under the eyes, acne, spots, large pores, skin texture and overall complexion.
[0008] In an optional embodiment of the present application, the above-mentioned step of inputting the user's facial image into an open set target detector and outputting the detection results of the face and facial features includes: inputting the user's facial image into an open set target detector, and the open set target detector detects the facial area and facial features area of the facial image as the detection results of the face and facial features; wherein the facial features area includes: forehead area, eye area, nose area, cheek area, mouth area and chin area.
[0009] In an optional embodiment of the present application, the above-mentioned open set object detector is fine-tuned and trained based on multiple face images with facial feature areas labeled.
[0010] In an optional embodiment of the present application, the above-mentioned step of inputting the detection results into a multimodal image-text comparator and outputting the verification results of the user's face and facial features includes: inputting the detection results and text into the multimodal image-text comparator, and outputting the similarity between the detection results and the text; wherein the text content includes: left eye, right eye, nose, left cheek, right cheek, mouth, chin and background; determining a target detection result whose similarity with the text with the content as the background is greater than a preset threshold; removing the target detection result from the detection result to obtain the verification results of the user's face and facial features.
[0011] In an optional embodiment of the present application, the above-mentioned multimodal image-text comparator is fine-tuned and trained based on multiple facial features images, face images and background images; wherein each facial features image, face image and background image is pre-annotated with corresponding text content.
[0012] In an optional embodiment of the present application, the above-mentioned step of performing an image search on the verification result based on the image search framework and determining the attributes corresponding to the verification result includes: inputting the verification result into a multimodal large model and outputting a text description of the verification result; and determining the attributes corresponding to the text description of the verification result based on the image search framework.
[0013] In an optional embodiment of the present application, the above method further includes: establishing a mapping database based on recipes constructed according to the nutritional elements of the ingredients and the relationship between the attributes and nutritional improvements.
[0014] In an optional embodiment of the present application, the above-mentioned step of inputting a facial image and at least one recipe into a multimodal big model and outputting a recipe recommended to the user includes: inputting a facial image and at least one recipe into the multimodal big model, and the multimodal big model determining a recipe recommended to the user from at least one recipe based on retrieval enhancement generation technology.
[0015] In a second aspect, an embodiment of the present invention further provides a recipe recommendation device, comprising: an open-set target detector processing module, configured to input a user's facial image into an open-set target detector and output detection results of the face and facial features; a multimodal image-text comparator processing module, configured to input the detection results into a multimodal image-text comparator and output verification results of the user's face and facial features; wherein the multimodal image-text comparator is configured to remove false detections in the detection results; an image search framework processing module, configured to perform an image search on the verification results based on the image search framework and determine attributes corresponding to the verification results; wherein the image search framework is equipped with a database of objects of interest, and the attributes are used to characterize features of different dimensions of the face and facial features; a mapping database processing module, configured to determine at least one recipe corresponding to the attributes based on a pre-established mapping database of attributes and recipes; The embodiments of the present invention bring the following beneficial effects: Embodiments of the present invention provide a recipe recommendation method and apparatus. These methods input a user's facial image into an open-set object detector, which outputs face and facial feature detection results. These detection results are then input into a multimodal image-text comparator, which outputs verification results of the user's facial and facial features. The multimodal image-text comparator is configured to remove false positives from the detection results. An image search is performed on the verification results based on an image search framework to determine attributes corresponding to the verification results. The image search framework includes a database of objects of interest, where attributes are used to characterize different dimensions of facial and facial features. At least one recipe corresponding to the attribute is determined based on a pre-established database of attribute-recipe mappings. The facial image and at least one recipe are then input into a multimodal large model, which outputs a recipe recommendation for the user. This method analyzes the user's real-time facial image to meet the user's real-time, personalized needs. Preliminary recipe recommendations are determined based on the facial and facial feature detection and verification results. The optimal recipe recommendation is then made based on the preliminary recipe recommendation information and the facial image information, achieving more personalized and dynamic recipe recommendations.
[0016] Other features and advantages of the present disclosure will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by practicing the above-mentioned technology of the present disclosure.
[0017] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 A flowchart of a recipe recommendation method provided by an embodiment of the present invention; Figure 2 A schematic diagram of the overall process of a recipe recommendation method provided by an embodiment of the present invention; Figure 3 A flowchart of another recipe recommendation method provided by an embodiment of the present invention; Figure 4 A schematic diagram of the structure of a recipe recommendation device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] Currently, recipe recommendations based on facial recognition analysis are a way for users to improve their health. However, existing technologies have certain limitations, such as: 1. Lack of personalization: Traditional recipe recommendation systems typically rely on users' historical data and preferences, lacking real-time analysis of their current status and needs, resulting in insufficiently personalized recommendations. 2. Inability to dynamically adjust: Existing recipe recommendation systems struggle to dynamically adjust to users' real-time status (e.g., mood, health status, etc.), resulting in limited accuracy and relevance of recommendations.
[0022] Based on this, embodiments of the present invention provide a recipe recommendation method and device, specifically providing a recipe recommendation algorithm based on facial recognition analysis. By combining facial recognition technology with a recipe recommendation system, more personalized and dynamic recipe recommendations can be achieved, specifically including: 1. Personalized recommendations, real-time analysis of user status, through face recognition technology, real-time analysis of facial features, given the user's health status, combined with the user's eating habits and preferences, to generate more personalized recipe recommendations; dynamic adjustment of recommendation results, according to the user's real-time status and needs, dynamically adjust the recipe recommendation results to ensure the accuracy and relevance of the recommended recipes.
[0023] 2. Improve recommendation accuracy by combining multi-dimensional information. By combining facial image information, facial features analysis, and the user's current dietary preferences, the accuracy and relevance of recipe recommendations can be improved from the three dimensions of images, text, and user needs. Reduce recommendation errors by analyzing the user's current health status in real time, reducing recommendation errors caused by historical data lags in traditional recommendation systems.
[0024] 3. Enhance user experience and provide real-time feedback. Analyze user feedback in real time through facial recognition technology, adjust recommendation results in a timely manner, and enhance user experience. Personalized interaction provides personalized interactive interfaces and recommended content based on the user's real-time health status and needs to improve user satisfaction.
[0025] In summary, by combining facial recognition analysis and a recipe recommendation system, the present invention can analyze the health status of a user's facial features, achieve more personalized and dynamic recipe recommendations, improve recommendation accuracy and user experience, and has broad application prospects and market value.
[0026] To facilitate understanding of this embodiment, a recipe recommendation method disclosed in an embodiment of the present invention is first introduced in detail.
[0027] Example 1: The present invention provides a recipe recommendation method. Figure 1 The flowchart of a recipe recommendation method shown in FIG. 1 includes the following steps: Step S102: Input the user's face image into an open set object detector and output the detection results of the face and facial features.
[0028] See Figure 2 The figure shows the overall process flow of a recipe recommendation method. In this embodiment, the user's facial image is first input into an open-set object detector to detect the face and corresponding facial features, and then output the face and facial features detection results. However, this detection result may contain redundant detection frames and false detections.
[0029] Open-set object detection is a technique that allows detection of dozens or even hundreds of target categories without limiting the scope of target categories. It can also identify new categories not seen during training. Unlike conventional object detectors, which can only detect a limited number of predefined categories, the open-set detection employed in this embodiment can simultaneously detect facial features.
[0030] In step S104, the detection results are input into a multimodal image-text comparator, which outputs verification results of the user's face and facial features; wherein the multimodal image-text comparator is used to remove false detections in the detection results.
[0031] like Figure 2As shown, in this embodiment, a multimodal image-text comparator can be applied to re-verify the detection results of the face and facial features, remove false detections in the detection results, and obtain the optimal face image and facial features image as the verification results of the user's face and facial features.
[0032] Step S106 , performing an image search on the verification result based on an image search framework to determine the attributes corresponding to the verification result; wherein the image search framework is equipped with a database of objects of interest, and the attributes are used to characterize features of different dimensions of the face and facial features.
[0033] like Figure 2 As shown, in this embodiment, an image search framework can be applied. First, the base database features of the target of interest are established, and then an image search is performed on the verification results. The attributes corresponding to the optimal target in the search result are used as the attributes corresponding to the verification results.
[0034] In some embodiments, the above attributes are used to characterize whether the user has blackheads, dark circles, bags under the eyes, acne, spots, large pores, skin texture and overall complexion.
[0035] Among them, interest refers to the facial attribute categories, which include eight dimensions of facial attributes (presence of blackheads, dark circles, eye bags, acne, spots, large pores, skin texture, and overall complexion).
[0036] The library refers to an image database. Each image represents one of the above facial attributes (such as whether there are dark circles). For example, one image shows the left eye with dark circles, and another shows the left eye without dark circles. And so on. The base library images include all facial attribute categories.
[0037] Step S108: determining at least one recipe corresponding to the attribute based on a pre-established mapping database of attributes and recipes.
[0038] like Figure 2 As shown, in this embodiment, a pre-built mapping database of attributes and recipes can be applied to output at least one recipe corresponding to the attribute recognition result.
[0039] Step S110: input the facial image and at least one recipe into the multimodal large model, and output the recipe recommended to the user.
[0040] like Figure 2 As shown, the multimodal large model in this embodiment can output the optimal recommended recipe based on the at least one recipe obtained in the previous step and combined with facial image information. The optimal recommended recipe output by the multimodal large model can be one of the at least one recipe obtained in the previous step, thereby achieving more personalized and dynamic recipe recommendations.
[0041] An embodiment of the present invention provides a recipe recommendation method. The method comprises inputting a user's facial image into an open-set object detector, which outputs face and facial feature detection results. The detection results are then input into a multimodal image-text comparator, which outputs verification results of the user's facial and facial features. The multimodal image-text comparator is configured to remove false positives from the detection results. An image search is then performed on the verification results based on an image search framework to determine attributes corresponding to the verification results. The image search framework is equipped with a database of objects of interest, where the attributes are used to characterize different dimensions of facial and facial features. At least one recipe corresponding to the attribute is determined based on a pre-established database of attribute-recipe mappings. The facial image and at least one recipe are then input into a multimodal large model, which outputs a recipe recommendation for the user. This method analyzes the user's real-time facial image to meet the user's real-time, personalized needs. Preliminary recipe recommendations are determined based on the facial and facial feature detection and verification results. The optimal recipe recommendation is then made based on the preliminary recipe recommendation information and the facial image information, achieving more personalized and dynamic recipe recommendations.
[0042] Example 2: This embodiment provides another recipe recommendation method, which is implemented on the basis of the above embodiment. Figure 3 A flowchart of another recipe recommendation method is shown, which includes the following steps: Step S302: Input the user's face image into an open set object detector and output the detection results of the face and facial features.
[0043] In some embodiments, the user's facial image can be input into an open-set target detector, which detects the facial area and facial feature areas of the facial image as the detection results of the face and facial features; wherein the facial feature areas include: forehead area, eye area, nose area, cheek area, mouth area and chin area.
[0044] This example focuses on the training optimization and inference application of an open-set object detector. The base model uses the Yolo-World open-set object detector, which primarily consists of the Yolov8 (You Only Look Oncev8) detector and the text branch of Clip (Contrastive Language-Image Pretraining, a pre-training method or model based on contrastive text-image pairs). In the inference application, a facial image can be fed into Yolo-World to detect the face and corresponding facial features, which include the forehead, eyes, nose, cheeks, mouth, and chin. However, this process may result in false detections.
[0045] In some embodiments, the open set object detector is fine-tuned and trained based on a plurality of facial images labeled with facial feature regions.
[0046] In actual Yolo-World applications, the detection of facial features in small areas is generally poor. This embodiment allows for fine-tuning and optimization. The steps are as follows: Step 1: Collect 1,000 facial images from different environments and angles; Step 2: Manually label the facial features using object detection and annotation tools (for example, the open-source image annotation tool labelme); Step 3: Fine-tune and optimize Yolo-World using this labeled facial feature data. This embodiment significantly improves Yolo-World's detection performance through fine-tuning and optimization.
[0047] In addition to the Yolo-World open set target detector, this embodiment can also use open set target detectors such as Grounding Dino and Dino-X.
[0048] In step S304, the detection results are input into a multimodal image-text comparator, which outputs the verification results of the user's face and facial features; wherein the multimodal image-text comparator is used to remove false detections in the detection results.
[0049] In some embodiments, the detection results and text can be input into a multimodal image-text comparator to output the similarity between the detection results and the text; wherein the text content includes: left eye, right eye, nose, left cheek, right cheek, mouth, chin and background; determine the target detection result whose similarity with the text with the content as the background is greater than a preset threshold; remove the target detection result from the detection result to obtain the verification result of the user's face and facial features.
[0050] Although the open set target detector in this embodiment can detect the face and facial features of a face image, there may be a small number of false detections in the detection results. These false detection images will have a negative impact on the final result when sent to subsequent steps. Taking this issue into consideration, in this embodiment, a multimodal image-text comparator can be used to remove false detections in the detection results and optimize the results of the facial features detection images. Among them, the text (including but not limited to text descriptions of eyes, nose, mouth, etc.) and the detection results can be sent to the multimodal image-text comparator BLIP (Bootstrapping Language-Image Pre-training) to remove false detections.
[0051] The above text content can include the left eye, right eye, nose, left cheek, right cheek, mouth, and chin, and these text contents correspond to the facial features image. This text can help eliminate false positives because BLIP is a multimodal image-text model, including an image branch, a text branch, and an image-text interaction branch. It can be used for image-text matching. The similarity between false positives and the text description of facial features is definitely lower than the similarity between the image and the text description of facial features. By matching the similarity, false positives can be eliminated. Furthermore, the text content can also include background, which corresponds to the false positives.
[0052] In some embodiments, the multimodal image-text comparator is fine-tuned and trained based on a plurality of facial features images, face images, and background images; wherein each facial features image, face image, and background image is pre-labeled with corresponding text content.
[0053] For example, for the application scenario of this embodiment, 1,000 images were initially collected, with correct facial features and text (eyes, mouth, etc.) and false positives and text (background) manually annotated. BLIP was then fine-tuned to improve the image-text comparison performance for this application scenario. BLIP training involves joint training of ITC (Image-Text Contrastive Learning) and ITM (Image-Text Matching).
[0054] Among them, image-text comparison means that the branches of image and text are opposite, and only contrast loss calculation training is performed at the end of the network; image-text matching refers to the image and text, which have fusion interaction in the middle layer of the network, and the last layer of the network is connected to a binary classification layer to determine whether the image and text match.
[0055] The ITC training first extracts the features of the image and text, and calculates the cosine similarity between the image and text. Finally, the contrastive loss is used for training. The training loss consists of two parts: the image-to-text contrast loss and the text-to-image contrast loss.
[0056] Image-to-text maps one image to many texts, allowing the model to learn how to match images with their corresponding text descriptions. This involves mapping from the image feature space to the text feature space and finding the most relevant text. Text-to-image maps one text to many images, a reverse process that allows the model to learn to find the corresponding image based on the text description. Bidirectional contrast loss allows for the image and text to be mapped in the feature space.
[0057] The contrastive loss formula for image to text is as follows:
[0058] The contrastive loss formula for text to image is as follows:
[0059] The total contrast loss is:
[0060] in, Is the temperature parameter, which is used to adjust the smoothness of the softmax function. is the total number of samples in the batch, Representing image features and text features The cosine similarity of is as follows:
[0061] in, and are the binary norms of image features and text features respectively.
[0062] The ITM part is consistent with the early ITC part. It extracts the features of the image and text, and then applies the Cross-Attention module to perform the fusion calculation of the image and text features. The image and text fusion features are then input into the linear layer and the cross entropy loss is used for binary classification training. The categories include mismatch and match.
[0063] During the inference application phase, the ITC image-text comparison branch can first be used to calculate the cosine similarity between the image and text. The ITM branch then combines the ITC similarity results to perform further ITM image-text matching, outputting the image-text matching similarity. Finally, the ITC and ITM similarity results are combined to select the facial features image with the greatest similarity as the test result, eliminating false positives from the test results. This embodiment uses a multimodal image-text aligner to select the optimal face and facial features images from the test results, ensuring the accuracy of subsequent attribute analysis.
[0064] Among them, ITC image-text comparison can obtain the image most similar to the text description; then, the above text description and the most similar image are sent to ITM for image-text matching to obtain the confidence of the image-text matching; finally, the image-text comparison similarity and image-text matching confidence are combined to obtain the final image-text similarity.
[0065] Step S306 , performing an image search on the verification result based on an image search framework to determine the attributes corresponding to the verification result; wherein the image search framework is equipped with a database of objects of interest, and the attributes are used to characterize features of different dimensions of the face and facial features.
[0066] In some embodiments, the verification result can be input into a multimodal large model, and a text description of the verification result can be output; and the attribute corresponding to the text description of the verification result can be determined based on the image search framework.
[0067] The image search framework in this embodiment primarily consists of the vision branch (Vision Transformer) of Clip-ReID (a method that applies a multimodal clip contrast loss pre-trained model to person re-ID tasks, improving image feature learning) and a feature base of interest. Before inference is applied, the model is trained using data consisting of eight-dimensional facial and facial feature images.
[0068] Detailed text descriptions, combined with image training, can help the network learn better and more generalizable image features. This embodiment also proposes that using more detailed text descriptions of images can help the image feature extraction network learn better and more generalizable image features. Therefore, the image feature extraction network of the image search framework is trained using the idea of Clip-ReID target re-identification. Target re-identification refers to accurately identifying the same target object under different camera perspectives or different time periods.
[0069] Specifically, the face and facial features images are first fed into a large multimodal model including but not limited to GPT4o, MiniCPM-V2.6, InternVL2 and other models to generate detailed text descriptions; then, the generated image-text pairs are input into the Clip-Reid network, the weights of the text feature extraction network are fixed, and the visual feature extraction branch is trained with the Contrastive Loss loss function to improve the image feature extraction capability of the visual branch; then the text branch is removed, and training data is constructed again from the eight-dimensional face and facial features images, with each attribute as a category ID, such as dark circles under the eyes for category one, large pores on the left side of the face for category two, and rosy complexion for category three, and so on, and the visual image feature extraction network is trained again with the cross-entropy loss function.
[0070] It's important to note that this example uses a large model, as it generates detailed and accurate image descriptions. Smaller models generate brief and inaccurate descriptions, which can affect training. Clip-reid has two branches: image and text. Detailed and accurate image text descriptions are fed into the text branch, which is then fixed, while only the image branch is trained. This training approach improves the image branch's ability to extract features, resulting in more generalized image features.
[0071] In the application inference phase, eight dimensions of facial and facial feature images are prepared as the base image database for image search. The attributes of these images are known. Features are extracted from the image to be identified and the base image database, and the images are sorted and output based on their cosine similarity. The attributes corresponding to the base image with the best match are used as the attributes of the image to be identified. The image search framework in this step outputs attribute results for the facial and facial feature images.
[0072] Step S308: determining at least one recipe corresponding to the attribute based on a pre-established attribute-recipe mapping database.
[0073] In some embodiments, a mapping database may be established based on recipes constructed according to the nutritional elements of ingredients and the relationship between attributes and nutritional improvements.
[0074] The aforementioned steps provide the relevant attributes of the face image and the corresponding facial features image through the image retrieval framework. This embodiment can establish a mapping database of facial features images and recipes in advance, which involves the knowledge system of medicine and nutrition. By investigating the nutritional elements of the ingredients to construct recipes with different functions, the connection between facial features and nutritional improvement is established. For example, if there are dark circles under the eyes, it is suitable to eat ingredients rich in vitamin C, vitamin K, and iron. The corresponding dish can be spinach, mushroom and chicken rolls for improvement. With the help of this facial features recipe mapping database, the attribute recognition results obtained in the aforementioned steps can be used to give relevant recipe recommendations from eight dimensions. Among them, the number of recipes given here can be multiple, that is, one or more recommended recipes can be given for each dimension.
[0075] Step S310: Input the facial image and at least one recipe into a multimodal macro model. The multimodal macro model determines a recipe to be recommended to the user from the at least one recipe based on retrieval enhancement generation technology.
[0076] Taking the example of eight recipe recommendations generated based on eight dimensions in the aforementioned steps, at the application level, eight-dimensional recommendations can lead to user selection issues, and the issue of recommendation priority warrants further resolution. Therefore, this embodiment introduces a multimodal macro model to provide optimal recipe recommendations through Retrieval-Augmented Generation (RAG). RAG is an AI technology that combines retrieval and generation, primarily used to enhance the performance of natural language processing (NLP) tasks. The RAG model retrieves relevant information before generating text, thereby improving the accuracy and relevance of the generated text. Drawing on RAG technology, the present invention uses the eight-dimensional recipe recommendations generated in the aforementioned steps as the RAG retrieval results, integrates facial and facial feature images, and uses them as the final prompt of the multimodal macro model, enhancing the effectiveness of the optimal recipe recommendation output by the multimodal macro model. For example, this embodiment can utilize the Qwen2-VL-7B model for the final recipe recommendation output. Therefore, the multimodal macro model can select the optimal recipe from the eight recommended recipes for recommendation.
[0077] The embodiments of the present invention specifically include: 1. This embodiment of the present invention provides an open set detection method for detecting faces and facial features. Using an open set detector, the entire face and facial features are simultaneously output for subsequent analysis, accelerating the algorithm's response speed.
[0078] 2. The present invention can use a multimodal image-text comparator to eliminate false detections of facial features. The image-text comparator can re-verify the accuracy of the detection results and improve the detection effect of the algorithm.
[0079] 3. This embodiment of the present invention designs an image search framework that builds a database of objects of interest and outputs the attributes of the objects to be identified. This framework can simultaneously analyze facial features and recognize recipes, achieving multiple tasks within a single framework.
[0080] 4. The present invention provides a multimodal image feature extractor that utilizes image-text comparison training to improve the generalization of image features.
[0081] 5. The present invention designs an overall algorithm similar to RAG search enhancement technology. By combining image and text information in a similar RAG-like manner, recipe recommendations can be made based on the level of coarseness and fineness.
[0082] In summary, this paper proposes a recipe recommendation algorithm based on facial recognition analysis. It recommends optimal recipes based on eight dimensions of facial attributes (presence of blackheads, dark circles, eye bags, acne, spots, enlarged pores, skin texture, and overall complexion). Specifically, this paper proposes an open-set detection algorithm to detect faces and facial features, a multimodal image-text comparator to eliminate false detections, and an image-to-image framework to output facial analysis results. This is combined with a face-recipe mapping database to output eight recipe recommendations. Finally, combined with image-text information, a technique similar to RAG (Retrieval Augmented Generation) is employed to leverage a large multimodal model to identify the optimal recipe.
[0083] Example 3: Corresponding to the above method embodiment, the present invention provides a recipe recommendation device, see Figure 4 The schematic diagram of the structure of a recipe recommendation device shown in FIG. 1 includes: An open-set object detector processing module 41 is configured to input a user's face image into an open-set object detector and output detection results of the face and facial features; The multimodal image-text comparator processing module 42 is used to input the detection results into the multimodal image-text comparator and output the verification results of the user's face and facial features; wherein the multimodal image-text comparator is used to remove false detections in the detection results; An image search framework processing module 43 is configured to perform an image search on the verification results based on an image search framework and determine attributes corresponding to the verification results; wherein the image search framework is equipped with a database of objects of interest, and the attributes are used to characterize different dimensions of human faces and facial features; a mapping database processing module 44 for determining at least one recipe corresponding to an attribute based on a pre-established mapping database of attributes and recipes; The multimodal large model processing module 45 is used to input the facial image and at least one recipe into the multimodal large model, and output the recipe recommended to the user.
[0084] An embodiment of the present invention provides a recipe recommendation device that inputs a user's facial image into an open-set object detector, which outputs face and facial feature detection results. The detection results are then input into a multimodal image-text comparator, which outputs verification results of the user's face and facial features. The multimodal image-text comparator is configured to remove false detections from the detection results. An image search is performed on the verification results based on an image search framework to determine attributes corresponding to the verification results. The image search framework is equipped with a database of objects of interest, where the attributes are used to characterize different dimensions of face and facial features. At least one recipe corresponding to the attribute is determined based on a pre-established database of attribute-recipe mappings. The facial image and at least one recipe are then input into a multimodal large model, which outputs a recipe recommendation for the user. This method analyzes the user's real-time facial image to meet the user's real-time, personalized needs. Preliminary recipe recommendations are determined based on the facial and facial feature detection and verification results. The optimal recipe recommendation is then made based on the preliminary recipe recommendation information and the facial image information, achieving more personalized and dynamic recipe recommendations.
[0085] The above attributes are used to characterize whether the user has blackheads, dark circles, eye bags, acne, spots, enlarged pores, skin texture, and overall complexion.
[0086] The above-mentioned open-set target detector processing module is used to input the user's facial image into the open-set target detector, and the open-set target detector detects the facial area and facial features area of the facial image as the detection results of the face and facial features; wherein the facial features area includes: forehead area, eye area, nose area, cheek area, mouth area and chin area.
[0087] The above open set object detector is fine-tuned and trained based on multiple face images with facial features annotated.
[0088] The above-mentioned multimodal image-text comparator processing module is used to input the detection results and text into the multimodal image-text comparator, and output the similarity between the detection results and the text; wherein the text content includes: left eye, right eye, nose, left cheek, right cheek, mouth, chin and background; determine the target detection result whose similarity with the text with the content as the background is greater than a preset threshold; eliminate the target detection result from the detection result to obtain the verification result of the user's face and facial features.
[0089] The multimodal image-text comparator is fine-tuned and trained based on multiple facial features images, face images, and background images, wherein each facial features image, face image, and background image is pre-labeled with corresponding text content.
[0090] The image search framework processing module is used to input the verification results into the multimodal large model and output a text description of the verification results; and determine the attributes corresponding to the text description of the verification results based on the image search framework.
[0091] The above-mentioned device also includes: a mapping database establishment module, which is used to establish a mapping database based on the recipes constructed according to the nutritional elements of the ingredients and the relationship between the attributes and nutritional improvement.
[0092] The multimodal large model processing module is used to input a facial image and at least one recipe into the multimodal large model, and the multimodal large model determines a recipe to recommend to the user from the at least one recipe based on retrieval enhancement generation technology.
[0093] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the recipe recommendation device described above can refer to the corresponding process in the aforementioned embodiment of the recipe recommendation method, and will not be repeated here.
[0094] Example 4: An embodiment of the present invention also provides an electronic device for running the above-mentioned recipe recommendation method; the electronic device includes a memory and a processor, wherein the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the above-mentioned recipe recommendation method.
[0095] Furthermore, the electronic device also includes a bus and a communication interface, and the processor, the communication interface and the memory are connected via the bus.
[0096] The memory may include high-speed random access memory (RAM) and non-volatile memory, such as at least one disk drive. The system network element communicates with at least one other network element via at least one communication interface (wired or wireless), which can be the Internet, a wide area network (WAN), a local area network (LAN), a metropolitan area network (MAN), etc. The bus may be an ISA bus, a PCI bus, or an EISA bus. Buses can be categorized as address buses, data buses, and control buses.
[0097] The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of the present invention may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or the like. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the method of the above embodiment in combination with its hardware.
[0098] An embodiment of the present invention also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned recipe recommendation method. The specific implementation can be found in the method embodiment, which will not be repeated here.
[0099] The computer program product of the recipe recommendation method and apparatus provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods in the previous method embodiments. For specific implementation, please refer to the method embodiments and will not be repeated here.
[0100] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the system and / or device described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0101] In addition, in the description of the embodiments of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0102] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0103] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0104] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A recipe recommendation method, characterized in that: The method comprises: Input the user's face image into the open set object detector and output the detection results of the face and facial features; Inputting the detection results into a multimodal image-text comparator to output verification results of the user's face and facial features; wherein the multimodal image-text comparator is used to remove false detections in the detection results; Performing an image search on the verification result based on an image search framework to determine attributes corresponding to the verification result; wherein the image search framework is equipped with a database of objects of interest, and the attributes are used to characterize features of different dimensions of a human face and facial features; determining at least one recipe corresponding to the attribute based on a pre-established mapping database of attributes and recipes; The facial image and the at least one recipe are input into a multimodal large model, and a recipe recommended to the user is output.
2. The method according to claim 1, characterized in that The attributes are used to characterize whether the user has blackheads, dark circles, eye bags, acne, spots, enlarged pores, skin texture, and overall complexion.
3. The method according to claim 1, characterized in that The steps of inputting the user's face image into the open set object detector and outputting the detection results of the face and facial features include: The user's facial image is input into an open set target detector, and the open set target detector detects the facial area and facial feature areas of the facial image as the detection results of the face and facial features; wherein the facial feature areas include: forehead area, eye area, nose area, cheek area, mouth area and chin area.
4. The method according to claim 3, characterized in that The open set object detector is fine-tuned and trained based on a plurality of face images with facial feature regions annotated thereon.
5. The method according to claim 1, wherein The step of inputting the detection result into a multimodal image-text comparator and outputting the verification result of the user's face and facial features includes: Inputting the detection result and text into a multimodal image-text comparator, and outputting the similarity between the detection result and the text; wherein the text content includes: left eye, right eye, nose, left cheek, right cheek, mouth, chin and background; Determine a target detection result whose similarity to the text as the background content is greater than a preset threshold; The target detection result is eliminated from the detection results to obtain the verification results of the user's face and facial features.
6. The method according to claim 5, characterized in that The multimodal image-text comparator is fine-tuned and trained based on a plurality of facial features images, face images and background images, wherein each facial features image, face image and background image is pre-labeled with corresponding text content.
7. The method according to claim 1, characterized in that The step of performing an image search on the verification result based on an image search framework to determine an attribute corresponding to the verification result includes: Inputting the verification results into the multimodal large model and outputting a text description of the verification results; An attribute corresponding to the text description of the verification result is determined based on an image search framework.
8. The method according to claim 1, characterized in that The method further comprises: The mapping database is established based on the recipes constructed according to the nutritional elements of the ingredients and the relationship between the attributes and nutritional improvement.
9. The method according to claim 1, characterized in that The step of inputting the facial image and the at least one recipe into a multimodal macro model and outputting a recipe recommended to the user comprises: The facial image and the at least one recipe are input into a multimodal big model, and the multimodal big model determines a recipe recommended to the user from the at least one recipe based on retrieval-enhanced generation technology.
10. A recipe recommendation device, characterized in that: The device comprises: An open-set object detector processing module, which is used to input the user's face image into the open-set object detector and output the detection results of the face and facial features; a multimodal image-text comparator processing module, configured to input the detection results into a multimodal image-text comparator and output verification results of the user's face and facial features; wherein the multimodal image-text comparator is configured to remove false detections in the detection results; An image search framework processing module, configured to perform an image search on the verification result based on an image search framework, and determine attributes corresponding to the verification result; wherein the image search framework is equipped with a database of objects of interest, and the attributes are used to characterize features of different dimensions of a human face and facial features; a mapping database processing module, configured to determine at least one recipe corresponding to the attribute based on a pre-established mapping database of attributes and recipes; The multimodal large model processing module is used to input the facial image and the at least one recipe into the multimodal large model and output a recipe recommended to the user.
Citation Information
Patent Citations
Information pushing method and related products
CN109146605A
Menu recommendation method, device, storage medium and cooking equipment
CN110797105A
Dish recommendation method and device based on face recognition and storage medium
CN111144967A
Dish recommendation method and device, server, electronic equipment and storage medium
CN111161035A
Key point positioning method and device based on open set target detection and storage medium
CN119169688A