Food value generation method, device, electronic device and computer-readable medium
The method addresses inaccuracies in food value determination by integrating image and semantic information through preprocessing and trained models, achieving precise food value generation.
Patent Information
- Application Number
- CN202411659581.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-20
AI Technical Summary
In the prior art, the image recognition model cannot logically match the food determination logic during the process of determining the value of food, resulting in errors in determining the value information and wasted resources, which requires manual coordination.
By obtaining food shooting images and semantic division information, the image encoding model, semantic encoding model and decoding model are used for feature extraction and semantic segmentation to generate accurate food value information.
It realizes accurate generation of food value information based on food semantic segmentation information, reduces resource waste and manual intervention, and improves the accuracy of determination.
Smart Images

Figure CN119152497B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a food value generation method, apparatus, electronic device, and computer-readable medium. Background Art
[0002] Currently, with the continuous development of the restaurant field, existing automatic services for meals are becoming increasingly intelligent. For determining the value information of the food corresponding to a user, the commonly adopted method is as follows: First, a food recognition is performed through an image recognition model. Then, based on the recognized food, the value information of the recognized food is determined through a pre-set food value determination logic. Finally, the value information of the food is sent to the user corresponding to the food.
[0003] However, when the above method is used to determine the value information, the following technical problems often exist:
[0004] There may be different determination logics for determining the value of food. The image recognition model can only realize the recognition of food and cannot be matched with the food determination logic, resulting in possible confusion in the selection of the final food value determination logic and errors in the determination of food value information. A large amount of recognition resources are wasted, and relevant technical personnel are required to participate in the coordination and processing.
[0005] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0006] This content part of the present disclosure is used to briefly introduce the concepts, which will be described in detail in the following detailed implementation part. This content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure propose a food value generation method, apparatus, electronic device, and computer-readable medium to solve one or more of the technical problems mentioned in the above background art section.
[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating food value, including: obtaining a currently acquired food captured image and food semantic segmentation information; performing image preprocessing on the food captured image to generate a first preprocessed image; inputting the first preprocessed image into a pre-trained image encoding model to generate first image encoding information; inputting the food semantic segmentation information into a pre-trained semantic encoding model to generate food semantic encoding information; splicing the food semantic encoding information and the first image encoding information to generate splicing information; inputting the splicing information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information; and generating food value information corresponding to the food captured image according to the food semantic segmentation information and the food semantic segmentation information.
[0009] In a second aspect, some embodiments of the present disclosure provide a food value generation device, including: an obtaining unit configured to obtain a currently acquired food captured image and food semantic segmentation information; a preprocessing unit configured to perform image preprocessing on the food captured image to generate a first preprocessed image; a first input unit configured to input the first preprocessed image into a pre-trained image encoding model to generate first image encoding information; a second input unit configured to input the food semantic segmentation information into a pre-trained semantic encoding model to generate food semantic encoding information; a splicing unit configured to splice the food semantic encoding information and the first image encoding information to generate splicing information; a third input unit configured to input the splicing information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information; and a generating unit configured to generate food value information corresponding to the food captured image according to the food semantic segmentation information and the food semantic segmentation information.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.
[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the food value generation method of some embodiments of the present disclosure, on the basis of accurately generating food semantic segmentation information, the value information of the corresponding food can be accurately generated. Specifically, the reason for the inaccurate value information of the relevant food is that there may be different determination logics for the value of the food. The image recognition model can only achieve the recognition of the food and cannot be matched with the determination logic of the food, resulting in possible confusion in the selection of the determination logic for the value of the food, and thus errors in the determination of the food value information. A large amount of recognition resources are wasted, and relevant technical personnel are required to participate in coordination and processing. Based on this, in the food value generation method of some embodiments of the present disclosure, first, the currently acquired food capture image and food semantic division information are obtained, so as to determine the food value information by obtaining food information and food plate information, etc. through the food capture image. The division of the food corresponding to the food capture image is determined through the food semantic division information, which is convenient for the determination of the food value information. Then, the above-mentioned food capture image is subjected to image preprocessing to generate a first preprocessed image to improve the image quality of the image, which is convenient for the subsequent extraction of image feature information. Next, the above-mentioned first preprocessed image is input into a pre-trained image encoding model to generate first image encoding information, and the feature semantic content related to the food in the image is extracted. Then, the above-mentioned food semantic division information is input into a pre-trained semantic encoding model to generate food semantic encoding information, and the feature semantic content related to how the food is divided in the food semantic division information is extracted. Secondly, the above-mentioned food semantic encoding information and the above-mentioned first image encoding information are spliced to generate splicing information to fuse the food semantic feature content and the image feature semantic content. Furthermore, the above-mentioned splicing information is input into a pre-trained decoding model for semantic segmentation to accurately generate food semantic segmentation information. Here, through the decoding model, the feature semantics corresponding to the food semantic encoding information and the first image encoding information can be fully considered to accurately perform food semantic segmentation processing. Finally, according to the above-mentioned food semantic segmentation information and the above-mentioned food semantic division information, the food value information corresponding to the above-mentioned food capture image can be accurately generated. In summary, through the image encoding model, the semantic encoding model and the decoding model, the extraction of the feature semantic content of the food capture image, the extraction of the feature of the food semantic content of the semantic encoding model, and the food semantic segmentation processing based on the fused feature information can be realized, so that accurate food value information can be obtained subsequently. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of a food value generation method according to the present disclosure;
[0015] Figure 2 is a schematic structural diagram of some embodiments of a food value generation apparatus according to the present disclosure;
[0016] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure;
[0017] Figure 4 is a schematic diagram of first semantic partition information, second semantic partition information, third semantic partition information, and fourth semantic partition information. Detailed implementation manners
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0019] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or mutual dependence relationship of the functions performed by these devices, modules, or units.
[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0024] Refer to Figure 1 , which shows a process 100 of some embodiments of a food value generation method according to the present disclosure. The food value generation method includes the following steps:
[0025] Step 101: Obtain the currently acquired food captured image and food semantic segmentation information.
[0026] In some embodiments, the execution subject of the above food value generation method (e.g., an electronic device) may obtain the currently acquired food captured image and food semantic segmentation information through a wired connection or a wireless connection. Among them, the food captured image may be an image captured for the corresponding food at the target time point. There is a corresponding serving plate (i.e., a plate) for the corresponding food. The food semantic segmentation information may be information characterizing how to perform semantic segmentation on the food. In practice, the food semantic segmentation information may be a classification rule characterizing how to classify the food in the corresponding situation. The target time point may be the time point corresponding to the target location. That is, when the food acquisition object brings the food to the target location, corresponding shooting processing will be performed to obtain the food captured image. The food semantic segmentation information may be information in text form.
[0027] In some alternative implementations of some embodiments, there is a corresponding food division label for each food plate in the above food captured images. Among them, the food division label can represent the division criteria for the food corresponding to the food plate. Among them, the food division label can be the standard information pasted on each food plate for determining how each food in the food plate is divided. For example, the food division label can be: the first division label, the second division label, the third division label, and the fourth division label. The first division label corresponds to the first semantic division information. The second division label corresponds to the second semantic division information. The third division label corresponds to the third semantic division information. The fourth division label corresponds to the fourth semantic division information. The first semantic division information can represent the division information determined based on the value of the food plate. Specifically, the first semantic division information can be the division information representing the value determination based on the shape of the food plate corresponding to the food plate. In practice, the food value can be the food price. The food value can be the food price that the food object should pay during the food purchase process. For example, for the first semantic division information, the food plate shapes include: square shape, triangular shape, and circular shape. The food price corresponding to the square shape can be 10 yuan. The food price corresponding to the triangular shape can be 15 yuan. The food price corresponding to the circular shape can be 20 yuan. The second semantic division information can represent the division information determined based on the food type without counting the quantity. Specifically, the second semantic division information can be the division information representing the value determination based on the food type without counting the quantity. The food type can be the type of food dishes. Different food dishes have different prices. And the second semantic division information does not calculate the quantity of the food dishes. For example, for the second semantic division information, the corresponding food types can include: scrambled eggs with tomatoes, stir-fried pork with peppers, and spicy chicken cubes. The price of scrambled eggs with tomatoes is 14 yuan. The price of stir-fried pork with peppers is 24 yuan. The price of spicy chicken cubes is 29 yuan. The third semantic division information can represent the division information determined based on the food type and by quantity. Specifically, the third semantic division information can represent the division information determined based on the food type and by quantity. For example, for the third semantic division information, the corresponding food types can include: scrambled eggs with tomatoes, stir-fried pork with peppers, spicy chicken cubes, 2 eggs, and 4 steamed buns. The price of scrambled eggs with tomatoes is 14 yuan. The price of stir-fried pork with peppers is 24 yuan. The price of spicy chicken cubes is 29 yuan. The price corresponding to each egg is 2 yuan, and the price corresponding to 2 eggs is 4 yuan. The price of 1 steamed bun is 0.5 yuan. The price corresponding to 4 steamed buns is 2 yuan. The fourth semantic division information can represent the division information determined based on the fourth semantic division information of the food ingredients. Specifically, the fourth semantic division information can represent the division information determined based on whether the food is a meat-based food or a vegetarian food. The food ingredients can include: food ingredients including meat items and food ingredients not including meat items.For example, for the fourth semantic partition information, the corresponding food components include meat dishes, with 1 meat dish and 2 vegetarian dishes, and the value is 14 yuan. Including meat dishes, 2 meat dishes and 1 vegetarian dish, the value is 18 yuan. Including meat dishes, 3 meat dishes and 1 vegetarian dish, the value is 25 yuan.
[0028] As Figure 4 shown, a schematic diagram showing the first semantic partition information, the second semantic partition information, the third semantic partition information, and the fourth semantic partition information is presented.
[0029] Optionally, the above food semantic partition information is generated through the following steps:
[0030] In the first step, obtain the respective label description feature information corresponding to each food partition label from the target database. Among them, the label description feature information includes: label shape feature information, label position feature information, and label font feature information. Among them, the target database can be a database storing all food partition labels. The all food partition labels can be all complete food partition labels. There is a one-to-one correspondence between the food partition labels in each food partition label and the label description feature information in each label description feature information. The label description feature information can be information in vector form. The label description feature information can represent the descriptive feature semantic content of the label. The label shape feature information can be information in vector form that represents the feature semantic content of the corresponding shape of the label. The label shape feature information can be information in vector form that represents the feature semantic content of the position of the label on the food plate. The label font feature information can be information in vector form that represents the feature semantic content of the corresponding font of the label.
[0031] In the second step, input the above food captured image into the target detection model to generate at least one detection feature information for at least one initial detection object. Among them, the target detection model can be a neural network model that detects labels as target detection objects. In practice, the target detection model can be the YOLO V5 model + convolutional layer. There is a one-to-one correspondence between the initial detection objects in at least one initial detection object and the detection feature information in at least one detection feature information. The initial detection object can be a detection object that the target detection model initially predicts as a label. The initial detection object can be an object to be determined as a label. The detection feature information can represent the feature semantic content corresponding to the initial detection object. The detection feature information can be information in vector form.
[0032] As an example, the above execution entity can input the above food captured image into the YOLO V5 model to generate at least one detection information corresponding to at least one initial detection object. Then, input at least one detection information into the convolutional layer to generate at least one detection feature information.
[0033] Step 3, for each of the at least one detection feature information above, perform the following first generation step:
[0034] Sub-step 1, input the above detection feature information into a pre-trained label shape feature extraction model to generate initial label shape feature information. Among them, the label shape feature extraction model can be a neural network model for extracting label shape feature information. In practice, the label shape feature extraction model can be the first predetermined number of convolutional layers + fully connected layers connected after the object detection model. The initial label shape feature information can be the feature information related to the label shape in the detection feature information. For example, the first predetermined number of layers can be 11 layers.
[0035] Sub-step 2, input the above detection feature information into a pre-trained label position feature extraction model to generate initial label position feature information. Among them, the label position feature extraction model can be a neural network model for extracting label position feature information. In practice, the label position feature extraction model can be the second predetermined number of convolutional layers + fully connected layers connected after the object detection model. The initial label position feature information can be the feature information related to the label position in the detection feature information. For example, the second predetermined number of layers can be 11 layers.
[0036] Sub-step 3, input the above detection feature information into a pre-trained label font feature extraction model to generate initial label font feature information. Among them, the label font feature extraction model can be a neural network model for extracting label font feature information. In practice, the label font feature extraction model can be the third predetermined number of convolutional layers + fully connected layers connected after the object detection model. The initial label font feature information can be the feature information related to the label font in the detection feature information. For example, the third predetermined number of layers can be 8 layers.
[0037] Sub-step 4, compare the above label shape feature information with each of the initial label shape feature information to generate first feature comparison information. Among them, the first feature comparison information can represent the feature similarity between the label shape feature information and the initial label shape feature information. For example, the first feature comparison information can be the cosine similarity. The first feature comparison information can be a value between 0 and 1. The higher the value, the higher the feature similarity between the label shape feature information and the corresponding initial label shape feature information.
[0038] Sub-step 5: Compare the above-mentioned label position feature information with each initial label position feature information to generate second feature comparison information. The second feature comparison information can characterize the feature similarity between the label position feature information and the initial label position feature information. For example, the second feature comparison information can be cosine similarity. The second feature comparison information can be a value between 0 and 1. The higher the value, the higher the feature similarity between the label position feature information and the corresponding initial label position feature information.
[0039] Sub-step 6: Compare the above-mentioned label font feature information with each initial label font feature information to generate third feature comparison information. The third feature comparison information can characterize the feature similarity between the label font feature information and the initial label font feature information. For example, the third feature comparison information can be cosine similarity. The third feature comparison information can be a value between 0 and 1. The higher the value, the higher the feature similarity between the label font feature information and the corresponding initial label font feature information.
[0040] Sub-step 7: Determine whether the initial detection object corresponding to the above-mentioned detection feature information is the target detection object according to the above-mentioned first feature comparison information, the above-mentioned second feature comparison information, and the above-mentioned third feature comparison information.
[0041] As an example, the above-mentioned execution entity can determine the average similarity corresponding to the first feature comparison information, the second feature comparison information, and the third feature comparison information. Finally, in response to determining that the average similarity is greater than or equal to the target value, it is determined that the initial detection object corresponding to the above-mentioned detection feature information is the target detection object. In response to determining that the average similarity is less than the target value, it is determined that the initial detection object corresponding to the above-mentioned detection feature information is not the target detection object.
[0042] Sub-step 8: Input the at least one target detection feature information corresponding to the above-mentioned at least one target detection object into a pre-trained food classification label content generation model to generate at least one food classification label content. The food classification label content generation model can be a neural network model for generating food classification label content. There is a one-to-one correspondence between the food classification label content in the at least one food classification label content and the target detection object in the at least one target detection object. The food classification label content can be the label content corresponding to the food classification label. The food classification label content generation model can be the decoding model of a multi-output model. In practice, the food classification label content generation model can include: a downsampling model + multiple cascaded convolutional layers. The downsampling model can be a pyramid model.
[0043] Fourth step: Perform deduplication processing on the above-mentioned at least one food classification label content to generate at least one semantic classification information as the above-mentioned food semantic classification information.
[0044] Step 102: Perform image preprocessing on the above food captured image to generate a first preprocessed image.
[0045] In some embodiments, the above execution subject may perform image preprocessing on the above food captured image to generate a first preprocessed image. Among them, the image preprocessing may be, but is not limited to, at least one of the following: initial image cropping, image size adjustment, and image sharpening processing.
[0046] As an example, the above execution subject may use image processing software to perform image preprocessing on the above food captured image to generate a first preprocessed image.
[0047] Step 103: Input the above first preprocessed image into a pre-trained image encoding model to generate first image encoding information.
[0048] In some embodiments, the above execution subject may input the above first preprocessed image into a pre-trained image encoding model to generate first image encoding information. Among them, the image encoding model may be a neural network model for content encoding processing of semantic content in an image. In practice, the image encoding model may be a series of convolutional layers. For example, the image encoding model may be 11 series-connected convolutional layers for downsampling processing. The first image encoding information may be information in vector form representing the corresponding feature semantic content of the corresponding first preprocessed image.
[0049] Step 104: Input the above food semantic division information into a pre-trained semantic encoding model to generate food semantic encoding information.
[0050] In some embodiments, the above execution subject may input the above food semantic division information into a pre-trained semantic encoding model to generate food semantic encoding information. Among them, the semantic encoding model may be a neural network model for information encoding processing of semantic division information. The food semantic encoding information may be information in vector form representing the information obtained by encoding the food semantic content. In practice, the semantic encoding model may be a Bert pre-trained model.
[0051] In some alternative implementation manners of some embodiments, the above food semantic division information includes at least one of the following: first semantic division information based on a food plate, second semantic division information based on food types without counting the quantity, third semantic division information based on food types and determined by quantity, and fourth semantic division information based on food ingredients.
[0052] Optionally, the above-mentioned execution entity may input at least one of the above-mentioned first semantic division information, the above-mentioned second semantic division information, the above-mentioned third semantic division information, and the above-mentioned fourth semantic division information into the above-mentioned semantic encoding model to generate food semantic encoding information.
[0053] Step 105: Concatenate the above-mentioned food semantic encoding information and the above-mentioned first image encoding information to generate concatenated information.
[0054] In some embodiments, the above-mentioned execution entity may concatenate the above-mentioned food semantic encoding information and the above-mentioned first image encoding information to generate concatenated information.
[0055] Step 106: Input the above-mentioned concatenated information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information.
[0056] In some embodiments, the above-mentioned execution entity may input the above-mentioned concatenated information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information. Among them, the food semantic segmentation information may be information after semantic segmentation processing of food. The decoding model may be a neural network for upsampling processing in multiple layers in series. For example, the decoding model may be a convolutional layer for upsampling processing in 16 layers in series. It should be noted that the image encoding model, the semantic encoding model, and the decoding model are trained synchronously.
[0057] In the process of adopting technical solutions to solve the technical problems mentioned in the background art, the following problems often accompany: The decoding model cannot implement feature processing for the concatenated information, resulting in limited feature information extracted by the decoding model and inaccurate food semantic segmentation information generated. To address these problems, combined with the advantages / technical status of the company where the inventor is located, the following implementation is adopted to enable the decoding model to fully utilize each feature information included in the concatenated information.
[0058] In some alternative implementation manners of some embodiments, inputting the above-mentioned concatenated information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information includes:
[0059] In the first step, input the above splicing information into the attention mechanism model included in the above decoding model to generate semantic splicing information representing the first image coding information and the target semantic segmentation feature information. Among them, the splicing information may include: the first image coding information, the first semantic segmentation feature information corresponding to the first semantic segmentation information, the second semantic segmentation feature information corresponding to the second semantic segmentation information, the third semantic segmentation feature information corresponding to the third semantic segmentation information, and the fourth semantic segmentation feature information corresponding to the fourth semantic segmentation information. The target semantic segmentation feature information may be the semantic segmentation feature information among the first semantic segmentation feature information, the second semantic segmentation feature information, the third semantic segmentation feature information, and the fourth semantic segmentation feature information that is most similar in features to the first image coding information. The attention mechanism model is used to filter out the semantic segmentation feature information that best matches the feature information of the first image coding information. In practice, the attention mechanism model may be a multi-head attention mechanism model.
[0060] In the second step, input the above semantic splicing information into the feature fusion model to generate the first fusion feature information. The feature fusion model may be a multi-layer cascaded convolutional layer, which is a model used for feature information fusion and feature information extension of food semantic coding information and target semantic segmentation feature information.
[0061] In the third step, input the above first fusion feature information into the food semantic segmentation layer included in the above decoding model to output the first initial food semantic segmentation information and the first segmentation score. Among them, the food semantic segmentation layer may be a network layer for semantic segmentation of food. In practice, the food semantic segmentation layer may be a U-net model. The first segmentation score may represent the segmentation accuracy corresponding to the first initial food semantic segmentation information. The higher the first segmentation score, the higher the corresponding segmentation accuracy.
[0062] In the fourth step, input the above first image coding information and the set of remaining semantic segmentation feature information into the feature fusion model to generate the second fusion feature information. Among them, the set of remaining semantic segmentation feature information may be the information set obtained by removing the target semantic segmentation feature information from the first semantic segmentation feature information, the second semantic segmentation feature information, the third semantic segmentation feature information, and the fourth semantic segmentation feature information.
[0063] In the fifth step, input the above second fusion feature information into the food semantic segmentation layer included in the above decoding model to output the second initial food semantic segmentation information and the second segmentation score. The second segmentation score may represent the segmentation accuracy corresponding to the second initial food semantic segmentation information. The higher the second segmentation score, the higher the corresponding segmentation accuracy.
[0064] In the sixth step, determine whether the above first initial food semantic segmentation information and the above second food semantic segmentation information are consistent.
[0065] Step 7, in response to determining the inconsistency, determine whether the subtraction score corresponding to subtracting the second segmentation score from the first segmentation score is greater than the target score. The target score can be a preset score. For example, the target score can be 40.
[0066] Step 8, in response to determining that it is greater than the target score, determine the first initial food semantic segmentation information as the food semantic segmentation information.
[0067] Step 9, in response to determining that it is not greater than the target score, input the splicing information corresponding to the first image coding information and the remaining semantic segmentation feature information set into the attention mechanism model to generate the second semantic splicing information representing the first image coding information and its secondary semantic segmentation feature information. The secondary semantic segmentation feature information can be the remaining semantic segmentation feature information in the remaining semantic segmentation feature information set with the highest feature similarity to the first image coding information.
[0068] Step 10, perform splicing information fusion on the second semantic splicing information and the semantic splicing information to generate splicing information.
[0069] Step 11, input the splicing information into the feature fusion model to generate the third fusion feature information. The relevant explanation of the third fusion feature information can refer to the explanation of the second fusion feature information.
[0070] Step 12, remove the secondary semantic segmentation feature information from the remaining semantic segmentation feature information set to generate the post-removal feature information set.
[0071] Step 13, input the post-removal feature information set and the first image coding information into the feature fusion model to generate the fourth fusion feature information.
[0072] Step 14, input the third fusion feature information into the food semantic segmentation layer to generate the third initial food semantic segmentation information and the third segmentation score.
[0073] Step 15, input the fourth fusion feature information into the food semantic segmentation layer to generate the fourth initial food semantic segmentation information and the fourth segmentation score.
[0074] Step 16, in response to determining that the third initial food semantic segmentation information is different from the fourth initial food semantic segmentation information and the third segmentation score is higher than the fourth segmentation score, determine the first initial food semantic segmentation information as the food semantic segmentation information.
[0075] The above technical solution and its related content, as an inventive point of an embodiment of the present disclosure, solve the technical problem that the decoding model cannot implement feature processing for the spliced information, resulting in limited feature information extracted by the decoding model and inaccurate food semantic segmentation information generated. Based on this, the present disclosure uses the attention mechanism model, food semantic segmentation layer, and feature fusion model included in the decoding model to extract more accurate semantic segmentation feature information, and then generates food semantic segmentation information, making the obtained food semantic segmentation information more accurate.
[0076] Step 107: Generate food value information corresponding to the food captured image according to the above food semantic segmentation information and the above food semantic division information.
[0077] In some embodiments, the above execution subject may generate food value information corresponding to the food captured image according to the above food semantic segmentation information and the above food semantic division information. Among them, the food value information may be the value information of the food corresponding to the food captured image. In practice, the food value information may be the payable amount for the corresponding food.
[0078] In some optional implementation manners of some embodiments, the above food semantic segmentation information includes: at least one sub-divided image. Among them, each sub-divided image may represent each sub-image after food semantic division of the target food image.
[0079] Optionally, the generating food value information corresponding to the food captured image according to the above food semantic segmentation information and the above food semantic division information includes:
[0080] First step: In response to determining that the above food semantic division information is the above first semantic division information, determine the image size information corresponding to the above at least one sub-divided image, and obtain at least one image size information. Among them, there is a one-to-one correspondence between the sub-divided image in the at least one sub-divided image and the image size information in the at least one image size information. The image size information may represent the image size corresponding to the sub-divided image.
[0081] Second step: Obtain a first value table from the target database, which represents the correspondence between the image size information and the value information for the above first semantic division information. In practice, for the dining scenario, the value information may be the price. The first value table may be an association table representing the correspondence between the image size and the price.
[0082] Third step: According to the above first value table, determine the value information corresponding to each image size information in the above at least one image size information, and obtain at least one first value information.
[0083] Step 4: Add up each of the at least one first value information in the above to generate added information as food value information.
[0084] Step 5: In response to determining that the above food semantic division information is the above second semantic division information, input each of the at least one divided sub-image into a pre-trained first food type information generation model to generate first food type information, obtaining at least one first food type information, where the first food type information includes: food type. The first food type information generation model can be a neural network model for generating food type information. The food type information can be food dish information. In practice, the first food type information generation model can be an object detection model with the food type (i.e., food dish) as the detection object. In practice, the first food type information generation model can include: a multi-layer cascaded convolutional layer + YOLO V5 model.
[0085] Step 6: Obtain a second value table from the above target database, and according to the second value table, determine the value information corresponding to each of the at least one first food type information in the above to obtain at least one second value information. Among them, the second value table can represent the corresponding relationship between the food type and the corresponding value information (for example, price). For example, the second value table can represent the corresponding relationship between each food dish and the corresponding price.
[0086] Step 7: Add up each of the at least one second value information in the above to generate added information as food value information.
[0087] Step 8: In response to determining that the above food semantic division information is the above third semantic division information, input each of the at least one divided sub-image into a pre-trained second food type information generation model to generate second food type information, obtaining at least one second food type information. Among them, the second food type information includes: food type and food quantity. Among them, the second food type information generation model can be a neural network model for generating food type information. The food type information can be food dish information. In practice, the second food type information generation model can be an object detection model with the food type (i.e., food dish) as the detection object. In practice, the second food type information generation model can include: a multi-layer cascaded convolutional layer + YOLO V5 model. The number of convolutional layers corresponding to the second food type information generation model and the number of convolutional layers corresponding to the first food type information generation model can be different.
[0088] Step 9: According to the above second value table, determine the value information corresponding to each second food type information in the at least one second food type information, and obtain at least one third value information. Among them, the third value information is the value information obtained by multiplying the value corresponding to the food type corresponding to the second food type information by the number of foods. The third value information can be the price corresponding to the food type.
[0089] Step 10: Add up the third value information in the at least one third value information to generate an added information, which is used as the food value information.
[0090] Step 11: In response to determining that the above food semantic segmentation information is the above fourth semantic segmentation information, input each sub-image in the at least one sub-image into a pre-trained food ingredient information generation model to generate food ingredient information, and obtain at least one food ingredient information. Among them, the food ingredient information generation model can be a meat food recognition model. In practice, the meat food recognition model can also be a neural network model with meat food as the recognition object. The food ingredient information can represent the cost information of whether each dish is a meat dish.
[0091] Step 12: Obtain from the above target database a pre-stored third value table that represents the relationship between food ingredient combinations and value information. The third value table can represent the association relationship between food ingredients and prices under various types. The food ingredients under various types can include: food ingredients representing one meat and two vegetables, food ingredients representing two meats and one vegetable, food ingredients representing two meats and two vegetables, food ingredients representing three meats, and food ingredients representing three vegetables.
[0092] Step 13: According to the above third value table, determine the value information corresponding to the food ingredient combination corresponding to the at least one food ingredient information, and use it as the above food value information.
[0093] As an example, the above execution entity can determine the value information corresponding to the food ingredient combination corresponding to the at least one food ingredient information by querying the third value table, and use it as the above food value information.
[0094] In some optional implementation manners of some embodiments, after step 107, the steps further include:
[0095] Step 1: Instruct the food acquisition object corresponding to the above food capture image to place the food corresponding to the above food capture image on the target placement platform. Among them, the food acquisition object can be the acquisition object of the food corresponding to the food capture image. That is, the food acquisition object can be the user who purchases the food. The target placement platform can be a platform for placing food.
[0096] Second step, in response to determining that the corresponding food has been placed on the above-mentioned target placement platform, obtain the historical food image sequence corresponding to the above-mentioned food captured image. Among them, the historical food captured image corresponds to a shooting time earlier than the shooting time corresponding to the above-mentioned food captured image. The historical food captured image can be an image captured for the food held by the food acquisition object at a historical time. There is a corresponding relationship between the historical food captured images in the historical food image sequence and the historical times in the historical time period.
[0097] Third step, remove the historical food captured images in the above-mentioned historical food image sequence whose corresponding food content does not meet the preset content conditions, and obtain the historical food image sequence after removal. Among them, the preset content condition can be that the pixel ratio of the food in the image is greater than the target ratio value.
[0098] Fourth step, according to the above-mentioned historical food image sequence after removal and the above-mentioned food semantic segmentation information, perform semantic segmentation verification on the above-mentioned food semantic segmentation information to obtain verification information.
[0099] Fifth step, in response to determining that the above-mentioned verification information indicates no error, display the food content corresponding to the above-mentioned food captured image on the display platform corresponding to the above-mentioned target placement platform, and display the above-mentioned food value information as a QR code on the display platform for the above-mentioned food acquisition object to transfer the corresponding value. The display platform can be a platform that displays QR code information, food value, and food information. The display platform is located above the above-mentioned target placement platform.
[0100] In some optional implementation manners of some embodiments, the above-mentioned performing semantic segmentation verification on the above-mentioned food semantic segmentation information according to the above-mentioned historical food image sequence after removal and the above-mentioned food semantic segmentation information to obtain verification information may include the following steps:
[0101] First step, for each historical food captured image in the above-mentioned historical food image sequence after removal, perform the following second generation steps:
[0102] Sub-step 1, perform image preprocessing on the above-mentioned historical food captured image to generate a second preprocessed image.
[0103] Sub-step 2, input the above-mentioned second preprocessed image into the above-mentioned image encoding model to generate second image encoding information.
[0104] Sub-step 3, input the above-mentioned food semantic encoding information and the above-mentioned second image encoding information into the above-mentioned decoding model to generate historical food semantic segmentation information.
[0105] Sub-step 4, determine the semantic segmentation image set corresponding to the above-mentioned historical food semantic segmentation information. The specific implementation manner will not be elaborated here.
[0106] In the second step, clustering is performed on the semantic segmentation images in the obtained sequence of semantic segmentation image sets to obtain a semantic segmentation image clustering cluster set.
[0107] As an example, first, image encoding processing is performed on each semantic segmentation image in the above-mentioned sequence of semantic segmentation image sets to generate image encoding information, and an image encoding information set is obtained. Then, the K-means algorithm is used to perform clustering processing on the image encoding information set to generate an image encoding information cluster set. Finally, a semantic segmentation image clustering cluster set corresponding to the image encoding information cluster set is generated.
[0108] In the third step, the similarity between the cluster center set corresponding to the above-mentioned semantic segmentation image clustering cluster set and at least one divided sub-image is determined. Among them, the similarity can represent the overall image content similarity between the cluster center set and at least one divided sub-image. Among them, the similarity can include multiple sub-similarities. The sub-similarity can represent the image similarity between the cluster center and the divided sub-image.
[0109] As an example, first, the execution subject can perform corresponding combination on the cluster centers in the cluster center set and the divided sub-images in at least one divided sub-image to generate a combination information set. Then, the image semantic similarity between the two images corresponding to each combination information in the combination information set is determined as the sub-similarity. Finally, the obtained sub-similarity set is determined as the similarity.
[0110] In the fourth step, in response to determining that the above-mentioned similarity is greater than or equal to the target similarity value, verification information indicating that the verification has passed is generated. Among them, the target similarity value can be a preset value.
[0111] Optionally, the above method further includes:
[0112] In the first step, in response to determining that the above-mentioned similarity is less than the above-mentioned target similarity value, verification information indicating that the verification has not passed is generated.
[0113] In the second step, a food image of the corresponding food placed on the above-mentioned target placement platform is re-taken as a real-time food image.
[0114] In the third step, corresponding food value information is generated according to the above-mentioned real-time food image. The specific implementation manner can refer to the implementation manner of generating corresponding food value information from the food captured image.
[0115] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the food value generation method of some embodiments of the present disclosure, based on accurately generating food semantic segmentation information, the value information of the corresponding food can be accurately generated. Specifically, the reason for the inaccurate value information of the relevant food is that there may be different determination logics for the value of the food. The image recognition model can only recognize the food and cannot be matched with the determination logic of the food, resulting in possible confusion in the selection of the determination logic for the food value finally, and errors in the determination of the food value information. A large amount of recognition resources are wasted, and relevant technical personnel are required to participate in coordination and processing. Based on this, in the food value generation method of some embodiments of the present disclosure, first, obtain the currently acquired food captured image and food semantic division information, so as to determine the food value information by obtaining food information and food plate information, etc. through the food captured image. The division and determination of the food corresponding to the food captured image are realized through the food semantic division information, which is convenient for the determination of the food value information. Then, perform image preprocessing on the above-mentioned food captured image to generate a first preprocessed image to improve the image quality of the image, which is convenient for the extraction of subsequent image feature information. Next, input the above-mentioned first preprocessed image into a pre-trained image encoding model to generate first image encoding information, and extract the feature semantic content related to the food in the image. Then, input the above-mentioned food semantic division information into a pre-trained semantic encoding model to generate food semantic encoding information, and extract the feature semantic content related to how the food is divided in the food semantic division information. Secondly, splice the above-mentioned food semantic encoding information and the above-mentioned first image encoding information to generate spliced information to fuse the food semantic feature content and the image feature semantic content. Furthermore, input the above-mentioned spliced information into a pre-trained decoding model for semantic segmentation to accurately generate food semantic segmentation information. Here, through the decoding model, the feature semantics corresponding to the food semantic encoding information and the first image encoding information can be fully considered to accurately perform food semantic segmentation processing. Finally, according to the above-mentioned food semantic segmentation information and the above-mentioned food semantic division information, the food value information corresponding to the above-mentioned food captured image can be accurately generated. In summary, through the image encoding model, the semantic encoding model and the decoding model, the extraction of the feature semantic content of the food captured image, the extraction of the feature of the food semantic content of the semantic encoding model, and the food semantic segmentation processing based on the fused feature information can be realized, so that accurate food value information can be obtained subsequently.
[0116] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a food value generation device. These device embodiments correspond to Figure 1 the method embodiments shown, and the food value generation device can be specifically applied to various electronic devices.
[0117] As Figure 2 shown, a food value generation device 200 includes: an acquisition unit 201, a preprocessing unit 202, a first input unit 203, a second input unit 204, a splicing unit 205, a third input unit 206, and a generation unit 207. Among them, the acquisition unit 201 is configured to acquire a currently acquired food capture image and food semantic segmentation information; the preprocessing unit 202 is configured to perform image preprocessing on the food capture image to generate a first preprocessed image; the first input unit 203 is configured to input the first preprocessed image into a pre-trained image encoding model to generate first image encoding information. The second input unit 204 is configured to input the food semantic segmentation information into a pre-trained semantic encoding model to generate food semantic encoding information; the splicing unit 205 is configured to splice the food semantic encoding information and the first image encoding information to generate splicing information; the third input unit 206 is configured to input the splicing information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information; the generation unit 207 is configured to generate food value information corresponding to the food capture image according to the food semantic segmentation information and the food semantic segmentation information.
[0118] It can be understood that the units described in the food value generation device 200 correspond to the respective steps in the method described in the reference Figure 1 description. Therefore, the operations, features, and beneficial effects described above for the method also apply to the food value generation device 200 and the units included therein, and will not be elaborated here.
[0119] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0120] As Figure 3 shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0121] Typically, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had. Figure 3 Each block shown in
[0122] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are executed.
[0123] It should be noted that in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0124] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0125] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain a currently acquired food captured image and food semantic segmentation information; perform image preprocessing on the above food captured image to generate a first preprocessed image; input the above first preprocessed image into a pre-trained image encoding model to generate first image encoding information; input the above food semantic segmentation information into a pre-trained semantic encoding model to generate food semantic encoding information; splice the above food semantic encoding information and the above first image encoding information to generate spliced information; input the above spliced information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information; generate food value information corresponding to the above food captured image according to the above food semantic segmentation information and the above food semantic segmentation information.
[0126] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0128] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a preprocessing unit, a first input unit, a second input unit, a splicing unit, a third input unit, and a generation unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit that acquires the currently captured food image and food semantic segmentation information".
[0129] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0130] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for generating food value, comprising: Obtaining a currently acquired food captured image and food semantic division information; Performing image preprocessing on the food captured image to generate a first preprocessed image; Inputting the first preprocessed image into a pre-trained image encoding model to generate first image encoding information; Inputting the food semantic division information into a pre-trained semantic encoding model to generate food semantic encoding information; Performing information splicing on the food semantic encoding information and the first image encoding information to generate spliced information; Inputting the spliced information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information, wherein Input the splicing information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information, including: Inputting the splicing information into the attention mechanism model included in the decoding model to generate semantic splicing information representing the first image encoding information and the target semantic segmentation feature information, where the splicing information includes: the first image encoding information, the first semantic segmentation feature information corresponding to the first semantic segmentation information, the second semantic segmentation feature information corresponding to the second semantic segmentation information, the third semantic segmentation feature information corresponding to the third semantic segmentation information, the fourth semantic segmentation feature information corresponding to the fourth semantic segmentation information, and the target semantic segmentation feature information is the semantic segmentation feature information among the first semantic segmentation feature information, the second semantic segmentation feature information, the third semantic segmentation feature information, and the fourth semantic segmentation feature information that is most similar in features to the first image encoding information. The attention mechanism model is used to select the semantic segmentation feature information that best matches the feature information of the first image encoding information; Input the semantic splicing information into the feature fusion model to generate the first fusion feature information. The feature fusion model is a model used to fuse the feature information of the food semantic encoding information and the target semantic segmentation feature information and extend the feature information; Input the first fusion feature information into the food semantic segmentation layer included in the decoding model to output the first initial food semantic segmentation information and the first segmentation score; Input the first image encoding information and the set of the remaining semantic segmentation feature information into the feature fusion model to generate the second fusion feature information; Input the second fusion feature information into the food semantic segmentation layer included in the decoding model to output the second initial food semantic segmentation information and the second segmentation score; Determine whether the first initial food semantic segmentation information and the second initial food semantic segmentation information are consistent; In response to determining that they are inconsistent, determine whether the subtraction score corresponding to subtracting the second segmentation score from the first segmentation score is greater than the target score; In response to determining that it is greater than the target score, determine the first initial food semantic segmentation information as the food semantic segmentation information; In response to determining that it is not greater than the target score, input the splicing information corresponding to the first image encoding information and the set of the remaining semantic segmentation feature information into the attention mechanism model to generate secondary semantic splicing information representing the first image encoding information and the secondary semantic segmentation feature information, where the secondary semantic segmentation feature information is the remaining semantic segmentation feature information in the set of the remaining semantic segmentation feature information that has the highest feature similarity to the first image encoding information; Perform splicing information fusion on the secondary semantic splicing information and the semantic splicing information to generate splicing information; Input the splicing information into the feature fusion model to generate the third fusion feature information; Remove the secondary semantic segmentation feature information from the set of the remaining semantic segmentation feature information to generate the removed feature information set; Input the removed feature information set and the first image encoding information into the feature fusion model to generate the fourth fusion feature information;Input the third fusion feature information into the food semantic segmentation layer to generate third initial food semantic segmentation information and a third segmentation score; input the fourth fusion feature information into the food semantic segmentation layer to generate fourth initial food semantic segmentation information and a fourth segmentation score; in response to determining that the third initial food semantic segmentation information is different from the fourth initial food semantic segmentation information and the third segmentation score is higher than the fourth segmentation score, determine the first initial food semantic segmentation information as the food semantic segmentation information; According to the food semantic segmentation information and the food semantic division information, generating food value information corresponding to the food captured image.
2. The method according to claim 1, wherein, The food semantic division information includes at least one of the following: first semantic division information based on a food plate, second semantic division information based on food types without counting the quantity, third semantic division information based on food types and determined by quantity, and fourth semantic division information based on food ingredients. The first semantic division information is division information representing value determination based on the corresponding plate shape of the food plate, the second semantic division information is division information representing value determination based on food types without counting the quantity, the third semantic division information represents division information for value determination based on food types and by quantity, and the fourth semantic division information represents division information for value determination based on whether the food is a meat-type food or a vegetarian-type food; and The step of inputting the food semantic division information into a pre-trained semantic encoding model to generate food semantic encoding information includes: Inputting at least one of the first semantic division information, the second semantic division information, the third semantic division information, and the fourth semantic division information into the semantic encoding model to generate food semantic encoding information.
3. The method according to claim 1, wherein, Each food plate in the food captured image has a corresponding food division label, where the food division label represents the division criterion of the food corresponding to the food plate, and the food division label is one of the following: first division label, second division label, third division label, and fourth division label. The first division label corresponds to the first semantic division information, the second division label corresponds to the second semantic division information, the third division label corresponds to the third semantic division information, and the fourth division label corresponds to the fourth semantic division information; and The food semantic division information is generated through the following steps: Obtaining various label description feature information corresponding to each food division label from a target database, where the label description feature information includes: label shape feature information, label position feature information, and label font feature information; Inputting the food captured image into a target detection model to generate at least one detection feature information for at least one initial detection object; For each detection feature information in the at least one detection feature information, performing the following first generation step: Input the detected feature information into a pre-trained label shape feature extraction model to generate initial label shape feature information; Input the detected feature information into a pre-trained label position feature extraction model to generate initial label position feature information; Input the detected feature information into a pre-trained label font feature extraction model to generate initial label font feature information; Compare the label shape feature information with the initial label shape feature information to generate first feature comparison information; Compare the label position feature information with the initial label position feature information to generate second feature comparison information; Compare the label font feature information with the initial label font feature information to generate third feature comparison information; Determine whether the initial detection object corresponding to the detected feature information is the target detection object according to the first feature comparison information, the second feature comparison information, and the third feature comparison information; Input the at least one target detection feature information corresponding to the obtained at least one target detection object into a pre-trained food classification label content generation model to generate at least one food classification label content, where the food classification label content is the label content corresponding to the food classification label; Perform deduplication processing on the at least one food classification label content to generate at least one semantic classification information as the food semantic classification information.
4. The method according to claim 1, wherein The method further includes: Instruct the food acquisition object corresponding to the food capture image to place the food corresponding to the food capture image on the target placement platform; In response to determining that the corresponding food has been placed on the target placement platform, obtain the historical food capture image sequence corresponding to the food capture image, where the historical food capture image corresponds to a capture time earlier than the capture time of the food capture image; Remove the historical food capture images in the historical food capture image sequence whose corresponding food content does not meet the preset content condition to obtain the filtered historical food capture image sequence; Perform semantic segmentation verification on the food semantic segmentation information according to the filtered historical food capture image sequence and the food semantic classification information to obtain verification information; In response to determining that the verification information indicates no error, display the food content corresponding to the food capture image on the display platform corresponding to the target placement platform, and display the food value information as a QR code on the display platform for the food acquisition object to transfer the corresponding value.
5. The method according to claim 4, wherein The performing semantic segmentation verification on the food semantic segmentation information according to the filtered historical food capture image sequence and the food semantic classification information to obtain verification information includes: For each historical food capture image in the filtered historical food capture image sequence, perform the following second generation steps: Perform image preprocessing on the historical food capture image to generate a second preprocessed image; Input the second preprocessed image into the image encoding model to generate second image encoding information; Input the food semantic coding information and the second image coding information into the decoding model to generate historical food semantic segmentation information; Determine the semantic segmentation image set corresponding to the historical food semantic segmentation information; Perform clustering processing on the semantic segmentation images in the obtained semantic segmentation image set sequence to obtain a semantic segmentation image clustering cluster set; Determine the similarity between the cluster center set corresponding to the semantic segmentation image clustering cluster set and at least one divided sub-image; In response to determining that the similarity is greater than or equal to the target similarity value, generate verification information indicating that the verification has passed; 6. The method according to claim 5, wherein, The method further includes: In response to determining that the similarity is less than the target similarity value, generate verification information indicating that the verification has not passed; Reshoot the food image of the corresponding food placed on the target placement platform as a real-time food image; Generate the corresponding food value information according to the real-time food image; 7. The method according to claim 2, wherein The food semantic segmentation information includes: at least one divided sub-image; and The generating the food value information corresponding to the food captured image according to the food semantic segmentation information and the food semantic division information includes: In response to determining that the food semantic division information is the first semantic division information, determine the image size information corresponding to the at least one divided sub-image to obtain at least one image size information; Obtain a first value table from the target database, which represents the corresponding relationship between the image size information and the value information for the first semantic division information; According to the first value table, determine the value information corresponding to each image size information in the at least one image size information to obtain at least one first value information; Add up the respective first value information in the at least one first value information to generate an added information as the food value information; In response to determining that the food semantic division information is the second semantic division information, input each of the at least one divided sub-images into a pre-trained first food type information generation model to generate first food type information, obtaining at least one first food type information, where the first food type information includes: food type; Obtain a second value table from the target database, and according to the second value table, determine the value information corresponding to each first food type information in the at least one first food type information to obtain at least one second value information; Add up the respective second value information in the at least one second value information to generate an added information as the food value information; In response to determining that the food semantic division information is the third semantic division information, input each of the at least one divided sub-images into a pre-trained second food type information generation model to generate second food type information, obtaining at least one second food type information, where the second food type information includes: food type and food quantity; According to the second value table, determine the value information corresponding to each second food type information in the at least one second food type information, to obtain at least one third value information, where the third value information is the value information obtained by multiplying the value corresponding to the food type corresponding to the second food type information by the number of foods; Add up the respective third value information in the at least one third value information to generate an added information as the food value information; In response to determining that the food semantic division information is the fourth semantic division information, input each sub-image in the at least one sub-image into a pre-trained food ingredient information generation model to generate food ingredient information, to obtain at least one food ingredient information; Obtain from the target database a pre-stored third value table representing the relationship between food ingredient combinations and value information; According to the third value table, determine the value information corresponding to the food ingredient combination corresponding to the at least one food ingredient information as the food value information.
8. A food value generation device, comprising: An acquisition unit configured to acquire a currently acquired food capture image and food semantic division information; A preprocessing unit configured to perform image preprocessing on the food capture image to generate a first preprocessed image; A first input unit configured to input the first preprocessed image into a pre-trained image encoding model to generate first image encoding information; A second input unit configured to input the food semantic division information into a pre-trained semantic encoding model to generate food semantic encoding information; A splicing unit configured to splice the food semantic encoding information and the first image encoding information to generate spliced information; A third input unit, configured to input the splicing information into a pre-trained decoding model for semantic segmentation to generate food semantic segmentation information, where inputting the splicing information into the pre-trained decoding model for semantic segmentation to generate food semantic segmentation information includes: inputting the splicing information into an attention mechanism model included in the decoding model to generate semantic splicing information representing first image encoding information and target semantic partition feature information, where the splicing information includes: first image encoding information, first semantic partition feature information corresponding to first semantic partition information, second semantic partition feature information corresponding to second semantic partition information, third semantic partition feature information corresponding to third semantic partition information, fourth semantic partition feature information corresponding to fourth semantic partition information, and the target semantic partition feature information is the semantic partition feature information among the first semantic partition feature information, the second semantic partition feature information, the third semantic partition feature information, and the fourth semantic partition feature information that is most similar in features to the first image encoding information, and the attention mechanism model is used to select the semantic partition feature information that best matches the feature information of the first image encoding information; inputting the semantic splicing information into a feature fusion model to generate first fusion feature information, where the feature fusion model is a model for fusing feature information and extending feature information of food semantic encoding information and target semantic partition feature information; inputting the first fusion feature information into a food semantic segmentation layer included in the decoding model to output first initial food semantic segmentation information and a first segmentation score; inputting the first image encoding information and the remaining semantic partition feature information set into the feature fusion model to generate second fusion feature information; inputting the second fusion feature information into the food semantic segmentation layer included in the decoding model to output second initial food semantic segmentation information and a second segmentation score; determining whether the first initial food semantic segmentation information and the second initial food semantic segmentation information are consistent; in response to determining that they are inconsistent, determining whether a subtraction score corresponding to subtracting the second segmentation score from the first segmentation score is greater than a target score; in response to determining that it is greater than the target score, determining the first initial food semantic segmentation information as the food semantic segmentation information; in response to determining that it is not greater than the target score, inputting the splicing information corresponding to the first image encoding information and the remaining semantic partition feature information set into the attention mechanism model to generate secondary semantic splicing information representing the first image encoding information and secondary semantic partition feature information, where the secondary semantic partition feature information is the remaining semantic partition feature information in the remaining semantic partition feature information set that has the highest feature similarity to the first image encoding information; performing splicing information fusion on the secondary semantic splicing information and the semantic splicing information to generate splicing information; inputting the splicing information into the feature fusion model to generate third fusion feature information; removing the secondary semantic partition feature information from the remaining semantic partition feature information set to generate a removed feature information set;Input the post-removal feature information set and the first image encoding information into a feature fusion model to generate fourth fused feature information; input the third fused feature information into a food semantic segmentation layer to generate third initial food semantic segmentation information and a third segmentation score; input the fourth fused feature information into the food semantic segmentation layer to generate fourth initial food semantic segmentation information and a fourth segmentation score; in response to determining that the third initial food semantic segmentation information is different from the fourth initial food semantic segmentation information and the third segmentation score is higher than the fourth segmentation score, determine the first initial food semantic segmentation information as the food semantic segmentation information; A generation unit configured to generate food value information corresponding to the food capture image according to the food semantic segmentation information and the food semantic division information.
9. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method and system based on attention regulation and control
CN115049919A