Food identification method and system based on multi-feature fusion
Through the food recognition method of multi-feature fusion, combined with shape, texture and knife features, food is identified step by step, and ultimately the accuracy of food recognition is achieved, solving the problem of insufficient food recognition accuracy in existing technologies.
Patent Information
- Application Number
- CN202510705189.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-05
AI Technical Summary
The existing methods based on food image recognition mainly rely on single-dimensional feature detection, resulting in low accuracy of food recognition results and inability to achieve accurate multi-level control of food.
A food recognition method based on multi-feature fusion is adopted. The image to be identified is generated through food image preprocessing, and the shape features, texture features and knife features are determined to generate a comprehensive feature vector. The outer contour and dynamic image are combined for step-by-step recognition to finally determine the final recognition result of the food.
The food recognition results are highly accurate. The precision and accuracy of food recognition are ensured through a multi-level recognition process. It is compatible with the overall consideration of shape, texture and knife characteristics, and improves the accuracy of recognition.
Smart Images

Figure CN120599602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of food recognition methods, and in particular to a food recognition method and system based on multi-feature fusion. Background Art
[0002] With the advancement of science and technology, food image recognition has become a core technology in the field of food computing. As a key research area for fine-grained visual classification in computer vision, its theoretical value and practical significance are significant. With the rapid development of artificial intelligence technology, food image recognition has shown tremendous application potential in multiple fields. Existing technologies rely on food images for recognition and output corresponding morphological features. These features are then used to determine food recognition results. This approach relies on single-dimensional feature detection, resulting in low accuracy in food recognition results and an inability to achieve precise, multi-level control of food recognition. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art, and the present invention provides a food recognition method and system based on multi-feature fusion.
[0004] An embodiment of the present invention provides a food recognition method based on multi-feature fusion, including:
[0005] outputting an image to be recognized according to image preprocessing of the food image;
[0006] Determining shape features, texture features, and knife features based on the recognition of the image to be recognized;
[0007] Generate a comprehensive feature vector based on the fusion of multiple features such as shape features, texture features, and knife features, and determine the primary recognition result of the food based on the comprehensive feature vector;
[0008] determining an advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food;
[0009] A review area is determined based on the high-level recognition result of the food and the food image, and a final recognition result of the food is determined based on the features corresponding to the review area, the high-level recognition result of the food, and the review mapping relationship.
[0010] An embodiment of the present invention provides a food recognition system based on multi-feature fusion. The food recognition system based on multi-feature fusion is applied to the above-mentioned food recognition method based on multi-feature fusion. The food recognition system based on multi-feature fusion includes:
[0011] An image to be identified module, configured to output an image to be identified based on image preprocessing of the food image;
[0012] A feature module, used to determine shape features, texture features, and knife features based on the recognition of the image to be recognized;
[0013] A primary recognition result module is used to generate a comprehensive feature vector based on the fusion of multiple features such as shape features, texture features, and knife features, and to determine the primary recognition result of the food based on the comprehensive feature vector;
[0014] an advanced recognition result module, configured to determine an advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food;
[0015] The final recognition result module is used to determine the review area based on the high-level recognition result of the food and the food image, and to determine the final recognition result of the food according to the features corresponding to the review area, the high-level recognition result of the food and the review mapping relationship.
[0016] Compared with the prior art, the present invention has the following beneficial effects:
[0017] In an embodiment of the present invention, through the method in the embodiment of the present invention, the image to be identified is output based on the image preprocessing of the food image; the shape features, texture features and knife features are determined based on the identification of the image to be identified; a comprehensive feature vector is generated based on the multi-feature fusion of shape features, texture features and knife features, and the primary identification result of the food is determined based on the comprehensive feature vector, which is compatible with the overall consideration of shape features, texture features and knife features, realizes the technical effect of multi-feature fusion, and ensures the accuracy of the primary identification result of the food.
[0018] Therefore, the advanced recognition result of the food is determined based on the outer contour of the food, multiple dynamic images and the primary recognition result of the food; the review area is determined based on the advanced recognition result of the food and the food image, and the final recognition result of the food is determined based on the features corresponding to the review area, the advanced recognition result of the food and the review mapping relationship, thereby realizing step-by-step control of the primary recognition result, advanced recognition result and final recognition result of the food, further ensuring the high accuracy of the final recognition result of the food, and realizing accurate multi-level control of food recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flowchart of a food recognition method based on multi-feature fusion in an embodiment of the present invention;
[0020] Figure 2 1 is a flow chart of step S11 in the food recognition method based on multi-feature fusion in an embodiment of the present invention;
[0021] Figure 3 is a flow chart of step S12 in the food recognition method based on multi-feature fusion in an embodiment of the present invention;
[0022] Figure 4 is a flow chart of step S13 in the food recognition method based on multi-feature fusion in an embodiment of the present invention;
[0023] Figure 5 is a flow chart of step S14 in the food recognition method based on multi-feature fusion in an embodiment of the present invention;
[0024] Figure 6 is a flow chart of step S15 in the food recognition method based on multi-feature fusion in an embodiment of the present invention;
[0025] Figure 7 Schematic diagram of the structure of a food recognition system based on multi-feature fusion in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0027] See also Figures 1 to 7 A food recognition method based on multi-feature fusion is applied to a food recognition scenario based on multi-feature fusion; the food recognition method based on multi-feature fusion includes:
[0028] Step S11: outputting an image to be recognized based on image preprocessing of the food image;
[0029] Step S12: determining shape features, texture features, and knife features based on the recognition of the image to be recognized;
[0030] Step S13: generating a comprehensive feature vector based on the fusion of the shape feature, texture feature, and knife feature, and determining a primary recognition result of the food according to the comprehensive feature vector;
[0031] Step S14: determining an advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food;
[0032] Step S15: determining a review area based on the high-level recognition result of the food and the food image, and determining a final recognition result of the food based on features corresponding to the review area, the high-level recognition result of the food, and the review mapping relationship;
[0033] refer to Figure 2 , in step S11, outputting an image to be recognized according to image preprocessing of the food image;
[0034] In the specific implementation process of the present invention, the specific steps are:
[0035] S111: generating a plurality of sub-images in different directions based on the circular photographing of the food, and determining a food image by synthesizing the plurality of sub-images, the food image being used as a panoramic image of the food;
[0036] S112: performing image preprocessing on the food image and outputting a corresponding grayscale image, determining a plurality of areas to be detected based on the detection of the grayscale image, and determining an image to be recognized based on the synthesis of the plurality of areas to be detected.
[0037] In an embodiment of the present application, multiple sub-images in different directions are generated based on the annular shooting of food, and a food image is determined according to the synthesis of the multiple sub-images. The food image is introduced as a panoramic image of the food.
[0038] At this time, capture the food's appearance features from different directions by shooting in a circular motion around the food; use a shooting device with an automatic rotation function, such as a camera with a rotating gimbal, or manually operate the camera to shoot around the food; ensure that the shooting parameters (such as exposure, focal length, white balance, etc.) are consistent so that subsequent image synthesis can be seamless; start from a starting point of the food and take a photo at a certain angle (such as 30 degrees, 45 degrees, etc.) until a full circle of shooting is completed; when shooting, ensure that the food is at the center of the shooting and that there is some overlap between each sub-image to facilitate subsequent image stitching.
[0039] Alternatively, suppose you want to capture and synthesize a panoramic image of a plate of fruit salad. Use a camera with a pan / tilt head to take a picture of the salad from above, every 45 degrees, for a total of eight sub-images. Ensure that the salad is centered in the frame and that the sub-images partially overlap. Perform denoising and contrast enhancement on each sub-image to make the color and texture of the fruit clearer.
[0040] The multiple sub-images captured are stitched together into a continuous, seamless panoramic image; each sub-image captured is pre-processed, such as denoising, contrast enhancement, etc., to improve image quality; image stitching algorithms (such as SIFT, SURF, ORB, etc.) are used to detect feature points in each sub-image and match them to find common features between adjacent sub-images; based on the feature point matching results, adjacent sub-images are transformed (such as affine transformation, perspective transformation, etc.) to enable them to be seamlessly connected; then, image fusion techniques (such as weighted averaging, multi-band fusion, etc.) are used to make the stitching smoother and more natural; after all sub-images are stitched together, a continuous, seamless panoramic image is generated, which will be used as the food image for subsequent processing.
[0041] Optionally, the SIFT algorithm is used to detect feature points in each sub-image and match them. Then, based on the matching results, adjacent sub-images are affine transformed to make them seamlessly connected. Finally, a weighted average method is used to smooth the joints. After all sub-images are stitched together, a complete, continuous, and seamless panoramic image of the fruit salad is obtained. This image clearly shows the shape, color, and texture characteristics of the salad, providing strong support for subsequent food recognition.
[0042] Furthermore, the food image is preprocessed and a corresponding grayscale image is output. Based on the detection of the grayscale image, multiple areas to be detected are determined, and the image to be identified is determined based on the synthesis of the multiple areas to be detected, thereby ensuring the accuracy of the image to be identified.
[0043] At this time, the food image is preprocessed to improve the quality of the food image and provide a clear image basis for subsequent processing; at the same time, filters (such as Gaussian filters, mean filters, etc.) are used to remove noise in the image to make the image smoother; the contrast of the image is adjusted to make the characteristics of the food more obvious for subsequent detection; the image is adjusted to a suitable size to meet the requirements of the subsequent processing algorithm.
[0044] Convert color images into grayscale images to simplify image information and improve processing efficiency; use common grayscale conversion formulas (such as weighted average method) to convert each pixel of the color image into a grayscale value; save the converted grayscale image for subsequent processing.
[0045] Detect the area containing food in the grayscale image as the area to be detected; use an edge detection algorithm (such as the Canny edge detector) to detect the edges of the food in the grayscale image; based on the edge detection results, use an image segmentation algorithm (such as threshold segmentation, region growing, etc.) to divide the image into multiple regions; based on the size, shape and other characteristics of the region, filter out the area to be detected containing food.
[0046] The screened areas to be detected are synthesized into a complete image to be identified. At this time, adjacent and similar areas to be detected are merged into a larger area; the merged area is cropped from the original food image to obtain the image to be identified; necessary adjustments (such as rotation, scaling, etc.) are made to the cropped image to be identified to ensure that its quality and features meet the requirements of subsequent processing.
[0047] Specifically, suppose there is a color image containing multiple fruits, and you want to extract the fruit area as the image to be recognized; denoise the original image to remove the noise points in the image; then, enhance the contrast of the image to make the color and texture of the fruit more vivid; finally, resize the image to an appropriate size.
[0048] The preprocessed color image is converted into a grayscale image; the converted grayscale image is more concise and only contains brightness information; the Canny edge detector is used on the grayscale image to detect the edges of the fruit; then, based on the edge detection results, the threshold segmentation algorithm is used to segment the image into multiple regions; according to the size and shape characteristics of the region, the detection areas containing fruits are screened out, which include complete fruits, partial fruits, and gaps between fruits; the screened detection areas are merged and adjusted; adjacent and similar areas are merged to form a larger fruit area; then, the merged fruit area is cropped from the original color image to obtain the image to be identified; finally, the cropped image to be identified is adjusted as necessary, such as rotation and scaling, to ensure that its quality and characteristics meet the requirements of subsequent food recognition.
[0049] In one embodiment of the present application, assume that there is a color food image containing apples and bananas; preprocessing operations such as denoising and contrast enhancement are performed on the original image, and a corresponding grayscale image is output.
[0050] In the grayscale image, features such as edges, textures, and shapes are extracted and compared with known food features in the matching table. Through comparison, two areas to be detected are determined: one is the circular apple area, and the other is the curved banana area. To quantify the degree of matching, a weight is assigned to each feature, and a score is calculated for each area. Assume that the score of the apple area is 0.85 (higher) and the score of the banana area is 0.75 (slightly lower).
[0051] Based on the weights and scores, the apple and banana regions with higher scores were selected for merging. During the merging process, the degree of overlap and shape similarity between the regions were taken into consideration to ensure that the merged image accurately reflects the key features of the food. Finally, the merged region was cropped from the original color image to obtain the image to be identified. This image to be identified clearly shows the shape, color, and texture features of the apple and banana, providing strong support for subsequent food recognition.
[0052] refer to Figure 3 In step S12, shape features, texture features, and knife features are determined based on the recognition of the image to be recognized;
[0053] In the specific implementation process of the present invention, the specific steps are:
[0054] S121: Determine a plurality of sub-regions based on the position division of the image to be recognized, and determine a shape region set, a texture region set, and a knife region set based on screening of the plurality of sub-regions;
[0055] S122: determining a plurality of sub-shape features according to the recognition of the shape region set, and determining a shape feature according to the morphologies of the plurality of sub-shape features, the spatial positions of the plurality of sub-shape features, and the shape mapping relationship;
[0056] S123: determining a plurality of sub-texture features based on the identification of the texture region set, and determining a texture feature based on the patterns of the plurality of sub-texture features, the spatial positions of the plurality of sub-texture features, and a texture mapping relationship;
[0057] S124: determining a plurality of sub-cutting features according to the identification of the cutting area set, and determining the cutting feature according to the tool paths of the plurality of sub-cutting features, the spatial positions of the plurality of sub-cutting features, and the cutting mapping relationship;
[0058] In an embodiment of the present application, multiple sub-regions are determined based on the position division of the image to be identified, and a shape region set, a texture region set, and a knife region set are determined based on the screening of the multiple sub-regions. The shape region set, the texture region set, and the knife region set are introduced.
[0059] At this time, the image to be identified is divided into multiple sub-regions according to different parts or features of the food, so that each sub-region can be subsequently analyzed and processed specifically; at the same time, the image to be identified is divided into multiple sub-regions using image segmentation algorithms (such as threshold segmentation, edge detection, region growing, superpixel segmentation, etc.). These algorithms can segment the image into different parts based on pixel similarity, edge information, etc.; in some cases, manual participation is required to mark key areas in the image to ensure the accuracy of the segmentation. This is usually performed after the preliminary algorithm segmentation to correct and optimize the segmentation results.
[0060] Regions related to shape, texture and knife marks are screened out from the multiple sub-regions obtained by segmentation, and they are respectively classified into corresponding sets; at this time, feature analysis is performed on each sub-region, including shape, texture and knife marks, which is accomplished by calculating the shape descriptors (such as perimeter, area, circularity, etc.), texture descriptors (such as gray-level co-occurrence matrix, local binary pattern, etc.) of the region and detecting cutting edges or patterns; based on the results of feature analysis and preset rules, the sub-regions are classified into shape regions, texture regions or knife marks regions; the basis of classification is the similarity of shape descriptors, the statistical characteristics of texture descriptors and the presence or absence of cutting edges; the classified sub-regions are respectively classified into shape region sets, texture region sets and knife marks region sets, which will be used for subsequent feature extraction and analysis.
[0061] Specifically, suppose there is an image to be identified containing sliced ham; a superpixel segmentation algorithm is used to divide the image into multiple sub-regions, which include the sliced part of the ham, the background area, and the cutting edge.
[0062] Shape regions: By calculating the shape descriptor of each subregion, we found that some subregions have regular rectangular or elliptical shapes, which are very similar to sliced parts of ham; therefore, these regions are classified into the shape region set;
[0063] Texture regions: Next, we compute the texture descriptor for each sub-region. We find that some sub-regions have unique texture patterns, such as the fibrous texture of ham, and these sub-regions are grouped into a set of texture regions.
[0064] Cutting area: Finally, the cutting edges in the image are detected; the edge detection algorithm can identify the cutting lines between the ham slices, and the areas where these cutting lines are located are classified into the cutting area set.
[0065] Through such steps and examples, we can more specifically understand how to determine multiple sub-regions based on the position of the image to be identified in step S121, and determine the shape region set, texture region set and knife region set based on the characteristics of the sub-regions.
[0066] Furthermore, multiple sub-shape features are determined based on the recognition of the shape region set. The shape feature is then determined based on the morphology, spatial location, and shape mapping relationships of the multiple sub-shape features. This takes into account the morphology, spatial location, and shape mapping relationships of the multiple sub-shape features, ensuring the accuracy of the determined shape features.
[0067] At this point, specific sub-shape features are extracted from the shape region set. These features can describe the key information of the food shape. For each region in the shape region set, its shape descriptor is calculated. Common shape descriptors include perimeter, area, circularity, rectangularity, aspect ratio, convex hull, etc. These descriptors can quantify the shape characteristics of the region. Sub-shape features are further extracted using feature extraction algorithms (such as shape context, contour matching, etc.). These algorithms can capture subtle differences in shape, such as the curvature of the contour, the key points of the shape, etc.
[0068] The morphology, spatial position and mutual mapping relationship of multiple sub-shape features are integrated to form a comprehensive description of the shape characteristics of food; at this time, the morphology of each sub-shape feature is analyzed, including its size, shape regularity, symmetry, etc., which helps to understand the overall structure and local characteristics of the food shape; the spatial position of the sub-shape features in the image is analyzed, including their relative distance, direction, arrangement, etc., which can reveal the spatial layout and hierarchy of the food shape; based on the morphology and spatial position relationship between the sub-shape features, a shape mapping relationship is constructed, which is achieved through shape matching algorithms (such as shape context matching, Hausdorff distance, etc.) to quantify the similarities and differences between different sub-shape features; the results of morphological analysis, spatial position analysis and shape mapping relationship are combined to determine the final shape features, which are weighted combinations of shape descriptors, encoded representations of shape contours or quantitative indicators of shape mapping relationships.
[0069] Specifically, assume that there is an image to be identified containing cut potatoes, and the shape area set has been determined through step S121; assume that there is an image to be identified containing cut potatoes, and the shape area set has been determined through step S121.
[0070] Morphological analysis shows that most of these cut areas are irregular polygonal shapes, but have a certain symmetry; spatial position analysis shows that the cut areas are evenly distributed in the image, without obvious aggregation or dispersion; at the same time, the shape context matching algorithm is used to compare the contour features of different cut areas; the results show that although the specific shapes of the cut areas are different, their contours show a similar pattern as a whole, indicating that these cuts have certain similarities in shape. The comprehensive shape characteristics of potato cuts are determined by combining the results of morphological analysis, spatial position analysis and shape mapping. These characteristics include the average perimeter, area, circularity of the cut area and the encoding representation of contour features.
[0071] Through such steps and examples, we can understand more specifically how to determine multiple sub-shape features based on the identification of the shape area set in step S122, and determine the final shape features based on the morphology, spatial position and mapping relationship of these sub-shape features. This process involves the calculation of shape descriptors, the application of feature extraction algorithms, and the comprehensive analysis of morphology, spatial position and mapping relationship to ensure accurate description of food shape features.
[0072] Furthermore, multiple sub-texture features are determined based on the identification of a texture area set, and texture features are determined based on the textures of the multiple sub-texture features, the spatial positions of the multiple sub-texture features, and the texture mapping relationship. This is compatible with the overall consideration of the textures of the multiple sub-texture features, the spatial positions of the multiple sub-texture features, and the texture mapping relationship, ensuring the accuracy of the determined texture features.
[0073] At this time, specific sub-texture features are extracted from the texture region set. These features can describe the key information of food texture. At the same time, for each region in the texture region set, its texture descriptor is calculated. Common texture descriptors include gray-level co-occurrence matrix (GLCM), local binary pattern (LBP), Gabor filter response, etc. These descriptors can quantify the texture characteristics of the region, such as roughness, directionality, contrast, etc.; specific feature extraction methods (such as wavelet transform, fractal analysis, etc.) are used to further extract sub-texture features. These methods can capture subtle differences in texture, such as texture periodicity, self-similarity, etc.
[0074] The texture, spatial position and mutual mapping relationship of multiple sub-texture features are integrated to form a comprehensive description of the food texture characteristics; at this time, the texture of each sub-texture feature is analyzed, including its directionality, periodicity, coarseness, etc., which helps to understand the overall structure and local characteristics of the food texture; the spatial position of the sub-texture features in the image is analyzed, including the relative distance and distribution pattern between them, which can reveal the spatial layout and hierarchical structure of the food texture; based on the texture and spatial position relationship between the sub-texture features, a texture mapping relationship is constructed, which is achieved through a texture matching algorithm (such as GLCM-based texture similarity measurement, LBP histogram comparison, etc.) to quantify the similarities and differences between different sub-texture features; the results of texture analysis, spatial position analysis and texture mapping relationship are combined to determine the final texture features, which are weighted combinations of texture descriptors, encoded representations of texture patterns or quantitative indicators of texture mapping relationships.
[0075] Specifically, assume that there is an image to be identified containing barbecue, and a texture region set has been determined through step S121; from the texture region set, multiple regions containing barbecue textures are identified; for each region, its GLCM descriptor is calculated, including statistics such as contrast, energy, homogeneity, and correlation, and the LBP algorithm is used to extract local texture features.
[0076] Texture analysis shows that the texture of the grilled meat area has obvious directionality and a certain periodicity, which is manifested as the fibrous structure of the meat and grill marks; spatial position analysis shows that these texture areas are evenly distributed in the image, but the texture in some areas is denser, which is related to the thickness or degree of grilling of the meat.
[0077] Using a GLCM-based texture similarity measurement algorithm, the similarities between different texture regions were compared. The results showed that although the texture details of different regions were different, they presented similar directional and periodic patterns as a whole, indicating that these regions have certain similarities in texture. The results of comprehensive texture analysis, spatial position analysis and texture mapping relationship were combined to determine the comprehensive texture characteristics of barbecue. These characteristics include a weighted combination of texture statistics such as directionality, periodicity, contrast, energy, and the encoding representation of texture patterns, which are used for subsequent food recognition or classification tasks.
[0078] Therefore, multiple sub-cutting features are determined based on the identification of the cutting area set, and the cutting features are determined based on the tool paths of multiple sub-cutting features, the spatial positions of multiple sub-cutting features and the cutting mapping relationship. This is compatible with the overall consideration of the tool paths of multiple sub-cutting features, the spatial positions of multiple sub-cutting features and the cutting mapping relationship, ensuring the accuracy of the cutting features.
[0079] At this time, specific sub-knife features are extracted from the knife area set. These features can describe the key information of food cutting; edge detection algorithms (such as Canny edge detection, Sobel operator, etc.) are used to identify the cutting edges in the knife area. These edges are usually manifested as a set of pixels with sudden changes in brightness or color in the image, corresponding to the position where the food is cut; features are extracted from the detected cutting edges, which include geometric features such as edge direction, length, curvature, continuity, and visual features such as color and brightness of pixels on both sides of the edge. These features together constitute the sub-knife feature set.
[0080] The tool paths (cutting trajectories), spatial positions and mapping relationships of multiple sub-knife features are integrated to form a comprehensive description of the food knife features; at this time, the tool path of each sub-knife feature is analyzed, including the cutting direction, depth, uniformity, etc. This information helps to understand the cutting method and degree of fineness; the spatial position of the sub-knife features in the image is analyzed, including their relative distance, arrangement, coverage area, etc. This information can reveal the overall layout and local details of the cutting; based on the tool path and spatial position relationship between the sub-knife features, a knife mapping relationship is constructed, which is achieved by comparing the geometric features, visual features and relative positions of different sub-knife features; the mapping relationship quantifies the similarities and differences between different cuts, thereby reflecting the knife style of the food; the results of the tool path analysis, spatial position analysis and knife mapping relationship are combined to determine the final knife features. These features are a weighted combination of the geometric features and visual features of the cutting edge, and are also quantitative indicators of the knife mapping relationship, which are used for subsequent food recognition, classification or quality assessment tasks.
[0081] Specifically, assume that there is an image to be identified containing finely cut vegetable salad, and a set of knife-cut areas has been determined through step S121; from the set of knife-cut areas, multiple areas containing cut edges are identified; these cut edges are accurately detected using the Canny edge detection algorithm; feature extraction is performed on the detected edges to obtain geometric features such as the direction, length, and curvature of the edges, as well as the color and brightness information of the pixels on both sides of the edges.
[0082] Tool path analysis shows that these cutting edges present uniform and fine cutting trajectories, indicating that the cutting process is very meticulous; spatial position analysis shows that the cutting edges are evenly distributed in the image, and the cutting depth is consistent, covering the entire vegetable salad area.
[0083] By comparing the geometric and visual features of different cutting edges, it was found that they had a high degree of similarity; further analysis of the relative position relationship between the cutting edges showed that they showed a regular arrangement, such as parallel or staggered arrangement; comprehensive knife-technique features of vegetable salad were determined by integrating the results of tool path analysis, spatial position analysis and knife-technique mapping. These features include statistics of geometric features such as the average length, curvature, and directional consistency of the cutting edges, as well as descriptions of visual features such as the color and brightness uniformity of the cutting area. These features together reflect the fine cutting style and high-quality knife-technique of the vegetable salad, which are used in subsequent food identification or quality assessment tasks.
[0084] refer to Figure 4 In step S13, a comprehensive feature vector is generated based on the fusion of multiple features including shape features, texture features, and knife features, and a primary recognition result of the food is determined based on the comprehensive feature vector;
[0085] In the specific implementation process of the present invention, the specific steps are:
[0086] S131: assigning corresponding weights to the shape features, texture features, and knife features, and generating multiple feature combinations based on the shape features, texture features, and knife features, wherein the weights of the multiple feature combinations are weighted results of the corresponding weights;
[0087] S132: The combined contents of multiple feature combinations and the weights of each feature combination are used to generate a comprehensive feature vector under multi-feature fusion; multiple phased results are determined based on the operation of the comprehensive feature vector, and the primary recognition result of the food is determined based on the multiple phased results and the corresponding result mapping relationship.
[0088] In the embodiment of the present application, corresponding weights are configured for shape features, texture features and knife features. At the same time, multiple feature combinations are generated based on shape features, texture features and knife features. The weights of the multiple feature combinations are the weighted results of the corresponding weights.
[0089] At this time, a weight value is assigned to each feature according to its importance in food recognition; at the same time, shape feature weight: shape feature is very critical information in food recognition, because many foods are easy to identify due to their unique shape; for example, round apples, rectangular bread, etc.; the weight of shape feature is usually higher.
[0090] Texture feature weight: Texture features describe the texture and pattern of food surfaces, such as smooth, rough, spotted, etc. These features are very useful for distinguishing foods with similar appearances; the weight of the texture feature depends on its importance in the specific food recognition task.
[0091] Knife feature weight: The knife feature reflects the way and degree of fineness of cutting of food, which is very important for evaluating the quality and preparation of food. For example, finely cut vegetable salad will look very different from casually cut vegetables. The weight of the knife feature depends on whether the task focuses on the fineness of cutting.
[0092] By combining different features, we explore feature interactions and complementarities to improve the accuracy of food recognition. We generate all feature combinations, including shape features + texture features, shape features + knife features, texture features + knife features, and a combination of the three. The number of combinations depends on the number of features and the limitations of computing resources. Too many combinations lead to low computational efficiency, while too few combinations fail to fully utilize the complementarity between features.
[0093] Assign a weight to each feature combination, which is the weighted result of the weights of each feature in the combination; at this time, common weighting methods include simple weighting (i.e. direct summation), normalized weighting (i.e. normalizing the weight of each feature and then summing it), etc.; the choice of weighting method depends on the importance of the feature and the requirements of the task; for each feature combination, the weight of each feature contained in it is weighted and summed to obtain the weight of the combination.
[0094] Specifically, suppose you are developing a food recognition system that needs to identify three fruits: apples, bananas, and oranges; you have extracted shape features, texture features, and cutter features, and decided to assign weights to them; shape feature weight: 0.5 (because shape is very important for identifying these three fruits); texture feature weight: 0.3 (because texture helps distinguish the smooth surface of apples from the rough surface of oranges); cutter feature weight: 0.2 (because in this task, you are mainly concerned with the overall shape and texture of the fruit, not the way it is cut).
[0095] Combination 1: shape features + texture features (weight: 0.5 + 0.3 = 0.8);
[0096] Combination 2: shape feature + knife feature (weight: 0.5 + 0.2 = 0.7);
[0097] Combination 3: texture feature + knife feature (weight: 0.3 + 0.2 = 0.5);
[0098] Combination 4: Shape + Texture + Knife Handling (Weight: 0.5 + 0.3 + 0.2 = 1.0, but considering computational efficiency and feature redundancy, this combination is not chosen in practice, or other combinations may be tried first, and then the decision to add a third feature is made based on the results.) In this example, the first three feature combinations were selected for further analysis and recognition; the weight of each combination reflects its importance in the food recognition task; in practice, the weights and combinations will be adjusted based on the performance of these combinations in the recognition model.
[0099] Furthermore, the combined content of multiple feature combinations and the weights of each feature combination are used to generate a comprehensive feature vector under multi-feature fusion; multiple stage results are determined based on the operation of the comprehensive feature vector, and the primary recognition result of the food is determined based on the multiple stage results and the corresponding result mapping relationship, which is compatible with the overall consideration of shape features, texture features and knife features, realizes the technical effect of multi-feature fusion, and ensures the accuracy of the primary recognition result of the food.
[0100] At this time, based on multi-feature fusion, the contents of multiple feature combinations and their weights are integrated into a comprehensive feature vector; according to the feature combinations generated in step S131, the specific numerical values or vector representations of the features in each combination are extracted; the weights of each feature combination are applied to their corresponding feature values or vectors, usually through weighted summation or weighted connection (concatenation with weights considered in subsequent steps); the weighted feature values or vectors are combined into a comprehensive feature vector, which contains information about all feature combinations, and the contribution of each feature combination is determined by its weight.
[0101] By performing operations on the comprehensive feature vector, multiple recognition results or phased outputs are obtained; the operation method depends on the recognition model or algorithm used, which can be a classifier (such as SVM, KNN, neural network, etc.), regression model, clustering algorithm, etc.; the result of the operation is a preliminary prediction of the food category, a similarity score, a probability distribution or other forms of output, which represent the model's different interpretations or classification tendencies of the input data.
[0102] Based on multiple phased results and preset result mapping relationships, the final primary recognition result of the food is determined; at this time, the result mapping relationship presents the mapping of the phased results to specific food categories, which is usually determined by labels in the training process or subsequent processing rules; according to the result mapping relationship, the output that best matches the phased results is selected as the primary recognition result of the food.
[0103] Specifically, suppose that a food recognition system is being developed. Three feature combinations have been generated according to step S131 and weights have been assigned to them. Now, these feature combinations and weights are used in step S132. Assume that there are the following feature combinations and their weights:
[0104] Combination 1 (shape feature + texture feature): weight 0.6;
[0105] Combination 2 (shape feature + knife feature): weight 0.3;
[0106] Combination 3 (texture feature + knife feature): weight 0.1;
[0107] For each combination, the corresponding feature vector is extracted and weighted summation is performed by applying weights (or weighted concatenation is performed and the weights are considered in subsequent steps); finally, a comprehensive feature vector is obtained, which contains the information of all feature combinations, and the contribution of each combination is determined by its weight.
[0108] A trained neural network model is used to operate on the comprehensive feature vector. The output of the neural network model is the probability distribution of three food categories (apple, banana, and orange). The interim results are: apple 0.7, banana 0.2, orange 0.1. Based on the resulting mapping relationship, the category with the highest probability is selected as the primary recognition result of the food. In this example, the probability of apple is the highest (0.7), so the food is identified as apple.
[0109] refer to Figure 5 , in step S14, determining an advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food;
[0110] In the specific implementation process of the present invention, the specific steps are:
[0111] S141: determining a plurality of dynamic images based on the acquisition of the dynamic position of the food, marking corresponding dynamic features on the plurality of dynamic images, and determining the outer contour of the food by contour processing of the food images;
[0112] S142: Determine multiple primary result combinations based on the outer contour of the food, the multiple dynamic features, and the primary recognition result of the food;
[0113] S143: Determine corresponding combination coefficients based on the recognition of multiple primary result combinations, and determine advanced food recognition results based on the multiple combination coefficients, the primary food recognition results, and the corresponding result optimization mapping relationship.
[0114] In an embodiment of the present application, multiple dynamic images are determined based on the acquisition of the dynamic position of food, and corresponding dynamic features are marked on the multiple dynamic images. At the same time, the outer contour of the food is determined based on the contour processing of the food image, and the outer contour of the food is introduced.
[0115] At this point, a series of dynamic images are obtained by capturing the position information of the food in space that changes over time. At this point, cameras, sensors, or other image acquisition devices are used to capture the movement of the food. These devices are fixed in specific positions and move with the food. The frequency of image acquisition is determined based on the movement speed of the food and the required recognition accuracy. Faster movements require a higher acquisition frequency to capture details. Ensure that the acquired images have sufficient resolution and clarity so that features and contours can be accurately extracted in subsequent steps.
[0116] Extract features related to food movement from each dynamic image for subsequent analysis; at this time, select features related to food movement, such as position, speed, acceleration, direction, etc., according to task requirements; use image processing algorithms or machine learning models to extract the selected features from each dynamic image, which involves image preprocessing (such as denoising, contrast enhancement, etc.), feature detection (such as edge detection, corner detection, etc.) and feature calculation (such as calculating the center of mass, velocity vector, etc.); associate the extracted features with the corresponding dynamic image for subsequent analysis and processing.
[0117] The outer contour of the food is extracted from the dynamic image for subsequent shape analysis and recognition. At this time, image processing algorithms (such as the Canny edge detector, Sobel operator, Laplacian operator, etc.) are used to detect the edges in the food image. The extracted contour is smoothed to remove noise and unnecessary details to obtain a more accurate contour representation. The optimized contour is represented as a series of points or curves for subsequent shape matching, recognition, and other processing.
[0118] Specifically, suppose you are developing an intelligent kitchen assistant system that can automatically identify and track food placed on the kitchen countertop. A high-definition camera fixed to the kitchen ceiling is used to capture the food on the countertop. The camera captures images at a speed of 30 frames per second to ensure that the rapid movement of food can be captured. The captured images are stored in the system's local storage for subsequent analysis.
[0119] From each dynamic image, the food's location features (such as center of mass coordinates), speed features (such as the distance moved per second), and direction features (such as the direction angle of movement) are extracted. These features are marked and associated with the corresponding image.
[0120] The Canny edge detection algorithm is used to extract the outer contour of the food from each dynamic image. The extracted contour is then smoothed to remove noise and unnecessary details. Ultimately, a point set representing the outer contour of the food is obtained, which will be used for subsequent shape analysis and recognition. Through this example, we can see the role of step S141 in the intelligent kitchen assistant system. By capturing the dynamic position information of the food and extracting dynamic features and outer contours, it provides an important data foundation for subsequent food recognition and tracking.
[0121] Furthermore, multiple primary result combinations are determined based on the outer contour of the food, multiple dynamic features and the primary recognition results of the food, which is compatible with the overall consideration of the outer contour of the food, multiple dynamic features and the primary recognition results of the food, and ensures the accuracy of the multiple primary result combinations.
[0122] At this point, the outer contour information of the food is integrated into the recognition process as an important basis for shape analysis; at this point, in step S141, the outer contour of the food has been extracted from the dynamic image. This step is to ensure the accuracy and completeness of the contour information; the contour is represented as a series of points or curves for subsequent shape matching and recognition, which is achieved through the contour representation function in the image processing library (such as OpenCV); features are extracted from the contour, such as perimeter, area, shape index (such as circularity, rectangularity, etc.), and these features will be used for subsequent analysis and recognition.
[0123] Multiple dynamic features (such as position, speed, acceleration, direction, etc.) are integrated into the recognition process to capture the dynamic behavior of food. At this time, according to the task requirements, features related to the dynamic behavior of food are selected, including position changes, speed changes, acceleration changes, and direction changes. The extracted features are standardized to ensure that they are on the same scale for subsequent analysis and comparison. Multiple dynamic features are fused into a comprehensive feature vector for subsequent recognition and analysis, which is achieved through simple feature concatenation or more complex feature fusion methods (such as PCA, LDA, etc.).
[0124] Combine the primary recognition result of food (such as preliminary classification based on static images) with dynamic features and contour information to improve the accuracy and robustness of recognition; at this time, in the previous step (such as S132), the primary recognition result of food has been obtained, which is a probability distribution, category label or other form of output; design a fusion strategy to combine the primary recognition result with dynamic features and contour information, which is achieved through weighted summation, Bayesian fusion, decision tree fusion and other methods; output the fused result, which is an updated probability distribution, a more reliable category label or other form of output.
[0125] Specifically, suppose that a smart restaurant system is being developed that can automatically identify the types of food on the customer's table; the outer contour of the food extracted from the dynamic image is represented as a series of points, and the contour's perimeter, area, and circularity are calculated. These features are used for subsequent shape matching and recognition.
[0126] Dynamic features such as the position, speed, and direction of the food are extracted from the dynamic images. These features are normalized and fused into a comprehensive feature vector; for example, the position feature is represented as a two-dimensional coordinate (x, y), the speed feature is represented as a velocity vector (vx, vy), and the direction feature is represented as a direction angle θ; then, these features are concatenated into a comprehensive feature vector [x, y, vx, vy, θ].
[0127] In the previous step, the food was preliminarily classified based on the static image, resulting in a probability distribution representing the probability of the food belonging to different categories. In step S142, this probability distribution is combined with the dynamic and contour features. Specifically, a weighted summation method is used to weight the probability distribution of the primary recognition result and the scores of the dynamic and contour features to obtain an updated probability distribution. This updated probability distribution is more accurate and reliable because it comprehensively considers the static, dynamic, and shape features of the food. Through this example, we can see the role of step S142 in the smart restaurant system; it improves the accuracy and robustness of food recognition by integrating the food's outer contour information, multiple dynamic features, and the primary recognition results.
[0128] Therefore, the corresponding combination coefficient is determined based on the identification of multiple primary result combinations, and the advanced recognition result of the food is determined based on the multiple combination coefficients, the primary recognition results of the food and the corresponding result optimization mapping relationship. This is compatible with the overall consideration of multiple combination coefficients, the primary recognition results of the food and the corresponding result optimization mapping relationship, ensuring the accuracy of the advanced recognition results of the food.
[0129] At this point, the recognition accuracy and reliability of each primary result combination are evaluated and a combination coefficient is assigned to it; at the same time, for each primary result combination, its recognition accuracy is evaluated, which is achieved by comparing the recognition results in the combination with the true label (if any), or using some form of validation set for evaluation; based on the evaluation results, a combination coefficient is calculated for each primary result combination, which reflects the recognition accuracy, consistency or other relevant indicators of the combination; the combination coefficient is usually a value between 0 and 1, where a higher value indicates higher accuracy and reliability; in order to ensure that the sum of all combination coefficients is 1 (or a fixed value), the calculated coefficients are normalized, which helps to fairly weight the contribution of each combination in subsequent steps.
[0130] Optionally, assume that an intelligent food recognition system is being developed that can automatically identify the types of food on supermarket shelves; assume that there are three primary result combinations, which are identified based on color features, shape features, and texture features respectively; use a validation set to evaluate the recognition accuracy of each combination and assign it a combination coefficient; for example, the color feature combination obtains a coefficient of 0.4, the shape feature combination obtains a coefficient of 0.3, and the texture feature combination obtains a coefficient of 0.3. These coefficients reflect the relative importance and accuracy of each combination in the recognition process.
[0131] Multiple primary recognition results and their corresponding combination coefficients are integrated together for final advanced recognition; at this time, each primary recognition result is represented as a vector or probability distribution for weighted summation with the combination coefficient; using the combination coefficient as a weight, each primary recognition result is weighted summed to obtain a comprehensive recognition result that combines the advantages of multiple primary results; the result after weighted summation is subjected to necessary processing, such as normalization, smoothing or denoising, to ensure its accuracy and reliability.
[0132] Optionally, assume that each primary recognition result is a probability distribution, representing the probability that the food belongs to different categories; these probability distributions are weighted and summed with the corresponding combination coefficients to obtain a comprehensive probability distribution, which combines the advantages of color, shape and texture features to provide more accurate recognition results.
[0133] Utilize the known result optimization mapping relationship to further optimize the comprehensive recognition result to determine the high-level recognition result of the food; at this point, in the previous step or training process, a result optimization mapping relationship has been established, which is a lookup table, decision tree, machine learning model or other form of mapping; input the comprehensive recognition result into the mapping relationship to obtain the optimized high-level recognition result, which is a more accurate category label, probability distribution or other form of output; if so, use the validation set or real label to verify the high-level recognition result to ensure its accuracy and reliability.
[0134] Optionally, in the previous step, a machine learning model has been trained to optimize the mapping relationship as a result. This model can output a more accurate category label based on the comprehensive probability distribution; the comprehensive probability distribution is input into the model to obtain an optimized high-level recognition result; for example, the model outputs "apple" as the final recognition result, which is more accurate and reliable than any single primary result.
[0135] In one embodiment of the present application, it is assumed that there are the following primary recognition results and combination coefficients:
[0136] Primary result 1 (color recognition): Apple, combination coefficient 0.6
[0137] Primary result 2 (shape recognition): pear, combination coefficient 0.3
[0138] Primary result 3 (texture recognition): apple, combination coefficient 0.1
[0139] The optimized advanced recognition result matching table is shown in Table 1:
[0140] Table 1. Matching table of optimized advanced recognition results
[0141] Primary recognition result combination Optimized advanced recognition results Apple + Apple apple Apple + Pear Apple (assuming apples are more common) pear + pear pear
[0142] In this example, the combination coefficient is used as the weight to calculate the weighted similarity; since "apple" appears twice in the primary results (once for color recognition and once for texture recognition), and the sum of its combination coefficients is high (0.6+0.1=0.7), "apple" is considered to be a more recognized result; according to the optimized advanced recognition result matching table, the advanced recognition result of the food is finally determined to be "apple".
[0143] refer to Figure 6 In step S15, a review area is determined based on the high-level recognition result of the food and the food image, and a final recognition result of the food is determined based on the features corresponding to the review area, the high-level recognition result of the food, and the review mapping relationship;
[0144] In the specific implementation process of the present invention, the specific steps are:
[0145] S151: Determine multiple core features based on the high-level recognition result of the food and the comparison of the food image, and determine a review area based on the relative positions of the multiple core features, the shapes of the core features, and the corresponding sides of the food;
[0146] S152: Determine the features corresponding to the review area based on the identification of the review area, and determine the final recognition result of the food based on the features corresponding to the review area, the high-level recognition result of the food, and the review mapping relationship.
[0147] In an embodiment of the present application, multiple core features are determined based on the high-level recognition results of food and the comparison of food images, and the review area is determined based on the relative positions of the multiple core features, the shapes of the core features and the corresponding sides of the food. This takes into account the overall consideration of the relative positions of the multiple core features, the shapes of the core features and the corresponding sides of the food to ensure the accuracy of the review area.
[0148] At this time, the high-level recognition results based on food (i.e., the output of step S143) are compared with the original food image to determine the core features that have a significant impact on the recognition results; at this time, it is necessary to understand and interpret the high-level recognition results, which are usually one or more category labels, probability distributions or other forms of output, representing the type or attribute of food; the high-level recognition results are compared with the original food image to find the feature areas in the image that correspond to the recognition results. These feature areas include color, shape, texture, edges, etc.; from the comparison results, the core features that have a significant impact on the recognition results are extracted. These features are usually those that can significantly distinguish different types or attributes of food.
[0149] Optionally, assume that an intelligent food recognition system is being developed that can automatically identify the types of food on supermarket shelves; the system's high-level recognition identifies a certain food as "apple" and outputs a probability distribution representing the probability that the food belongs to different categories; at the same time, the system obtains the original image of the food; by comparing with the high-level recognition result, it is found that the color features (red) and shape features (circle) of the apple in the image are highly consistent with the recognition result.
[0150] Further analysis of the relative positional relationship of the core features in the food image and their respective morphological characteristics will provide a basis for the subsequent determination of the review area. At this point, the position of each core feature in the food image and the relative positional relationship between them will be determined, which will help to understand the spatial distribution and mutual correlation between the features. The morphology of each core feature will be analyzed in detail, including size, shape, texture, color, etc. These features reflect the physical properties and appearance characteristics of the food.
[0151] Optionally, the relative positions and morphological features of the color and shape features of the apple in the image are further analyzed; it is found that the color features are mainly distributed on the surface of the apple, while the shape features appear as a complete circular outline.
[0152] Based on the relative position and shape of the core features and the corresponding sides of the food (such as top, side, bottom, etc.), the image areas that require in-depth review are determined. At this point, the high-level recognition results and image information of the food are used to identify the side of the food shown in the image. This requires combining the 3D model of the food and the perspective information in the image.
[0153] Based on the relative position and morphological characteristics of the core features, as well as the side of the food display, image areas that require in-depth review are identified. These areas are usually those where the features are not obvious, there are recognition errors, or further confirmation is needed.
[0154] Optionally, based on the side of the apple shown in the image (assuming it is the top), and the relative position and morphological features of the color and shape features, the bottom area of the apple is delineated as the review area. This is because the color features of the bottom area of the apple are not obvious due to lighting, occlusion or viewing angle, and the shape features are also affected to a certain extent; by reviewing this area, the accuracy of the recognition is further confirmed, and potential recognition errors are discovered; through this example, we can see the role of step S151 in the intelligent food recognition system; it determines the core features by analyzing the high-level recognition results and image information of the food, and further analyzes the relative position and morphological features of these features, and finally determines the image area that needs to be reviewed in depth. This process helps to improve the accuracy and reliability of food recognition.
[0155] Furthermore, the features corresponding to the review area are determined based on the identification of the review area, and the final identification result of the food is determined based on the features corresponding to the review area, the advanced identification result of the food, and the review mapping relationship. This is compatible with the overall consideration of the features corresponding to the review area, the advanced identification result of the food, and the review mapping relationship, ensuring the accuracy of the final identification result of the food. At the same time, it realizes the step-by-step control of the primary identification result, advanced identification result, and final identification result of the food, further ensuring the high accuracy of the final identification result of the food, and realizing the precise multi-level control of the identification of the food.
[0156] At this point, the review area is carefully analyzed to identify and extract the features of the area, which will be used for subsequent comprehensive evaluation. At the same time, the review area determined in step S151 is carefully analyzed, which requires the use of higher resolution images or more advanced image processing technology to enhance details. Within the review area, key features are identified and extracted, including color, texture, shape, edges, etc., depending on the food type and recognition requirements. The extracted features are recorded and compared with the core features extracted previously to evaluate the consistency of the review area with the overall food image.
[0157] Optionally, assume that an intelligent food recognition system is being developed to identify apples on supermarket shelves; in step S151, the bottom area of the apple is determined as the review area; in step S152, the area is carefully analyzed, and features such as color (slightly green) and shape (slightly irregular) are extracted, which are slightly different from the overall color (red) and shape (round) features of the apple.
[0158] A comprehensive evaluation is conducted on the review area features, the advanced recognition results of the food, and the pre-set review mapping relationship to determine the accuracy of the final recognition result; at this time, the review mapping relationship is a mapping table or model constructed based on historical data, expert knowledge or machine learning algorithm, which is used to evaluate the consistency between the review area features and the advanced recognition results; the review area features are compared with the advanced recognition results to check whether there are obvious inconsistencies or conflicts between them; combined with the review mapping relationship, a comprehensive evaluation is conducted on the review area features and the advanced recognition results, which involves weighting, scoring or sorting the features to determine the reliability of the final recognition result.
[0159] Optionally, the review area features are comprehensively evaluated with the advanced recognition results (apple) and the review mapping relationship; the review mapping relationship tells us that the color of the bottom area of the apple varies due to different lighting or maturity, but the shape should generally remain round; in this example, there is a slight inconsistency between the color features of the review area and the advanced recognition results, but the shape features, although slightly irregular, are still within an acceptable range.
[0160] Based on the results of the comprehensive evaluation, the final recognition result of the food is determined; at this time, based on the results of the comprehensive evaluation, decision rules are formulated to determine the final recognition result, which includes setting thresholds, majority voting, weighted average and other methods; according to the decision rules, the final recognition result of the food is output, which is a category label, probability distribution or other form of output, depending on the design and requirements of the system.
[0161] Optionally, a decision rule is formulated based on the results of the comprehensive evaluation; in this example, a threshold is set, and if the degree of inconsistency between the review area features and the advanced recognition results is lower than the threshold, the advanced recognition results are maintained; otherwise, the recognition algorithm needs to be re-evaluated or adjusted; in this example, since the degree of inconsistency between the review area features and the advanced recognition results is low, the final recognition result of the food is determined to be "apple".
[0162] Through this example, we can see the role of step S152 in the intelligent food recognition system; it conducts a comprehensive evaluation of the review area characteristics, combined with the advanced recognition results and the review mapping relationship, and finally determines the final recognition result of the food. This process helps to improve the accuracy and reliability of recognition, while being able to deal with uncertainties and errors in the recognition process.
[0163] In one embodiment of the present application, assume that a fruit is being identified, and the advanced recognition result is initially determined to be "apple." In step S151, the bottom of the fruit is determined as the review area, and the features of this area are extracted: the color is light green and the shape is slightly flat. Now, the final recognition result matching table is used to determine the final recognition result. The final recognition result matching table is shown in Table 2:
[0164] Table 2 Final recognition result matching table
[0165] Review regional characteristics Final recognition results Color: red, shape: round apple Color: green, shape: round Green Apple Color: light green, shape: flat pear
[0166] In this example, the matching table was searched and it was found that the features of the review area were closest to the description of "pear" (light green in color and slightly flat in shape); although the advanced recognition result was initially judged to be "apple", based on the review area features and the matching table, the final recognition result of the food was determined to be "pear".
[0167] In another embodiment of the present application, data enhancement operations such as illumination enhancement, angle transformation, and noise processing are performed on the input food image to improve the image quality; the shape features, texture features, and knife skills features of the food image are extracted respectively through the shape recognition module, the texture recognition module, and the knife skills recognition module; the above three types of features are weightedly fused through the feature fusion layer to generate a comprehensive feature vector; based on the comprehensive feature vector, the cuisine judgment and cooking method inference are performed, and the final classification result is output.
[0168] In the shape recognition module, the methods of contour extraction and shape feature calculation are adopted to complete the food shape classification (such as square, strip, sheet, block, roll and sphere, etc.) through the analysis of shape features such as area, perimeter, roundness and rectangularity.
[0169] In the texture recognition module, the texture features of food images are extracted through gray-level co-occurrence matrix (GLCM) analysis, local binary pattern (LBP) feature extraction and directional gradient calculation, and the texture patterns are classified (such as cross-cut texture, oblique texture, longitudinal texture and composite texture).
[0170] In the knife skill recognition module, based on edge detection and knife skill feature extraction, the cutting shape of the food (such as slender strips, cubes, thin slices and short columns, etc.) is analyzed, and the knife skill quality is rated in combination with size measurement and uniformity analysis.
[0171] In the feature fusion layer, by assigning different weights to shape features, texture features and knife features, weighted fusion of features is achieved, thereby generating a more recognizable comprehensive feature vector; in the comprehensive analysis stage, the comprehensive feature vector is processed by a multi-dimensional analysis model, and the model weight is dynamically adjusted based on user feedback to improve the accuracy of cuisine judgment and cooking method inference; the present invention introduces multi-feature fusion and attention mechanism, combined with transfer learning and ensemble learning strategies of deep learning models, to achieve high-precision and high-robustness food image recognition in complex scenarios, providing reliable technical support for fields such as smart catering and smart food equipment.
[0172] Furthermore, the shape recognition module is implemented as follows:
[0173] The shape recognition module calculates the shape features of food images, including area, perimeter, roundness, and rectangularity, through image segmentation and contour extraction. The calculated results of the shape features are vectorized for further classification. This module can classify food into categories such as square, strip, slice, block, roll, and sphere, providing important geometric information support for subsequent feature fusion.
[0174] Implementation of texture recognition module:
[0175] After grayscale conversion, the texture recognition module uses gray-level co-occurrence matrix (GLCM), local binary pattern (LBP) and directional gradient analysis to extract texture features. The texture features are fused and classified to identify food texture patterns (such as cross-cut texture, oblique texture, longitudinal texture and composite texture). This multi-dimensional texture feature extraction method can effectively cope with interference from lighting changes and complex backgrounds.
[0176] Implementation of knife recognition module:
[0177] The knife-cutting recognition module uses edge detection technology to extract food knife-cutting features, such as slender strips, cubic blocks, thin slices, and short columns; it rates the quality of knife-cutting through size measurement and uniformity analysis; this module not only reflects the food processing technology, but also provides an important basis for judging cuisine and inferring cooking methods.
[0178] In the feature fusion module, shape, texture and knife features are integrated into a comprehensive feature vector through weighted fusion; during the fusion process, the feature weights are optimized using a multidimensional analysis method to ensure that the contribution of each feature to the final classification result is optimal; the comprehensive analysis module receives the fused comprehensive feature vector and, combined with the multidimensional analysis model, makes cuisine judgments and cooking method inferences on food images; this module can optimize the performance of the classification model according to actual application scenarios through a dynamic weight adjustment mechanism; in addition, the user feedback module can collect users' evaluations of the classification results, further adjust the feature weights and update the model, thereby improving the intelligence and adaptability of the system; in order to improve the generalization ability of the model, the present invention introduces optimization strategies such as transfer learning, attention mechanism and ensemble learning; in terms of data enhancement, the diversity of the training data set is expanded through means such as illumination enhancement, angle transformation and noise processing, thereby enhancing the model's adaptability to complex scenarios.
[0179] See also Figure 7 , Figure 7 : is a schematic diagram of the structural composition of a food recognition system based on multi-feature fusion in an embodiment of the present invention; the food recognition system based on multi-feature fusion includes:
[0180] An image to be identified module 21 is configured to output an image to be identified based on image preprocessing of the food image;
[0181] A feature module 22 is used to determine shape features, texture features, and knife features based on the recognition of the image to be recognized;
[0182] A primary recognition result module 23 is configured to generate a comprehensive feature vector based on the fusion of multiple features including shape features, texture features, and knife-cut features, and to determine a primary recognition result of the food based on the comprehensive feature vector;
[0183] an advanced recognition result module 24 for determining an advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food;
[0184] The final recognition result module 25 is used to determine the review area based on the high-level recognition result of the food and the food image, and determine the final recognition result of the food according to the features corresponding to the review area, the high-level recognition result of the food and the review mapping relationship.
[0185] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. A food recognition method based on multi-feature fusion, characterized in that: include: outputting an image to be recognized according to image preprocessing of the food image; Determining shape features, texture features, and knife features based on the recognition of the image to be recognized; Generate a comprehensive feature vector based on the fusion of multiple features such as shape features, texture features, and knife features, and determine the primary recognition result of the food based on the comprehensive feature vector; determining an advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food; A review area is determined based on the high-level recognition result of the food and the food image, and a final recognition result of the food is determined based on the features corresponding to the review area, the high-level recognition result of the food, and the review mapping relationship.
2. The food recognition method based on multi-feature fusion according to claim 1 is characterized in that: The step of outputting an image to be recognized based on image preprocessing of the food image includes: generating a plurality of sub-images in different directions based on the circular photographing of the food, and determining a food image by synthesizing the plurality of sub-images, the food image being used as a panoramic image of the food; The food image is preprocessed and a corresponding grayscale image is output. A plurality of areas to be detected are determined based on the detection of the grayscale image, and an image to be identified is determined based on the synthesis of the plurality of areas to be detected.
3. The food recognition method based on multi-feature fusion according to claim 1 is characterized in that: The determining of shape features, texture features, and knife features based on the recognition of the image to be recognized includes: Based on the position division of the image to be identified, multiple sub-regions are determined, and based on the screening of the multiple sub-regions, a shape region set, a texture region set, and a knife-work region set are determined; regions related to shape, texture, and knife-work are screened out from the multiple sub-regions obtained by segmentation, and they are respectively classified into corresponding sets; at this time, feature analysis is performed on each sub-region, including shape, texture, and knife-work traces, and the sub-regions are classified into shape regions, texture regions, or knife-work regions based on the results of the feature analysis and preset rules; the classification is based on the similarity of shape descriptors, the statistical characteristics of texture descriptors, and the presence or absence of cutting edges; the classified sub-regions are respectively classified into the shape region set, the texture region set, and the knife-work region set; Determining a plurality of sub-shape features based on the recognition of the shape region set, and determining a shape feature based on the morphology of the plurality of sub-shape features, the spatial positions of the plurality of sub-shape features, and the shape mapping relationship; Determining a plurality of sub-texture features based on the identification of the texture region set, and determining a texture feature based on the grain of the plurality of sub-texture features, the spatial positions of the plurality of sub-texture features, and a texture mapping relationship; A plurality of sub-cutting features are determined according to the identification of the cutting area set, and the cutting features are determined according to the tool paths of the plurality of sub-cutting features, the spatial positions of the plurality of sub-cutting features and the cutting mapping relationship.
4. The food recognition method based on multi-feature fusion according to claim 1 is characterized in that: The multi-feature fusion based on shape features, texture features and knife features generates a comprehensive feature vector, and determines the primary recognition result of the food according to the comprehensive feature vector, including: Corresponding weights are configured for shape features, texture features and knife features. At the same time, multiple feature combinations are generated based on the shape features, texture features and knife features, and the weights of the multiple feature combinations are the weighted results of the corresponding weights; the combination methods include: shape features + texture features, shape features + knife features, texture features + knife features, and a combination of the three; for each feature combination, the weight of each feature it contains is weightedly summed to obtain the weights of the multiple feature combinations.
5. The food recognition method based on multi-feature fusion according to claim 4 is characterized in that: Based on the fusion of multiple features such as shape features, texture features, and knife features, a comprehensive feature vector is generated, and the primary recognition result of the food is determined based on the comprehensive feature vector, which also includes: The combined content of multiple feature combinations and the weights of each feature combination generate a comprehensive feature vector under multi-feature fusion; multiple interim results are determined based on the operation of the comprehensive feature vector, and the primary recognition result of the food is determined based on the multiple interim results and the corresponding result mapping relationship; the comprehensive feature vector contains information on all feature combinations, and the contribution of each feature combination is determined by its weight; at this time, a trained neural network model is used to operate on the comprehensive feature vector, and the output of the neural network model is the probability distribution of the three food categories; the interim results correspond to the probability distribution of the three food categories, and according to the result mapping relationship, the category with the highest probability is selected as the primary recognition result of the food.
6. The food recognition method based on multi-feature fusion according to claim 1 is characterized in that: The step of determining the advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food includes: Based on the acquisition of the dynamic position of the food, multiple dynamic images are determined, and corresponding dynamic features are marked on the multiple dynamic images. At the same time, the outer contour of the food is determined based on the contour processing of the food image; by capturing the position information of the food in space that changes with time, a series of dynamic images are obtained.
7. The food recognition method based on multi-feature fusion according to claim 6 is characterized in that: The step of determining the advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food further includes: determining a plurality of primary result combinations based on an outer contour of the food, a plurality of dynamic features, and a primary recognition result of the food; The corresponding combination coefficient is determined based on the recognition of multiple primary result combinations, and the advanced recognition result of the food is determined based on the multiple combination coefficients, the primary recognition results of the food and the corresponding result optimization mapping relationship.
8. The food recognition method based on multi-feature fusion according to claim 1, characterized in that: The step of determining a review area based on the high-level recognition result of the food and the food image, and determining a final recognition result of the food according to features corresponding to the review area, the high-level recognition result of the food, and the review mapping relationship, includes: Multiple core features are determined based on the high-level recognition results of the food and the comparison of the food image, and the review area is determined based on the relative positions of the multiple core features, the shapes of the core features and the corresponding sides of the food.
9. The food recognition method based on multi-feature fusion according to claim 8, characterized in that: The step of determining a review area based on the high-level recognition result of the food and the food image, and determining a final recognition result of the food according to features corresponding to the review area, the high-level recognition result of the food, and the review mapping relationship, further includes: The features corresponding to the review area are determined based on the identification of the review area, and the final recognition result of the food is determined based on the features corresponding to the review area, the high-level recognition result of the food, and the review mapping relationship.
10. A food recognition system based on multi-feature fusion, characterized in that: The food recognition system based on multi-feature fusion is applied to the food recognition method based on multi-feature fusion according to any one of claims 1 to 9, and the food recognition system based on multi-feature fusion includes: An image to be identified module, configured to output an image to be identified based on image preprocessing of the food image; A feature module, used to determine shape features, texture features, and knife features based on the recognition of the image to be recognized; A primary recognition result module is used to generate a comprehensive feature vector based on the fusion of multiple features such as shape features, texture features, and knife features, and to determine the primary recognition result of the food based on the comprehensive feature vector; an advanced recognition result module, configured to determine an advanced recognition result of the food based on the outer contour of the food, the multiple dynamic images, and the primary recognition result of the food; The final recognition result module is used to determine the review area based on the high-level recognition result of the food and the food image, and to determine the final recognition result of the food according to the features corresponding to the review area, the high-level recognition result of the food and the review mapping relationship.