Feature extraction and recognition system and method for smart restaurant dish images
By proposing a dish image feature extraction and recognition system in a smart restaurant, using feature extraction and clustering modules such as color, texture, and shape for classification, the problem of inaccurate dish image recognition in the existing technology is solved, and efficient and accurate dish recognition and management is achieved.
Patent Information
- Application Number
- CN202510045192.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has the problem of inaccurate identification of dishes in smart restaurants, especially under factors such as complex background, light changes and dish placement angles. The prior art has limited the extraction of dishes image features to one dimension, resulting in inaccurate analysis.
A feature extraction and recognition system for smart restaurant dishes images is proposed. The dish images are obtained through the acquisition module. The feature extraction module extracts and analyzes the image in color, texture, shape and other features, constructs the dish feature vector, and classifies the feature vectors through the clustering module, and finally uses the pre-trained dish recognition model to identify them.
It improves the accuracy and efficiency of dish image recognition, can adapt to the identification needs of multiple dishes, has good versatility and expansion, helps smart restaurants better manage and analyze dishes, and improve service quality and operational efficiency.
Smart Images

Figure CN120071331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular to a system and method for feature extraction and recognition of dish images in a smart restaurant. Background Art
[0002] With the continuous acceleration of the pace of life and the continuous improvement of people's living standards, the rise of the service industry has been brought about. The emergence of smart restaurants has enabled people to enjoy the convenience brought by technology.
[0003] A smart restaurant is a new business form that deeply integrates the catering industry with modern information technology. It introduces cutting-edge technologies such as cloud computing, big data, the Internet of Things, and artificial intelligence to intelligently transform and upgrade traditional catering services, making catering services more convenient, efficient, and personalized.
[0004] In the prior art, when people dine in a smart restaurant and use self-checkout, the recognition of dishes often appears inaccurate. Especially when facing various factors such as complex backgrounds, light changes, and dish placement angles, the recognition often appears inaccurate. Secondly, the feature extraction of dish images is limited to one dimension, resulting in inaccurate analysis of dishes.
[0005] For these reasons, the existing technologies that can accurately recognize dishes are very limited at present, and these situations have all increased the difficulty of accurately recognizing dish images. Therefore, there is an urgent need for a system and method for feature extraction and recognition of dish images in a smart restaurant to solve the above problems. Summary of the Invention
[0006] The present invention aims to solve at least one of the technical problems in the above technologies to some extent. To this end, the first aspect of the present invention aims to propose a system for feature extraction and recognition of dish images in a smart restaurant, which extracts and analyzes features such as the color, texture, and shape of dish images, constructs dish feature vectors, and classifies the feature vectors to achieve accurate recognition of dishes.
[0007] The second aspect of the present invention aims to propose a method for feature extraction and recognition of dish images in a smart restaurant.
[0008] To achieve the above object, an embodiment of the present invention proposes a system for feature extraction and recognition of dish images in a smart restaurant, including:
[0009] An acquisition module, configured to acquire a dish image in a smart restaurant as an image to be recognized;
[0010] A feature extraction module, configured to extract features from the image to be recognized, construct dish feature vectors, and obtain a plurality of dish feature vectors;
[0011] A clustering module, configured to cluster the plurality of dish feature vectors to obtain a plurality of sets of dish feature vectors;
[0012] An identification module, configured to respectively input the plurality of sets of dish feature vectors into a pre-trained dish identification model for identification to obtain an identification result of the image to be identified.
[0013] Preferably, the feature extraction module includes:
[0014] A texture feature extraction sub-module, configured to:
[0015] Perform grayscale processing on the image to be identified to obtain a grayscale image;
[0016] Obtain a gray-level co-occurrence matrix corresponding to the grayscale image;
[0017] Determine texture features corresponding to the image to be identified based on the gray-level co-occurrence matrix;
[0018] A color feature extraction sub-module, configured to:
[0019] Convert the image to be identified into the HSV color space;
[0020] Divide the HSV color space into a plurality of intervals, traverse all pixel points in the image to be identified, and determine a color histogram corresponding to the image to be identified;
[0021] Determine color features corresponding to the image to be identified based on the color histogram;
[0022] A shape feature extraction sub-module, configured to perform edge contour detection on the image to be identified to obtain a plurality of contours corresponding to the image to be identified; determine shape features corresponding to the image to be identified based on the plurality of contours;
[0023] A construction sub-module, configured to construct dish feature vectors based on the texture features, color features, and shape features corresponding to the image to be identified.
[0024] Preferably, the clustering module includes:
[0025] A calculation sub-module, configured to respectively calculate the density of the plurality of dish feature vectors to obtain a density value corresponding to each dish feature vector; determine a minimum Euclidean distance corresponding to each dish feature vector based on the density value corresponding to each dish feature vector; calculate the mean of the minimum Euclidean distances corresponding to all dish feature vectors to obtain a mean Euclidean distance;
[0026] A first determination sub-module, configured to compare the minimum Euclidean distance corresponding to each feature vector with the mean Euclidean distance, and use the dish feature vectors whose minimum Euclidean distance is greater than or equal to the mean Euclidean distance as initial clustering centers to obtain a plurality of initial clustering centers;
[0027] A second determination sub-module, configured to:
[0028] Based on the first density values and the minimum Euclidean distances corresponding to the several initial clustering centers, determine the clustering evaluation values of the several initial clustering centers;
[0029] Compare the clustering evaluation values of the several initial clustering centers with a preset clustering evaluation threshold respectively, and use the initial clustering centers corresponding to the clustering evaluation values greater than or equal to the preset clustering evaluation threshold as target clustering centers;
[0030] A clustering sub-module, configured to cluster several dish feature vectors based on the target clustering centers to obtain several dish feature vector sets.
[0031] Preferably, the calculation sub-module includes:
[0032] A first calculation unit, configured to:
[0033] Calculate the density of each of the several dish feature vectors respectively to obtain the density value corresponding to each dish feature vector;
[0034] Arbitrarily select a dish feature vector as a first target feature vector; use the density value corresponding to the first target feature vector as the target density value;
[0035] Compare the target density value with the density values corresponding to the other dish feature vectors except the first target feature vector;
[0036] Use the dish feature vectors corresponding to the density values of the other dish feature vectors except the first target feature vector that are greater than or equal to the target density value as second target feature vectors to obtain several second target feature vectors;
[0037] A second calculation unit, configured to:
[0038] Calculate the Euclidean distances between the first target feature vector and the several second target feature vectors respectively to obtain several Euclidean distances;
[0039] Obtain the minimum value among the several Euclidean distances as the minimum Euclidean distance corresponding to the first target feature vector;
[0040] Traverse all the dish feature vectors to obtain the minimum Euclidean distance corresponding to each dish feature vector;
[0041] A third calculation unit, configured to calculate the mean value of the minimum Euclidean distances corresponding to all the dish feature vectors to obtain the mean Euclidean distance.
[0042] Preferably, the first calculation unit is configured to calculate the density of each of the plurality of dish feature vectors respectively to obtain the density value corresponding to each dish feature vector, including:
[0043]
[0044] where ρ i represents the density value corresponding to the i-th dish feature vector; m represents the total number of dish feature vectors; d ij represents the Euclidean distance between the i-th dish feature vector and the j-th dish feature vector; d c represents a preset truncation distance.
[0045] Preferably, it further includes:
[0046] A preprocessing module, configured to preprocess the image to be recognized before feature extraction of the image to be recognized.
[0047] Preferably, the preprocessing module includes:
[0048] An enhancer module, configured to perform image enhancement on the image to be recognized to obtain an enhanced image to be recognized;
[0049] A determination sub-module, configured to use the enhanced image to be recognized as the preprocessed image to be recognized.
[0050] Preferably, the enhancer module includes:
[0051] A first acquisition unit, configured to:
[0052] Arbitrarily select an image to be recognized as the first image;
[0053] Divide the first image evenly into a plurality of sub-images;
[0054] A third calculation unit, configured to calculate the fuzzy evaluation value corresponding to each sub-image;
[0055] A first determination unit, configured to:
[0056] Compare the fuzzy evaluation value with a preset fuzzy evaluation threshold, and use the sub-image corresponding to the fuzzy evaluation value greater than or equal to the preset fuzzy evaluation threshold as the sub-image to be enhanced, to obtain a plurality of sub-images to be enhanced;
[0057] A second acquisition unit, configured to:
[0058] Arbitrarily select a sub-image to be enhanced as the second image;
[0059] Arbitrarily select a pixel point in the second image as the target pixel point;
[0060] Taking the target pixel as the center and the preset step length as the radius, determine the target area;
[0061] The fourth calculation unit is used for:
[0062] Calculate the grayscale mean value of all pixel points in the target area as the target grayscale mean value;
[0063] Calculate the absolute value of the difference between the grayscale value of the target pixel and the target grayscale mean value to obtain the target absolute difference;
[0064] Compare the target absolute difference with the preset absolute difference threshold;
[0065] The second determination unit is used to use the target pixel point corresponding to the target absolute difference being greater than or equal to the preset absolute difference threshold as the pixel point to be enhanced;
[0066] The enhancement unit is used for:
[0067] Replace the grayscale value of the pixel point to be enhanced based on the target grayscale mean value to obtain the enhanced pixel point;
[0068] Traverse all pixel points in the second image to obtain the enhanced second image;
[0069] Traverse all the sub-images to be enhanced to obtain the enhanced image to be recognized.
[0070] Preferably, the first calculation unit is used to calculate the blur evaluation value corresponding to each sub-image, including:
[0071] Arbitrarily select a sub-image and perform grayscale processing to obtain a grayscale sub-image;
[0072] Obtain the grayscale values of each pixel point in the grayscale sub-image;
[0073] Calculate the gradient amplitude of each pixel point in the grayscale sub-image;
[0074] Determine the gradient amplitude mean value corresponding to the grayscale sub-image based on the gradient amplitude of each pixel point;
[0075] Compare the gradient amplitude of each pixel point with the gradient amplitude mean value, and use the pixel points with the gradient amplitude greater than or equal to the gradient amplitude mean value as the first pixel points to obtain the first pixel point set;
[0076] Calculate the grayscale mean value of all pixel points in the grayscale sub-image;
[0077] Compare the grayscale value of each pixel point in the grayscale sub-image with the grayscale mean value, and use the pixel points with the grayscale value greater than or equal to the grayscale mean value as the second pixel points to obtain the second pixel point set;
[0078] Determine the fuzzy evaluation value of the grayscale sub-image based on the first set of pixel points and the second set of pixel points;
[0079] Traverse each sub-image to obtain the fuzzy evaluation value corresponding to each sub-image.
[0080] To achieve the above object, an embodiment of the present invention proposes a method for feature extraction and recognition of dish images in a smart restaurant, including:
[0081] Obtain the dish image of the smart restaurant as the image to be recognized;
[0082] Extract features from the image to be recognized, construct a dish feature vector, and obtain a number of dish feature vectors;
[0083] Cluster the number of dish feature vectors to obtain a number of dish feature vector sets;
[0084] Input the number of dish feature vector sets into a pre-trained dish recognition model for recognition respectively to obtain the recognition result of the image to be recognized.
[0085] The present invention provides a feature extraction and recognition system and method for dish images in a smart restaurant. The acquisition module quickly acquires images, the feature extraction module efficiently constructs feature vectors by extracting and analyzing features such as the color, texture, and shape of the dish image, and the clustering module performs classification processing. The overall process helps to improve the processing efficiency; the clustering module can classify vectors with similar features into one category, providing more targeted processing for subsequent recognition and improving the recognition accuracy; using a pre-trained dish recognition model for recognition can adapt to the recognition needs of various dishes, has good versatility and scalability, helps the smart restaurant better manage and analyze dishes, and improves the service quality and operation efficiency.
[0086] Other features and advantages of the present invention will be described in the subsequent description, and, in part, will be obvious from the description or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the written description and the drawings.
[0087] The technical solutions of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0088] The drawings are used to provide a further understanding of the present invention and constitute a part of the description. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0089] Figure 1 is a block diagram of a feature extraction and recognition system for dish images in a smart restaurant according to an embodiment of the present invention;
[0090] Figure 2 is a block diagram of a feature extraction module according to an embodiment of the present invention;
[0091] Figure 3 is a flowchart of a method for feature extraction and recognition of dish images in an intelligent restaurant according to an embodiment of the present invention. Detailed implementation manners
[0092] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0093] Embodiment 1
[0094] As Figure 1 shown, a feature extraction and recognition system for dish images in an intelligent restaurant includes:
[0095] An acquisition module, configured to acquire dish images in an intelligent restaurant as images to be recognized;
[0096] A feature extraction module, configured to extract features from the images to be recognized, construct dish feature vectors, and obtain a plurality of dish feature vectors;
[0097] A clustering module, configured to cluster the plurality of dish feature vectors to obtain a plurality of dish feature vector sets;
[0098] A recognition module, configured to input the plurality of dish feature vector sets into a pre-trained dish recognition model respectively for recognition, and obtain a recognition result of the images to be recognized.
[0099] In this embodiment, features are extracted from the images to be recognized, and the features include features such as the color, texture, and shape of the images to be recognized.
[0100] In this embodiment, the construction method of the dish recognition model includes:
[0101] Obtain a dish training data set;
[0102] Input the dish training data set into a neural network model for iterative training to obtain an initial dish recognition model;
[0103] Obtain a dish test data set;
[0104] Based on the dish test data set, test the initial dish recognition model, and when it is determined that the test result meets the requirements, obtain the final dish recognition model.
[0105] The beneficial effects of the above technical solution are as follows: The acquisition module quickly acquires images, the feature extraction module efficiently constructs feature vectors by extracting and analyzing features such as the color, texture, and shape of the dish images, and the clustering module performs classification processing. The overall process helps improve the processing efficiency; the clustering module can classify vectors with similar features into one category, providing more targeted processing for subsequent recognition and improving the accuracy of recognition; using a pre-trained dish recognition model for recognition can meet the recognition needs of various dishes, has good versatility and scalability, and helps the smart restaurant better manage and analyze dishes, improving service quality and operation efficiency.
[0106] Embodiment 2
[0107] As Figure 2 shown, the feature extraction module includes:
[0108] The texture feature extraction sub-module is used to:
[0109] Perform grayscale processing on the image to be recognized to obtain a grayscale image;
[0110] Obtain the gray-level co-occurrence matrix corresponding to the grayscale image;
[0111] Determine the texture feature corresponding to the image to be recognized based on the gray-level co-occurrence matrix;
[0112] The color feature extraction sub-module is used to:
[0113] Convert the image to be recognized into the HSV color space;
[0114] Divide the HSV color space into several intervals, traverse all pixel points in the image to be recognized, and determine the color histogram corresponding to the image to be recognized;
[0115] Determine the color feature corresponding to the image to be recognized based on the color histogram;
[0116] The shape feature extraction sub-module is used to perform edge contour detection on the image to be recognized to obtain several contours corresponding to the image to be recognized; determine the shape feature of the image to be recognized based on the several contours;
[0117] The construction sub-module is used to construct a dish feature vector based on the texture feature, color feature, and shape feature corresponding to the image to be recognized.
[0118] In this embodiment, the texture features include features such as contrast, energy, entropy, and uniformity.
[0119] In this embodiment, the color features include color distribution, contrast, and the mean, variance, median, and standard deviation of pixel points.
[0120] In this embodiment, the shape features include contour length, contour area, center distance, and eccentricity.
[0121] In this embodiment, the gray-level co-occurrence matrix is a statistical tool for describing the texture features of an image and is commonly used for texture feature extraction in image processing. The gray-level co-occurrence matrix is defined as the joint probability distribution of pixel pairs and is a symmetric matrix. It reflects the comprehensive information of the gray levels of the image in adjacent directions, adjacent intervals, and change amplitudes, and also reflects the positional distribution characteristics between pixels with the same gray level. It is the basis for calculating texture features. Suppose there is a gray-scale image with a maximum gray level of L. Arbitrarily take a point (x, y) in the image and a point (x + a, y + b) deviating from it (where a and b are integers, defined artificially) to form a point pair. Let the gray values of this point pair be (f1, f2), and then let the point (x, y) move on the entire image, then different (f1, f2) values will be obtained. For the entire image, count the number of occurrences of each (f1, f2) value, then arrange them into a square matrix, and then normalize them with the total number of occurrences of (f1, f2) to obtain the occurrence probability P(f1, f2). The resulting matrix is the gray-level co-occurrence matrix. Set the direction of the gray-level co-occurrence matrix (such as 0°, 45°, 90°, 135°, etc.) and the pixel interval d. In this example, the 0° direction and the pixel interval d = 1 are selected. For each pixel, count the frequency of occurrence of the gray value pairs between it and its neighboring pixels at an interval of 1 in the 0° direction. Arrange the statistical results into a square matrix and perform normalization. A series of texture features can be extracted from the normalized gray-level co-occurrence matrix. Common ones include: Energy: The sum of the squares of the elements of the gray-level co-occurrence matrix, which reflects the stability of the gray-level change of the image texture. A large energy value indicates that the current texture is a texture with relatively stable regular changes. Entropy: A measure of the randomness of the information content contained in the image. The entropy value indicates the complexity of the gray-level distribution of the image. The larger the entropy value, the more complex the image; Maximum probability: Represents the texture feature that appears most frequently in the image. Contrast: Measures how the values of the matrix are distributed and the amount of local change in the image, reflecting the clarity of the image and the depth of the texture grooves. The deeper the texture grooves, the greater the contrast and the clearer the effect; conversely, if the contrast value is small, the grooves are shallow and the effect is blurred. Inverse difference moment: Reflects the homogeneity of the image texture and measures the amount of local change in the image texture. A large value indicates that there is little change between different regions of the image texture and the local area is very uniform. Correlation: Reflects the consistency of the image texture. Measures the similarity degree of the elements of the spatial gray-level co-occurrence matrix in the row or column direction. The size of the correlation value reflects the local gray-level correlation in the image. When the matrix element values are uniformly equal, the correlation value is large; on the contrary, if the matrix pixel values vary greatly, the correlation value is small.
[0122] In this embodiment, the HSV color space consists of three components: hue, saturation, and value. To simplify the problem, each component can be divided into several intervals. For example: Hue (H): 0 - 360 degrees, which can be divided into 8 intervals, each interval being 45 degrees. Saturation (S): 0 - 100%, which can be divided into 5 intervals, each interval being 20%. Value (V): 0 - 100%, which can also be divided into 5 intervals, each interval being 20%. Assume that the image to be recognized is a two-dimensional matrix, where each element represents a pixel. Traverse all the pixels in the image. For each pixel, convert it from the RGB color space to the HSV color space. According to the divided HSV intervals, count the HSV interval to which each pixel belongs and calculate the number of pixels in each interval. In this way, a three-dimensional color histogram can be obtained, where the horizontal axis represents the hue interval, the vertical axis represents the combination of the saturation and value intervals, and the value of each interval represents the number of pixels in that interval. To simplify the representation, the three-dimensional color histogram can be projected onto a two-dimensional plane. For example, the saturation and value intervals can be combined into one dimension, or the histograms of hue, saturation, and value can be plotted separately. Based on the obtained color histogram, the color features of the image to be recognized can be determined. Common color features include: Dominant color: The color corresponding to the interval with the largest number of pixels in the color histogram. Color distribution: The distribution of the number of pixels in each color interval, which reflects the richness and diversity of colors in the image. Color contrast: The difference in the number of pixels between different color intervals, which reflects the degree of color contrast in the image. Suppose there is a simple image that only contains two colors, red and blue, and red and blue are evenly distributed in the image. After dividing the HSV color space into the above intervals and traversing all the pixels in the image, the following color histogram is obtained: Hue interval: The number of pixels in the red (0 - 45 degrees) and blue (210 - 255 degrees) intervals is relatively large, and the number of pixels in other hue intervals is relatively small. Saturation interval: Since both red and blue are pure colors, the number of pixels in the higher saturation intervals is relatively large. Value interval: Since there are no particularly bright or dark areas in the image, the value is evenly distributed in each interval. Based on the above color histogram, the color features of this image can be determined as: dominated by red and blue, with a relatively even color distribution and moderate color contrast.
[0123] In this embodiment, the specific method for performing edge contour detection on the image to be recognized can be an edge detection algorithm or an edge detection model.
[0124] The beneficial effects of the above technical solution are as follows: Different aspects of features of the image are extracted by the texture feature extraction sub-module, the color feature extraction sub-module, and the shape feature extraction sub-module respectively, making the description of the dish more comprehensive and accurate, covering key information such as texture, color, and shape. Detailed processing steps, such as determining texture by the gray-level co-occurrence matrix, analyzing color in the HSV color space, and determining shape by edge contour detection, etc., help to improve the accuracy of feature extraction, thus better identifying and distinguishing different dishes; The construction sub-module integrates various features into a dish feature vector, providing a concise and effective data representation for subsequent tasks such as recognition and classification, facilitating further analysis and processing; It is applicable to various types and styles of dish images and can handle complex and diverse dish appearances.
[0125] Embodiment 3
[0126] The clustering module includes:
[0127] A calculation sub-module, which is used to calculate the density of each of the several dish feature vectors respectively to obtain the density value corresponding to each dish feature vector; determine the minimum Euclidean distance corresponding to each dish feature vector based on the density value corresponding to each dish feature vector; calculate the mean value of the minimum Euclidean distances corresponding to all dish feature vectors to obtain the mean Euclidean distance; A first determination sub-module, which is used to compare the minimum Euclidean distance corresponding to each feature vector with the mean Euclidean distance, and use the dish feature vectors whose minimum Euclidean distance is greater than or equal to the mean Euclidean distance as the initial clustering centers to obtain several initial clustering centers;
[0128] A second determination sub-module, which is used for:
[0129] Determine the clustering evaluation values of several initial clustering centers based on the first density value and the minimum Euclidean distance corresponding to the several initial clustering centers;
[0130] Compare the clustering evaluation values of the several initial clustering centers with a preset clustering evaluation threshold respectively, and use the initial clustering centers corresponding to the clustering evaluation values greater than or equal to the preset clustering evaluation threshold as the target clustering centers;
[0131] A clustering sub-module, which is used to cluster several dish feature vectors based on the target clustering centers to obtain several sets of dish feature vectors.
[0132] In this embodiment, determining the clustering evaluation values of several initial clustering centers based on the first density value and the minimum Euclidean distance corresponding to the several initial clustering centers includes:
[0133] Taking the product of the first density value corresponding to each initial clustering center and the minimum Euclidean distance as the clustering evaluation value of each initial clustering center;
[0134] Traverse all the initial clustering centers to obtain the clustering evaluation values of all the initial clustering centers.
[0135] The working principle of the above technical solution is as follows: The calculation sub-module measures the density of the vector distribution around each dish feature vector by calculating its density, and then obtains the density value of each vector. Based on the density value, the minimum Euclidean distance is determined. Then, the mean of all the minimum Euclidean distances is calculated, which is the mean Euclidean distance. The first determination sub-module compares the minimum Euclidean distance of each feature vector with the mean Euclidean distance, and takes those feature vectors with larger distances as the initial clustering centers, which represent the possible clustering cores. The second determination sub-module further evaluates these initial clustering centers, combines the first density value and the minimum Euclidean distance to determine the clustering evaluation value, and filters out more appropriate target clustering centers after comparing with the preset clustering evaluation threshold. Finally, the clustering sub-module performs clustering operations on all the dish feature vectors according to the determined target clustering centers, classifies similar vectors into one category, and forms several dish feature vector sets, thus realizing the classification and clustering of the dish feature vectors.
[0136] The beneficial effects of the above technical solution are as follows: By means of complex calculations and comparisons, the initial clustering centers and target clustering centers are determined, which can classify the dish feature vectors more accurately and improve the accuracy of clustering. Considering various factors such as density values and Euclidean distances, the solution can adapt to dish feature vector sets with different distributions and characteristics. Through a series of evaluation mechanisms, appropriate target clustering centers are screened out to ensure the quality and reliability of clustering. A reasonable clustering process helps to quickly classify relevant vectors into one category and improve the processing efficiency. This modular design is convenient for adjustment and expansion according to actual needs and has good flexibility.
[0137] Embodiment 4
[0138] The calculation sub-module includes:
[0139] The first calculation unit is used for:
[0140] Calculate the density of each of the several dish feature vectors respectively to obtain the density value corresponding to each dish feature vector;
[0141] Arbitrarily select a dish feature vector as the first target feature vector; take the density value corresponding to the first target feature vector as the target density value;
[0142] Compare the target density value with the density values corresponding to the other dish feature vectors except the first target feature vector;
[0143] When the density values corresponding to the other dish feature vectors except the first target feature vector are greater than or equal to the target density value, the corresponding dish feature vectors are used as the second target feature vectors, and a number of second target feature vectors are obtained;
[0144] A second calculation unit, configured to:
[0145] Calculate the Euclidean distances between the first target feature vector and a number of second target feature vectors respectively, and obtain a number of Euclidean distances;
[0146] Obtain the minimum value among the number of Euclidean distances as the minimum Euclidean distance corresponding to the first target feature vector;
[0147] Traverse all the dish feature vectors to obtain the minimum Euclidean distance corresponding to each dish feature vector;
[0148] A third calculation unit, configured to calculate the mean value of the minimum Euclidean distances corresponding to all the dish feature vectors to obtain the mean Euclidean distance.
[0149] The working principle of the above technical solution is as follows: The first calculation unit first calculates the density of all dish feature vectors to obtain their respective density values, then selects a vector as the first target feature vector, sets its density value as the target density value, and then compares it with the density values of other vectors to find those vectors whose density values are not less than the target density value as the second target feature vectors; The second calculation unit calculates the Euclidean distance between each first target feature vector and a number of second target feature vectors, and selects the minimum value as the minimum Euclidean distance of the first target feature vector. By traversing all dish feature vectors, the minimum Euclidean distance corresponding to each vector can be obtained; The third calculation unit aggregates the minimum Euclidean distances of all dish feature vectors and calculates their mean value to obtain the mean Euclidean distance. Generally speaking, through a series of calculations and comparisons of density and distance, the calculation sub-module provides basic data and basis for subsequent operations such as determining the clustering center.
[0150] The beneficial effects of the above technical solution are as follows: By calculating density and making detailed comparisons to determine relevant vectors, the basis for subsequent clustering analysis is made more reliable, thereby improving the accuracy of clustering results; It can adapt to the distribution of different dish feature vectors, find more representative target feature vectors and distance relationships, and enhance the adaptability of the solution to various data situations; Accurately calculate the Euclidean distance and obtain the minimum value and mean value, providing an accurate quantitative basis for the subsequent determination of the clustering center and clustering operations; It helps to more reasonably divide the dish feature vector set, making the clustering results more in line with the actual needs and feature distributions, and optimizing the overall clustering effect; It provides key basic data such as density values and distances for other sub-modules of the subsequent clustering module, ensuring the smooth progress of the entire clustering process.
[0151] Example 5
[0152] The first calculation unit is used to calculate the density of each of the several dish feature vectors respectively to obtain the density value corresponding to each dish feature vector, including:
[0153]
[0154] where ρ i represents the density value corresponding to the i-th dish feature vector; m represents the total number of dish feature vectors; d ij represents the Euclidean distance between the i-th dish feature vector and the j-th dish feature vector; d c represents a preset truncation distance.
[0155] The beneficial effects of the above technical solution are as follows: Through specific function calculations, it can more accurately reflect the density of each dish feature vector in the whole; combined with the Euclidean distance between dish feature vectors, the density calculation is made more reasonable and relevant; by presetting the truncation distance, the range and sensitivity of density calculation can be adjusted according to the actual situation to adapt to different dish feature distributions; the obtained density value provides an important basis for subsequent determination of target feature vectors and further analysis.
[0156] Example 6
[0157] It further includes:
[0158] A preprocessing module, which is used to preprocess the image to be recognized before extracting features from the image to be recognized.
[0159] Example 7
[0160] The preprocessing module includes:
[0161] An enhancer module, which is used to enhance the image to be recognized to obtain an enhanced image to be recognized;
[0162] A determination sub-module, which is used to use the enhanced image to be recognized as the preprocessed image to be recognized.
[0163] Example 8
[0164] The enhancer module includes:
[0165] A first acquisition unit, which is used for:
[0166] Arbitrarily select an image to be recognized as the first image;
[0167] Divide the first image evenly into several sub-images;
[0168] A third calculation unit for calculating a blurriness evaluation value corresponding to each sub-image;
[0169] A first determination unit for:
[0170] Comparing the blurriness evaluation value with a preset blurriness evaluation threshold, and taking the sub-image corresponding to the blurriness evaluation value being greater than or equal to the preset blurriness rating threshold as a sub-image to be enhanced, thereby obtaining a plurality of sub-images to be enhanced;
[0171] A second acquisition unit for:
[0172] Arbitrarily selecting one sub-image to be enhanced as the second image;
[0173] Arbitrarily selecting one pixel point in the second image as the target pixel point;
[0174] Taking the target pixel point as the center and the preset step length as the radius to determine a target area;
[0175] A fourth calculation unit for:
[0176] Calculating the gray mean value of all pixel points in the target area as the target gray mean value;
[0177] Calculating the absolute value of the difference between the gray value of the target pixel point and the target gray mean value to obtain the target absolute difference;
[0178] Comparing the target absolute difference with a preset absolute difference threshold;
[0179] A second determination unit for taking the target pixel point corresponding to the target absolute difference being greater than or equal to the preset absolute difference threshold as a pixel point to be enhanced;
[0180] An enhancement unit for:
[0181] Replacing the gray value of the pixel point to be enhanced based on the target gray mean value to obtain an enhanced pixel point;
[0182] Traversing all pixel points in the second image to obtain an enhanced second image;
[0183] Traversing all sub-images to be enhanced to obtain an enhanced image to be recognized.
[0184] The working principle of the above technical solution is as follows: The enhancer module first evenly divides the image to be recognized into several sub-images, then calculates the fuzzy evaluation value of each sub-image, and determines the sub-images to be enhanced by comparing with a preset fuzzy evaluation threshold. For each sub-image to be enhanced, taking one pixel point as the target pixel point, determining the target area and calculating the average gray value of the pixel points within the area, calculating the absolute value of the difference between the gray value of the target pixel point and the average value, comparing with the preset absolute difference threshold to determine the pixel points to be enhanced, and then replacing the gray value of the pixel points to be enhanced based on the target gray average value. After traversing all the pixel points in a sub-image, the enhanced sub-image is obtained. Furthermore, after processing all the sub-images to be enhanced, the enhanced image to be recognized is obtained. The determination module directly determines the enhanced image to be recognized as the preprocessed image to be recognized.
[0185] The beneficial effects of the above technical solution are as follows: Through the targeted enhancement of the image by the enhancer module, the clarity, contrast, etc. of the image can be improved, making the image features more obvious, which is beneficial to the accuracy of subsequent recognition and other operations; it can automatically determine the sub-images to be enhanced and the pixel points to be enhanced according to the characteristics and blur degree of different images, and has good adaptability; by dividing the image into sub-images and processing them one by one, and conducting detailed analysis and enhancement on the pixel points, more refined image processing is achieved; it helps to highlight the key features in the image to be recognized, reduce the influence of interference and blurred parts, and improve the recognition efficiency and reliability; the preprocessed image can better adapt to the processing requirements of subsequent modules, laying a foundation for the performance improvement of the entire image recognition or analysis system.
[0186] Example 9
[0187] The first calculation unit is used to calculate the fuzzy evaluation value corresponding to each sub-image, including:
[0188] Arbitrarily select a sub-image and perform grayscale processing to obtain a grayscale sub-image;
[0189] Obtain the gray values of each pixel point in the grayscale sub-image;
[0190] Calculate the gradient magnitude of each pixel point in the grayscale sub-image;
[0191] Determine the average gradient magnitude corresponding to the grayscale sub-image based on the gradient magnitude of each pixel point;
[0192] Compare the gradient magnitude of each pixel point with the average gradient magnitude, and take the pixel points with gradient magnitude greater than or equal to the average gradient magnitude as the first pixel points to obtain a set of first pixel points;
[0193] Calculate the average gray value of all pixel points in the grayscale sub-image;
[0194] Compare the gray value of each pixel in the grayscale sub-image with the average gray value, and take the pixels with gray values greater than or equal to the average gray value as the second pixels to obtain a set of second pixels;
[0195] Determine the fuzzy evaluation value of the grayscale sub-image based on the first set of pixels and the second set of pixels;
[0196] Traverse each sub-image to obtain the fuzzy evaluation value corresponding to each sub-image.
[0197] In this embodiment, calculate the gradient magnitude of each pixel in the grayscale sub-image. The specific implementation method: for each pixel, calculate its gradient components in the horizontal and vertical directions respectively. Use the convolution kernels corresponding to the Sobel operator to perform convolution operations with the sub-image to obtain the horizontal and vertical gradient values. Calculate the gradient magnitude of each pixel. The gradient magnitude of the pixel can be obtained by adding the square of the horizontal gradient component and the square of the vertical gradient component and then taking the square root. Repeat the above steps for all pixels in the entire grayscale sub-image to obtain the gradient magnitude of each pixel.
[0198] In this embodiment, determining the fuzzy evaluation value of the grayscale sub-image based on the first set of pixels and the second set of pixels includes:
[0199] Obtain the intersection of the first set of pixels and the second set of pixels;
[0200] Obtain the union of the first set of pixels and the second set of pixels;
[0201] Take the ratio of the intersection to the union as the fuzzy evaluation value of the grayscale sub-image.
[0202] The working principle of the above technical solution is as follows: First, perform grayscale processing on any selected sub-image to obtain grayscale information. Calculate the gradient magnitude of each pixel to reflect the severity of the gray change in the image, and then determine the average gradient magnitude. Compare the gradient magnitude of the pixel with the average value to obtain the first set of pixels, which helps to screen out the parts with large gradient changes. At the same time, calculate the average gray value of all pixels, and obtain the second set of pixels according to the comparison with the average value, which can reflect the parts with relatively high gray values. Combine the first set of pixels and the second set of pixels to comprehensively evaluate the characteristics of the grayscale sub-image in terms of gradient change and gray distribution, so as to determine a fuzzy evaluation value that can reflect the blur degree of the sub-image. Repeat this process by traversing all sub-images to obtain the fuzzy evaluation value corresponding to each sub-image, so as to judge whether enhancement is needed and how to enhance the sub-image according to these values in the future.
[0203] The beneficial effects of the above technical solution are as follows: By comprehensively considering the gradient magnitude and gray value distribution of pixel points, the blur degree of the sub-image can be measured more comprehensively and accurately, providing a reliable basis for subsequent processing; After determining the blur evaluation value of each sub-image, more appropriate processing strategies can be adopted for sub-images with different blur degrees, improving the efficiency and effect of image processing; It helps to analyze image features more carefully, so as to make more accurate adjustments and optimizations during the image processing process, improving the image quality; This method can adapt to different types and characteristics of images and has good applicability for various image analysis and processing tasks.
[0204] To achieve the above object, as Figure 3 shown, an embodiment of the present invention provides a method for feature extraction and recognition of dish images in a smart restaurant, including steps S1 - S4:
[0205] S1: Obtain a dish image of a smart restaurant as an image to be recognized;
[0206] S2: Extract features from the image to be recognized, construct a dish feature vector, and obtain a number of dish feature vectors;
[0207] S3: Cluster the number of dish feature vectors to obtain a number of dish feature vector sets;
[0208] S4: Input the number of dish feature vector sets into a pre-trained dish recognition model for recognition respectively to obtain the recognition result of the image to be recognized.
[0209] The beneficial effects of the above technical solution are as follows: The acquisition module quickly acquires the image, the feature extraction module efficiently constructs the feature vector by extracting and analyzing features such as the color, texture, and shape of the dish image, and the clustering module performs classification processing. The overall process helps to improve the processing efficiency; The clustering module can group vectors with similar features into one category, providing more targeted processing for subsequent recognition and improving the recognition accuracy; Using a pre-trained dish recognition model for recognition can adapt to the recognition needs of various dishes, has good versatility and scalability, and helps the smart restaurant better manage and analyze dishes, improving the service quality and operation efficiency.
[0210] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A feature extraction and recognition system for food images in a smart restaurant, characterized in that: include: An acquisition module is used to acquire images of dishes in a smart restaurant as images to be recognized; A feature extraction module is used to extract features from the image to be identified, construct a dish feature vector, and obtain a plurality of dish feature vectors; A clustering module, used for clustering the plurality of dish feature vectors to obtain a plurality of dish feature vector sets; The recognition module is used to input several dish feature vector sets into the pre-trained dish recognition model for recognition, and obtain the recognition result of the image to be recognized.
2. The feature extraction and recognition system for smart restaurant dish images according to claim 1, characterized in that: Feature extraction module, including: Texture feature extraction submodule, used for: Performing grayscale processing on the image to be identified to obtain a grayscale image; Get the gray level co-occurrence matrix corresponding to the gray level image; Determine the texture features corresponding to the image to be identified based on the gray level co-occurrence matrix; The color feature extraction submodule is used to: Convert the image to be recognized into HSV color space; Divide the HSV color space into a number of intervals, traverse all pixels in the image to be identified, and determine a color histogram corresponding to the image to be identified; Determining a color feature corresponding to the image to be identified based on the color histogram; The shape feature extraction submodule is used to perform edge contour detection on the image to be identified to obtain a plurality of contours corresponding to the image to be identified; and determine the shape features of the image to be identified based on the plurality of contours; The construction submodule is used to construct a dish feature vector based on the texture features, color features and shape features corresponding to the image to be identified.
3. The feature extraction and recognition system for food images in a smart restaurant as claimed in claim 1, characterized in that: Clustering modules, including: A calculation submodule is used to perform density calculation on the several dish feature vectors respectively to obtain the density value corresponding to each dish feature vector; determine the minimum Euclidean distance corresponding to each dish feature vector based on the density value corresponding to each dish feature vector; calculate the mean of the minimum Euclidean distances corresponding to all dish feature vectors to obtain the mean Euclidean distance; a first determination submodule is used to compare the minimum Euclidean distance corresponding to each feature vector with the mean Euclidean distance, and use the dish feature vector whose minimum Euclidean distance is greater than or equal to the mean Euclidean distance as the initial cluster center to obtain several initial cluster centers; The second determination submodule is used to: Determining clustering evaluation values of the plurality of initial clustering centers based on the first density values and the minimum Euclidean distances corresponding to the plurality of initial clustering centers; The cluster evaluation values of the plurality of initial cluster centers are compared with the preset cluster evaluation thresholds respectively, and the initial cluster center corresponding to the cluster evaluation value being greater than or equal to the preset cluster evaluation threshold is taken as the target cluster center; The clustering submodule is used to cluster a plurality of dish feature vectors based on the target cluster center to obtain a plurality of dish feature vector sets.
4. The feature extraction and recognition system for smart restaurant dish images as claimed in claim 3, characterized in that: Computing submodule, including: The first computing unit is configured to: Performing density calculation on the plurality of dish feature vectors respectively to obtain a density value corresponding to each dish feature vector; Randomly select a dish feature vector as the first target feature vector; the density value corresponding to the first target feature vector is used as the target density value; Comparing the target density value with density values corresponding to other dish feature vectors except the first target feature vector; The dish feature vectors corresponding to the density values corresponding to the other dish feature vectors except the first target feature vector are greater than or equal to the target density value as the second target feature vectors, to obtain a plurality of second target feature vectors; The second computing unit is configured to: Calculating the Euclidean distances between the first target feature vector and a plurality of second target feature vectors respectively to obtain a plurality of Euclidean distances; Obtaining a minimum value among the plurality of Euclidean distances as the minimum Euclidean distance corresponding to the first target feature vector; Traverse all dish feature vectors and obtain the minimum Euclidean distance corresponding to each dish feature vector; The third calculation unit is used to calculate the mean of the minimum Euclidean distances corresponding to all dish feature vectors to obtain the mean Euclidean distance.
5. The feature extraction and recognition system for food images in a smart restaurant as claimed in claim 4, characterized in that: The first calculation unit is used to perform density calculation on the plurality of dish feature vectors respectively to obtain a density value corresponding to each dish feature vector, including: Among them, ρ i represents the density value corresponding to the i-th dish feature vector; m represents the total number of dish feature vectors; d ij represents the Euclidean distance between the i-th dish feature vector and the j-th dish feature vector; d c Indicates the preset cutoff distance.
6. The feature extraction and recognition system for food images in a smart restaurant as claimed in claim 1, characterized in that: Also includes: The preprocessing module is used to preprocess the image to be identified before extracting features from the image to be identified.
7. The feature extraction and recognition system for food images in a smart restaurant as claimed in claim 6, characterized in that: The preprocessing module comprises: An enhancement submodule, used for performing image enhancement on the image to be identified to obtain an enhanced image to be identified; A submodule is determined, which is used to use the enhanced image to be recognized as the preprocessed image to be recognized.
8. The feature extraction and recognition system for food images in a smart restaurant as claimed in claim 7, characterized in that: Enhancer modules, including: A first acquisition unit is used for: Take any image to be identified as the first image; Evenly divide the first image into a plurality of sub-images; A third calculation unit, used to calculate a fuzzy evaluation value corresponding to each sub-image; A first determining unit is configured to: Comparing the fuzzy evaluation value with a preset fuzzy evaluation threshold, taking the sub-image corresponding to the fuzzy evaluation value being greater than or equal to the preset fuzzy evaluation threshold as the sub-image to be enhanced, and obtaining a plurality of sub-images to be enhanced; The second acquisition unit is used for: Randomly select a sub-image to be enhanced as the second image; Randomly select a pixel point in the second image as the target pixel point; Determine the target area with the target pixel as the center and the preset step size as the radius; The fourth computing unit is configured to: Calculate the grayscale mean of all pixels in the target area as the target grayscale mean; Calculate the absolute value of the difference between the grayscale value of the target pixel and the target grayscale mean to obtain the target absolute difference; Comparing the target absolute difference with a preset absolute difference threshold; A second determining unit is used to take the target pixel point corresponding to the target absolute difference being greater than or equal to a preset absolute difference threshold as the pixel point to be enhanced; Enhancement unit for: Replacing the grayscale value of the pixel to be enhanced based on the target grayscale mean value to obtain an enhanced pixel; Traversing all pixel points in the second image to obtain an enhanced second image; Traverse all sub-images to be enhanced to obtain the enhanced image to be recognized.
9. The feature extraction and recognition system for food images in a smart restaurant as claimed in claim 8, characterized in that: The first calculation unit is used to calculate the fuzzy evaluation value corresponding to each sub-image, including: Take any sub-image and grayscale it to get a grayscale sub-image; Get the gray value of each pixel in the gray sub-image; Calculate the gradient amplitude of each pixel in the grayscale sub-image; Determine the gradient amplitude mean corresponding to the grayscale sub-image based on the gradient amplitude of each pixel point; Compare the gradient amplitude of each pixel point with the gradient amplitude mean, and take the pixel point whose gradient amplitude is greater than or equal to the gradient amplitude mean as the first pixel point, to obtain a first pixel point set; Calculate the grayscale mean of all pixels in the grayscale sub-image; Compare the grayscale value of each pixel in the grayscale sub-image with the grayscale mean, and take the pixel whose grayscale value is greater than or equal to the grayscale mean as the second pixel to obtain a second pixel set; Determine a fuzzy evaluation value of the grayscale sub-image based on the first pixel point set and the second pixel point set; Traverse each sub-image and obtain the fuzzy evaluation value corresponding to each sub-image.
10. A method for extracting and identifying features of dishes in a smart restaurant, characterized in that: include: Obtaining images of dishes in a smart restaurant as images to be recognized; Extracting features of the image to be identified, constructing a dish feature vector, and obtaining a plurality of dish feature vectors; Clustering the plurality of dish feature vectors to obtain a plurality of dish feature vector sets; Several dish feature vector sets are respectively input into the pre-trained dish recognition model for recognition to obtain the recognition results of the image to be recognized.
Citation Information
Patent Citations
Fast food positioning method
CN114627279A
Vegetable checking method and system based on image recognition
CN117475240A
Part surface quality detection method and system based on machine vision
CN118608504A