A dish cooking maturity grading method based on unsupervised learning

By constructing a four-order progressive clustering framework and combining physicochemical property indicators with clustering algorithms, the problems of subjectivity and accuracy in the classification of the cooking maturity level of dishes are solved, realizing the objective and accurate classification of unsupervised dish maturity levels and providing reliable data support and theoretical basis.

CN121456698BActive Publication Date: 2026-07-21SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2025-09-25
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies lack a universal and objective method for classifying the cooking maturity of dishes. Sensory evaluation is subjective, the database construction process is vague and lacks authoritative standards, and existing clustering methods are susceptible to outlier interference, resulting in insufficient accuracy.

Method used

A four-order progressive clustering framework consisting of principal component analysis, k-means clustering, DBSCAN clustering, and hierarchical clustering was adopted. Combined with physicochemical property indicators, PCA and LDA were used to classify the cooking maturity level of dishes in an unsupervised manner, and the results were visualized for verification and optimization.

Benefits of technology

This method achieves an objective and accurate classification of the maturity levels of dishes, improves the acceptance and accuracy of unsupervised maturity levels, solves the ambiguity and subjectivity problems in database construction, and enhances the anti-interference ability and result stability of the clustering method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456698B_ABST
    Figure CN121456698B_ABST
Patent Text Reader

Abstract

The application discloses a dish cooking maturity grading method based on unsupervised learning, and the method comprises the following steps: S1, acquiring recipe data of a target dish to determine a cooking time interval, reproducing a cooking process and controlling variables, and setting an experimental group according to an expected maturity classification number; S2, selecting physicochemical property indexes according to the type of the dish, detecting physicochemical data under different cooking times, and constructing a database; S3, using a four-order progressive clustering framework composed of principal component analysis, k-means clustering, DBSCAN clustering and system clustering methods to realize unsupervised dish cooking maturity grade division based on physicochemical properties; and S4, realizing visualization of the unsupervised dish cooking maturity grade division result based on PCA and LDA, and verifying and optimizing the result. The application constructs a four-order progressive collaborative framework according to the maturity grade division characteristics in the cooking process from the multiple physicochemical properties, and realizes the unsupervised dish maturity grade division.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of food testing technology, specifically relating to a method for grading the cooking maturity of dishes based on unsupervised learning. Background Technology

[0002] Currently, apart from steak, which has a clearly defined unsupervised standard for classifying doneness based on internal temperature, the doneness classification of most dishes lacks a unified unsupervised standard. Most chefs and researchers still rely on sensory evaluation methods to classify doneness. However, sensory evaluation of dishes is considered subjective and experiential. With the development of artificial intelligence, intelligent cooking has raised the need for intelligent perception of dish doneness. However, related research has focused on rapid, non-destructive determination of doneness and algorithmic innovation, while using relatively vague definitions in the initial sample standards and database construction.

[0003] Currently, related research focuses on the development of backend models and non-destructive judgment methods, while largely neglecting the accuracy and correctness of the basic database itself. Patent CN202510223925.0 discloses a "method, computer equipment, and system for detecting the maturity of steamed eggs," which uses only internal temperature as an unsupervised reference and lacks authoritative standard support when constructing its database of steamed egg dishes. Patent CN201910918784.9 discloses a "method, device, kitchen appliance, and server for determining the maturity of ingredients," which constructs a maturity judgment system applicable to various cooking appliances. However, in its database construction process, the unsupervised determination of ingredient maturity reaching a set maturity level is achieved by the chef observing the ingredients, a process highly susceptible to subjective influence and exhibiting high instability. Patent CN202311357797.6 discloses a "method for evaluating the cooking effect of food," which quickly judges maturity based on image information acquired during the cooking process and provides a judgment reference table; however, the source and correctness of this table have not been rigorously proven.

[0004] Therefore, there is an urgent need to design a universal, objective, and easily implementable unsupervised method for classifying the maturity levels of dishes, so as to provide a rigorous theoretical basis and methodological reference for the construction of various maturity databases in the field of intelligent cooking. Summary of the Invention

[0005] The main objective of this invention is to overcome the shortcomings and deficiencies of the prior art and propose a method for classifying the maturity of dishes based on unsupervised learning. Starting from multidimensional physicochemical properties, a four-order progressive collaborative framework is constructed based on the characteristics of maturity level classification during the cooking process to realize the classification of dish maturity levels without supervision.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A method for grading the cooking maturity of dishes based on unsupervised learning includes the following steps:

[0008] S1. Obtain the recipe data of the target dish to determine the cooking time interval, reproduce the cooking process and control variables, and set up experimental groups according to the expected maturity classification number;

[0009] S2. Select physicochemical property indicators according to the type of dish, detect physicochemical data under different cooking times, and construct a database;

[0010] S3. A four-order progressive clustering framework consisting of principal component analysis, k-means clustering, DBSCAN clustering, and hierarchical clustering methods is used to realize the unsupervised classification of the cooking maturity level of dishes based on physicochemical properties.

[0011] S4. Based on PCA and LDA, visualize the results of unsupervised cooking maturity level classification of dishes, and verify and optimize the results.

[0012] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0013] 1. Existing unsupervised methods for classifying the maturity level of dishes have serious shortcomings. Existing unsupervised maturity level classification methods based on sensory evaluation are highly subjective and not recognized by the academic community. To address these issues, the present invention uses objective physicochemical properties as a benchmark and constructs a four-order progressive clustering framework with the help of universally recognized conventional clustering methods, thereby improving the objectivity and acceptance of unsupervised maturity level classification.

[0014] 2. The method of this invention effectively solves the problems of unclear original sample standards, lack of literature support, and strong subjectivity in the process of constructing databases for various artificial intelligence-based rapid and non-destructive identification methods, thus providing reliable data support for the development of the field of intelligent cooking.

[0015] 3. Existing technologies use dimensionality reduction and clustering methods in a fragmented manner, failing to fully realize their potential. Furthermore, single clustering algorithms are susceptible to interference from outliers and isolated experimental groups, making it impossible to complete unsupervised maturity level classification. If too few experimental groups are set, the resulting maturity levels lack persuasiveness and accuracy. To address these issues, this invention integrates commonly used clustering techniques with the characteristics of the unsupervised food cooking maturity level classification process. Utilizing the information required for maturity level classification, different dimensionality reduction and clustering methods are employed to construct a four-order progressive clustering framework suitable for maturity level classification, resulting in accurate and persuasive maturity level classification results. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method of the present invention;

[0017] Figure 2 This is a schematic diagram illustrating the determination of the physicochemical properties of stir-fried pork with chili peppers in the examples;

[0018] Figure 3 This is a schematic diagram of the original data after PCA dimensionality reduction and clustering in the embodiment;

[0019] Figure 4 This is a schematic diagram of the k-means elbow analysis rule for the original data;

[0020] Figure 5 This is a schematic diagram of the k-means elbow analysis rule for preliminary data purification;

[0021] Figure 6 This is a schematic diagram of the original data system cluster analysis;

[0022] Figure 7 This is a schematic diagram of DBSCAN cluster analysis for clean data;

[0023] Figure 8 This is a schematic diagram of the final LDA clustering result of the purified data. Detailed Implementation

[0024] Explanation of relevant terms:

[0025] PCA: Principal Component Analysis;

[0026] LDA: Linear Discriminant Analysis.

[0027] like Figure 1 As shown, this invention provides a method for grading the cooking maturity of dishes based on unsupervised learning, comprising the following steps:

[0028] S1. Obtain the recipe data of the target dish to determine the cooking time interval, reproduce the cooking process and control variables, and set up experimental groups according to the expected maturity classification number;

[0029] The recipe data for the target dish is obtained using Python web scraping. Specifically, it involves sending network requests and parsing web pages to scrape recipe data from static web pages. The implementation approach is as follows:

[0030] a. Identify the target website: Select a food-related website (such as Xiachufang, Douguo Food, etc.) where the page contains the recipe for the identified dish.

[0031] b. Sending a request: Simulate a browser sending a request to the same Uniform Resource Locator (URL) to obtain the Hypertext Markup Language (HTML) content of the webpage.

[0032] c. Parse the webpage: Locate the key information of the recipe (dish name, ingredients, steps, etc.) through webpage tags (such as div, ul, li, etc.).

[0033] d. Extract data: Extract specific content from the parsed HTML, such as ingredient list, cooking steps, tips, etc.

[0034] e. Save data: Save the extracted information as a local file (such as CSV or JSON) for later use.

[0035] After retrieving the recipes for the corresponding dishes, the cooking process was reproduced. The cooking time was determined after strictly controlling for potential cooking variables, including ingredient quantities, heat, and utensils. Physicochemical analysis experimental groups were set up based on the time and the approximate number of maturity levels, with the number of experimental groups being at least twice the number of maturity levels.

[0036] S2. Select physicochemical property indicators based on the type of dish, detect physicochemical data at different cooking times, and construct a database; specifically:

[0037] Physicochemical properties were selected based on the type of dish. For example, for meat, indicators could include cooking loss, protein content, textural parameters, color difference, and moisture phase distribution. For vegetables, indicators could include moisture content, vitamin C retention rate, chlorophyll content, firmness (textural characteristics), total phenolic content, and pH value. Subsequently, these physicochemical indicators were tested using conventional techniques in the field. To ensure the effectiveness and reliability of the constructed unsupervised method for classifying the cooking maturity of dishes, the sample size for each indicator in each experimental group in the constructed database should not be less than 15.

[0038] S3. A four-order progressive clustering framework, consisting of principal component analysis, k-means clustering, DBSCAN clustering, and hierarchical clustering methods, is used to implement unsupervised classification of the cooking maturity levels of dishes based on physicochemical properties; specifically including:

[0039] S31. Principal component analysis (PCA) was used to reduce the dimensionality of the high-dimensional original data and analyze the number of experimental groups corresponding to the experimental group samples contained in the confidence ellipse of each experimental group. The mean μ, standard deviation σ, and threshold of the number of experimental groups corresponding to the experimental group samples contained in the confidence ellipse of all experimental groups were calculated. By comparing the number of experimental groups corresponding to the experimental group samples contained in the confidence ellipse of each experimental group with the calculated threshold, the experimental groups with the fastest changes in maturity level were identified and eliminated, forming a preliminary clean database. The confidence ellipse was a 95% confidence interval constructed based on the first two principal components of PCA. The threshold was set as the mean of the number of experimental groups contained in the confidence ellipse of all experimental groups plus 1.5 times the standard deviation, i.e., μ + 1.5σ. The threshold was dynamically adjusted for different types of dishes: for meat dishes, the threshold was reduced by about 10% because the maturity transition was more significant; for vegetables, the threshold was increased by about 10% because the moisture loss was faster and the number of maturity levels was smaller. The specific adjustment coefficients were obtained by fitting the pre-experimental data.

[0040] The criterion for determining the "riding the wall" experimental group is that the number of samples from other experimental groups included in the confidence ellipse is greater than or equal to a threshold.

[0041] S32. To address the subjectivity issue in determining the total number of categories, the k-means elbow analysis method is innovatively applied simultaneously to both the original and pre-cleaned databases. The optimal total number of maturity categories is determined by comparing the inflection points of the clustering errors of the two databases. The criteria for determining the optimal total number of maturity categories are as follows:

[0042] When the difference between the k-value corresponding to the inflection point of the elbow analysis clustering error curve of the original data and the preliminary clean database is ≤1, the k-value corresponding to the inflection point of the elbow analysis clustering error curve of the preliminary clean database is selected as the total number of optimal maturity categories; when the difference is >1, different preliminary clean databases that meet the standard are constructed and the above steps are repeated. In addition, the total number of optimal maturity categories must be less than half of the number of experimental groups set in step S1; otherwise, the number of experimental groups in step S1 needs to be increased and step S2 needs to be executed again.

[0043] S33. Guided by the total number of optimal maturity categories, perform systematic clustering on the original data. Utilize the tree-like structure of systematic clustering to identify experimental groups whose samples span two maturity categories and further eliminate them, forming a purified final database. This achieves precise correction of category boundaries to maximize the effectiveness of the final maturity level classification. Systematic clustering uses Euclidean distance, and each physicochemical index participates in the calculation with equal weight after being normalized by the maximum and minimum values.

[0044] S34. Based on the final cleaned database, the DBSCAN clustering algorithm is used to complete the final division of maturity levels. The parameters of the DBSCAN algorithm are obtained by traversing the total number of optimal maturity categories determined in step S32. During the traversal, parameter combinations that make the clustering results consistent with the total number of optimal maturity categories determined in step S32 are selected first. During the traversal, the neighborhood radius is set to 1-5 with a step size of 0.1, the minimum number of samples is set to 1-20 with a step size of 1, and the parameter combination with the smallest proportion of outliers is selected as the optimal parameters.

[0045] This invention effectively avoids the category shift problem that easily occurs when single clustering algorithms process nonlinear and noisy data such as dish maturity levels. The process is not a simple superposition of existing algorithms, but rather a progressive collaborative design involving "quantitative identification of 'riding the wall' points, objective calibration of the number of categories, data purification, and optimization of clustering parameters." This method focuses on key nodes in the unsupervised classification of dish cooking maturity levels, specifically addressing the inherent shortcomings of traditional methods such as low clustering accuracy, weak anti-interference ability, and poor result stability through logical collaboration and parameter passing. It significantly improves the accuracy and reliability of dish maturity level judgment. The logical connections between algorithms and the scenario-based customized processing fully demonstrate a profound understanding and innovative application of dish maturity data characteristics. This method effectively fills the gap in unsupervised dish cooking maturity level classification methods needed for database construction in the context of the current booming development of deep learning, improving the accuracy and reliability of related research.

[0046] S4. Visualize the results of unsupervised cooking maturity level classification of dishes based on PCA and LDA, and verify and optimize the results; specifically including:

[0047] S41. Obtain the results of the dish maturity level classification and the corresponding physicochemical data;

[0048] S42. PCA is used to reduce the dimensionality of the physicochemical data and generate a two-dimensional scatter plot that shows the global cluster distribution of the samples, intuitively presenting the spatial clustering pattern of each maturity level.

[0049] S43. Introduce category information from the segmentation results to guide linear discriminant analysis to focus on maximizing the discriminant boundary between maturity levels and generate a visualization result with clear category separation characteristics; the category information must satisfy the outlier ratio <0.05; if it does not meet the standard, backtrack to optimize the secondary straddle point removal parameters in step S33.

[0050] S44. Calculate the inter-class separation distance and intra-class clustering degree of each maturity level in the LDA space to quantitatively evaluate the reliability of the division results; the preset threshold for inter-class separation distance is twice the average intra-class clustering degree of each level.

[0051] S45. Determine whether the deviation between the visualization results of PCA and LDA exceeds the preset range, or whether the inter-class separation distance reaches the threshold. If not, backtrack and optimize until the result meets the standard. For the two visualization schemes, if the sample overlap rate is >30%, re-execute S31 and S33, and adjust the judgment threshold to increase the strictness of the exclusion of the straddle point and the experimental group, such as reducing the judgment threshold that spans two classes from 3 samples to 2 samples;

[0052] If the sample overlap rate is >50%, then S32 (class number calibration) is re-executed, and the optimal total number of classes is ±1 before re-clustering.

[0053] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0054] Example

[0055] This embodiment takes the classification of the doneness grades of pork slices in stir-fried pork with chili peppers as an example, and includes the following steps:

[0056] S1. Crawl recipe data for the target dish to determine the cooking time interval, reproduce the cooking process and control variables, and set up experimental groups according to the expected maturity classification number. The number of experimental groups should be at least twice the number of classifications.

[0057] We selected Xiachufang and Douguo Food as the core data crawling platforms, filtering recipes with the keyword "stir-fried pork with chili peppers," prioritizing recipes with ≥500 favorites and ≥100 comments for high credibility. The plan was to crawl 30-50 valid data entries from each platform. The crawling process involved first using Python's requests library, setting a specific User-Agent to simulate a browser, and then crawling recipe URLs paginated from the search results pages for "stir-fried pork with chili peppers" on platforms like Xiachufang to circumvent anti-crawling measures. Next, we used the BeautifulSoup4 library to parse the HTML, extracting stir-frying time, the amount of core ingredients, and key step descriptions from specific tags (such as "cooking time, ingredients, steps"). Then, invalid data with unclear times or ambiguous steps was removed, and the valid data was formatted into a standardized format. Finally, we used the pandas library to store the cleaned data in a local folder for subsequent analysis.

[0058] Based on the results obtained from web crawling, relevant cooking operation specifications were specified. On the day of slaughter, pork hind leg meat, after removing connective tissue and fat, was sliced ​​into thin slices of 3cm × 4cm × 2mm. 100g of sliced ​​pork hind leg meat, 0.5g of salt, 0.5g of MSG, 4g of soy sauce, and 20g of water were placed in a bowl and stirred until the seasonings were completely absorbed by the meat. Then, the induction cooker was set to 600℃, and an infrared thermometer was placed above it. When the surface temperature of the induction cooker reached 200±5℃, a wok containing 40ml of cooking oil was placed on top. When the bottom temperature of the wok reached 80±2℃, the marinated meat slices were added and stir-fried. According to the recipe information, the time for the meat slices to go from raw to overcooked under these operating conditions is approximately 2-3 minutes. The above standardized cooking steps were repeated 3 times, and the actual time for the meat slices to go from raw to overcooked (burnt) was recorded each time. The average cooking time was calculated to be 2.5±0.3 minutes. Therefore, 150 seconds was defined as the cooking time required for the meat slices to go from raw to overcooked under the conditions of this embodiment. Based on this, this embodiment aims to classify the doneness of the meat slices into five categories, thus constructing 10 experimental groups, and cooking was performed for 15 seconds, 30 seconds, 45 seconds, 60 seconds, 75 seconds, 90 seconds, 105 seconds, 120 seconds, 135 seconds, and 150 seconds respectively. Preliminary experiments verified the feasibility of this time span, confirming that it fully covers the entire cooking process from raw meat to overcooked meat slices.

[0059] S2. Select physicochemical property indicators according to the type of dish, detect physicochemical data under different cooking times and construct a database. The sample size of each experimental group for each indicator in the database shall not be less than 15.

[0060] Meat slices cooked for different times were obtained based on the cooking conditions and times set in S1, and their physicochemical properties were tested. For meat, this embodiment selects representative physicochemical properties according to its characteristics. Cooking loss, texture, moisture, myofibrillar protein, sarcoplasmic protein, color, malondialdehyde, and thiol groups were selected as the physicochemical properties to be tested. Figure 2 The diagram shown is a schematic diagram for determining the physicochemical properties of stir-fried pork with chili peppers.

[0061] Cooking loss was determined as follows: The total weight of the pork slices mixed with various seasonings was weighed and recorded as M1. After cooking for the appropriate time, the cooking oil on the pork slices was drained using a sieve, and then weighed and recorded as M2. Cooking loss was calculated using the following formula, and the average value was taken from three parallel trials under each experimental condition:

[0062]

[0063] M1 and M2 represent the weight of the pork before and after cooking.

[0064] Texture determination was performed using a texture analyzer: samples were cut perpendicularly along the muscle fiber direction, and texture data were acquired using an HDP / BSK probe. The prediction speed was set to 2 mm / s, the testing speed was also set to 2 mm / s, the trigger force was set to 5 g, and the testing distance was 40 mm. Three parallel tests were performed under each experimental condition, and the average value was taken as the final result.

[0065] The color of the meat slices was measured using a colorimeter. The colorimeter was used to measure the red-green hue, yellow-blue hue, and brightness of the original meat slices and the meat slices from each experimental group at five locations: top left, top right, bottom left, bottom right, and center. The color difference between the experimental group meat slices and the original meat slices was calculated and denoted as Δa*, Δb*, and ΔL*, respectively. For each sample at each maturity level, three arbitrary cross-sections were measured, and the average of the three measurements was taken.

[0066] The sarcoplasmic protein content was determined as follows: Pork slices at different degrees of doneness were minced using a meat grinder. 2 grams of pork and 25 ml of a prepared 0.01 mol / L sodium phosphate buffer solution (K₂HPO₄:KH₂PO₄ = 1:1 molar ratio) were added to a 50 ml centrifuge tube. The sample was then homogenized at 7160 × g for 1 min using a homogenizer, and subsequently stored at 4°C for 1 h. Next, the sample was centrifuged at 4°C and 6800 × g for 15 min in a refrigerated centrifuge, and the supernatant was diluted at a 2:3 ratio. Finally, the protein content was determined using the Coomassie Brilliant Blue method and recorded as M3.

[0067] Myofibrillar protein content was determined as follows:

[0068] 0.75 g of minced meat was added to a centrifuge tube along with 25 mL of a prepared 0.1 mol / L phosphate buffer (K₂HPO₄:KH₂PO₄ = 1:1 molar ratio) (containing 1.1 mol / L potassium iodide). The mixture was then homogenized at 7160 × g for 2 min and stored overnight. The next day, the sample was centrifuged at 4 °C and 6800 × g for 15 min, and the supernatant was diluted 2:3. The total soluble protein content was determined using the Coomassie Brilliant Blue method and recorded as the M value. Myofibrillar protein content was calculated using the following formula:

[0069] Myofibrillar protein content (mg / kg) = M - M3

[0070] M3 and M represent the sarcoplasmic protein content and the total soluble protein content, respectively.

[0071] This embodiment designed a series of physicochemical property and micro-index experiments from both macroscopic and microscopic perspectives. The methods used are all common in the industry, and therefore will not be elaborated upon further. Specific measurement data are shown in Table 1 below. The table shows that the physicochemical properties of pork slices changed rapidly in the early stages, gradually leveling off in the middle stages. Some physicochemical properties, represented by cooking loss and maximum shear force, exhibited abrupt changes in the later stages. No single physicochemical property can be used to classify the maturity level of the dish. This illustrates the complexity and rapid nature of changes during pork cooking. More comprehensive tools and data analysis methods are needed to classify maturity. The measurement process of the physicochemical properties in Table 1 was repeated for pork slice samples at different cooking times. The obtained data was stored in an Excel spreadsheet labeled with cooking time for subsequent establishment of a maturity level classification method. In this embodiment, each physicochemical property was repeatedly tested 21 times at each cooking time, therefore the constructed database contains a total of 210 sample points (10 experimental groups × 21 repeated tests).

[0072]

[0073] Table 1

[0074] S3. A four-order progressive clustering framework, which integrates principal component analysis, k-means clustering, DBSCAN clustering, and hierarchical clustering methods, is used to realize the unsupervised classification of the cooking maturity level of dishes based on physicochemical properties.

[0075] To achieve initial purification of the physicochemical property database, improve the accuracy of the machine learning-based ripeness classification method, and determine the appropriate number of ripeness levels for the samples in the current database, Principal Component Analysis (PCA) was used to reduce the 11 dimensions of the physicochemical properties of pork slices based on cooking time. The dimensionality reduction results were visualized using OriginPro 2022. The visualization results of the initial data are shown below. Figure 3 As shown. Analysis of the initial PCA clustering results revealed that the number of species included in the confidence ellipse corresponding to each cooking time was 1, 1, 2, 3, 4, 4, 4, 3, 2, 1, respectively. The mean and standard deviation of this group were then calculated to be 2.5 and 1.2, respectively, resulting in (μ + 1.5σ) × 0.9 = 3.87. Therefore, the experimental group with 4 species included in the confidence ellipse was deleted. It is hypothesized that the meat slice properties undergo abrupt changes at cooking times of 90s and 120s, and attempts were made to remove these two groups of data to obtain a preliminary cleaned database.

[0076] Elbow analysis is a data mining technique based on k-means clustering, used to determine the optimal number of clusters for unknown samples. When the slope of the within-cluster error square curve becomes relatively flat, it corresponds to the optimal number of clusters for the dataset under study. This embodiment uses the scikit-learn library to perform elbow analysis to process the initial and pre-cleaned data. The results of k-means elbow analysis on the original database and the pre-cleaned database are as follows: Figure 4 , Figure 5 As shown, the intra-cluster error squared curves of both databases tend to flatten out when k=5, and the number of maturity levels reflected by the initial database and the preliminary cleaned database is consistent. Therefore, it can be considered that the maturity of the dishes should be divided into 5 levels in this embodiment.

[0077] Hierarchical clustering is based on the progressive merging or splitting of clusters according to the similarity of a series of sample points until the final number of clusters matches the input number. This embodiment uses the number of maturity levels obtained from k-means elbow analysis to perform hierarchical clustering on the initial data. By analyzing the hierarchical classification results, experimental groups where samples are divided into two maturity levels are selected to examine the similarity and abrupt changes between groups in the initial data. In this embodiment, the number of hierarchical clustering partitions is set to 5. Hierarchical clustering of the original database is performed using OriginPro 2022, and the results are visualized as a dendrogram. The hierarchical clustering results of the original database are as follows: Figure 6 As shown in the figure (where letter A corresponds to 15s, B to 30s, C to 45s, and so on), the clustering results under the 5-cluster condition show that samples of cooking for 90s and 120s span two categories. Therefore, the points in the 90s and 120s experimental groups are designated as "cross-category points." After deleting these cross-category points, the final purified database is obtained after secondary purification.

[0078] DBSCAN is an unsupervised clustering method based on an improved K-means algorithm. By utilizing density information, this algorithm can identify the local density distribution of data points, discover cluster structures of arbitrary shapes, and automatically label noise points. It can not only reclassify processed data but also automatically label noise points. This embodiment uses it to demonstrate the correctness of the number of maturity levels obtained by the aforementioned PCA, k-means clustering, and hierarchical clustering in relation to the clean database. This embodiment uses the scikit-learn library to perform DBSCAN-based maturity level clustering on pork slices with different cooking times and visualizes the results. The visualization results of the clean database after DBSCAN clustering are shown below. Figure 7As shown in the figure, the confidence ellipses for the five maturity levels do not overlap at all, and each sample is assigned to only one maturity level. Furthermore, the DBSCAN clustering results did not identify any outliers, and the proportion of outliers was 0 < 0.05. This meets the requirements for further LDA analysis.

[0079] S4. Based on PCA and LDA, visualize the results of unsupervised cooking maturity level classification of dishes, and verify and optimize the results.

[0080] Since the cluster ratio of DBSCAN clustering in step S3 is 0 < 0.05, satisfying the LDA input condition, the 5 maturity levels determined in step S3 are used as category labels to perform LDA dimensionality reduction on the 11-dimensional physicochemical data. The maturity level classification method is analyzed based on the dimensionality reduction results to determine if it meets the preset standard. It is also determined whether the obtained maturity levels need to be recalculated. In this embodiment, the samples are specifically divided into 5 categories (Level 1: 15s raw meat group, Level 2: 30s, 45s early-ripened group, Level 3: 60s, 75s semi-ripe group, Level 4: 105s moderately ripe group, Level 5: 135s, 150s overripe group). The cleaned database is then subjected to LDA dimensionality reduction based on this maturity level classification method. This embodiment uses the scikit-learn library to perform LDA dimensionality reduction on pork slices with different cooking times based on DBSCAN clustering results and visualizes the results. Specific LDA dimensionality reduction results are as follows: Figure 8 As shown in the figure. The final results show that the cumulative contribution rate of the PCA algorithm is 88.5%, and the cumulative contribution rate of the LDA algorithm is 96.1%. Both are greater than the preset standard of 80%. The sample overlap rate of the PCA algorithm is 1.98%, and the sample overlap rate of the LDA algorithm is zero, both less than the preset standard of 30%. The intra-class clustering degree of the LDA algorithm is 0.96, and the inter-class separation distance is 10.51, satisfying the threshold that the inter-class separation distance is ≥2 times the intra-class clustering degree. Furthermore, the class consistency between the LDA and PCA algorithms is 100%. Therefore, the maturity level classification results in this embodiment meet all the verification criteria established by this method, and the obtained results do not require backtracking. The dish maturity level obtained based on the multi-clustering method linkage method has a certain degree of accuracy.

[0081] To verify the important role of multiple clustering methods in improving the accuracy of unsupervised maturity level classification, under the same experimental environment and data acquisition conditions as in Example 1, only PCA dimensionality reduction and DBSCAN clustering were used to process the obtained physicochemical property database.

[0082] The database used in the experiment was exactly the same as in Example 1, therefore the results after PCA dimensionality reduction were also identical. Since the optimal number of maturity levels to be divided was not determined when using the DBSCAN clustering algorithm, the two parameters of the DBSCAN algorithm—neighborhood radius and minimum sample size—were calculated and iterated through the data in the experiment. The neighborhood radius was set to 1-5 with a step size of 0.1. The minimum sample size was set to 1-20 with a step size of 1. The results showed that the DBSCAN clustering algorithm could not directly distinguish the five maturity levels. During the iteration, the experimental group with a cooking time of 15s was separated first, with a neighborhood radius of 1.8 and a minimum sample size of 2. Subsequently, the experimental group with a cooking time of 30s was separated, with a neighborhood radius of 1.6 and a minimum sample size of 4. Then, the experimental groups with cooking times of 135s and 150s were separated, with a neighborhood radius of 2.54 and a minimum sample size of 17. Subsequently, within the parameter search range, no suitable parameter combination was found that could effectively cluster without outliers. At this point, the lowest outlier ratio is 0.3125 > 0.05. This indicates that the data itself has a complex distribution, making it difficult to cluster using the DBSCAN algorithm without generating outliers.

[0083] This experiment demonstrates that the multi-clustering system constructed in this invention, which combines principal component analysis, k-means clustering, DBSCAN clustering, and hierarchical clustering, is not a simple superposition of existing algorithms. Rather, it is a groundbreaking solution tailored to the complex characteristics of dish maturity data (high-dimensional coupling, interference from transitional samples, and difficulty in determining the number of categories). Its core value and irreplaceability are fully demonstrated here. Only the multi-clustering linkage logic of this invention can effectively handle complex, nonlinear, and noisy data such as dish maturity, breaking through the application bottleneck of single algorithms in unsupervised maturity level classification. It significantly improves the accuracy, stability, and reliability of the results, providing a practical technical path for unsupervised detection of dish maturity and a highly valuable scenario-based solution for unsupervised classification of complex data in related fields. It provides reliable data support and a solid theoretical foundation for various maturity algorithms in intelligent cooking driven by artificial intelligence, effectively promoting the healthy and positive development of the industry.

[0084] This invention constructs a four-order progressive collaborative framework based on multidimensional physicochemical properties and the characteristics of maturity level classification during cooking. It successfully achieves unsupervised classification of dish maturity levels, effectively solving the problems of relying on a single indicator and subjective experience in dish maturity classification. This provides a reliable foundational method and data support for the development of intelligent cooking driven by artificial intelligence. This invention is particularly suitable for Chinese complex cooking (such as stir-frying, quick-frying, and other short-time high-temperature processes combined with multiple seasonings), solving the problem of asynchronous maturity of multiple component ingredients.

[0085] It should also be noted that, in this specification, terms such as "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0086] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for grading the cooking maturity of dishes based on unsupervised learning, characterized in that, Includes the following steps: S1. Obtain the recipe data of the target dish to determine the cooking time interval, reproduce the cooking process and control variables, and set up experimental groups according to the expected maturity classification number; S2. Select physicochemical property indicators according to the type of dish, detect physicochemical data under different cooking times, and construct a database; S3. A four-order progressive clustering framework, consisting of principal component analysis, k-means clustering, DBSCAN clustering, and hierarchical clustering methods, is used to implement unsupervised classification of the cooking maturity levels of dishes based on physicochemical properties. Specifically, this includes: S31. Principal component analysis was used to reduce the dimensionality of the high-dimensional original data and the number of experimental groups corresponding to the experimental group samples contained in the confidence ellipse of each experimental group was analyzed. The mean μ, standard deviation σ, and threshold of the number of experimental groups corresponding to the experimental group samples contained in the confidence ellipse of all experimental groups were calculated. By comparing the number of experimental groups corresponding to the experimental group samples contained in the confidence ellipse of each experimental group with the calculated threshold, the experimental groups with the fastest changes in maturity level were identified and eliminated, forming a preliminary clean database. S32. Apply the k-means elbow analysis method to the original data and the obtained preliminary clean database respectively, and determine the total number of optimal maturity categories by comparing the inflection points of clustering error. S33. Based on the total number of determined optimal maturity categories, perform systematic clustering, identify and remove the fence points and experimental groups that span two maturity categories through a tree structure, and form the final clean database. S34. Based on the final cleaned database, the DBSCAN clustering algorithm is used to complete the final division of maturity levels. The parameters of the DBSCAN algorithm are obtained by traversing the total number of optimal maturity categories determined in step S32. During the traversal, parameter combinations that make the clustering results consistent with the total number of optimal maturity categories determined in step S32 are selected first. S4. Based on PCA and LDA, visualize the results of unsupervised cooking maturity level classification of dishes, and verify and optimize the results.

2. The method for grading the cooking maturity of dishes based on unsupervised learning according to claim 1, characterized in that, In step S1, obtaining the recipe data for the target dish is specifically implemented using a Python web crawler, as follows: Select recipe pages from food websites, simulate browser requests to obtain HTML content, crawl cooking steps and time information, and save them as local files; In step S1, cooking variables include the amount of ingredients, heat level, and utensils; In step S1, the number of experimental groups is at least twice the expected number of maturity categories.

3. The method for grading the cooking maturity of dishes based on unsupervised learning according to claim 1, characterized in that, In step S2, the physicochemical property indicators selected according to the type of dish are as follows: For meat, the physicochemical properties include cooking loss, protein content, texture parameters, color difference, and moisture phase distribution. For vegetables, the physicochemical properties include moisture content, vitamin C retention rate, chlorophyll content, firmness, total phenol content, and pH value. In step S2, the sample size for each experimental group of each indicator in the constructed database is no less than 15.

4. The method for grading the cooking maturity of dishes based on unsupervised learning according to claim 1, characterized in that, In step S31, the confidence ellipse is a 95% confidence interval constructed based on the first two principal components of PCA; The threshold is specifically the mean of the number of other experimental groups included in the confidence ellipse of all experimental groups plus 1.5 times the standard deviation, i.e., μ+1.5σ, and is dynamically adjusted by ±10% according to the characteristics of different types of dishes in the rate of change of maturity. The criterion for determining the "riding the wall" experimental group is that the number of samples from other experimental groups included in the confidence ellipse is greater than or equal to the threshold.

5. The method for grading the cooking maturity of dishes based on unsupervised learning according to claim 1, characterized in that, In step S32, the criterion for determining the total number of optimal maturity categories is: When the difference between the k-value corresponding to the inflection point of the elbow analysis clustering error curve of the original data and the preliminary clean database is ≤1, the k-value corresponding to the inflection point of the elbow analysis clustering error curve of the preliminary clean database is selected as the total number of optimal maturity categories; when the difference is >1, different preliminary clean databases that meet the standards are constructed and the above steps are repeated. In addition, the total number of optimal maturity categories must be less than half of the number of experimental groups set in step S1; otherwise, the number of experimental groups in step S1 must be increased and step S2 must be executed again.

6. The method for grading the cooking maturity of dishes based on unsupervised learning according to claim 1, characterized in that, In step S33, Euclidean distance is used for system clustering, and each physicochemical index participates in the calculation with equal weight after being normalized by the maximum-minimum value. In step S34, the neighborhood radius is set to 1-5 with a step size of 0.1 during the traversal process, and the minimum number of samples is set to 1-20 with a step size of 1. The optimal parameters are selected by choosing the parameter combination that minimizes the proportion of outliers.

7. The method for grading the cooking maturity of dishes based on unsupervised learning according to claim 1, characterized in that, Step S4 specifically includes: S41. Obtain the results of the dish maturity level classification and the corresponding physicochemical data; S42. PCA is used to reduce the dimensionality of the physicochemical data and generate a two-dimensional scatter plot that shows the global cluster distribution of the samples, intuitively presenting the spatial clustering pattern of each maturity level. S43. Introduce category information from the classification results to guide linear discriminant analysis to focus on maximizing the discriminant boundary between maturity levels and generate visualization results with clear category separation characteristics; S44. Calculate the inter-class separation distance and intra-class clustering degree of each maturity level in the LDA space to quantitatively evaluate the reliability of the partitioning results; S45. If the deviation between the visualization results of PCA and LDA exceeds the preset range or the inter-class separation distance does not reach the threshold, backtrack to optimize the straddling point and experimental group elimination steps or the total number of categories calibration steps until the results meet the standards.

8. The method for grading the cooking maturity of dishes based on unsupervised learning according to claim 7, characterized in that, In step S43, the category information must satisfy the outlier ratio <0.05; if the standard is not met, the secondary fence point removal parameters in step S33 are backtracked and optimized. In step S44, the preset threshold for inter-class separation distance is twice the average intra-class aggregation degree of each level.

9. A method for grading the cooking maturity of dishes based on unsupervised learning according to claim 7, characterized in that, In step S45, backtracking optimization includes: If the overlap rate is >30%, then repeat steps S31 and S33, and adjust the judgment threshold to increase the strictness of the exclusion of the "riding the wall" point and the experimental group. If the overlap rate is greater than 50%, then step S32 is executed again to re-cluster the optimal total number of categories by ±1.