Food cooking state detection method and food cooking equipment

By combining the feature fusion processing of maximum pooling and average pooling, the problems of information loss and noise impact in the prior art are solved, detailed description and accurate detection of food cooking status are realized, and automation and intelligence of intelligent cooking equipment are improved.

CN120260032APending Publication Date: 2025-07-04NINGBO FOTILE KITCHEN WARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234318.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the maximum pooling discards non-maximum information, resulting in the loss of detailed information, the average pooling is sensitive to outliers and is susceptible to noise, affecting the accuracy and robustness of food cooking status detection.

Method used

Using a method that combines maximum pooling and average pooling, the texture features and background information of the image are obtained through feature fusion processing, and the advantages of the two pooling are combined to improve detection accuracy.

Benefits of technology

It realizes a more comprehensive and detailed description of the cooking status of food, improves the accuracy and robustness of detection, facilitates users to understand the cooking process in real time, and improves the automation and intelligence level of intelligent cooking equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260032A_ABST
    Figure CN120260032A_ABST
Patent Text Reader

Abstract

The invention discloses a food cooking state detection method and food cooking equipment, and the method comprises the steps: obtaining an image feature extraction model which comprises a feature extraction layer, a maximum pooling layer and an average pooling layer; inputting a target image of cooking food into a feature extraction layer of the initial image feature extraction model to obtain initial image features corresponding to the target image; respectively inputting the initial image features into a maximum pooling layer and an average pooling layer to obtain a maximum pooling feature matrix corresponding to the maximum pooling layer and an average pooling feature matrix corresponding to the average pooling layer; performing feature fusion processing on the maximum pooling feature matrix and the average pooling feature matrix to obtain a target image feature matrix corresponding to the target image; and analyzing the target image based on the target image feature matrix, and determining the cooking state of the cooked food. According to the method, two pooling modes are integrated, texture features and background information are considered, the food detection capability is enhanced, and the cooking performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of household appliances, and particularly to a method for detecting the cooking state of food and a food cooking device. Background Art

[0002] During the working process of a food cooking device, it often detects the cooking state of food in order to better control the cooking time and temperature and cook more delicious food. Therefore, the detection effect of the cooking state of food will directly affect the success or failure of the food cooking device. With the development of artificial intelligence, more and more image detection algorithms are applied to the detection of steam-convection-oven food. The intelligent algorithm first extracts features of an image in different dimensions, and then analyzes and judges these features. Therefore, the method of feature extraction of the image will directly affect the detection effect of the image. The current mainstream methods are average pooling or max pooling, which are used to strengthen or weaken the features of a certain dimension of the picture. However, max pooling discards the non-maximum information, which may lose some detailed information. Continuously using max pooling in a multi-layer network may lead to over-abstraction of information, causing some subtle features to be ignored; average pooling is sensitive to outliers, may be affected by noise, and it cannot highlight the most significant features in the image like max pooling. Summary of the Invention

[0003] The present invention aims to at least solve one of the technical problems existing in the prior art. For this purpose, the first aspect of the present invention proposes a method for detecting the cooking state of food, including:

[0004] Obtaining an image feature extraction model, where the image feature extraction model includes a feature extraction layer, a max pooling layer, and an average pooling layer;

[0005] Inputting a target image of the cooking food into the feature extraction layer of the initial image feature extraction model to obtain the initial image features corresponding to the target image;

[0006] Inputting the initial image features into the max pooling layer and the average pooling layer respectively to obtain a max pooling feature matrix corresponding to the max pooling layer and an average pooling feature matrix corresponding to the average pooling layer;

[0007] Performing feature fusion processing on the max pooling feature matrix and the average pooling feature matrix to obtain a target image feature matrix corresponding to the target image;

[0008] Analyzing the target image based on the target image feature matrix to determine the cooking state of the cooking food.

[0009] Further, the process of performing feature fusion on the max-pooling feature matrix and the average-pooling feature matrix to obtain the target image feature matrix corresponding to the target image includes:

[0010] Performing arithmetic processing based on the max-pooling feature matrix and the average-pooling feature matrix to obtain a first weight matrix corresponding to the max-pooling feature matrix and a second weight matrix corresponding to the average-pooling feature matrix;

[0011] Based on the first weight matrix and the second weight matrix, perform weighted calculations on the average-pooling feature matrix and the max-pooling feature matrix respectively to obtain the target image feature matrix.

[0012] Further, the matrix dimension of the max-pooling feature matrix is the same as the matrix dimension of the average-pooling feature matrix;

[0013] The process of performing arithmetic processing based on the max-pooling feature matrix and the average-pooling feature matrix to obtain a first weight matrix corresponding to the max-pooling feature matrix and a second weight matrix corresponding to the average-pooling feature matrix includes:

[0014] Obtain an initial weight matrix, and the matrix dimension of the initial weight matrix is the same as the matrix dimension of the max-pooling feature matrix;

[0015] Perform ratio calculation based on the elements at each position of the max-pooling feature matrix and the elements at the corresponding positions of the average-pooling feature matrix to obtain an element calculation result;

[0016] Based on the element calculation result, update the elements at the corresponding positions of the initial weight matrix to obtain the first weight matrix;

[0017] Based on the difference between the identity matrix and the first weight matrix, obtain the second weight matrix.

[0018] Further, the process of performing weighted calculations on the average-pooling feature matrix and the max-pooling feature matrix respectively based on the first weight matrix and the second weight matrix to obtain the target image feature matrix includes:

[0019] Obtain a preset noise threshold matrix and an initial target image feature matrix; the matrix dimension of the preset noise threshold matrix is the same as the matrix dimension of the first weight matrix;

[0020] Traverse each element in the first weight matrix, and compare the element at the current position in the first weight matrix with the element at the corresponding position in the preset noise threshold matrix;

[0021] When the element at the current position in the first weight matrix is less than the element at the corresponding position in the preset noise threshold matrix, the elements at the corresponding positions in the max-pooling feature matrix are weighted based on the elements at each position in the first weight matrix to obtain a first weighted result;

[0022] The elements at the corresponding positions in the average-pooling feature matrix are weighted based on the elements at the corresponding positions in the second weight matrix to obtain a second weighted result;

[0023] A first feature element is determined based on the sum of the first weighted result and the second weighted result;

[0024] The elements at the corresponding positions in the initialized target image feature matrix are updated based on the first feature element to obtain the target image feature matrix.

[0025] Further, the method further includes:

[0026] When the element at the current position in the first weight matrix is greater than or equal to the element at the corresponding position in the preset noise threshold matrix, the elements at the corresponding positions in the average-pooling feature matrix are weighted based on the elements at the corresponding positions in the second weight matrix to obtain the second weighted result;

[0027] A second feature element is determined based on the difference between the element at the corresponding position in the second weight matrix and the second weighted result;

[0028] The elements at the corresponding positions in the initialized target image feature matrix are updated based on the second feature element to obtain the target image feature matrix.

[0029] Further, after performing feature fusion processing on the max-pooling feature matrix and the average-pooling feature matrix to obtain the target image feature matrix corresponding to the target image, the method further includes:

[0030] When there is at least one target element in the target image feature matrix, the at least one target element is approximated to obtain an approximated target image feature matrix; the target element is an element that does not meet the preset data accuracy.

[0031] Further, the method further includes:

[0032] Repeat the steps: During the food cooking process, obtain the original image of the food at a preset time interval; perform preprocessing on the original image to obtain the target image, where the preprocessing includes at least one of cropping, rotation, and padding; obtain an image feature extraction model, where the image feature extraction model includes a feature extraction layer, a max pooling layer, and an average pooling layer; until analyzing the target image based on the target image feature matrix to determine the cooking state of the cooked food;

[0033] Until the cooking state is a preset state.

[0034] Further, the step of respectively inputting the initial image features into the max pooling layer and the average pooling layer to obtain the max pooling feature matrix corresponding to the max pooling layer and the average pooling feature matrix corresponding to the average pooling layer includes:

[0035] Determine the max pooling window;

[0036] Traverse the initial image features based on the max pooling window according to a preset stride;

[0037] Within each max pooling window, determine the maximum value among the elements within the max pooling window;

[0038] Determine the max pooling feature matrix based on the maximum value of the elements within each max pooling window.

[0039] Further, the method further includes:

[0040] Determine the average pooling window;

[0041] Traverse the initial image features based on the average pooling window according to a preset stride;

[0042] Within each average pooling window, calculate the arithmetic mean of all the elements within the average pooling window;

[0043] Determine the average pooling feature matrix based on the arithmetic mean of the elements within each average pooling window.

[0044] A second aspect of the present invention provides a food cooking device, including: an image acquisition device and a controller;

[0045] The image acquisition device is used to obtain the original image of the food during the food cooking process;

[0046] The controller is used to obtain an image feature extraction model, where the image feature extraction model includes a feature extraction layer, a max pooling layer, and an average pooling layer;

[0047] Input the target image of the cooked food into the feature extraction layer of the initial image feature extraction model to obtain the initial image features corresponding to the target image; the target image is obtained by preprocessing the original image.

[0048] Input the initial image features into the max pooling layer and the average pooling layer respectively to obtain the max pooling feature matrix corresponding to the max pooling layer and the average pooling feature matrix corresponding to the average pooling layer.

[0049] Perform feature fusion processing on the max pooling feature matrix and the average pooling feature matrix to obtain the target image feature matrix corresponding to the target image.

[0050] Analyze the target image based on the target image feature matrix to determine the cooking state of the cooked food.

[0051] For the food cooking state detection method and food cooking device provided by the present invention as described above, the beneficial effects of the present invention are as follows: By obtaining an image feature extraction model, which includes a feature extraction layer, a max pooling layer, and an average pooling layer, the max pooling layer and the average pooling layer are introduced, thereby obtaining the texture features and background information of the image, realizing a more comprehensive and detailed description of the image features, and helping to improve the accuracy of food cooking state detection; By inputting the target image of the cooked food into the feature extraction layer of the initial image feature extraction model to obtain the initial image features corresponding to the target image, it provides the initial image features for subsequent pooling operations and lays the foundation for accurately extracting the deep features of the image; By inputting the initial image features into the max pooling layer and the average pooling layer respectively to obtain the max pooling feature matrix corresponding to the max pooling layer and the average pooling feature matrix corresponding to the average pooling layer, it captures the significant features in the image and retains the overall information of the image. The combination of the two provides richer feature information and provides multi-dimensional basic data for feature fusion; By performing feature fusion processing on the max pooling feature matrix and the average pooling feature matrix to obtain the target image feature matrix corresponding to the target image, it combines the advantages of max pooling and average pooling, can more comprehensively reflect various features of the image, improves the model's ability to capture details and global information, and makes the detection results more robust and accurate; By analyzing the target image based on the target image feature matrix to determine the cooking state of the cooked food, it can accurately evaluate the current cooking state of the food, facilitate the user to understand the cooking process of the food in real time, and improve the automation and intelligence level of intelligent cooking devices.

[0052] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0054] Figure 1 It is a flowchart of a method for detecting the cooking state of food provided by an embodiment of the present invention;

[0055] Figure 2 It is a flowchart of a maximum pooling processing and average pooling processing method provided by an embodiment of the present invention;

[0056] Figure 3 It is a schematic diagram of a maximum pooling feature extraction method provided by an embodiment of the present invention;

[0057] Figure 4 It is a schematic diagram of an average pooling feature extraction method provided by an embodiment of the present invention;

[0058] Figure 5 It is a flowchart of a feature fusion processing method provided by an embodiment of the present invention;

[0059] Figure 6 It is a flowchart of a first weight matrix calculation method provided by an embodiment of the present invention;

[0060] Figure 7 It is a flowchart of a weighted calculation method provided by an embodiment of the present invention;

[0061] Figure 8 It is a schematic diagram of a weighted pooling calculation method provided by an embodiment of the present invention;

[0062] Figure 9 It is a structural block diagram of a food cooking device provided by an embodiment of the present invention. Detailed implementation manners

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout.

[0064] Embodiment

[0065] In view of the problem that the existing max pooling discards non-maximum information, which may result in the loss of some detailed information, and the continuous use of max pooling in a multi-layer network may lead to over-abstraction of information, causing some subtle features to be ignored; and the average pooling is sensitive to outliers, may be affected by noise, and it cannot highlight the most significant features in the image like max pooling. The present invention proposes a method for detecting the cooking state of food, which provides a solution for a food cooking device that combines the two types of pooling, takes into account texture features and background information, enhances the food detection ability, and improves the cooking performance. This technical solution can not only enable the extracted image features to retain some key texture features and well retain the background information, but also improve the cooking performance of the device through the cooking performance of the device. Figure 1 It is a flowchart of a method for detecting the cooking state of food provided by an embodiment of the present invention. This specification provides the method operation steps such as in the embodiment or flowchart, but based on routine or non-creative labor, there may be more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps, and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order shown in the embodiment or the drawing or in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 1 shown, the controller of the device executes the following steps:

[0066] S101: Obtain an image feature extraction model, which includes a feature extraction layer, a max pooling layer, and an average pooling layer;

[0067] Specifically, the feature extraction layer extracts corresponding image features from the image, providing image features for subsequent pooling operations; the max pooling layer obtains the max pooling feature matrix corresponding to the image features from the image features, which can capture the significant features in the image and retain the overall information of the image; the average pooling layer obtains the average pooling feature matrix corresponding to the image features from the image features, which can retain the overall information of the image features, provide basic smooth features for the recognition of the image, and enhance the comprehensiveness of the analysis.

[0068] By obtaining an image feature extraction model, which includes a feature extraction layer, a max pooling layer, and an average pooling layer, the max pooling layer and the average pooling layer are introduced, thereby obtaining the texture features and background information of the image, realizing a more comprehensive and detailed description of the image features, and helping to improve the accuracy of food cooking state detection;

[0069] S102: Input the target image of the cooking food into the feature extraction layer of the initial image feature extraction model to obtain the initial image features corresponding to the target image.

[0070] Specifically, the feature extraction layer extracts the initial image features corresponding to the target image from the target image of the cooked food, where the target image is obtained by preprocessing the original image, and the preprocessing includes at least one of clipping, rotation, and padding.

[0071] By inputting the target image of the cooked food into the feature extraction layer of the initial image feature extraction model, the initial image features corresponding to the target image are obtained, providing the initial image features for subsequent pooling operations and laying the foundation for accurately extracting the deep features of the image.

[0072] S103: Input the initial image features into the max pooling layer and the average pooling layer respectively to obtain the max pooling feature matrix corresponding to the max pooling layer and the average pooling feature matrix corresponding to the average pooling layer;

[0073] Specifically, Figure 2 is the flowchart of the max pooling process and the average pooling process provided by the embodiments of the present invention. Specifically, as Figure 2 shown, step S103 includes the following steps:

[0074] S201: Determine the max pooling window;

[0075] Specifically, in one embodiment, the input image or feature map is divided into a series of overlapping or non - overlapping small regions, and the size of the max pooling window (e.g., 2*2) is selected, as Figure 3 shown in

[0076] S202: Based on the preset stride, traverse the initial image features based on the max pooling window;

[0077] Specifically, in one embodiment, slide the pooling window on the input image or feature map and move it according to the specified preset stride until the entire initial image features are traversed, as Figure 3 shown in Feature one Feature two

[0078] S203: In each max pooling window, determine the maximum value among the elements in the max pooling window;

[0079] Specifically, in one embodiment, as Figure 3 shown in

[0080] the maximum value 7 in feature one, the maximum value 8 in feature two, and the maximum value 9 in feature three are obtained respectively.

[0081] Specifically, in one embodiment, asFigure 3 As shown, the maximum pooling feature matrix maxpool is [7 8 9].

[0082] Through the precise operation of the max pooling layer, max pooling helps to retain the significant features in the picture, that is, the maximum value within the pooling region. Such an operation is particularly useful for extracting important visual features such as textures and edges. It can also increase the invariance of the network to small position changes, rotations, and scale transformations. In addition, by reducing the number of features, max pooling helps to prevent the model from overfitting to the training data and improves the generalization ability of the model.

[0083] S205: Determine the average pooling window;

[0084] Specifically, in one embodiment, the input image or feature map is divided into a series of overlapping or non - overlapping small regions, and the size of the average pooling window (e.g., 2*2) is selected, as Figure 4 shown.

[0085] S206: Traverse the initial image features based on the average pooling window according to the preset stride;

[0086] Specifically, in one embodiment, the pooling window is slid on the input image or feature map and moved according to the specified preset stride until the entire initial image features are traversed, as Figure 4 shown, to obtain Feature One Feature Two Feature Three

[0087] S207: Calculate the arithmetic mean of all elements within each average pooling window;

[0088] Specifically, in one embodiment, as Figure 4 shown, the arithmetic mean of all elements in Feature One is 4, the arithmetic mean of all elements in Feature Two is 4, and the arithmetic mean of all elements in Feature Three is 5.

[0089] S208: Determine the average pooling feature matrix based on the arithmetic mean of the elements within each average pooling window.

[0090] Specifically, in one embodiment, as Figure 4 shown, the average pooling feature matrix meanpool is [4 4 5].

[0091] Through the precise operation of average pooling, all pixel points within the pooling window are considered, and the features are averaged, which helps to retain the background information of the image and reflects the overall statistical characteristics of the input, contributing to expressing the overall information of the image, smooth textures, etc. Since all values are processed smoothly, there will be no excessive feature deviation due to individual extreme points, providing a basic smooth feature for the recognition of cooking states and enhancing the comprehensiveness of the analysis.

[0092] S104: Perform feature fusion processing on the max-pooling feature matrix and the average-pooling feature matrix to obtain the target image feature matrix corresponding to the target image;

[0093] Specifically, Figure 5 is the flowchart of the feature fusion processing provided by the embodiment of the present invention. Specifically, as Figure 5 shown, step S104 includes the following steps:

[0094] S501: Based on the max-pooling feature matrix and the average-pooling feature matrix, perform arithmetic processing to obtain the first weight matrix corresponding to the max-pooling feature matrix and the second weight matrix corresponding to the average-pooling feature matrix;

[0095] Specifically, Figure 6 is the flowchart of the calculation of the first weight matrix provided by the embodiment of the present invention. Specifically, as Figure 6 shown, step S501 includes the following steps:

[0096] S601: Obtain the initialized weight matrix, and the matrix dimension of the initialized weight matrix is the same as that of the max-pooling feature matrix;

[0097] Specifically, in one embodiment, the obtained initialized weight matrix is [0 0 0].

[0098] S602: Based on the elements at each position of the max-pooling feature matrix and the elements at the corresponding positions of the average-pooling feature matrix, perform ratio calculation to obtain the element calculation result;

[0099] Specifically, the calculation formula is:

[0100] A i,j =(maxpool i,j -meanpool i,j ) / meanpool i,j (Equation 1)

[0101] where A i,j is the element in the i-th row and j-th column of the first weight matrix, maxpool i,j is the element corresponding to the i-th row and j-th column in the max-pooling matrix, and meanpool i,jis the element corresponding to the \(i\)-th row and \(j\)-th column in the average pooling matrix.

[0102] Specifically, in one embodiment, it is calculated that:

[0103] A 1,1 =(7 - 4) / 4 = 0.75

[0104] A 1,2 =(8 - 4) / 4 = 1

[0105] A 1,3 =(9 - 5) / 5 = 0.8

[0106] S603: Based on the element calculation results, update the elements at the corresponding positions in the initialized weight matrix to obtain the first weight matrix;

[0107] Specifically, in one embodiment, the updated first weight matrix is [0.75 1 0.8].

[0108] S604: Based on the difference between the identity matrix and the first weight matrix, obtain the second weight matrix;

[0109] Specifically, the calculation formula is:

[0110] B = ones - A (Equation 2)

[0111] where A is the first weight matrix, B is the second weight matrix, ones is the identity matrix, and the dimension of ones is the same as that of A and V.

[0112] Specifically, in one embodiment, it is calculated that:

[0113] B = [1 1 1] - [0.75 1 0.8] = [0.25 0 0.2]

[0114] By initializing the weight matrix and calculating based on the ratio of these two pooling feature matrices to determine the first and second weight matrices, not only the information of the maximum value and the average value is utilized, but also the weight distribution between them is optimized, thereby better retaining the image feature information and improving the accuracy, efficiency and stability of food cooking state detection.

[0115] S502: Based on the first weight matrix and the second weight matrix, perform weighted calculations on the average pooling feature matrix and the max pooling feature matrix respectively to obtain the target image feature matrix;

[0116] Specifically, Figure 7 is the flowchart of the weighted calculation provided by the embodiment of the present invention. Specifically, as Figure 7 shown, step S502 includes the following steps:

[0117] S701: Obtain a preset noise threshold matrix and initialize a target image feature matrix; the matrix dimension of the preset noise threshold matrix is the same as that of the first weight matrix;

[0118] The calculation formula for obtaining the preset noise threshold matrix is:

[0119] D = 0.8ones(Equation 3)

[0120] where D is the noise threshold matrix, 0.8 is a fixed value, ones is the identity matrix, and the dimension of ones is the same as that of A and B.

[0121] Specifically, in one embodiment, the initialized target image feature matrix C obtained is [0 0 0], and the preset noise threshold matrix D is [0.8 0.8 0.8]

[0122] S702: Traverse each element in the first weight matrix, and compare the element at the current position in the first weight matrix with the element at the corresponding position in the preset noise threshold matrix;

[0123] Specifically, compare A i,j with D i,j where A i,j is the element in the i-th row and j-th column of the first weight matrix, and D i,j is the element in the corresponding i-th row and j-th column of the noise threshold matrix.

[0124] S703: When the element at the current position in the first weight matrix is less than the element at the corresponding position in the preset noise threshold matrix, weight the element at the corresponding position in the max-pooling feature matrix based on each element in the first weight matrix to obtain a first weighted result;

[0125] S704: Weight the element at the corresponding position in the average-pooling feature matrix based on the element at the corresponding position in the second weight matrix to obtain a second weighted result.

[0126] S705: Determine a first feature element based on the sum of the first weighted result and the second weighted result;

[0127] Specifically, the calculation formulas for steps S703 to S705 are:

[0128] If A i,j < D i,j :

[0129] C i,j = A i,j * maxpool i,j + B i,j * meanpool i,j (Equation 4)

[0130] Among them, C i,j is the element at the i-th row and j-th column in the initialized target image feature matrix, A i,j is the element corresponding to the i-th row and j-th column in the first weight matrix, maxpool i,j is the element corresponding to the i-th row and j-th column in the max pooling matrix, B i,j is the element corresponding to the i-th row and j-th column in the second weight matrix, meanpool i,j is the element corresponding to the i-th row and j-th column in the average pooling matrix.

[0131] Specifically, in one embodiment, since: A 1,1 = 0.75 < D 1,1 = 0.8

[0132] Then it is calculated that: C 1,1 = 0.75 * 7 + 0.25 * 4 = 6.25

[0133] S706: Update the element at the corresponding position in the initialized target image feature matrix based on the first feature element to obtain the target image feature matrix;

[0134] Specifically, in one embodiment, the updated target image feature matrix C obtained is [6.25 0 0].

[0135] Specifically, Figure 8 contains the weighted pooling calculation schematic diagram in steps S701 to S706 provided by the embodiments of the present invention when the element at the current position in the first weight matrix is less than the element at the corresponding position in the preset noise threshold matrix. Specifically, as Figure 8 shown.

[0136] By setting the noise threshold matrix and judging the relationship between each element in the weight matrix and the noise threshold based on this, to select a suitable calculation method, it is possible to select the most suitable calculation method according to the actual situation, optimize the extraction process of the target image features, reduce the influence of noise, and improve the accuracy of feature extraction.

[0137] S707: When the element at the current position in the first weight matrix is greater than or equal to the element at the corresponding position in the preset noise threshold matrix, weight the element at the corresponding position in the average pooling feature matrix based on the element at the corresponding position in the second weight matrix to obtain the second weighted result.

[0138] S708: Determine the second feature element based on the difference between the element at the corresponding position in the second weight matrix and the second weighted result;

[0139] Specifically, the calculation formulas for steps S707 to S708 are:

[0140] If A i,j ≥D i,j :

[0141] C i,j = meanpool i,j -B i,j *meanpool i,j (Equation Five)

[0142] Wherein, C i,j is the element at the i-th row and j-th column in the initialized target image feature matrix, B i,j is the element corresponding to the i-th row and j-th column in the second weight matrix, and meanpool i,j is the element corresponding to the i-th row and j-th column in the average pooling matrix.

[0143] Specifically, in one embodiment, since: A 1,2 = 1 > D 1,2 = 0.8,

[0144] Then it is calculated that: C 1,2 = 4 - 0 * 4 = 4;

[0145] Since: A 1,3 = 0.8 = D 1,2 = 0.8,

[0146] Then it is calculated that: C 1,3 = 5 - 0.2 * 5 = 4;

[0147] S709: Update the elements at the corresponding positions in the initialized target image feature matrix based on the second feature element to obtain the target image feature matrix;

[0148] Specifically, in one embodiment, the updated target image feature matrix C is [6.25 4 4].

[0149] Specifically, Figure 8 includes the weighted pooling calculation schematic diagrams in steps S707 to S709 provided by the embodiments of the present invention when the element at the current position in the first weight matrix is less than or equal to the element at the corresponding position in the preset noise threshold matrix, specifically as Figure 8 shown.

[0150] By using different calculation methods to determine the element values in the target feature matrix, it is allowed to select different processing methods according to the importance of the features, so that both the key features in the image can be emphasized and the unimportant noises can be weakened, improving the accuracy of the final cooking state judgment.

[0151] S503: When there is at least one target element in the target image feature matrix, perform approximation processing on the at least one target element to obtain the target image feature matrix after approximation processing; the target element is an element that does not meet the preset data accuracy.

[0152] Specifically, the approximation processing can be at least one of rounding, ceiling, floor, or other approximation methods.

[0153] Specifically, in one embodiment, the rounding approximation method can be adopted, and the final target image feature matrix is [6 4 4].

[0154] By performing approximation processing on the element values in the feature matrix to obtain the final target image feature matrix, it helps to quickly extract the features of the target image and meet the requirements of real-time or near-real-time food cooking monitoring.

[0155] S105: Analyze the target image based on the target image feature matrix to determine the cooking state of the cooked food.

[0156] Specifically, the extracted image features are classified and regressed through a neural network to obtain the detection result.

[0157] By analyzing the target image based on the target image feature matrix to determine the cooking state of the cooked food, it can accurately evaluate the current cooking state of the food, facilitate the user to understand the cooking process of the food in real time, and improve the automation and intelligence level of the intelligent cooking device.

[0158] In the actual food cooking process, the controller of the food cooking device repeats the steps: during the food cooking process, obtain the original image of the food at a preset time interval; perform preprocessing on the original image to obtain the target image, and the preprocessing includes at least one of cropping, rotating, and filling; obtain the image feature extraction model, and the image feature extraction model includes a feature extraction layer, a max pooling layer, and an average pooling layer; until analyzing the target image based on the target image feature matrix to determine the cooking state of the cooked food; until the cooking state is the preset state to achieve real-time detection of the cooking state of the cooked food.

[0159] By setting a fixed time interval to obtain the images of the cooked food and performing standardized preprocessing steps on these images, it ensures the consistency and quality of the image data input to the image feature extraction model; by cycling through the steps, it improves the dynamic tracking accuracy of the cooking state detection, and by continuously obtaining the latest cooking data, it provides the possibility of timely response for the control system of the cooking device, enabling quick adjustment of cooking parameters according to the real-time cooking situation, realizing continuous and real-time monitoring of the food cooking process, and increasing the adaptability of the cooking device to various ingredients and cooking methods.

[0160] Figure 9 It is a structural block diagram of a food cooking device provided by an embodiment of the present invention. Specifically, as Figure 9 shown, an embodiment of the present invention also provides a food cooking device, and this device includes the following parts:

[0161] An image acquisition device 901, configured to acquire an original image of food during the food cooking process;

[0162] A controller 902, configured to acquire an image feature extraction model, and the image feature extraction model includes a feature extraction layer, a max pooling layer, and an average pooling layer;

[0163] Input the target image of the cooked food into the feature extraction layer of the initial image feature extraction model to obtain the initial image features corresponding to the target image; the target image is obtained by preprocessing the original image;

[0164] Input the image features into the max pooling layer and the average pooling layer respectively to obtain a max pooling feature matrix corresponding to the max pooling layer and an average pooling feature matrix corresponding to the average pooling layer;

[0165] Perform feature fusion processing on the max pooling feature matrix and the average pooling feature matrix to obtain a target image feature matrix corresponding to the target image;

[0166] Analyze the target image based on the target image feature matrix to determine the cooking state of the cooked food.

[0167] It should be noted that without departing from the scope of the present disclosure, the food cooking state detection method provided by the embodiment of the present invention can be applied to stove steam baking devices, cooking machines, food processors, or any other type of food cooking detection.

[0168] As can be seen from the embodiments of the food cooking state detection method and the food cooking device provided by the present invention described above, in the embodiments of the present invention, by obtaining an image feature extraction model, which includes a feature extraction layer, a max pooling layer, and an average pooling layer, the max pooling layer and the average pooling layer are introduced, so as to obtain the texture features and background information of the image, realize a more comprehensive and detailed description of the image features, and help improve the accuracy of food cooking state detection; by inputting the target image of the cooked food into the feature extraction layer of the initial image feature extraction model, the initial image features corresponding to the target image are obtained, which provides the initial image features for subsequent pooling operations and lays a foundation for accurately extracting the deep features of the image; by inputting the initial image features into the max pooling layer and the average pooling layer respectively, the max pooling feature matrix corresponding to the max pooling layer and the average pooling feature matrix corresponding to the average pooling layer are obtained, capturing the significant features in the image and retaining the overall information of the image. The combination of the two provides richer feature information and provides multi-dimensional basic data for feature fusion; by performing feature fusion processing on the max pooling feature matrix and the average pooling feature matrix, the target image feature matrix corresponding to the target image is obtained, which combines the advantages of max pooling and average pooling, can more comprehensively reflect various features of the image, improves the model's ability to capture details and global information, and makes the detection results more robust and accurate; by analyzing the target image based on the target image feature matrix, the cooking state of the cooked food is determined, the current cooking state of the food can be accurately evaluated, which is convenient for users to understand the cooking process of the food in real time and improves the automation and intelligence level of the intelligent cooking device;

[0169] It should be noted that: the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0170] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0171] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting the cooking state of food, characterized in that, The method includes: Obtain an image feature extraction model, where the image feature extraction model includes a feature extraction layer, a max pooling layer, and an average pooling layer; Input the target image of the cooked food into the feature extraction layer of the initial image feature extraction model to obtain the initial image features corresponding to the target image; Input the initial image features into the max pooling layer and the average pooling layer respectively to obtain the max pooling feature matrix corresponding to the max pooling layer and the average pooling feature matrix corresponding to the average pooling layer; Perform feature fusion processing on the max pooling feature matrix and the average pooling feature matrix to obtain the target image feature matrix corresponding to the target image; Analyze the target image based on the target image feature matrix to determine the cooking state of the cooked food.

2. The method according to claim 1, characterized in that, The performing feature fusion processing on the max pooling feature matrix and the average pooling feature matrix to obtain the target image feature matrix corresponding to the target image includes: Perform arithmetic processing based on the max pooling feature matrix and the average pooling feature matrix to obtain the first weight matrix corresponding to the max pooling feature matrix and the second weight matrix corresponding to the average pooling feature matrix; Based on the first weight matrix and the second weight matrix, perform weighted calculations on the average pooling feature matrix and the max pooling feature matrix respectively to obtain the target image feature matrix.

3. The method according to claim 2, characterized in that, The matrix dimension of the max pooling feature matrix is the same as the matrix dimension of the average pooling feature matrix; The performing arithmetic processing based on the max pooling feature matrix and the average pooling feature matrix to obtain the first weight matrix corresponding to the max pooling feature matrix and the second weight matrix corresponding to the average pooling feature matrix includes: Obtain an initialized weight matrix, where the matrix dimension of the initialized weight matrix is the same as the matrix dimension of the max pooling feature matrix; Perform ratio calculation based on the elements at each position of the max pooling feature matrix and the elements at the corresponding positions of the average pooling feature matrix to obtain the element calculation results; Update the elements at the corresponding positions of the initialized weight matrix based on the element calculation results to obtain the first weight matrix; Obtain the second weight matrix based on the difference between the identity matrix and the first weight matrix.

4. The method according to claim 2, characterized in that, The performing weighted calculations on the average pooling feature matrix and the max pooling feature matrix respectively based on the first weight matrix and the second weight matrix to obtain the target image feature matrix includes: Obtain a preset noise threshold matrix and an initialized target image feature matrix; the matrix dimension of the preset noise threshold matrix is the same as the matrix dimension of the first weight matrix; Traverse each element in the first weight matrix, and compare the element at the current position in the first weight matrix with the element at the corresponding position in the preset noise threshold matrix. When the element at the current position in the first weight matrix is less than the element at the corresponding position in the preset noise threshold matrix, the elements at the corresponding positions in the max-pooling feature matrix are weighted based on the elements at each position in the first weight matrix to obtain a first weighted result; The elements at the corresponding positions in the average-pooling feature matrix are weighted based on the elements at the corresponding positions in the second weight matrix to obtain a second weighted result; A first feature element is determined based on the sum of the first weighted result and the second weighted result; The elements at the corresponding positions in the initialized target image feature matrix are updated based on the first feature element to obtain the target image feature matrix.

5. The method according to claim 4, wherein The method further includes: When the element at the current position in the first weight matrix is greater than or equal to the element at the corresponding position in the preset noise threshold matrix, the elements at the corresponding positions in the average-pooling feature matrix are weighted based on the elements at the corresponding positions in the second weight matrix to obtain the second weighted result; A second feature element is determined based on the difference between the element at the corresponding position in the second weight matrix and the second weighted result; The elements at the corresponding positions in the initialized target image feature matrix are updated based on the second feature element to obtain the target image feature matrix.

6. The method according to claim 4, wherein After performing feature fusion processing on the max-pooling feature matrix and the average-pooling feature matrix to obtain the target image feature matrix corresponding to the target image, the method further includes: When there is at least one target element in the target image feature matrix, the at least one target element is approximated to obtain an approximated target image feature matrix; the target element is an element that does not meet the preset data accuracy.

7. The method according to claim 1, characterized in that The method further includes: Repeatedly execute the steps: during the food cooking process, obtain the original image of the food at a preset time interval; preprocess the original image to obtain the target image, and the preprocessing includes at least one of preprocessing methods such as cropping, rotation, and padding; obtain an image feature extraction model, and the image feature extraction model includes a feature extraction layer, a max-pooling layer, and an average-pooling layer; until analyzing the target image based on the target image feature matrix to determine the cooking state of the cooked food; Until the cooking state is a preset state.

8. The method according to claim 1, characterized in that The step of inputting the initial image features into the max-pooling layer and the average-pooling layer respectively to obtain the max-pooling feature matrix corresponding to the max-pooling layer and the average-pooling feature matrix corresponding to the average-pooling layer includes: Determine the max-pooling window; Based on the preset stride, traverse the initial image features based on the max-pooling window; Within each max-pooling window, determine the maximum value among the elements within the max-pooling window; Determine the max-pooling feature matrix based on the maximum values of the elements within each max-pooling window.

9. The method according to claim 8, wherein The method further includes: Determine the average-pooling window; Based on the preset stride, traverse the initial image features based on the average-pooling window; Within each of the average pooling windows, calculate the arithmetic mean of all elements within the average pooling window; Determine the average pooling feature matrix based on the arithmetic mean of the elements within each average pooling window.

10. A food cooking device, characterized in that, The device includes: an image acquisition device and a controller; The image acquisition device is configured to acquire an original image of food during the food cooking process; The controller is configured to obtain an image feature extraction model, which includes a feature extraction layer, a max pooling layer, and an average pooling layer; Input the target image of the cooked food into the feature extraction layer of the initial image feature extraction model to obtain the initial image features corresponding to the target image; the target image is obtained by preprocessing the original image; Input the initial image features into the max pooling layer and the average pooling layer respectively to obtain the max pooling feature matrix corresponding to the max pooling layer and the average pooling feature matrix corresponding to the average pooling layer; Perform feature fusion processing on the max pooling feature matrix and the average pooling feature matrix to obtain the target image feature matrix corresponding to the target image; Analyze the target image based on the target image feature matrix to determine the cooking state of the cooked food.