A food cooking state detection method and system based on cross-modal deep learning
Through cross-modal deep learning methods, combined with improved Yolo model and TCN model, a mapping relationship between food appearance and center maturity is established, which solves the problem of difficulty in comprehensively evaluating the maturity in the food cooking process in the existing technology, real-time monitoring and evaluation of food cooking status is realized, and the intelligence of the cooking system in the food industry is improved.
Patent Information
- Application Number
- CN202411666916.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The prior art is difficult to comprehensively evaluate the appearance maturity and central maturity in the cooking process of food, resulting in insufficient intelligence in the cooking system of the food industry, and the detection of traditional temperature probes is time-consuming and labor-intensive and destroys the appearance of food.
The cross-modal deep learning method is adopted to extract image features through the improved Yolo model, combine the UMAP algorithm dimensionality reduction and the TCN model to extract temperature-time series features, and use the lightweight multimodal attention mechanism and the CPO optimization algorithm of the ViT model to optimize the hyperparameters of the ViT model to establish a mapping relationship between the appearance of food and the center maturity, and realize real-time detection of food cooking status.
Real-time monitoring and evaluation of food cooking status is realized, detection accuracy is improved, food maturity is comprehensively evaluated, and the intelligence of the food industry's cooking system is improved.
Smart Images

Figure CN119600592B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and in particular to a method and system for detecting the cooking state of food based on cross-modal deep learning. Background Art
[0002] In recent years, food safety issues have attracted much attention. During the food processing process, in addition to the freshness of the ingredients and the hygiene of the processing environment, the maturity of food during cooking is also a key concern.
[0003] Maturity includes comprehensive indicators such as the external maturity degree of food, the central maturity degree of food, and the nutrient content under different maturity states of food. Generally, the standard definition of maturity is determined by manually judging whether the food has reached the maturity that meets the usage standard.
[0004] Traditional methods mainly use temperature probes to detect the central maturity of food; however, manual detection is time-consuming and laborious, while temperature probes will damage the appearance of food and the cleaning of probe devices also poses a problem.
[0005] In response to the above problems, the existing technology uses machine vision to detect the maturity of food, and the discrimination of the maturity during the food cooking process is handed over to the object detection algorithm to implement the maturity grading and classification task. However, the existing food image recognition only judges the maturity from the external feature information of the image, and cannot comprehensively evaluate the maturity of food. How to establish the mapping relationship between the external maturity and the central maturity of food, so as to comprehensively evaluate the optimal solution of the maturity during the food processing process, which will contribute to the development of a more intelligent cooking system in the food industry. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides a method and system for detecting the cooking state of food based on cross-modal deep learning. The present invention establishes the mapping relationship between the external maturity and the central maturity of food, so as to comprehensively evaluate the maturity of food during the cooking process.
[0007] The technical solution of the present invention is: a method for detecting the cooking state of food based on cross-modal deep learning, including the following steps:
[0008] S1), collect a food cooking data set, preprocess the food cooking data set and then divide it into a training set and a test set;
[0009] S2), construct an improved Yolo model and train it, and use the trained improved Yolo model to extract features from the preprocessed image data set to obtain image features containing rich feature information;
[0010] S3), Use the UMAP algorithm to reduce the dimension of the extracted image features, visualize the data after dimension reduction, and classify the external maturity of the food cooking process according to the visualized data distribution map;
[0011] S4), Construct a TCN time series model, and extract temperature-time series features from the food temperature series dataset through the TCN time series model;
[0012] S5), Fusion the image features and temperature-time series features through a multi-modal fusion layer to obtain multi-modal features;
[0013] S6), Use the fused multi-modal features to train the ViT model, and use the Crested Porcupine Optimization Algorithm CPO improved by an adaptive mechanism to intelligently find the optimal solution for the setting of the hyperparameters of the trained ViT model;
[0014] S7), Construct a mapping relationship between the food surface maturity value and the food center maturity value according to the food appearance image feature information and time series;
[0015] S8), Use the trained ViT model to identify the cooking state in the food image to be detected, and realize the real-time detection of the state change during the food cooking process and feedback the corresponding maturity according to the mapping relationship.
[0016] Preferably, in step S2), the improved Yolo model is: replacing the Feature Pyramid Network FPN of the Neck network of the Yolo model with the Path Aggregation Network PANet; and replacing the ordinary convolution blocks of the Path Aggregation Network PANet with depthwise separable convolution blocks; at the same time, inserting the CBAM attention mechanism module in the upsampling and downsampling of the Path Aggregation Network PANet.
[0017] Preferably, in step S2), the improved Yolo model includes an image input layer, a backbone network Backbone, a neck network Neck, and an output layer; wherein the backbone network Backbone uses CSPDarknet53.
[0018] Preferably, in step S3), use the UMAP algorithm to reduce the dimension of the extracted image feature map, visualize the data after dimension reduction, and classify the external maturity of the food cooking process according to the visualized data distribution map, which specifically includes the following steps:
[0019] S31), Construct a graph in the high-dimensional space, where each point is connected to its nearest neighbor points; and the UMAP algorithm describes the similarity between points through a probability distribution. For each point x i , calculate the k-nearest neighbors, and use the probability distribution p ij to represent the point xi and x j the similarity between, and use the Gaussian kernel function to calculate the probability distribution p ij :
[0020]
[0021] wherein, d(x i , x j ) represents the distance between the point x i and x j , σ i is the bandwidth parameter related to the point x i , and is determined by binary search, such that the sum of the probabilities of the k-nearest neighbors of the point x i is close to a preset value;
[0022] S32) Construct a graph in the low-dimensional space, use probabilities to describe the similarity between points, for each point y i , calculate its k-nearest neighbors, and use the Student's t-distribution to calculate the probability distribution q ij to represent the similarity between the point y i and y j , and the calculation formula of the probability distribution q ij is:
[0023]
[0024] S33) Optimize the low-dimensional embedding through the fuzzy cross-entropy loss function between the minimum high-dimensional graph and the low-dimensional graph, and the fuzzy cross-entropy loss function C is:
[0025]
[0026] wherein, C represents the fuzzy cross-entropy loss function;
[0027] S34) Use Stochastic Gradient Descent (SGD) to optimize the fuzzy cross-entropy loss function C to approximately obtain the projection of the manifold structure of the high-dimensional data on the low-dimensional space. The same kind of food forms different clusters in the two-dimensional image in different cooking state dimensions, and accordingly, the different cooking states of various foods are visually judged, and grading is carried out in combination with the actual ripening state.
[0028] Preferably, in step S4), the TCN time series model is composed of multiple Temporal blocks, and each Temporal block includes causal convolution, dilated convolution, residual connection, and normalization.
[0029] Preferably, in step S4), the causal convolution is expressed as:
[0030]
[0031] In the formula, y t represents the output of causal convolution, w i is the convolution kernel, b is the bias term, and x t-i represents the time series at time point t - i; k is the convolution kernel size;
[0032] The dilated convolution is expressed as:
[0033]
[0034] In the formula, Z t is the output of the dilated convolution, and d is the dilation factor.
[0035] Preferably, in step S5), the fusion of the image features and the temperature - time series features includes feature alignment, feature weight adjustment, and splicing operation.
[0036] Preferably, in step S5), the feature alignment uses dynamic time warping (DTW) to align the image features and the temperature - time series features.
[0037] Preferably, in step S5), the feature weight adjustment is performed by inserting a lightweight multi - modal attention module before the multi - modal fusion layer to adjust the optimal weights of the image features and the time series features and perform weighted fusion.
[0038] Preferably, in step S5), the expression of the weighted fusion is:
[0039] F fused = α(A I + C IT )+(1 - α)(A T + C TI );
[0040] In the formula, F fused represents the feature after weighted fusion, α is a learnable weight parameter, A I represents the self - attention weight matrix of the image features; C IT represents the cross - attention weight matrix of the image features to the time series features; A T represents the self - attention weight matrix of the time series features; C TI represents the cross - attention weight matrix of the time series features to the image features.
[0041] Preferably, the present invention also provides a food cooking state detection system based on cross - modal deep learning, including:
[0042] A data acquisition module for acquiring a food cooking dataset, where the food cooking dataset includes a food status appearance image dataset and a food temperature sequence dataset;
[0043] An image feature extraction module for extracting features from the preprocessed image dataset using an improved Yolo model to obtain image features containing rich feature information;
[0044] A maturity grading module for reducing the dimension of the image features extracted by the image feature extraction module, visualizing the data after dimension reduction, and grading the appearance maturity of the food cooking process based on the visualized data distribution map;
[0045] A temperature-time series feature extraction module for extracting temperature-time series features from the food temperature sequence dataset through a TCN time series model;
[0046] A feature fusion module for fusing image features and temperature-time series features to obtain multi-modal features;
[0047] A cooking state detection module for identifying the cooking state in a food image.
[0048] Preferably, the improved Yolo model is: replacing the Feature Pyramid Network (FPN) of the Neck network of the Yolo model with the Path Aggregation Network (PANet); replacing the ordinary convolution blocks of the Path Aggregation Network (PANet) with depthwise separable convolution blocks; and inserting the CBAM attention mechanism module during the upsampling and downsampling of the Path Aggregation Network (PANet).
[0049] Preferably, the cooking state detection module uses a Vision Transformer (ViT) model to identify the cooking state.
[0050] Preferably, the Vision Transformer (ViT) model is trained using multi-modal features, and the Crown Porcupine Optimization (CPO) algorithm improved by an adaptive mechanism is used to intelligently find the optimal solution for setting the hyperparameters of the trained Vision Transformer (ViT) model.
[0051] Preferably, the cooking state detection module constructs a mapping relationship between the food surface maturity value and the food center maturity value based on the food appearance image information and the temperature-time information, realizes real-time detection of the state change during the food cooking process, and feedbacks the corresponding maturity.
[0052] The beneficial effects of the present invention are:
[0053] 1. The present invention obtains the image and temperature information during the food cooking process, uses the improved YOLO model and TCN model to extract image features and time series features respectively, and uses the UMAP algorithm to perform dimensionality reduction processing on the image features, thereby establishing a food maturity evaluation system;
[0054] 2. The present invention realizes the optimal weight adjustment and weighted fusion of image features and time series features through a lightweight multi-modal attention mechanism module;
[0055] 3. The present invention also introduces a hyperparameter optimization strategy for the ViT model based on the Crest Porcupine Optimization Algorithm (CPO), which improves the model detection accuracy and realizes the real-time monitoring and evaluation of the food cooking state. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a schematic flow chart of the method of the present invention;
[0057] Figure 2 is a framework structure diagram of the system of the present invention;
[0058] Figure 3 is a framework structure diagram of the improved Yolo model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0059] The following further describes the specific embodiments of the present invention with reference to the drawings:
[0060] Embodiment 1
[0061] As Figure 1 shown, this embodiment provides a method for detecting the food cooking state based on cross-modal deep learning, including the following steps:
[0062] S1). Collect a food cooking data set, preprocess the food cooking data set, and divide it into a training set and a test set;
[0063] In this embodiment, the food cooking data set includes a food state appearance image data set and a food temperature sequence data set; in this embodiment, the preprocessing of the food state appearance image data includes, but is not limited to, annotating labels, flipping, tilting, and partially cropping the image to generate a preprocessed image data set, and normalizing the food temperature sequence data set. The preprocessed image data set is divided into a training set and a test set according to a ratio of 7:3.
[0064] S2). Construct an improved Yolo model and train it, and use the trained improved Yolo model to extract features from the preprocessed image data set to obtain image features containing rich feature information;
[0065] In this embodiment, the improved Yolo model is as follows: the feature pyramid network FPN of the Neck network of the Yolo model is replaced by the path aggregation network PANet; the ordinary convolution blocks of the path aggregation network PANet are replaced by depthwise separable convolution blocks; meanwhile, the CBAM attention mechanism module is inserted in the upsampling and downsampling of the path aggregation network PANet.
[0066] In this embodiment, the improved Yolo model extracts the image feature F of the food state appearance image I The expression of which can be represented as:
[0067] F I = Yolo(I);
[0068] In the formula, I represents the food state appearance image.
[0069] In this embodiment, as Figure 3 shown, the improved Yolo model includes an image input layer, a backbone network Backbone, a neck network Neck, and an output layer; wherein the backbone network Backbone adopts CSPDarknet53.
[0070] In this embodiment, since the forward calculation path of the improved Yolo model is relatively long, the image feature information will be diluted once every time it passes through a convolutional layer. However, during the food cooking process, the feature information that needs to be extracted in this embodiment is capable of characterizing different maturity levels of the food appearance. Therefore, in this embodiment, by introducing the path aggregation network PANet into the neck network Neck, in addition to the top-down network structure, an additional bottom-up path aggregation module is added, and the standard convolution blocks are replaced by depthwise separable convolution blocks. Moreover, after upsampling and downsampling, the CBAM attention mechanism is used for the output feature map to optimize the PANet pyramid. Depthwise separable convolution can reduce the computational complexity and the number of parameters of the model, and CBAM is a lightweight mechanism module that can improve the performance while ignoring its overhead.
[0071] S3), Use the UMAP algorithm to reduce the dimension of the extracted image features, visualize the dimension-reduced data, and classify the appearance maturity of the food cooking process according to the visualized data distribution map; specifically, it includes the following steps:
[0072] S31), Construct a graph in the high-dimensional space, where each point is connected to its nearest neighbor points; and the UMAP algorithm describes the similarity between points through probability distribution. For each point x i , calculate the k-nearest neighbors, and use the probability distribution p ij to represent the point x i and x jThe similarity between them, and use the Gaussian kernel function to calculate the probability distribution p ij :
[0073]
[0074] In the formula, d(x i , x j ) represents the distance between points x i and x j , σ i is the bandwidth parameter related to point x i , which is determined by binary search, so that the sum of the probabilities of the k-nearest neighbors of point x i is close to a preset value;
[0075] S32), construct a graph in the low-dimensional space, use probability to describe the similarity between points, for each point y i , calculate its k-nearest neighbors, and use the Student's t-distribution to calculate the probability distribution q ij to represent the similarity between points y i and y j , and the calculation formula of the probability distribution q ij is:
[0076]
[0077] S33), optimize the low-dimensional embedding through the fuzzy cross-entropy loss function between the minimum high-dimensional graph and the low-dimensional graph, and the fuzzy cross-entropy loss function C is:
[0078]
[0079] In the formula, C represents the fuzzy cross-entropy loss function, which is used to measure the difference between the similarity p ij between data points in the high-dimensional space and the similarity q ij between these points in the low-dimensional space;
[0080] S34), use Stochastic Gradient Descent (SGD) to optimize the fuzzy cross-entropy loss function C to approximately obtain the projection of the manifold structure of high-dimensional data on the low-dimensional space. The same kind of food forms different clusters on the two-dimensional image in different cooking state dimensions. Based on this, the different cooking states of various foods can be intuitively judged, and combined with the actual maturity state for grading. In this embodiment, the maturity of the food is divided into three grades: undercooked, cooked, and overcooked.
[0081] S4), construct a TCN time series model, and extract the temperature-time series feature F T from the food temperature sequence dataset through the TCN time series model; the extraction of the temperature-time series feature F T can be expressed as:
[0082] F T = TCN(T);
[0083] T represents the food temperature sequence data set; it is a data set regarding heating time and ingredient temperature;
[0084] The TCN time series model of this embodiment is composed of multiple temporal blocks, and each temporal block includes causal convolution, dilated convolution, residual connection, and normalization.
[0085] Among them, the causal convolution is expressed as:
[0086]
[0087] In the formula, y t represents the output of the causal convolution, w i is the convolution kernel, b is the bias term, x t-i represents the time series at time point t - i; k is the convolution kernel size;
[0088] Correspondingly, the dilated convolution is expressed as:
[0089]
[0090] In the formula, Z t is the output of the dilated convolution, and d is the dilation factor.
[0091] S5), fuse the image features and temperature - time series features through a multi - modal fusion layer to obtain multi - modal features;
[0092] In this embodiment, the fusion of the image features and temperature - time series features includes feature alignment, feature weight adjustment, and splicing operation. Among them, the feature alignment refers to the process of mapping features of different modalities to the same reference framework. Through this process, the relationships and similarities between different features can be better captured; in this embodiment, the dynamic time warping (DTW) is used for feature alignment of the image features and time series features.
[0093] The feature weight adjustment is to insert a lightweight multi - modal attention module before the multi - modal fusion layer to adjust the optimal weights of the image features and temperature - time series features and perform weighted fusion. The expression of the weighted fusion is:
[0094] F fused = α(A I + C IT )+(1 - α)(A T + C TI );
[0095] In the formula, F fused represents the feature after weighted fusion, α is a learnable weight parameter, and A I represents the self-attention weight matrix of the image feature; C IT represents the cross-attention weight matrix of the image feature to the time series feature; A T represents the self-attention weight matrix of the time series feature; C TI represents the cross-attention weight matrix of the time series feature to the image feature, indicating the influence of the time series feature on the image feature.
[0096] S6). Use the fused multi-modal features to train the ViT model, and use the Crowned Porcupine Optimization algorithm CPO improved by the adaptive mechanism to intelligently find the optimal solution for the setting of the hyperparameters of the trained ViT model;
[0097] In this embodiment, the Crowned Porcupine Optimization algorithm CPO realizes the optimization process by simulating four defense behaviors of the crowned porcupine. When facing a predator, the crowned porcupine will adopt different defense strategies, and these strategies can be divided into two categories:
[0098] Global Exploration: Visual intimidation and sound intimidation.
[0099] Local Exploitation: Odor attack and physical attack.
[0100] To improve the convergence speed, an adaptive mechanism is added to the Crowned Porcupine Optimization algorithm CPO in this embodiment. The adaptive mechanism is a method of dynamically adjusting parameters through feedback, which can effectively adjust the different weights of global exploration and local exploitation in the algorithm.
[0101] Use the Crowned Porcupine Optimization algorithm CPO improved by the adaptive mechanism to automatically find the optimal setting of the hyperparameters of the ViT model, so that the effect of the ViT model during training for different food cooking state images can reach the optimal solution, thereby obtaining a better detection model.
[0102] Specifically, it includes the following steps:
[0103] S61) Initialize the algorithm parameter settings: Initialize the population size N, that is, randomly generate N individuals (solutions), and each individual X i is randomly initialized within the search range; the maximum number of iterations T max ; the initial search range [L, U]; the initial values and adjustment strategies of the adaptive parameters. Here, the parameters mainly include the learning rate α, the exploration probability p e and the exploitation probability p d ;
[0104] After setting the initial values of the algorithm-related parameters, calculate the fitness value f(X i ) of each individual in the current population, and find the best individual X best in the current population;
[0105] S63) For each individual X i , the CPO optimization algorithm will perform global exploration and local exploitation respectively with the exploration probability p e and the exploitation probability p d , and keep the probability 1 - p e - p d unchanged;
[0106] Perform global exploration with probability p e : X′ i = X i + α × rand × (U - L);
[0107] Perform local exploitation with probability p d : X′ i = X i + α × rand × (X best - X i );
[0108] S64) After the algorithm has explored each individual to varying degrees, limit the updated individual X i ′ to the search range: X′ i = min[max(X′ i , L), U], and at the same time calculate the fitness value f(X′ i ) of the updated individual X′ i ). If f(X′ i ) is better than f(X i ), then accept the new individual X′ i and update the best individual X best in the current population;
[0109] S65) Dynamically adjust the adaptive parameters according to the current iteration number t and the change of the fitness value;
[0110] Perform adaptive adjustment on the learning rate α:
[0111] S66) When the iteration number reaches the maximum T max or the fitness value does not improve significantly, stop the iteration and output the best individual X best and its corresponding fitness value f(X i ′)
[0112] S7), construct a mapping relationship between the surface maturity value and the center maturity value of the food based on the food appearance image feature information and the time series;
[0113] S8), use the trained ViT model to identify the cooking state in the food image to be detected, and based on the mapping relationship, realize real-time detection of the state change during the food cooking process and feedback the corresponding maturity.
[0114] Embodiment 2
[0115] As Figure 2 shown, provide a food cooking state detection system based on cross-modal deep learning, including:
[0116] A data acquisition module for acquiring a food cooking data set, where the food cooking data set includes a food state appearance image data set and a food temperature sequence data set;
[0117] An image feature extraction module for extracting features from the preprocessed image data set by using an improved Yolo model to obtain image features containing rich feature information;
[0118] A maturity grading module for reducing the dimension of the image features extracted by the image feature extraction module, visualizing the data after dimension reduction, and grading the appearance maturity of the food cooking process based on the visualized data distribution map;
[0119] A temperature-time series feature extraction module for extracting temperature-time series features from the food temperature sequence data set through a TCN time series model;
[0120] A feature fusion module for fusing the image features and the temperature-time series features to obtain multi-modal features;
[0121] A cooking state detection module for identifying the cooking state in the food image.
[0122] Preferably in this embodiment, the improved Yolo model is: using the Path Aggregation Network (PANet) to replace the Feature Pyramid Network (FPN) of the Neck network of the Yolo model; and replacing the ordinary convolution blocks of the Path Aggregation Network (PANet) with depthwise separable convolution blocks; at the same time, inserting the CBAM attention mechanism module in the upsampling and downsampling of the Path Aggregation Network (PANet).
[0123] Preferably in this embodiment, the maturity grading module uses the UMAP algorithm to reduce the dimension of the extracted image features.
[0124] Preferably in this embodiment, the feature fusion of the feature fusion module includes feature alignment, feature weight adjustment, and splicing operation.
[0125] Preferably in this embodiment, the cooking state detection module uses a ViT model to identify the cooking state.
[0126] Preferably in this embodiment, the ViT model is trained with multi-modal features, and the improved Crested Porcupine Optimization Algorithm CPO with an adaptive mechanism is used to intelligently find the optimal solution for setting the hyperparameters of the trained ViT model.
[0127] Preferably in this embodiment, the cooking state detection module constructs a mapping relationship between the surface maturity value and the center maturity value of the food based on the food appearance image information and the temperature-time information, so as to realize the real-time detection of the state change during the cooking process of the food and feedback the corresponding maturity.
[0128] The above embodiments and descriptions in the specification only illustrate the principles and the best embodiments of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A food cooking state detection method based on cross-modal deep learning, characterized in that, It includes the following steps: S1), collect a food cooking dataset, preprocess the food cooking dataset and divide it into a training set and a test set; S2), construct an improved Yolo model and train it, and use the trained improved Yolo model to extract features from the preprocessed image dataset to obtain image features containing rich feature information; S3), use the UMAP algorithm to reduce the dimension of the extracted image features, visualize the data after dimension reduction, and classify the external maturity of the food cooking process according to the visualized data distribution map; S4), construct a TCN time series model, and extract temperature-time series features from the food temperature series dataset through the TCN time series model; S5), fuse the image features and temperature-time series features through a multimodal fusion layer to obtain multimodal features; S6), use the fused multimodal features to train a ViT model, and use the improved Crested Porcupine Optimization Algorithm CPO with an adaptive mechanism to intelligently find the optimal solution for setting the hyperparameters of the trained ViT model; S7), construct a mapping relationship between the food surface maturity value and the food center maturity value according to the food external image feature information and time series; S8), use the trained ViT model to identify the cooking state in the food image to be detected, and realize real-time detection of the state change during the food cooking process and feedback the corresponding maturity according to the mapping relationship.
2. The method for detecting the cooking state of food based on cross-modal deep learning according to claim 1, characterized in that: In step S2), the improved Yolo model is: replace the Feature Pyramid Network FPN of the Neck network of the Yolo model with the Path Aggregation Network PANet; replace the ordinary convolution blocks of the Path Aggregation Network PANet with depthwise separable convolution blocks; at the same time, insert the CBAM attention mechanism module in the upsampling and downsampling of the Path Aggregation Network PANet.
3. The food cooking state detection method based on cross-modal deep learning according to claim 1, wherein: In step S3), use the UMAP algorithm to reduce the dimension of the extracted image feature map, visualize the data after dimension reduction, and classify the external maturity of the food cooking process according to the visualized data distribution map. The specific steps are as follows: S31), construct a graph in a high-dimensional space where each point is connected to its nearest neighbor; and the UMAP algorithm describes the similarity between points through a probability distribution. For each point x i , calculate the k-nearest neighbors and use a probability distribution p ij to represent the similarity between point x i and x j , and use a Gaussian kernel function to calculate the probability distribution p ij : where d(x i , x j ) represents the distance between points x i and x j , σ i is the bandwidth parameter related to point x i , which is determined by binary search such that the sum of the probabilities of the k-nearest neighbors of point x i approaches a preset value; S32), construct a graph in the low-dimensional space, use probabilities to describe the similarity between points, for each point y i , calculate its k-nearest neighbors, and use the Student's t-distribution to calculate the probability distribution q ij to represent the point y i and y j the similarity between them, the probability distribution q ij is calculated by the formula: S33), optimize the low-dimensional embedding through the fuzzy cross-entropy loss function between the minimum high-dimensional graph and the low-dimensional graph. The fuzzy cross-entropy loss function C is: In the formula, C represents the fuzzy cross-entropy loss function; S34), use Stochastic Gradient Descent SGD to optimize the fuzzy cross-entropy loss function C to approximately obtain the projection of the manifold structure of the high-dimensional data in the low-dimensional space. The same kind of food forms different clusters in the two-dimensional image in different cooking state dimensions. Based on this, intuitively judge the different cooking states of various foods and classify them in combination with the actual maturity state.
4. A method for detecting the cooking state of food based on cross-modal deep learning according to claim 1, characterized in that: In step S4), the TCN time series model is composed of multiple Temporal blocks. Each Temporal block includes causal convolution, dilated convolution, residual connection and normalization.
5. A method for detecting the cooking state of food based on cross-modal deep learning according to claim 1, characterized in that: In step S5), the fusion of the image features and temperature-time series features includes feature alignment, feature weight adjustment, and splicing operation.
6. The food cooking state detection method based on cross-modal deep learning according to claim 1, characterized in that: In step S5), the feature alignment uses dynamic time warping (DTW) to align image features and temperature-time series features; The feature weight adjustment is performed by inserting a lightweight multi-modal attention module before the multi-modal fusion layer to adjust the optimal weights of image features and time series features and perform weighted fusion; the expression of the weighted fusion is: F fused = α(A I + C IT ) + (1 - α)(A T + C TI ); where F fused represents the feature after weighted fusion, α is a learnable weight parameter, and A I represents the self-attention weight matrix of the image feature; C IT represents the cross-attention weight matrix of the image feature to the time series feature; A T represents the self-attention weight matrix of the time series feature; C TI represents the cross-attention weight matrix of the time series feature to the image feature.
7. A food cooking state detection system based on cross-modal deep learning, characterized in that The system uses the detection method described in any one of claims 1-6 to detect the food cooking state. The system includes: A data acquisition module for acquiring a food cooking data set, where the food cooking data set includes a food state appearance image data set and a food temperature sequence data set; An image feature extraction module for using an improved Yolo model to extract features from the preprocessed image data set to obtain image features containing rich feature information; A maturity grading module for reducing the dimension of the image features extracted by the image feature extraction module, visualizing the dimension-reduced data, and grading the appearance maturity of the food cooking process based on the visualized data distribution map; A temperature-time series feature extraction module for extracting temperature-time series features from the food temperature sequence data set through a TCN time series model; A feature fusion module for fusing image features and temperature-time series features to obtain multi-modal features; A cooking state detection module for identifying the cooking state in the food image.
8. The food cooking state detection system based on cross-modal deep learning according to claim 7, characterized in that: The improved Yolo model is: replacing the feature pyramid network (FPN) of the Neck of the Yolo model with a path aggregation network (PANet); replacing the ordinary convolution blocks of the path aggregation network (PANet) with depthwise separable convolution blocks; and inserting a CBAM attention mechanism module in the upsampling and downsampling of the path aggregation network (PANet).
9. The food cooking state detection system based on cross-modal deep learning according to claim 7, characterized in that: The cooking state detection module uses a ViT model to identify the cooking state; The ViT model is trained using multi-modal features, and the intelligent search for the optimal solution of the hyperparameter settings for training the ViT model is performed using the crown porcupine optimization algorithm (CPO) improved by an adaptive mechanism.
10. A food cooking state detection system based on cross-modal deep learning according to claim 7, characterized in that: The cooking state detection module constructs a mapping relationship between the food surface maturity value and the food center maturity value based on the food appearance image information and the temperature-time information, realizes the real-time detection of the state change during the food cooking process, and feeds back the corresponding maturity.
Citation Information
Patent Citations
Deep cross-mode correlation learning-based image retrieval method for free-hand sketch
CN108595636A
A human-certificate integrated verification terminal, system and method based on multi-mode face recognition
CN109902780A