Chef machine accurate food material identification and processing method based on computer vision
Through the parallel convolutional neural network architecture and staged training method, the accuracy and adaptability problems of intelligent chef machines in food recognition and processing are solved, and high-precision recognition and adaptive processing of stacked food ingredients are realized, improving the automation level and user experience of chef machines.
Patent Information
- Application Number
- CN202510318884.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
The existing smart chef machines have problems with low recognition accuracy and poor ability to adapt to complex environments in terms of food identification and processing, especially in the case of food stacking, deformation or occlusion, which is difficult to accurately identify and handle.
The parallel convolutional neural network architecture is adopted, combining multi-scale feature extraction, centerline positioning regression, reversible transformation network and attention mechanism, and optimized neural networks through phased training to achieve high-precision recognition and adaptive processing strategies for stacked food ingredients.
It improves the accuracy of food identification and the automation level of the chef machine, enhances the adaptability to complex food scenes, optimizes processing accuracy and system response speed, and reduces user manual intervention.
Smart Images

Figure CN120260033A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to an accurate ingredient recognition and processing method for a cooking machine based on computer vision. Background Art
[0002] With the development of intelligent kitchen appliance technology, more and more households and commercial kitchens are starting to use intelligent cooking machines to improve the automation level of ingredient processing, reduce manual intervention, and enhance cooking efficiency and consistency. In existing intelligent cooking machine products, they mainly rely on users to manually set the processing mode or make processing decisions based on simple physical parameters such as weight sensors and color recognition. However, due to the wide variety and diverse shapes of ingredients, traditional methods still have many technical problems in terms of recognition accuracy, processing of stacked ingredients, understanding of ingredient arrangement, and adaptive processing strategies, which limit the adaptability of intelligent cooking machines in complex environments.
[0003] Currently, existing ingredient recognition technologies mainly include analysis methods based on optical sensors, traditional computer vision methods based on specific color and shape features, and intelligent recognition methods based on deep learning that have emerged in recent years. For example, some intelligent cooking machines integrate spectral analysis technology to irradiate ingredients with light of a specific wavelength and analyze their spectral reflection characteristics to determine the type of ingredient. This method has a high recognition accuracy in specific scenarios, especially suitable for classifying certain specific types of fruits and vegetables, such as using near-infrared spectroscopy to distinguish apples or bananas with different maturities. However, this spectral-based recognition method is often affected by the surface state of the ingredient. For example, if there is moisture or oil residue on the surface, it may lead to misjudgment. In addition, the cost of spectral analysis equipment is relatively high and is not suitable for embedding in consumer-grade cooking machine products.
[0004] Another common ingredient recognition method is a computer vision algorithm based on color and shape features. This method usually uses traditional image processing techniques such as edge detection, color segmentation, and contour analysis to extract features from the ingredient images captured by the camera and match them with predefined templates in the database. This method can achieve relatively accurate ingredient recognition in certain specific scenarios, especially when a single ingredient is neatly arranged. For example, in supermarket checkout systems, many automatic weighing devices can recognize different fruits and vegetables based on color and shape. However, the limitation of this method is that when the ingredient is deformed (such as sliced or shredded), partially occluded, or ingredients are stacked together, the recognition accuracy of the system will drop significantly. In addition, since color information is easily affected by changes in lighting, using color features for ingredient classification often fails to ensure stability in an actual kitchen environment. For example, the same ingredient may exhibit different color distributions under different lighting conditions, resulting in inconsistent recognition results. Summary of the Invention
[0005] The embodiment of the present application provides a method for accurate ingredient recognition and processing of a cooking machine based on computer vision. By adopting a parallel convolutional neural network architecture, combining multi-scale feature extraction, centerline positioning regression, reversible transformation network and attention mechanism, high-precision recognition of stacked ingredients is achieved. By training and optimizing the neural network in stages and combining an adaptive processing strategy, the system can accurately judge the ingredient category, arrangement method and relative stacking relationship, and match the optimal processing method according to the recognition result. The present invention improves the accuracy of ingredient recognition, enhances the adaptability of the cooking machine to complex ingredient scenarios, optimizes the processing accuracy and system response speed, reduces user manual intervention, and significantly improves the automation level and user experience of the intelligent cooking machine.
[0006] The embodiment of the present application provides a method for accurate ingredient recognition and processing of a cooking machine based on computer vision, and the method includes:
[0007] Step 1: The cooking machine acquires an image of the ingredients to be processed that are stacked.
[0008] Step 2: The image of the ingredients to be processed is subjected to image recognition through a convolutional neural network model to obtain an image recognition result; the convolutional neural network model includes two convolutional neural networks, namely a first convolutional neural network and a second convolutional neural network; both the first convolutional neural network and the second convolutional neural network are convolutional neural networks with a parallel sub fully connected-convolution replacement structure, and are constructed through the following process: the fully connected layer is divided into two independent sub fully connected layers, and the two sub fully connected layers are combined in parallel and replaced with a single convolutional layer;
[0009] Step 3: According to the image recognition result, the ingredients to be processed are processed according to a preset processing method corresponding to the image recognition result.
[0010] Further, in step 2, the process of performing image recognition on the to-be-processed food ingredient image through a convolutional neural network model to obtain the image recognition result specifically includes: performing at least one of multi-scale pyramid decomposition, random morphological transformation, and phase consistency feature extraction on the to-be-processed food ingredient image to generate preprocessed images with different resolutions or different feature versions; inputting the preprocessed images into the convolutional neural network model, and using the multi-path parallel convolutional layer composed of the first convolutional neural network and the second convolutional neural network in the convolutional neural network model to respectively perform feature extraction on the preprocessed images with different feature versions to obtain multiple feature maps; using an attention mechanism to perform weighted fusion on the multiple feature maps to obtain a fused feature map with enhanced representation; performing position regression on the fused feature map to infer the vertical centerline positions of multiple recognition objects in the to-be-processed food ingredient image corresponding to the stacked to-be-processed food ingredients; outputting the inferred centerline positions as the vertical direction position information of the recognition objects; performing image template matching recognition according to the centerline positions of each recognition object to obtain the recognition result corresponding to each recognition object; and using all the recognition results as elements in a set to form the image recognition result.
[0011] Further, step 3 specifically includes: obtaining the food ingredient combination method of the stacked to-be-processed food ingredients according to each recognition result in the image recognition result, and using the preset processing method corresponding to the food ingredient combination method to process the to-be-processed food ingredients.
[0012] Further, the method further includes: using training data including multiple training images and corresponding ground truth images to perform staged training on the first convolutional neural network or the second convolutional neural network of the parallel sub fully connected-convolutional replacement structure, specifically including:
[0013] First-stage training: Only train the first convolutional neural network or the second convolutional neural network on the training images processed by random morphological transformation or phase consistency feature to enhance the robustness to different noise patterns or local structures.
[0014] Second-stage training: Using another batch of training images and corresponding ground truth images containing the true labeled centerline positions to further jointly fine-tune the feature extraction layer and the output layer of the first convolutional neural network or the second convolutional neural network to improve the accuracy of centerline position prediction.
[0015] Furthermore, the process of staged training further includes: using a loss function based on cross-entropy to measure the classification or discrete position annotation accuracy of the network for centerline localization; simultaneously using a regression error function based on Euclidean distance or Huber loss to measure the numerical difference between the predicted centerline position and the ground truth annotation; weighting and combining the regression error function and the loss function, and performing backpropagation using the combined total loss function to update the weights of each convolutional layer and sub fully connected layer in the first convolutional neural network or the second convolutional neural network; stopping the update when the combined total loss converges to a preset threshold or reaches the minimum value; and determining the weights of the first convolutional neural network or the second convolutional neural network at this time as the final weights for inference.
[0016] Furthermore, during the staged training process, randomly close some convolutional channels in the first convolutional neural network and / or the second convolutional neural network, and resample the truncated channels using noise distribution or adjacent channel information; after completing inference based on the resampling, calculate the confidence score of the recognition result, and when the confidence is lower than the threshold, automatically increase the resolution of the preprocessed image or perform additional morphological operations and then input it into the first convolutional neural network and / or the second convolutional neural network again to improve the stability and accuracy of the recognition result.
[0017] Furthermore, in step 2, a reversible transformation network is used to perform feature decomposition and reconstruction on the fused feature map to generate representations in different feature spaces; based on the reconstruction results in the different feature spaces and the preprocessed image, multi-dimensional matching verification is performed, and automatic correction is implemented for inconsistent local regions, thereby further improving the centerline localization stability and accuracy of each recognition object.
[0018] Furthermore, when the image of the ingredient to be processed is recognized through a convolutional neural network model, it first passes through the first convolutional neural network and then through the second convolutional neural network.
[0019] Furthermore, all convolutional layers of the first convolutional neural network and all convolutional layers of the second convolutional neural network are in a parallel relationship.
[0020] The computer vision-based precise ingredient recognition and processing method provided by this application has the following beneficial effects: The present invention significantly improves the system's ability to recognize stacked ingredients. Traditional object detection methods often rely on bounding box detection or instance segmentation techniques when facing the problem of ingredient stacking, but these methods are prone to misjudgment when dealing with highly overlapping objects. The present invention proposes a method for analyzing the stacking relationship based on centerline prediction, using deep learning regression to calculate the centerline position of each ingredient, thereby determining its relative height in the entire stacked ingredient structure. This strategy effectively solves the problem that traditional methods are difficult to correctly segment in the case of multi-layer ingredient arrangement, enabling the system to accurately identify which ingredients are on the upper layer and which are on the lower layer, thereby optimizing the processing order and ensuring the correct ingredient processing process. In addition, the present invention uses a reversible transformation network to perform feature decomposition and reconstruction on the fused feature map, enabling the system to independently model the morphological information of ingredients in different feature spaces and perform error correction through multi-dimensional matching verification, which further improves the stability of stacked ingredient recognition. The present invention has achieved a significant improvement in the accuracy of ingredient recognition. Traditional cookers usually rely on fixed sensors, such as color sensors, pressure sensors, or simple shape detection methods, to determine the ingredient category. These methods have poor adaptability to complex ingredient combinations, especially in the case of ingredient stacking, partial occlusion, or morphological changes, and are prone to incorrect recognition. The present invention introduces a dual convolutional neural network with a parallel sub fully connected-convolutional replacement structure, where the first convolutional neural network is mainly responsible for global feature extraction to identify the general outline and morphology of ingredients, while the second convolutional neural network focuses on local feature analysis, such as texture, glossiness, surface structure, etc. This parallel architecture ensures that the system can capture the key features of ingredients at different scales, improving the accuracy and stability of recognition. In addition, through data augmentation means such as multi-scale pyramid decomposition, morphological transformation, and phase consistency feature extraction, the adaptability of the model to different lighting, noise interference, and background complexity is further enhanced, enabling ingredient recognition to maintain a high accuracy rate under various environmental conditions. Description of the Drawings
[0021] The following, in conjunction with the drawings, through a detailed description of the specific embodiments of this application, will make the technical solutions and other beneficial effects of this application obvious.
[0022] Figure 1 It is a schematic flowchart of the method for the computer vision-based precise ingredient recognition and processing method provided by the embodiments of the present invention. Specific Embodiments
[0023] Next, in conjunction with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0024] Embodiment 1: Refer to Figure 1 , a method for accurate ingredient recognition and processing of a cooking machine based on computer vision, the method comprising:
[0025] Step 1: The cooking machine acquires an image of the ingredients to be processed in a stacked state.
[0026] Traditional cooking machines often rely on users to manually select the processing mode or make judgments based on simple physical parameters such as weight sensors and color sensors when processing ingredients. They cannot accurately identify complex ingredient information in a stacked state, resulting in inaccurate processing methods. In contrast, the present invention introduces computer vision technology, uses a camera to collect high-quality ingredient images, and combines a variety of preprocessing strategies to ensure that the input image contains complete ingredient features and has high recognizability, thus laying a foundation for subsequent deep learning-based analysis.
[0027] First, before the cooking machine performs the ingredient recognition task, it needs to use a high-precision camera to perform real-time imaging of the working area to ensure that complete ingredient information can be collected. This imaging process needs to fully consider the working environment of the cooking machine, including issues such as lighting conditions, reflection effects, and the placement of ingredients. Since the colors, textures, and shapes of different ingredients vary greatly, and in actual use, multiple ingredients may overlap, block, or even partially spill out of the field of view, simply relying on ordinary RGB images for ingredient recognition is prone to environmental interference, resulting in the accumulation of recognition errors. Therefore, during the image acquisition process, the cooking machine can adopt a variety of imaging enhancement strategies, such as automatic exposure control, white balance adjustment, high dynamic range (HDR) imaging, etc., to enhance the detail expression of the ingredient image. In addition, to reduce image deviation caused by environmental light changes, the cooking machine can integrate auxiliary light sources around the camera and combine an adaptive brightness adjustment algorithm to ensure that images collected under different lighting conditions are consistent.
[0028] To further improve the distinguishability of ingredients, the present invention can combine depth imaging technology or multispectral imaging to enhance the ability to analyze the structure of complex ingredients. A depth camera can calculate depth information by actively projecting infrared light or using binocular imaging, thereby providing additional geometric information about the three-dimensional shape of the ingredients, enabling the system to not only rely on two-dimensional texture features when identifying ingredients, but also use features such as height and volume to further distinguish different types of ingredients. In addition, multispectral imaging can capture spectral information invisible to the human eye, such as the absorption characteristics in the near-infrared band, thereby effectively distinguishing ingredients with similar appearances but different internal compositions. For example, when distinguishing raw meat from cooked meat, different types of vegetables, or detecting the freshness of ingredients, multispectral imaging can provide more reliable auxiliary information.
[0029] Since the goal of the present invention is to accurately identify stacked ingredients, and traditional single-view imaging often has difficulty obtaining complete spatial information, therefore, during the image acquisition process, a multi-view imaging scheme can be introduced. For example, a cooking machine can use a robotic arm to control the camera to take pictures of the same batch of ingredients from multiple angles, thereby obtaining images from multiple perspectives and performing multi-view fusion through computer vision algorithms to reconstruct a more complete ingredient shape. In addition, if a fixed camera is used, a rotating tray can be combined to enable the ingredients to be captured by the camera in turn at different angles, thereby avoiding occlusion problems caused by limited perspectives. This method can not only improve the ability to distinguish stacked ingredients, but also enhance the system's understanding of the arrangement of complex ingredients, providing more accurate input information for subsequent processing steps.
[0030] After the image acquisition is completed, to ensure the quality of the input data, the system needs to preprocess the collected original images to reduce noise interference and enhance the feature expression of the ingredients. During the preprocessing process, denoising algorithms such as Gaussian filtering and median filtering can be used to eliminate random noise in the images. In addition, considering that the distribution of different ingredients in the color space may vary greatly, the system can convert RGB images to color spaces more suitable for object segmentation and recognition, such as HSV and Lab, to enhance the ability to distinguish specific ingredient categories. At the same time, to improve the extraction effect of edge features, adaptive edge enhancement algorithms such as the Laplacian operator and Sobel operator can be used to ensure that the contour information of the ingredients can be clearly captured. In addition, due to possible distortion in different cameras, especially in the case of wide-angle lenses, obvious geometric deformation may occur at the image edges. Therefore, during the preprocessing stage, it is also necessary to correct the distortion of the images using camera calibration parameters to ensure the authenticity of the ingredient shape.
[0031] Step 2: Perform image recognition on the to-be-processed food ingredient image through a convolutional neural network model to obtain an image recognition result; the convolutional neural network model includes two convolutional neural networks, namely the first convolutional neural network and the second convolutional neural network; both the first convolutional neural network and the second convolutional neural network are convolutional neural networks with a parallel sub fully-connected-convolution replacement structure, and are constructed through the following process: divide the fully-connected layer into two independent sub fully-connected layers, and combine and replace the two sub fully-connected layers in parallel with a single convolutional layer;
[0032] First of all, when processing the input image, a series of feature enhancements are required to ensure that the convolutional neural network can fully extract food ingredient information at different scales and levels. Since stacked food ingredients may present complex boundary and texture relationships, it is difficult for a single-resolution image to effectively capture all detailed information. Therefore, the present invention adopts a multi-scale pyramid decomposition method to decompose the input image into different-resolution versions in order to extract different-level features in the convolutional neural network respectively. For example, a high-resolution image can provide more refined boundary information, while a low-resolution image can capture the overall shape contour. In addition, in order to enhance the perception ability of local features, morphological transformation can be used to preprocess the image. For example, dilation and erosion operations can be used to strengthen boundary information, or noise interference can be removed through opening and closing operations. The phase consistency feature extraction method is also an important part of the present invention. This method extracts the structural information of food ingredients based on local frequency domain features and is not affected by changes in illumination intensity, thereby enhancing the expression ability of the internal texture of food ingredients and improving the recognition effect of weak boundary targets.
[0033] After the input data is preprocessed, the system will input it into the parallel convolutional neural network model, which consists of the first convolutional neural network and the second convolutional neural network, and adopts a sub fully-connected-convolution replacement structure to reduce redundant calculations and improve the feature expression ability. Traditional convolutional neural networks usually rely on fully-connected layers for classification decisions, but fully-connected layers often have problems such as parameter redundancy, large computational volume, and easy overfitting. The present invention splits the fully-connected layer into two sub fully-connected layers, combines them in parallel and replaces them with a convolutional layer, so that the model can reduce the computational overhead and improve the generalization ability while maintaining a high feature expression ability. In the specific calculation process, the first convolutional neural network and the second convolutional neural network process different versions of the input image respectively. The former mainly focuses on global shape features, while the latter focuses on extracting local texture information. Through this dual-path parallel structure, the model can take into account both the overall contour and detailed features of food ingredients, improving the adaptability to complex stacking scenarios.
[0034] To further optimize the feature extraction results, the present invention adopts an attention mechanism to perform weighted fusion on feature maps of different channels, so as to enhance the network's attention to the features of key ingredients. In a traditional convolutional neural network, the feature contribution of each channel is fixed, while the attention mechanism allows the model to dynamically adjust the weights of different channels, enabling the feature channels that contribute more to the recognition task to receive higher attention, while irrelevant or redundant information is suppressed. For example, when the ingredient to be recognized has a high edge contrast, the attention mechanism can automatically enhance the response of the convolutional layer to edge features, and when the ingredient is mainly distinguished by texture information, the system can strengthen the role of the local texture channel. In addition, the attention mechanism of the present invention not only acts at the feature channel level, but also applies to the spatial distribution level, enabling the model to more effectively focus on the core area of the ingredient rather than being interfered by background noise.
[0035] After completing the feature extraction, the system needs to determine the relative positional relationship of the stacked ingredients for subsequent processing decisions. The present invention proposes a method based on centerline positioning to infer the upper and lower hierarchical structure of the ingredients. Specifically, the system calculates the centerline position of the ingredients through regression to determine which ingredients are located in the upper layer and which are in the lower layer of the stack. In traditional object detection tasks, bounding boxes or instance segmentation methods are usually used to determine the position of objects, but these methods are prone to misjudgment when dealing with stacked objects, especially when some ingredients are occluded. The centerline positioning method can directly provide the relative height information of the ingredients, making it more suitable for judging the stacking relationship. In the actual calculation process, the system will use a regression network to predict the centerline position of each recognized object and output it as part of the vertical position information. Subsequently, the system will perform template matching based on the centerline position to further confirm the category of each recognized object and finally obtain the complete ingredient recognition result.
[0036] Step 3: According to the image recognition result, process the ingredient to be processed according to the preset processing method corresponding to the image recognition result.
[0037] When the system obtains the recognition result of the ingredient, it first needs to match the category of each ingredient with the processing parameter library of the chef machine. The processing parameter library is a pre-set knowledge database that contains the optimal processing methods corresponding to different types of ingredients. For example, for meat, the system can decide whether to stir, slice or mince according to the fat content, cutting direction and thickness information, while for vegetables, it may involve operations such as washing, cutting or mixing. Since the same type of ingredient may have different processing requirements, for example, carrots can be sliced, shredded or pureed, the system not only needs to rely on the recognized category information, but also needs to combine the user's recipe settings or default recommended processing methods to ensure that the processing process meets the desired cooking requirements.
[0038] Before actually performing the processing task, the system also needs to perform a spatial adjustment on the identified food ingredients to ensure that the cutting tools, mixing paddles or other processing components of the cooking machine can accurately act on the target food ingredients. Since the food ingredients may be randomly arranged, directly performing processing may result in some parts not being fully processed. Therefore, the system can rearrange the food ingredients through a robotic arm or a vibrating platform. For example, if multiple food ingredients are detected to be stacked on top of each other, the system can first perform a layering operation, giving priority to processing the topmost food ingredient, while the lower-layer food ingredients are processed after the first batch of processing is completed. If the shape of the food ingredient is relatively complex, the system can also adjust the working angles of the cutting tools or mixing paddles inside the cooking machine based on the contour information of the target to adapt to food ingredients of different shapes.
[0039] After determining the processing strategy, the system will start the corresponding cooking machine module to perform the processing operation. For example, if the identified food ingredient needs to be cut, the system will control the cutting depth, speed and direction of the cutting tool according to the thickness and shape parameters of the food ingredient to ensure the uniformity of the slices. For stirring or mixing operations, the system can adjust the rotation speed and time of the mixing paddle based on the texture of the food ingredient to ensure that the stirring effect meets the expectations. In addition, if heat treatment operations such as heating, steaming or others need to be performed, the system can adjust the heating power and duration according to the moisture content and heat conduction characteristics of the food ingredient to avoid overcooking or uneven heating.
[0040] Embodiment 2: In step 2, the process of obtaining the image recognition result by performing image recognition on the image of the food ingredient to be processed through a convolutional neural network model specifically includes: performing at least one of multi-scale pyramid decomposition, random morphological transformation, and phase consistency feature extraction on the image of the food ingredient to be processed to generate preprocessed images with different resolutions or different feature versions; inputting the preprocessed images into the convolutional neural network model, and using the multi-path parallel convolutional layer composed of the first convolutional neural network and the second convolutional neural network in the convolutional neural network model to respectively perform feature extraction on the preprocessed images with different feature versions to obtain multiple feature maps; using an attention mechanism to perform weighted fusion on the multiple feature maps to obtain a fused feature map with enhanced representation; performing position regression on the fused feature map to infer the vertical centerline positions of multiple recognition objects in the image of the food ingredient to be processed corresponding to the stacked food ingredients to be processed; outputting the inferred centerline positions as the vertical direction position information of the recognition objects; performing image template matching recognition according to the centerline positions of each recognition object to obtain the recognition result corresponding to each recognition object; all the recognition results are used as elements in a set to form the image recognition result.
[0041] Specifically, first, before the input image of the ingredient to be processed enters the convolutional neural network model, the system will perform a series of preprocessing operations on the original image to enhance the expression ability of different scales and different features. This preprocessing process mainly includes multi-scale pyramid decomposition, random morphological transformation, and phase consistency feature extraction. The purpose of multi-scale pyramid decomposition is to construct images with different resolutions, enabling the model to capture both the local details and the overall contour information of the ingredient simultaneously. During this process, the input image will be downsampled layer by layer to generate multiple versions at different scales, each retaining different levels of features. For example, high-resolution images can provide more detailed edge information, while low-resolution images can retain global structural information. This strategy ensures that the convolutional neural network can take into account both the global shape and local details during feature extraction. At the same time, random morphological transformation effectively reduces background noise and enhances the boundary contrast between different categories of ingredients by introducing morphological operations such as erosion, dilation, opening, and closing. In addition, the phase consistency feature extraction method is an enhancement strategy based on frequency domain analysis. Its core idea is to calculate the structural information of local edges through phase consistency filtering, regardless of the change in illumination intensity. This method is particularly effective in low-contrast or complex illumination environments, which can enhance the contour information of the ingredient and improve the robustness of recognition.
[0042] After completing the preprocessing, the system will input the images of different feature versions generated into the convolutional neural network model. The convolutional neural network of the present invention adopts a multi-path parallel structure, in which the first convolutional neural network and the second convolutional neural network respectively constitute two independent feature extraction paths, and process different versions of the preprocessed images in parallel. The advantage of this parallel architecture is that it can simultaneously extract feature information of different scales and different modalities, and fuse them in the subsequent stage to improve the recognition accuracy. The first convolutional neural network is mainly responsible for extracting global shape features, such as the contour, main direction, and overall shape of the ingredient, while the second convolutional neural network focuses on local texture features, such as surface details, color changes, and local structures. This division of labor enables the system to simultaneously focus on the overall structure and local differences when facing various types of ingredients, thereby improving the recognition accuracy. In addition, due to the possible deformation, rotation, or partial occlusion of the ingredient, the parallel structure can also enhance the system's adaptability to complex scenarios through information complementarity of different paths.
[0043] After the feature extraction is completed, the system uses the attention mechanism to perform weighted fusion on the feature maps of multiple channels to obtain a fused feature map with enhanced representation. In traditional convolutional neural network models, the feature maps of all channels are usually treated equally, while the attention mechanism allows the system to dynamically adjust the weights of different channels according to the task requirements, so that the features that play a decisive role in the recognition of key ingredients receive higher attention, while irrelevant or redundant features are suppressed. For example, if the recognition of a certain ingredient mainly depends on texture features, the system will automatically increase the weight of the corresponding texture feature channel, while for ingredients that mainly rely on shape features, the global contour information will be enhanced. In addition, the attention mechanism not only acts at the feature channel level, but also at the spatial level, enabling the network to more accurately locate the key areas of the ingredients and improve the stability of recognition.
[0044] Based on the fused feature map, the system further performs position regression calculation to infer the position of the center line in the up-down direction of the stacked ingredients. Since the ingredients may be at different heights and in some cases partially occlude other ingredients, simple object detection methods are difficult to accurately judge the relative stacking relationship of the ingredients. The present invention uses a regression model based on deep learning to perform center line positioning on multiple recognition objects in the fused feature map and predict their relative positions in the vertical direction. The concept of the center line comes from the geometric center point of the object, which can effectively reflect the stacking order of the ingredients and avoid misjudgments that may occur in the case of bounding box detection under occlusion. Specifically, the regression network will output the center line coordinates of each ingredient and use them as the vertical position information of the recognition object to help the system correctly understand which ingredients are on the upper layer and which are on the lower layer. This information is crucial for subsequent processing decisions, because the cooking machine needs to process the upper layer ingredients first to ensure that the processing quality is not affected by the wrong order.
[0045] After obtaining the center line position information, the system further performs image template matching recognition to finally determine the specific category of each ingredient. Template matching is a method that uses existing ingredient templates for comparison. Its core idea is to calculate the similarity between the ingredient to be recognized and the standard ingredient template in the feature space. Since the system of the present invention has extracted rich feature information through the convolutional neural network, it can perform matching in the feature space instead of the traditional pixel space matching. This method has stronger robustness and can effectively reduce the influence of factors such as illumination changes and angle deviations on the recognition results. During the matching process, the system will perform feature alignment based on the center line position to ensure the accuracy of template matching, and at the same time use a multi-scale matching strategy so that even if the size of the ingredient changes, its category can still be correctly recognized.
[0046] Embodiment 3: Step 3 specifically includes: obtaining the food combination mode of the stacked to-be-processed ingredients according to each recognition result in the image recognition result, and processing the to-be-processed ingredients using a preset processing method corresponding to the food combination mode.
[0047] Specifically, first, after the ingredient recognition is completed, the system analyzes each ingredient object in the recognition result and determines the combination mode of the current to-be-processed ingredients according to their spatial distribution relationship. Since the ingredients processed by the cooking machine may not be of a single category, but multiple ingredients are randomly stacked together, simply processing each ingredient independently may affect the final ingredient processing effect. For example, when preparing a salad, lettuce leaves, tomatoes, cucumbers, etc. may be mixed and stacked together, and when making meat products, different parts of the meat may have different fat contents and texture distributions and need to be cut, stirred, or marinated separately. Traditional methods usually rely on users to manually sort and input processing instructions, while the cooking machine of the present invention can automatically complete this classification and processing decision-making process through computer vision technology.
[0048] In the specific implementation process, the system first calculates the spatial combination relationship of the ingredients according to the recognition result of each ingredient, including their relative positions, stacking order, and possible contact areas. To ensure the calculation accuracy, the system can infer the upper and lower layer relationships of the ingredients based on the recognized center line position and use image segmentation technology to further refine the ingredient boundaries to more accurately determine the arrangement mode of the ingredients. For example, if the system detects that some ingredients have a large contact area and regular shapes, it can be inferred that these ingredients may belong to the same group. For example, cut potato pieces or carrot pieces may need to be stir-fried together, and if a large height difference is detected between the ingredients, it may mean that different ingredients need to be processed in batches. For example, when making a composite dish, some ingredients need to be processed first and then subsequent steps are carried out.
[0049] Once the ingredient combination method is determined, the system will query the preset processing method database and select a processing strategy that matches the current combination method. The processing method database is a knowledge base that stores the best processing methods for different categories of ingredients, including different operation parameters such as cutting, stirring, crushing, heating, freezing, etc. For example, when it is recognized that the ingredient combination contains flour, eggs, and milk, the system will automatically match the baking processing strategy, while when it is recognized that the ingredient combination includes beef, onions, and peppers, cutting and marinating operations may be recommended. In addition, the system can also be adjusted in combination with the user's custom settings or cooking preferences to provide a processing plan that better suits personal taste. For example, in different cuisines, the same ingredient combination may have different processing methods. For example, in Asian cuisines, shredded potatoes are usually cut finer and stir-fried quickly, while in Western cuisines, potato chunks are usually used for stewing or baking. Therefore, the system can automatically optimize the processing parameters based on the user's historical usage data to make the processing results more in line with the user's habits.
[0050] In the processing execution stage, the system will precisely control the hardware components such as the cutter, stirrer, heating element, etc. of the cooking machine according to the matched processing method. For example, if it is detected that the ingredients need to be cut, the system will automatically adjust the angle, cutting speed, and pressure of the blade according to the recognized size and shape of the ingredients to ensure the uniformity of cutting; if the ingredients need to be stirred, the system will adjust the rotation speed and time of the stirring paddle to ensure that the stirring effect meets the expectations. For cases involving multi-step processing, such as ingredients that need to be cut first and then heated, the system will set a task queue according to the processing flow and control different processing modules to execute tasks in sequence. In addition, for some special ingredients, such as ingredients with high water content (such as fruits and vegetables), the system can adjust the cutting method according to the recognized moisture characteristics to prevent extrusion damage or affect the taste.
[0051] Example 4: The method further includes: using training data including a plurality of training images and corresponding correct solution images to perform staged training on the first convolutional neural network or the second convolutional neural network of the parallel sub fully connected-convolutional replacement structure, specifically including:
[0052] First stage training: Only train the first convolutional neural network or the second convolutional neural network for the training images processed by random morphological transformation or phase consistency feature to enhance the robustness to different noise forms or local structures;
[0053] Second stage training: Use another batch of training images and corresponding correct solution images containing the true marked centerline positions to further jointly fine-tune the feature extraction layer and the output layer of the first convolutional neural network or the second convolutional neural network to improve the accuracy of centerline position prediction.
[0054] Specifically, in the first-stage training, the training data mainly consists of images processed by random morphological transformations or phase consistency features, aiming to enhance the model's robustness to different noise patterns or local structural variations. In actual application scenarios, the ingredients to be processed may have blurred local boundaries, abnormal color distributions, or distorted texture features due to external environmental interferences (such as uneven illumination, surface contamination, reflection interference, etc.). If the model is directly trained based on the original images, it is prone to overfitting to these interference factors, thus reducing the generalization ability. Therefore, in the first-stage training, the system performs various morphological enhancement operations on the training data, such as using erosion and dilation to enhance the boundary contrast, using opening and closing operations to eliminate small noise regions, or adopting the phase consistency feature extraction method to enhance the structural information, enabling the convolutional neural network to learn more stable edge and shape features at the feature level. The training in this stage mainly targets the low-level feature extraction modules of the model, namely the convolutional layer and the local pooling layer, so that the network can still extract stable features when facing different environmental conditions, thereby improving the robustness of the model.
[0055] In the second-stage training, the system uses a training image dataset containing the true annotated centerline positions to jointly fine-tune the feature extraction layer and the output layer of the first convolutional neural network or the second convolutional neural network to improve the accuracy of centerline position prediction. Due to the complex stacking relationship of the ingredients, traditional object detection methods (such as bounding box localization) are prone to misjudgment when dealing with occluded objects, while the centerline prediction method can more effectively describe the vertical position relationship of the ingredients, thus improving the accuracy of ingredient stacking analysis. In this training stage, the system will use a large number of manually annotated training data, where the centerline positions of each ingredient have been accurately annotated to ensure that the neural network can form correct spatial perception ability during the learning process. Specifically, the training in this stage not only optimizes the feature extraction ability of the convolutional layer but also adjusts the parameters of the fully connected layer and the output layer to minimize the error of centerline prediction as much as possible. During the training process, the system adopts the supervised learning method and optimizes by calculating the Euclidean distance or Huber loss between the predicted centerline and the annotated centerline, so that the final model can more accurately predict the stacking order and spatial distribution of the ingredients during inference.
[0056] The phased training method adopted by the present invention has significant advantages compared to the traditional end-to-end training method. In the training process of traditional deep learning models, all parameters are usually optimized in the same phase. Although this method can converge to a relatively optimal solution when the data volume is large enough, for complex tasks (such as ingredient stacking recognition), single-phase training is often affected by the selection of initial parameters, data noise, and gradient update strategies, resulting in the model being difficult to converge to the global optimal solution. The present invention first performs robustness optimization (the first phase) and then precision optimization (the second phase), enabling the neural network to converge more stably and maintain a high recognition accuracy under different environmental conditions.
[0057] In addition, the phased training strategy of the present invention can also improve the transferability of the model. In different application scenarios, the types, arrangement methods, and environmental lighting conditions of ingredients may change. Therefore, the generalization ability of the model is crucial. By adding morphological enhancement and phase consistency feature extraction in the first-phase training, the model can learn stable ingredient features under diverse input conditions, thereby reducing the overfitting phenomenon. In the second-phase training, the model can be finely tuned according to the requirements of specific tasks to adapt to specific recognition scenarios. For example, if applied to different kitchen appliances or different lighting environments, the model parameters obtained from the first-phase training can remain unchanged, and only the parameters of the second phase are fine-tuned to quickly adapt to the new scenario without having to retrain the entire model completely, thus significantly reducing the computational cost and data requirements.
[0058] Example 5: The process of the phased training further includes: using a loss function based on cross-entropy to measure the classification or discrete position annotation accuracy of the network for centerline positioning; simultaneously using a regression-type error function based on Euclidean distance or Huber loss to measure the numerical difference between the predicted centerline position and the labeled ground truth; weighting and combining the regression-type error function and the loss function, and using the combined total loss function for backpropagation to update the weights of each convolutional layer and sub fully connected layer in the first convolutional neural network or the second convolutional neural network; when the combined total loss converges to a preset threshold or reaches the minimum value, stop the update; determine the weights of the first convolutional neural network or the second convolutional neural network at this time as the final weights for inference.
[0059] Specifically, first, the system uses a cross-entropy-based loss function to measure the classification accuracy of the network for centerline localization. This means that during training, the system regards the centerline position of the food ingredients as an object to be classified, and by adjusting the model parameters, makes the predicted centerline position as close as possible to the truly annotated centerline. In this way, the system can more effectively learn the features of the food ingredient stacking structure and improve the recognition accuracy of the model under different food ingredient arrangements. The core idea of the cross-entropy loss is to maximize the prediction probability of the correct class, enabling the model to continuously optimize its own parameters during the training process to improve the classification accuracy.
[0060] Meanwhile, the system also uses a regression error function based on the Euclidean distance or Huber loss to measure the numerical difference between the predicted centerline position and the true annotation. The role of this regression error is to ensure that when the model predicts the centerline position of the food ingredients, it is not just a discrete classification task, but can perform fine-grained numerical optimization on the specific position of the centerline. During the optimization process of the regression error, the system can adjust the feature extraction ability of the model, enabling the convolutional neural network to not only identify the categories of the food ingredients but also accurately judge the stacking relationship of the food ingredients in space, thereby improving the accuracy of the processing decision.
[0061] To make the training more efficient, the present invention adopts a weighted combination strategy to combine the classification loss and the regression loss to form a comprehensive total loss function. This total loss function can optimize both the classification ability and the regression ability of the model, enabling the model to correctly judge the food ingredient categories and accurately predict the spatial positions of the food ingredients when identifying the food ingredients. During the training process of the system, the backpropagation algorithm is used to continuously adjust the weights of each convolutional layer and sub fully connected layer in the first convolutional neural network and the second convolutional neural network according to the combined loss value, making the model gradually converge to the optimal state.
[0062] When the combined total loss converges to a preset threshold or reaches the minimum value, the system will stop updating the model and determine the weights of the first convolutional neural network or the second convolutional neural network at this time as the final weights for inference. This process ensures that after the model is trained, it can perform inference with the optimal parameter configuration, enabling the cooking machine to accurately identify the stacking structure of the food ingredients to be processed during actual operation and providing a reliable basis for subsequent processing operations.
[0063] Compared with traditional neural network training methods, the loss optimization strategy of the present invention has significant advantages in multiple aspects. First, the combination of classification loss and regression error enables the model to not only identify the categories of ingredients but also accurately predict spatial positions, thereby improving the overall recognition accuracy. Second, the optimization method of the weighted combined loss function makes the training process more stable and can effectively avoid the local optimum problem caused by a single loss term. In addition, by setting a convergence threshold to automatically stop training, the present invention can ensure that the model will not overfit after reaching the optimal state, enabling the final ingredient recognition system to have stronger generalization ability and adapt to environments with different lighting conditions, different ingredient arrangements, and complex background interferences.
[0064] Example 6: During the staged training process, randomly close some convolutional channels in the first convolutional neural network and / or the second convolutional neural network, and resample the truncated channels using noise distribution or adjacent channel information; after completing the inference based on the resampling, calculate the confidence score of the recognition result. When the confidence is lower than the threshold, automatically increase the resolution of the preprocessed image or perform additional morphological operations and then input it into the first convolutional neural network and / or the second convolutional neural network again to improve the stability and accuracy of the recognition result.
[0065] Specifically, during the staged training process, the system will randomly close some convolutional channels in the first convolutional neural network or the second convolutional neural network. The core idea of this method is to simulate the feature loss situation in the real environment. For example, due to lighting changes, ingredient occlusion, or camera perspective problems, some key information may be missing in the image. By randomly closing some convolutional channels during the training process, the system can effectively improve the adaptability of the model in the case of feature loss, thereby reducing the overfitting phenomenon that depends on specific feature channels. In addition, to avoid excessive information loss caused by closing the channels, the system will resample using noise distribution or adjacent channel information to compensate for the features of the truncated channels to a certain extent. The resampling strategy of noise distribution can introduce Gaussian noise or Poisson noise, enabling the neural network to maintain good recognition ability when facing different ingredient features, while the resampling of adjacent channel information is based on spatial relationships, interpolating adjacent feature maps to fill the information gap caused by the closed channels.
[0066] After completing inference based on resampling, the system evaluates the recognition result and calculates the confidence score of the recognition result. The confidence score is an indicator that measures the credibility of the model for the current recognition result and is usually calculated based on the output probability distribution of the neural network. If the confidence is high, it means that the result of the model's recognition of the current ingredient is relatively reliable and can be directly used as the final result. However, in some cases, due to the complex arrangement of ingredients, large feature interference, or insufficient image quality, the confidence may be low, which means that there may be a high degree of uncertainty in the recognition result of the model on the current input data.
[0067] When the system detects that the confidence of the recognition result is lower than the preset threshold, it will automatically adjust the input data to improve the recognition accuracy of the model. First, the system increases the resolution of the preprocessed image so that the input image contains more detailed information. This strategy is particularly effective when dealing with small ingredients or ingredients with complex textures. For example, when recognizing mixed grains or chopped vegetables, increasing the resolution can help the neural network extract more key features. In addition, the system also performs additional morphological operations, such as dilation, erosion, or adaptive edge enhancement, to optimize the contour information of the ingredients and improve the feature extraction ability of the recognition model. After this adjustment, the system re-enters the optimized image into the first convolutional neural network or the second convolutional neural network for a second round of recognition, thereby obtaining a more stable recognition result.
[0068] This recognition optimization strategy based on adaptive adjustment has stronger adaptability and higher stability compared to the traditional static model inference method. In the traditional neural network recognition method, the model usually adopts the same processing method for all input images. Even when encountering low-confidence samples, it will not make targeted adjustments, resulting in an increase in the recognition error rate. In contrast, the present invention dynamically adjusts the input data during the inference stage, enabling the system to perform adaptive optimization according to different ingredient arrangements, thereby significantly improving the stability of recognition.
[0069] In addition, this strategy of the present invention can also effectively reduce the waste of computing resources. Traditional high-precision image recognition methods usually default to using high-resolution images for training and inference. However, high-resolution images not only increase the computational overhead but also may cause unnecessary noise interference. The system of the present invention adopts an intelligent adjustment mechanism, using standard resolution for recognition in most cases, and only increasing the resolution or performing additional morphological operations when the confidence is low, thereby improving the computational efficiency while ensuring the recognition accuracy.
[0070] Example 7: In step 2, the fused feature map is subjected to feature decomposition and reconstruction using a reversible transformation network to generate representations in different feature spaces; based on the reconstruction results in the different feature spaces and the preprocessed image, multi-dimensional matching verification is performed, and automatic correction is implemented for inconsistent local regions, thereby further improving the positioning stability and accuracy of the centerlines of each recognition object.
[0071] Specifically, first, after completing the initial feature extraction and fusion, the system inputs the fused feature map into a reversible transformation network to further decompose and reconstruct the features. The core idea of the reversible transformation network is to map high-dimensional features into different sub-feature spaces through a set of reversible mappings, and reconstruct multiple representations with different feature distributions while maintaining the integrity of the original information. Such a transformation can effectively reduce feature redundancy and retain the most critical information for ingredient recognition in different feature spaces. For example, for a stack of ingredients, some feature channels may be better at capturing shape information, while others may be more suitable for extracting texture or color information. Through the reversible transformation network, the system can enhance this information separately in different feature spaces and further optimize the recognition results in the subsequent verification process.
[0072] After completing the feature decomposition and reconstruction, the system uses the representations in these different feature spaces to perform multi-dimensional matching verification with the preprocessed image to detect and correct possible recognition errors. In traditional deep learning methods, the prediction results of neural networks are usually not actively subjected to secondary verification, so once the model produces incorrect recognition results, the system often cannot effectively correct them. In the solution of the present invention, through the method of multi-dimensional matching verification, the recognition process can have the ability of self-correction, improving the adaptability to complex ingredient arrangement situations.
[0073] Specifically, the system calculates the matching degree between the reconstruction results of different feature spaces and the original preprocessed image, and detects the inconsistent regions. For example, if the representation of a certain region in a certain feature space has a large deviation from the edge information in the original image, it indicates that there may be an error in the recognition of this region. To further improve the recognition accuracy, the system automatically corrects these inconsistent regions, adjusts the feature weights of this region, or re-performs feature extraction when necessary to reduce error accumulation. This method is particularly suitable for dealing with severely occluded ingredient stacking situations, because in traditional methods, the features of the occluded regions are often misrecognized as the background, while the combination of the reversible transformation network and multi-dimensional matching verification in the present invention enables the system to automatically discover and correct these problems, thereby improving the overall recognition accuracy.
[0074] In addition, the feature decomposition method based on the reversible transformation network can also improve the stability of centerline positioning. In the scenario of stacked food ingredients, the positioning accuracy of the centerline is crucial for the decision-making of the processing sequence. Especially in the case of layer-by-layer processing, incorrect centerline prediction may lead to confusion in processing steps. The method of the present invention performs feature reconstruction in different feature spaces, making the prediction of the centerline not rely on a single feature channel, but optimized by combining the results of multi-dimensional matching verification, thereby significantly reducing the recognition deviation caused by local feature loss or abnormal data. This mechanism not only enhances the robustness of the model, but also enables the system to maintain a high centerline prediction accuracy when facing complex situations such as light changes and food ingredient shape changes.
[0075] Compared with traditional food ingredient recognition methods, Embodiment 7 of the present invention provides a more intelligent and stable feature optimization scheme. Traditional convolutional neural networks often rely on single-path feature extraction when facing complex scenarios and are unable to dynamically adjust their own feature expression methods. The present invention enables the system to perform information reconstruction in multiple feature spaces through a reversible transformation network and combines multi-dimensional matching verification to achieve automatic error correction. This not only improves the accuracy of food ingredient recognition, but also significantly enhances the adaptive ability of the system, enabling it to better cope with complex food ingredient arrangements and processing requirements.
[0076] Embodiment 8: When the image of the food ingredient to be processed is recognized through a convolutional neural network model, it first passes through the first convolutional neural network and then through the second convolutional neural network.
[0077] Specifically, first, in the first stage of image recognition, the first convolutional neural network is mainly responsible for the extraction of low-level features, such as basic information such as edges, textures, colors, and shapes. The goal of this stage is to preprocess the original image and extract key features that can characterize the basic attributes of the food ingredient. In traditional computer vision tasks, the extraction of low-level features is often affected by noise, light changes, and background interference. The first convolutional neural network of the present invention is specially trained to enable it to maintain a relatively stable feature extraction ability in complex environments. For example, in the case of uneven lighting conditions, the first convolutional neural network can enhance edge information through techniques such as local contrast normalization, thereby reducing the interference of the external environment on food ingredient recognition. In addition, the network can also use pooling layers to reduce the data dimension while maintaining the overall structural features of the food ingredient for subsequent in-depth analysis.
[0078] After being processed by the first convolutional neural network, the preliminary features extracted will be input into the second convolutional neural network for further analysis. The main task of the second convolutional neural network is to perform high-level feature extraction, that is, to construct more complex semantic information based on the low-level features extracted in the first stage. For example, based on the edge information extracted in the first stage, the second convolutional neural network can identify the complete outline of the ingredients, determine the category of the ingredients, and analyze their positions in the stacked structure. In addition, the second convolutional neural network can also combine the information of multiple feature maps to achieve the fusion of local details and global features of the ingredients, which is crucial for dealing with the complex relationships of stacked ingredients. For example, in the case of multiple overlapping ingredients, the first convolutional neural network may only be able to identify some visible edges, while the second convolutional neural network can infer the shape of the occluded part through context information, thereby improving the overall recognition accuracy.
[0079] An important advantage of adopting a cascaded two-stage convolutional neural network architecture is that it can hierarchically process different types of features, enabling the model to focus on different information levels at different stages. This approach not only improves the generalization ability of the model but also reduces computational redundancy and enhances the recognition efficiency. In the traditional single convolutional neural network architecture, all feature extraction tasks are usually completed in a unified network structure, which often makes it difficult for the model to distinguish the different roles of low-level and high-level features, thus affecting the recognition accuracy. The two-stage processing method adopted in the present invention enables the first convolutional neural network to focus on the stable extraction of low-level features, while the second convolutional neural network focuses on the semantic analysis of high-level features, thereby optimizing the accuracy and reliability of the entire recognition process.
[0080] In addition, this cascaded structure can also adapt to ingredient recognition tasks of different complexities. In simple ingredient classification tasks, the basic features extracted by the first convolutional neural network may be sufficient to complete the recognition. For more complex scenarios, such as dealing with multiple stacked ingredients or cases with partial occlusion, the second convolutional neural network can further optimize the recognition results to ensure the recognition accuracy. For example, when identifying a group of overlapping ingredients, the first convolutional neural network can extract the boundary information of each ingredient, while the second convolutional neural network can infer through deep learning which parts belong to the same ingredient and use global information to optimize the final classification result.
[0081] Compared with traditional single-stage neural networks, the two-stage convolutional neural network method of the present invention improves the adaptability to complex scenarios, enabling the intelligent cooking machine to maintain a high recognition accuracy when facing different types of ingredient arrangements. This method is particularly suitable for tasks that require precise classification and localization. For example, in the process of automatic cooking, different ingredients require different processing methods, so ensuring the accuracy of ingredient recognition is crucial for subsequent processing decisions. By optimizing the neural network structure, the present invention enables the ingredient recognition system to better meet the actual application requirements and improves the intelligence level and user experience of the cooking machine.
[0082] Embodiment 9: All convolutional layers of the first convolutional neural network and all convolutional layers of the second convolutional neural network are in a parallel relationship.
[0083] Specifically, in traditional deep learning image processing methods, convolutional neural networks usually use a serial manner for feature extraction, that is, the output of one network is used as the input of the next network. This method can gradually extract higher-level features, but it may also cause information to be gradually lost during multiple transformations. Especially in tasks that require simultaneous attention to global morphology and local details, the serial structure may lead to the attenuation or neglect of some key information. The parallel convolutional neural network architecture of the present invention can solve this problem, enabling the first convolutional neural network and the second convolutional neural network to independently learn in different feature dimensions and finally fuse the different features extracted, thereby improving the robustness of ingredient recognition.
[0084] In the parallel network structure, the first convolutional neural network focuses on global feature extraction, such as the overall shape, contour structure, main color distribution, and spatial arrangement information of the ingredients. Through a larger receptive field, this network can capture the relative position relationship between the ingredients and establish discriminative ability in the overall structure. This plays an important role in dealing with complex scenarios, such as when multiple ingredients are stacked together or some ingredients are occluded. For example, in the case of stacked ingredients, it may be difficult to distinguish which boundaries belong to the same ingredient relying only on local information, while global features can provide more complete context information, thereby improving the accuracy of segmentation and classification.
[0085] At the same time, the second convolutional neural network focuses on the extraction of local detail features, such as the texture on the surface of the ingredients, tiny color changes, and the complexity of the surface structure. Through a smaller receptive field and deeper feature layers, this network can distinguish different ingredients with similar shapes. For example, when identifying carrots and sweet potatoes, they may have a similar overall shape, but their surface textures, smoothness, and subtle color differences may be different. The second convolutional neural network can effectively capture these differences through its more refined feature extraction ability and improve the accuracy of classification.
[0086] Since these two neural networks run in parallel, they can learn different levels of features separately without interfering with each other, and at the same time can avoid information loss during the hierarchical transmission. Finally, the features extracted by these two networks will be integrated in the fusion layer to form a complete feature representation. In the fusion stage, the system can use the attention mechanism to assign different weights to different features to ensure more accurate final recognition results. For example, when the shape information of the food ingredient is more prominent, the fusion layer can rely more on the global features of the first convolutional neural network, and when the detailed information of the food ingredient is more important, it can rely more on the local features of the second convolutional neural network. Through this adaptive feature fusion strategy, the system can maintain a high recognition accuracy in different environments and food ingredient arrangements.
[0087] Another advantage of the parallel network architecture compared to the serial network is the improvement in computational efficiency. Since the two convolutional neural networks perform parallel computations, the overall inference time can be optimized without computational delays caused by deep stacking. This is particularly important for real-time food ingredient recognition tasks. For example, during the automatic processing of a cooking machine, it is necessary to complete food ingredient classification as quickly as possible to perform operations such as cutting, stirring, or heating in a timely manner. The parallel structure of the present invention can effectively reduce the inference latency, enabling the cooking machine to maintain a high response speed when processing complex food ingredient combinations.
[0088] In addition, the design of the parallel network also enhances the scalability of the system. In practical applications, users may encounter different types of food ingredients, and the features of these food ingredients may vary. Through the parallel structure, the system can easily expand new convolutional neural networks to adapt to more diverse food ingredient recognition tasks. For example, dedicated convolutional neural networks can be added for specific categories of food ingredients. For different categories of food ingredients such as meat, vegetables, and fruits, a set of independent feature extraction models can be trained respectively and unifiedly processed in the final fusion layer. This scalability enables the system to be continuously optimized as new food ingredients are added without having to retrain the entire network.
[0089] The embodiments of the present invention have been introduced in detail above. Specific examples are used in this article to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, based on the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for accurate ingredient recognition and processing of a cooking machine based on computer vision, characterized in that, The method includes: Step 1: The cooking machine acquires an image of the ingredients to be processed stacked; Step 2: The image of the ingredients to be processed is subjected to image recognition through a convolutional neural network model to obtain an image recognition result; the convolutional neural network model includes two convolutional neural networks, namely the first convolutional neural network and the second convolutional neural network; both the first convolutional neural network and the second convolutional neural network are convolutional neural networks with a parallel sub fully connected-convolution replacement structure, and are constructed through the following process: the fully connected layer is divided into two independent sub fully connected layers, and the two sub fully connected layers are combined in parallel and replaced with a single convolutional layer; Step 3: According to the image recognition result, the ingredients to be processed are processed according to a preset processing method corresponding to the image recognition result.
2. The method for accurate ingredient recognition and processing of a chef machine based on computer vision according to claim 1, wherein, In Step 2, the process of subjecting the image of the ingredients to be processed to image recognition through a convolutional neural network model to obtain an image recognition result specifically includes: performing at least one of multi-scale pyramid decomposition, random morphological transformation, and phase consistency feature extraction on the image of the ingredients to be processed to generate preprocessed images with different resolutions or different feature versions; inputting the preprocessed images into the convolutional neural network model, and using the multi-path parallel convolutional layer composed of the first convolutional neural network and the second convolutional neural network in the convolutional neural network model to respectively perform feature extraction on the preprocessed images with different feature versions to obtain multiple feature maps; using an attention mechanism to perform weighted fusion on the multiple feature maps to obtain a fused feature map with enhanced representation; performing position regression on the fused feature map to infer the vertical centerline positions of multiple recognition objects in the image of the ingredients to be processed stacked corresponding to the image of the ingredients to be processed; outputting the inferred centerline positions as the vertical direction position information of the recognition objects; performing image template matching recognition according to the centerline positions of each recognition object to obtain the recognition result corresponding to each recognition object; all the recognition results are used as elements in a set to form the image recognition result.
3. The method for accurate ingredient recognition and processing of a cooking machine based on computer vision according to claim 2, characterized in that, Step 3 specifically includes: according to each recognition result in the image recognition result, obtaining the ingredient combination method of the ingredients to be processed stacked, and according to the ingredient combination method, using a preset processing method corresponding to the ingredient combination method to process the ingredients to be processed.
4. The method for precise ingredient recognition and processing of a chef machine based on computer vision according to claim 3, wherein The method further includes: using training data including a plurality of training images and corresponding correct images to perform staged training on the first convolutional neural network or the second convolutional neural network with a parallel sub fully connected-convolution replacement structure, specifically including: First-stage training: Only train the first convolutional neural network or the second convolutional neural network for the training images processed by random morphological transformation or phase consistency feature to enhance the robustness to different noise forms or local structures; Second-stage training: Use another batch of training images and corresponding correct images including real annotated centerline positions to further perform joint fine-tuning on the feature extraction layer and the output layer of the first convolutional neural network or the second convolutional neural network to improve the accuracy of centerline position prediction.
5. The method for accurate ingredient recognition and processing of a cooking machine based on computer vision according to claim 4, wherein The process of staged training further includes: using a loss function based on cross-entropy to measure the classification or discrete position annotation accuracy of the network for centerline localization; simultaneously using a regression error function based on Euclidean distance or Huber loss to measure the numerical difference between the predicted centerline position and the annotation ground truth; weighting and combining the regression error function and the loss function, and using the combined total loss function for backpropagation to update the weights of each convolutional layer and sub fully connected layer in the first convolutional neural network or the second convolutional neural network; when the combined total loss converges to a preset threshold or reaches the minimum value, stop the update; and determining the weights of the first convolutional neural network or the second convolutional neural network at this time as the final weights for inference.
6. The method for accurate ingredient recognition and processing of a chef machine based on computer vision according to claim 5, characterized in that, During the staged training process, randomly close some convolutional channels in the first convolutional neural network and / or the second convolutional neural network, and resample the truncated channels using noise distribution or adjacent channel information; after completing inference based on the resampling, calculate the confidence score of the recognition result, and when the confidence is lower than the threshold, automatically increase the resolution of the preprocessed image or perform additional morphological operations and then input it into the first convolutional neural network and / or the second convolutional neural network again to improve the stability and accuracy of the recognition result.
7. The method for accurate ingredient recognition and processing of a chef machine based on computer vision according to claim 6, characterized in that, In step 2, the fused feature map is subjected to feature decomposition and reconstruction using a reversible transformation network to generate representations in different feature spaces; based on the reconstruction results in the different feature spaces and the preprocessed image, perform multi-dimensional matching verification, and automatically correct the inconsistent local regions, thereby further improving the centerline localization stability and accuracy of each recognition object.
8. The method for accurate ingredient recognition and processing of a chef machine based on computer vision according to claim 7, characterized in that, When the image of the ingredient to be processed is recognized through a convolutional neural network model, it first passes through the first convolutional neural network and then through the second convolutional neural network.
9. The method for accurate ingredient recognition and processing of a cooking machine based on computer vision according to claim 8, wherein, All convolutional layers of the first convolutional neural network and all convolutional layers of the second convolutional neural network are in a parallel relationship.
Citation Information
Cited By
Signal identification system of neural network based on biological superficial brain architecture
CN120910783A
Heterogeneous cauliflower multi-class classification method based on color features
CN120932017A