Food detection method and device, electronic equipment and computer readable medium
By training a lightweight food detection model using object segmentation and knowledge distillation techniques, the problems of high cost of manual annotation and large model size are solved, achieving high-precision and low-latency online food quality detection.
Patent Information
- Application Number
- CN202511085463.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies for food testing rely on a large amount of manually labeled data to train models, which is costly and prone to subjective errors. Furthermore, the large size of food testing models makes them difficult to deploy on restaurant production lines for real-time testing.
Image segmentation masks are generated using object segmentation processing. A lightweight food detection model is trained using knowledge distillation technology, and semantic text binding pseudo-labels are combined to achieve automated training and deployment of the model.
It improves the speed and accuracy of defective food detection, reduces reliance on manual labeling, and enables real-time defective food detection at the restaurant production line.
Smart Images

Figure CN120997147A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular, to a food detection method and device, an electronic device and a computer readable medium. BACKGROUND
[0002] In the key areas of the back kitchen of a restaurant or a central kitchen, it is necessary to quickly determine the types of food appearing in these areas to pick out defective food. To improve efficiency, a food recognition model is generally used for food recognition. For food detection, the commonly used method is to train a food detection model using manually annotated data, and deploy the model to a cloud server for real-time detection and analysis of food.
[0003] However, when the above method is used to detect food, the following technical problems often exist:
[0004] First, the traditional method relies on a large amount of manually annotated data to train the model, which requires professional personnel to manually annotate food images, which is costly and prone to subjective errors.
[0005] Second, the food detection model is too large to be deployed to the monitoring camera device at the end of the restaurant production line to realize real-time detection of defective food.
[0006] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present inventive concept, and therefore, it can include information that does not form the prior art known to those of ordinary skill in the art in the country. SUMMARY
[0007] The summary section is provided to introduce the concepts briefly in a simplified form, which will be described in detail in the specific embodiments section. The summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0008] Some embodiments of the present disclosure propose a food detection method, device, electronic device and computer readable medium to solve one or more of the technical problems mentioned in the background section.
[0009] In a first aspect, some embodiments of the present disclosure provide a food detection method, comprising: for each food image in a set of acquired food images, performing the following generation steps: performing object segmentation processing on the food image to obtain a set of at least one image segmentation mask corresponding to at least one segmented object; performing segmentation mask fusion on the set of image segmentation masks corresponding to each segmented object in the at least one segmented object according to a preset semantic text library to generate an actual image segmentation mask; binding the obtained at least one actual image segmentation mask with a corresponding semantic text to obtain at least one binary tuple pseudo label; training an initial food detection model according to the obtained set of binary tuple pseudo labels by a knowledge distillation technique to obtain a food detection model; and deploying the food detection model to a video detection module in a restaurant monitoring terminal corresponding to a food production line, wherein the restaurant monitoring terminal further comprises: an image acquisition module; and a food detection unit configured to input a real-time food image acquired by the image acquisition module to the food detection model deployed in the video detection module to obtain a food category in response to receiving the real-time food image.
[0010] In a second aspect, some embodiments of the present disclosure provide a food detection apparatus, comprising: a generation unit configured to, for each food image in a set of acquired food images, perform the following generation steps: perform object segmentation processing on the food image to obtain a set of at least one image segmentation mask corresponding to at least one segmented object; perform segmentation mask fusion on the set of image segmentation masks corresponding to each segmented object in the at least one segmented object according to a preset semantic text library to generate an actual image segmentation mask; and bind the obtained at least one actual image segmentation mask with a corresponding semantic text to obtain at least one binary tuple pseudo label; a model training unit configured to train an initial food detection model according to the obtained set of binary tuple pseudo labels by a knowledge distillation technique to obtain a food detection model; and a model deployment unit configured to deploy the food detection model to a video detection module in a restaurant monitoring terminal corresponding to a food production line, wherein the restaurant monitoring terminal further comprises: an image acquisition module; and a food detection unit configured to input a real-time food image acquired by the image acquisition module to the food detection model deployed in the video detection module to obtain a food category in response to receiving the real-time food image.
[0011] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; and a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0012] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method as described in any implementation manner of the first aspect.
[0013] The above various embodiments of the present disclosure have the following beneficial effects: the food detection model obtained by the food detection method of some embodiments of the present disclosure improves the speed of detecting defective food. Specifically, the reason for slow detection of defective food is that manual annotation of defective food data is required to train the food detection model. Based on this, the food detection method of some embodiments of the present disclosure first, for each food image in the obtained food image set, performs the following generation step: performing object segmentation processing on the food image to obtain at least one image segmentation mask set corresponding to at least one segmented object. Thus, the mask corresponding to the food object can be automatically generated instead of manual annotation, providing visual information for subsequent pseudo-labels. Second, according to a preset semantic text library, the image segmentation mask set corresponding to each segmented object in the at least one segmented object is fused to generate an actual image segmentation mask. Thus, the semantic information is fused to optimize the mask boundary, solve the segmentation fragmentation problem, and improve the semantic consistency and integrity of the mask. Then, the obtained at least one actual image segmentation mask is bound with the corresponding semantic text to obtain at least one binary pseudo-label. Thus, structured training samples (mask + text) are constructed, and a strong association between visual objects and semantic descriptions is established. Next, according to the obtained binary pseudo-label set, the initial food detection model is trained by a knowledge distillation technology to obtain a food detection model. Thus, the generalization ability of the dual-teacher model (vision + semantics) is utilized to transfer the cross-modal knowledge in the pseudo-label to the lightweight student model, significantly reducing the dependence on manual annotation of defective data. After that, the food detection model is deployed to a video detection module in a restaurant monitoring terminal corresponding to a food production line, wherein the restaurant monitoring terminal further comprises an image acquisition module. Thus, the model is lightweight and landed, meeting the low latency requirement of the production line terminal and providing hardware support for real-time defect detection. Finally, in response to receiving a real-time food image collected by the image acquisition module, the real-time food image is input into the food detection model deployed in the video detection module to obtain a food category. Thus, the classification result is output online and in real time, triggering a defect alarm mechanism and enabling defective food to be filtered out. In summary, through the pseudo-label automation and lightweight distillation technology chain, the problem of slow manual annotation and high cost in defective food detection is solved, and high-precision, low-latency online food quality detection and monitoring are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0014] The above and other features, aspects and advantages of various embodiments of the present disclosure will become more apparent with reference to the following specific embodiments described exemplary and illustrated in the accompanying drawings. Identical or similar components shown throughout the figures are identified by identical or similar reference numerals. It is to be understood that the drawings are schematic and elements and features are not necessarily to scale.
[0015] Figure 1 is a flowchart of some embodiments of a food detection method according to the present disclosure;
[0016] Figure 2 is a structural schematic diagram of some embodiments of a food detection device according to the present disclosure;
[0017] Figure 3 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in detail with reference to the drawings, wherein the same or similar components are denoted by the same or similar reference numerals. It is to be understood that the drawings are diagrammatic and elements and features are not necessarily to scale.
[0019] It is further noted that, for the sake of brevity, only some of the features of the embodiments are described or shown in the drawings and / or the specification. It is to be understood that not all of the features of the embodiments are required, that the embodiments can be implemented or utilized with respect to only some of the features, and that specific embodiments can be implemented or utilized with respect to additional or different features.
[0020] It is to be noted that the terms "first", "second", and the like in the present disclosure are used only for differentiating different devices, modules or units, and do not imply the order or sequence of the functions of the devices, modules or units.
[0021] It is to be noted that the terms "one", "multiple", etc. are used in the present disclosure only for illustrative purpose and should not be construed as limiting. It is to be understood by those skilled in the art that, unless otherwise explicitly stated in the context, "one" should be understood as "one or more".
[0022] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are used only for illustrative purpose and should not be construed as limiting the scope of the messages or information.
[0023] The present disclosure will be described in detail with reference to the drawings and embodiments.
[0024] Reference Figure 1FIG. 10 shows a flow 100 of some embodiments of the food detection method according to the present disclosure. The food detection method comprises the following steps:
[0025] At step 101, for each food image in the set of acquired food images, the following generation steps are performed:
[0026] At step 1011, the food image is subjected to object segmentation processing to obtain a set of at least one image segmentation mask corresponding to at least one segmented object.
[0027] In some embodiments, the execution subject of the food detection method described above (e.g., an electronic device) can be a high-performance industrial computer. Among them, the high-performance industrial computer described above can be a computer with image processing capability and hardware interaction capability. The food image described above can be a static picture of one or more foods. The object segmentation processing described above can be a computer vision technology for identifying the static contour of each independent object in the image to distinguish different object instances. The segmented object described above can be a single independent object separated from the image (e.g., a meat patty and a leaf in a hamburger). The set of image segmentation masks described above can be a collection of all image segmentation masks corresponding to the segmented object. The image segmentation mask described above can be a binary matrix with the same size as the original image (visualized as a black and white contour map).
[0028] As an example, first, the food image can be preprocessed to obtain a standard food image. Second, the standard food image is input into a pre-trained segmentation model to obtain object instances and a plurality of image segmentation masks corresponding thereto. Finally, the plurality of image segmentation masks are determined as the set of image segmentation masks for the object instances.
[0029] In some optional implementations of some embodiments, the execution subject can perform object segmentation processing on the food image to obtain a set of at least one image segmentation mask corresponding to at least one segmented object, which can comprise the following steps:
[0030] First, the feature map of the food image is extracted. Among them, the feature map can be an abstract feature extracted from the image by a pre-trained convolutional neural network. The convolutional neural network described above can be ResNet (Residual Network) and VGG (Visual Geometry Group).
[0031] Second, a feature pyramid is constructed from the feature map. Among them, the feature pyramid can be a hierarchical structure composed of feature maps of different scales.
[0032] As an example, first, multi-level features (e.g., shallow feature layers, middle feature layers, deep feature layers) can be selected from a backbone network (a pre-trained convolutional neural network). Second, elements of deep features and shallow features are added using a top-down method to obtain feature maps. Finally, the number of channels of each feature map is adjusted using convolution (e.g., 1x1 convolution) to obtain a feature pyramid.
[0033] Third, according to a preset sliding window, the feature maps corresponding to the feature pyramid are slid to obtain a candidate box set. The preset sliding window can be a rectangular region used to traverse all possible object positions in the feature map. The candidate box set can be a set of object bounding boxes with food objects.
[0034] Fourth, for each candidate box in the candidate box set, the following mask segmentation steps are performed:
[0035] First sub-step, determine the object class in the original image corresponding to the candidate box to obtain at least one segmented object. The object class can be the main object class (e.g., leaf) in the corresponding region of the original image within the candidate box. The segmented object can be the object class within the candidate box.
[0036] As an example, the feature region in the original image corresponding to the candidate box can be intercepted. The object class is obtained by classifying the output class probability of the fully connected layer of the pre-trained convolutional neural network model (e.g., VGG) as a segmented object.
[0037] Second sub-step, according to the at least one segmented object, segment the food image to obtain an initial image segmentation mask. The initial image segmentation mask can be the region of the exact pixels of the object within the candidate box in the initial food image.
[0038] As an example, a pre-trained convolutional neural network can be used to extract features of the candidate region of the segmented object in the food image and generate an initial image segmentation mask.
[0039] Fifth, for each segmented object, the obtained initial image segmentation mask set is integrated to obtain an image segmentation mask set corresponding to the segmented object. The initial image segmentation mask set can be a set of multiple image segmentation masks of the same object.
[0040] As an example, initial image segmentation masks of the same segmented object class can be grouped together. Then, according to the classification confidence, the masks in the group that overlap in the candidate region are weighted and fused to generate an image segmentation mask corresponding to the segmented object.
[0041] In the process of adopting the technical solutions to solve the above technical problem one, the following problems often accompany, the object in the food image exists occlusion, adhesion and boundary blur, leading to the decline of segmentation accuracy. At the same time, the training of high-precision segmentation model consumes increases, it is difficult to meet the real-time detection requirements of defective food. In view of these problems, the general solution is: using a complex segmentation network to increase the receptive field, or introducing artificial assistance to correct the annotation. However, the inventors consider that increasing the network model structure will increase the training overhead, and adding artificial assistance will add subjective errors, so we decide to use the following solution.
[0042] In some optional implementations of some embodiments, the above execution subject can perform object segmentation processing on the above food image to obtain at least one image segmentation mask set corresponding to at least one segmented object, which can include the following steps:
[0043] First, the image height and image width corresponding to the food image are obtained. The image height can be the number of vertical pixels of the image. The image width can be the number of horizontal pixels of the image.
[0044] As an example, an image processing library can be used to read the food image to obtain the size information of the food image, and obtain the image height and image width. The image processing library can be OpenCV (Open Source Computer Vision Library, cross-platform computer vision library).
[0045] Second, according to the preset grid density parameter, the image height and the image width, the column coordinate list and the row coordinate list are determined. The preset grid density parameter can be a custom parameter that controls the grid point spacing. The column coordinate list can be a sequence of grid point positions in the horizontal axis direction. The row coordinate list can be a sequence of grid point positions in the vertical axis direction.
[0046] As an example, the height and width of the image can be separated according to the preset grid density parameter to form the coordinate list. For example, the image height can be 900. The image width can be 1200. The preset grid density parameter can be 20. Then the column coordinate list can be [0, 20, …, 1180]. The row coordinate list can be [0, 20, …, 880].
[0047] Third, the Cartesian product of the column coordinate list and the row coordinate list is determined to obtain the initial grid point set. The Cartesian product can be a combination of ordered pairs formed by all elements in the column coordinate list and the row coordinate list. The initial grid point set can be a set formed by the grid coordinates combined by the column coordinate list and the row coordinate list.
[0048] In the fourth step, the range correction processing is performed on each initial grid point in the initial grid point set based on the size of the food image to generate grid points, thereby obtaining a grid point set. The range correction processing can be a processing method of restricting the grid point coordinate range within the size boundary of the food image. The grid point set can be a set of final grid coordinates after correction.
[0049] In the fifth step, the following entity confirmation step is performed for each grid point in the grid point set:
[0050] In the first sub-step, an image block in the food image corresponding to the grid point is intercepted. The image block can be a local image region of the food image centered on the grid point.
[0051] In the second sub-step, feature extraction is performed on the image block to obtain an image block feature vector. The image block feature vector can be a numerical vector describing the image content.
[0052] As an example, a pre-trained convolutional neural network (e.g., ResNet and VGG) can be used to extract feature information in the image block and represent it in the form of a feature vector.
[0053] In the third sub-step, the image block feature vector is input into a pre-trained classification model to obtain at least one segmentation object. The pre-trained classification model can be a pre-trained convolutional neural network. The convolutional neural network can be obtained by using commonly seen food images in restaurants as a data set to train an initial network model of ResNet.
[0054] In the fourth sub-step, the grid point coordinates of the grid point corresponding to the at least one segmentation object are determined. The grid point coordinates can be the position coordinates of the grid point in the original image.
[0055] In the fifth sub-step, a mask segmentation point is constructed according to the grid point coordinates and the at least one segmentation object. The mask segmentation point can be a data structure containing grid point coordinates and segmentation object categories.
[0056] As an example, the grid point coordinates covered by the mask corresponding to each segmentation object in the image can be paired and combined with the segmentation category label (e.g., the object category to which it belongs) corresponding to the segmentation object, thereby generating a structured representation.
[0057] In the sixth sub-step, a mask region of the food image is determined according to the grid point coordinates. The mask region can be a set of grid points belonging to the same food object.
[0058] As an example, first, the grid point coordinates are converted into pixel positions. Then, the region corresponding to the pixel position is located in the food image to obtain the mask region.
[0059] A seventh sub-step is to determine a region of interest in the mask region to obtain a set of regions of interest. The region of interest can be a rectangular bounding box including a specific object or feature.
[0060] An eighth sub-step is to generate a set of original segmentation masks according to the set of regions of interest and the mask segmentation points using a pre-trained segmentation model. The pre-trained segmentation model can be a model with a pre-trained semantic segmentation network. The semantic segmentation network can be a U-Net (Convolutional Networks for Biomedical Image Segmentation). The set of original segmentation masks can be a set of preliminary binary segmentation masks generated by the pre-trained segmentation model according to the mask segmentation points and the region of interest images.
[0061] A ninth sub-step is to perform edge optimization processing on each original segmentation mask in the set of original segmentation masks corresponding to the set of regions of interest to obtain at least one image segmentation mask. The edge optimization processing can be a processing procedure for optimizing the boundary of the original segmentation mask within the region of interest by applying morphological operations (e.g., opening and closing operations) and CRF (Conditional Random Field).
[0062] As an inventive point of the above operation steps, the technical problem mentioned in the background is solved, that is, the need for professional manual labeling of food images is costly and prone to subjective errors. In practice, complete reliance on manual labeling is inefficient. When a simple full-image segmentation model is applied to handle occlusion, adhesion, and boundary ambiguity of food objects, the accuracy is often insufficient, and the use of high-precision models is costly and difficult to meet the response requirements of real-time detection on the production line. Conventional complex network or manual correction schemes can significantly increase model training and inference time, and cannot eliminate the problems of high cost and subjective bias. The present disclosure designs a scheme based on grid-based local object recognition and mask post-optimization automatic segmentation, which locuses on local image block recognition through grid points. The potential objects are located and the mask segmentation points are generated, and then a pre-trained segmentation model is used to generate the original mask in the region of interest (e.g., ROI), and finally the boundary of the segmentation mask is corrected by edge optimization techniques (e.g., morphological operations and conditional random fields). Therefore, the need for professional manual labeling is reduced, and subjective errors introduced by manual work are avoided. Through local processing and edge optimization, the problems of food occlusion, adhesion, and boundary ambiguity are effectively solved. While maintaining high segmentation accuracy, the segmentation efficiency is reduced, and real-time detection of defective food on the food production line can be achieved to select defective food.
[0063] At step 1012, the image segmentation mask set corresponding to each of the at least one segmented object is fused according to a preset semantic text library to generate an actual image segmentation mask.
[0064] In some embodiments, the execution subject can fuse the image segmentation mask set corresponding to each of the at least one segmented object according to a preset semantic text library to generate an actual image segmentation mask. The preset semantic text library can be a database of predefined food category names and corresponding semantic text feature information. The actual image segmentation mask can be an accurate pixel region representing the food type after fusion.
[0065] As an example, first, in the preset semantic text library, the type label in the mask set corresponding to each segmented object is queried. Second, the image segmentation masks are grouped according to the type label. Finally, the grouped segmentation masks are fused (e.g., weighted average) to obtain the actual image segmentation mask.
[0066] In some optional implementations of some embodiments, the execution subject can fuse the image segmentation mask set corresponding to each of the at least one segmented object according to a preset semantic text library to generate an actual image segmentation mask, which can include the following steps:
[0067] In a first step, feature information of each image region corresponding to each image segmentation mask in the image segmentation mask set is extracted to generate mask image feature information, to obtain a mask image feature information set. The mask image feature information can be a feature vector extracted from the original image region corresponding to each mask.
[0068] As an example, first, the corresponding image region is cropped from the original image using the image segmentation mask. Finally, the image region is subjected to feature extraction (e.g., color histogram and texture feature), and the extracted feature information is saved in vector form to form a set.
[0069] In a second step, for each mask image feature information in the mask image feature information set, the following semantic matching steps are performed:
[0070] In a first sub-step, the similarity information between the mask image feature information and each text feature information in the preset semantic text library is determined to obtain a similarity information matrix. The text feature information can be the feature information corresponding to each category in the semantic text library. The similarity information can be a similarity value between the mask image feature information and the text feature information. The similarity information matrix can be a vector matrix composed of the similarity values between the mask image feature information and all text feature information in the preset semantic text library. The rows of the vector matrix are the feature information of each mask image, the columns are each text feature information in the preset semantic text library, and the elements are the similarity information at the intersection of each text feature information in the preset semantic text library and the feature information of each mask image, so as to store the similarity information between the mask image feature information and each text feature information in the preset semantic text library, and facilitate calculation by calling a linear algebra library.
[0071] As an example, the similarity information can be determined by determining the cosine similarity between the mask image feature information and each text feature information in the preset semantic text library.
[0072] In a second sub-step, the main category label and the weight value of the mask image feature information are determined according to the similarity information matrix. The main category label can be the category label with the highest similarity in the similarity information matrix. The weight value can be the similarity value corresponding to the similarity information of the image segmentation mask in the corresponding main category label.
[0073] In a third step, each image segmentation mask in the image segmentation mask set is weighted and fused according to the obtained main category label set and the obtained weight value set to obtain an actual image segmentation mask. The actual image segmentation mask can be a segmentation mask obtained by weighted fusion of image segmentation masks and weights of the same category.
[0074] As an example, the actual image segmentation mask can be obtained by weighted fusion of image segmentation masks that are the same as the label.
[0075] In some optional implementations of some embodiments, the execution subject can perform segmentation mask fusion on the image segmentation mask set corresponding to each of the at least one segmentation object according to the preset semantic text library to generate an actual image segmentation mask, which can include the following steps:
[0076] First, the image segmentation mask set corresponding to the segmentation object is fused in the overlapping region to obtain an initial fusion mask. The overlapping region fusion can be to fuse the regions with more overlaps in the multiple image segmentation masks of the segmentation object. The initial fusion mask can be an image segmentation mask obtained after the overlapping region fusion.
[0077] As an example, first, the frequency of each pixel position in the original image being marked in all image segmentation masks is determined to obtain a frequency map. Second, the frequency value of each pixel position in the frequency map is divided by the total number of image segmentation masks to obtain a probability map representing whether the pixel belongs to the foreground. Finally, a threshold (for example, 0.5) is applied to the probability map, and the probability exceeding the threshold is set to the foreground in the blank initial fusion mask, and other positions are set to the background, thereby obtaining the initial fusion mask.
[0078] Second, the geometric feature information of the initial fusion mask is extracted. The geometric feature information can be a feature (for example, area, perimeter, and circular radian) describing the shape of the initial fusion mask.
[0079] As an example, an image processing library (for example, OpenCV) can be used to determine the connected domain of the initial fusion mask to obtain the geometric feature information.
[0080] Third, a semantic query key is generated according to the geometric feature information. The semantic query key can be an index key (for example, a vector and an index string) used to retrieve the semantic text library.
[0081] As an example, the geometric feature information can be classified, and the classified name information can be used as the semantic query key.
[0082] Fourth, the semantic value corresponding to the semantic query key is retrieved from the preset semantic text library to obtain a category query vector. The category query vector can be a vector obtained by encoding the semantic value by a word embedding method.
[0083] The fifth step involves determining the similarity information between the pixel positions corresponding to the initial fusion mask and the category query vector, resulting in a semantic attention map. The pixel positions can be the spatial coordinates of each pixel in the initial fusion mask. These spatial coordinates can be the pixel positions corresponding to color or texture features in the original image. The semantic attention map can be a single-channel matrix of the same size as the original image, where the value at each position represents the similarity between the feature at that position and the category query vector.
[0084] As an example, first, features are extracted from the initial fusion mask regions in the original image. Second, the cosine similarity between the feature vector of each pixel within these regions and the category query vector is determined. Finally, the obtained similarity is filled into the corresponding pixel positions to obtain the semantic attention map.
[0085] Step 6: Perform confidence filtering on each image segmentation mask in the above image segmentation mask set to obtain the image segmentation mask set to be fused. The image segmentation mask set to be fused can be the set of all image segmentation masks that satisfy the confidence condition.
[0086] Step 7: Use the semantic attention map described above to perform weighted fusion of each image segmentation mask in the set of image segmentation masks to be fused, and obtain the actual image segmentation mask.
[0087] As an example, firstly, for each mask in the set of masks to be fused, its binary matrix is multiplied element-wise by the semantic attention map to obtain a weighted mask. Secondly, all weighted masks are summed and normalized. Finally, edge optimization is performed on the normalized fused mask to obtain the actual image segmentation mask. The Canny (computational theory of edge detection) algorithm can be used for edge optimization.
[0088] Step 1013: Bind at least one actual image segmentation mask to the corresponding semantic text to obtain at least one binary pseudo-label.
[0089] In some embodiments, the execution entity can bind at least one actual image segmentation mask to corresponding semantic text to obtain at least one binary pseudo-label. The semantic text can be natural language labels describing food from a pre-defined semantic text library. The binding can create a mapping relationship between the segmentation mask and the text (e.g., key-value pairs and tuples).
[0090] In some optional implementations of certain embodiments, the execution entity may bind at least one actual image segmentation mask to the corresponding semantic text to obtain at least one binary pseudo-label, which may include the following steps:
[0091] In a first step, a semantic text corresponding to the actual image segmentation mask is obtained.
[0092] As an example, a feature analysis (e.g., shape analysis) can be performed on the actual image segmentation mask. According to the feature analysis result, a corresponding semantic text is matched in a preset semantic text library.
[0093] In a second step, the actual image segmentation mask and the semantic text are bound to obtain a binary tuple pseudo label. The binary tuple pseudo label can be a data pair including visual information of the actual image segmentation mask and semantic description of the semantic text.
[0094] As an example, the actual image segmentation mask and the corresponding semantic text can be bound in the form of a key-value pair, and a unique identifier is generated as the binary tuple pseudo label.
[0095] In step 102, an initial food detection model is trained according to the obtained binary tuple pseudo label set by a knowledge distillation technology to obtain a food detection model.
[0096] In some embodiments, the execution subject can train an initial food detection model according to the obtained binary tuple pseudo label set by a knowledge distillation technology to obtain a food detection model. The knowledge distillation technology can be a machine learning method for transferring knowledge of a complex model to a simple model. The initial food detection model can be a lightweight untrained neural network model. The neural network model can be YOLO (You Only Look Once). The food detection model can be a model that can perform food image reasoning after training the initial food detection model using the knowledge distillation technology.
[0097] As an example, first, the binary tuple pseudo label set can be converted into a commonly used format annotation file. The commonly used format can be COCO (Common Objects in Context) format. Second, a teacher model (e.g., YOLOv8) is trained using the annotation file. Finally, the initial food detection model (student model) is trained using the trained teacher model by using the knowledge distillation technology to obtain the food detection model.
[0098] In the process of using the technical solutions to solve the above technical problem two, the following problems often occur: using a lightweight model to deploy on a hardware device, and the precision of detection and recognition that can be caused by the lightweight model. To solve these problems, the conventional solution is generally to use model pruning or directly design a lightweight network architecture. However, the inventors consider that these methods are difficult to compress the model volume while maintaining high precision of food recognition. We decided to use the following solution.
[0099] In some optional implementations of some embodiments, the above execution subject can train an initial food detection model according to the obtained set of pseudo-labels of binary tuples to obtain a food detection model by a knowledge distillation technology, which can include the following steps:
[0100] Firstly, a teacher model in the knowledge distillation technology is obtained. The teacher model includes a first teacher model and a second teacher model. The teacher model can be a pre-trained model for providing a supervision signal. The first teacher model can be a model for processing visual data (for example, a ResNet model for image classification). The second teacher model can be a model for processing semantic data. The model for processing semantic data can be a BERT (Bidirectional Encoder Representation from Transformers) model for text feature extraction.
[0101] Secondly, the regions of interest of the actual image segmentation masks corresponding to the pseudo-labels of binary tuples in the set of pseudo-labels of binary tuples are integrated to generate visual data and obtain a set of visual data. The region of interest can be the region of the original image corresponding to the segmentation mask. The visual data can be the region of the original image corresponding to the segmentation mask extracted and processed into unified visual data. The set of visual data can be a set composed of all visual data.
[0102] Thirdly, according to the set of visual data and the set of pseudo-labels of binary tuples, the following model training steps are performed:
[0103] Firstly, at least one visual data in the set of visual data is input into the first teacher model to obtain visual feature information corresponding to the at least one visual data, wherein the at least one visual data is determined according to the number of model training times. The visual feature information can be image feature information extracted by the first teacher model.
[0104] Secondly, at least one pseudo-label of binary tuples in the set of pseudo-labels of binary tuples is input into the second teacher model to obtain semantic feature information corresponding to the at least one pseudo-label of binary tuples, wherein the at least one pseudo-label of binary tuples is determined according to the number of model training times. The semantic feature information can be semantic text feature information extracted by the second teacher model.
[0105] A third sub-step, inputting the visual feature information corresponding to the at least one visual data and the semantic feature information corresponding to the at least one binary pseudo label into the initial food detection model to obtain at least one student visual feature information projected into a visual teacher space and at least one student semantic feature information projected into a text teacher space. The initial food detection model can be a student model to be trained. The student visual feature information can be feature information extracted by the student model after processing the visual data. The student semantic feature information can be feature information extracted by the student model after processing the text data.
[0106] A fourth sub-step, comparing the visual feature information corresponding to the at least one visual data and the student visual feature information to obtain a visual knowledge distillation loss value. The visual knowledge distillation loss value can be an index value for measuring the difference loss between the student visual feature and the teacher visual feature.
[0107] As an example, the visual knowledge distillation loss value can be determined by the mean square error of the visual feature information and the student visual feature information.
[0108] A fifth sub-step, comparing the semantic feature information corresponding to the at least one binary pseudo label and the student semantic feature information to obtain a text knowledge distillation loss value. The text knowledge distillation loss value can be an index for measuring the loss between the teacher semantic feature information and the student semantic feature.
[0109] A sixth sub-step, weighting and fusing the visual knowledge distillation loss value and the text knowledge distillation loss value to obtain a comprehensive distillation loss value. The comprehensive distillation loss value can be a total training loss for determining the training state of the student model.
[0110] A seventh sub-step, determining whether the initial food detection model reaches a preset optimization target according to the comprehensive distillation loss value. The preset optimization target can be a convergence condition (e.g., loss threshold) of the preset model.
[0111] An eighth sub-step, in response to the initial food detection model reaching the preset optimization target, taking the initial food detection model as a food detection model trained.
[0112] A fourth step, in response to the initial food detection model not reaching the preset optimization target, adjusting the training parameters of the initial food detection model, and using the adjusted initial food detection model as the initial food detection model to execute the model training step again. The training parameters can be the learning rate and batch size of the student model training.
[0113] As an inventive point of the present disclosure, the above operation steps solve the second technical problem mentioned in the background art, i.e., the food detection model is too large in size and difficult to be deployed on the camera device at the production line end to realize real-time detection of food. In practice, directly compressing a high-performance model (e.g., YOLOv8) or using a lightweight model may result in a significant decrease in the accuracy of the model, which cannot meet the real-time detection requirements of defective food on the food production line. There may be a problem of excessive loss of accuracy and difficulty in effectively fusing multi-modal information (e.g., the defective food is not food, but a model of food). The present disclosure designs a scheme based on visual-semantic dual-teacher knowledge distillation, which guides the student model to learn high-quality visual features and semantic features through a pre-trained visual teacher model (e.g., ResNet) and a semantic teacher model (e.g., BERT), and fuses the knowledge into a single student model (food detection model). Therefore, the food detection model (student model) obtained by training can maintain a small size and low computational complexity (convenient for deployment and real-time detection on food production line monitoring devices), and can also inherit the high-precision multi-modal information processing capability of the teacher model. The lightweight food detection model can be deployed on the restaurant monitoring device to perform real-time detection of food on the production line and identify defective food.
[0114] Step 103, deploying the food detection model to a video detection module in the restaurant monitoring terminal corresponding to the food production line.
[0115] In some embodiments, the above execution subject can deploy the food detection model to a video detection module in the restaurant monitoring terminal corresponding to the food production line. The food production line can be a physical workflow area for restaurant food preparation. The physical workflow area can include a food material storage area, a food material preprocessing area, a food material cooking area, a food material assembly area, and a food material packaging area. The restaurant monitoring terminal can be a computer device for managing and controlling restaurant monitoring. The video detection module can be a software system for real-time processing of restaurant monitoring video streams. As an example, an interface can be used to realize the connection with the restaurant monitoring terminal.
[0116] Step 104, in response to receiving the real-time food image collected by the image collection module, inputting the real-time food image into the food detection model deployed in the video detection module to obtain the food category.
[0117] In some embodiments, the execution subject described above can input the real-time food image collected by the image collection module into a food detection model deployed in the video detection module to obtain a food category in response to receiving the real-time food image. The image collection module described above can be a hardware system (e.g., a camera) deployed on a production line to obtain real-time images. The real-time food image described above can be a food image captured by the camera. The food category described above can be the type and category of the food identified by the food detection model after the food image is input into the food detection model.
[0118] The above various embodiments of the present disclosure have the following beneficial effects: the food detection model obtained by the food detection method of some embodiments of the present disclosure improves the speed of detecting defective food. Specifically, the reason for slow detection of defective food is that manual annotation of defective food data is required to train the food detection model. Based on this, the food detection method of some embodiments of the present disclosure first, for each food image in the obtained food image set, performs the following generation step: performing object segmentation processing on the food image to obtain at least one image segmentation mask set corresponding to at least one segmented object. Thus, the mask corresponding to the food object can be automatically generated instead of manual annotation, providing visual information for subsequent pseudo-labels. Second, according to a preset semantic text library, the image segmentation mask set corresponding to each segmented object in the at least one segmented object is fused to generate an actual image segmentation mask. Thus, the semantic information is fused to optimize the mask boundary, solve the segmentation fragmentation problem, and improve the semantic consistency and integrity of the mask. Then, the obtained at least one actual image segmentation mask and the corresponding semantic text are bound to obtain at least one binary pseudo-label. Thus, the structured training sample (mask + text) is constructed, and the strong association between the visual object and the semantic description is established. Next, according to the obtained binary pseudo-label set, the initial food detection model is trained by a knowledge distillation technology to obtain a food detection model. Thus, the generalization ability of the dual-teacher model (vision + semantics) is used to transfer the cross-modal knowledge in the pseudo-label to the lightweight student model, significantly reducing the dependence on manually annotated defect data. Then, the food detection model is deployed to a video detection module in a restaurant monitoring terminal corresponding to a food production line, wherein the restaurant monitoring terminal further comprises an image acquisition module. Thus, the model is lightweight and landed, meeting the low delay requirement of the production line terminal and providing hardware support for real-time defect detection. Finally, in response to receiving a real-time food image collected by the image acquisition module, the real-time food image is input into the food detection model deployed in the video detection module to obtain a food category. Thus, the classification result is output online and in real time, triggering a defect alarm mechanism and enabling the defective food to be filtered out. In summary, through the pseudo-label automation and lightweight distillation technology chain, the problem of slow manual annotation and high cost in defective food detection is solved, and high-precision, low-delay online food quality detection and monitoring are achieved.
[0119] Further reference Figure 2 As an implementation of the method shown in the above figures, the present disclosure provides some embodiments of a food detection device, which correspond to the method embodiments shown in Figure 1 The food detection device can be applied to various electronic devices.
[0120] As Figure 2As shown, a food detection device 200 includes a generation unit 201, a model training unit 202, a model deployment unit 203, and a food detection unit 204. The generation unit 201 is configured to, for each food image in a set of acquired food images, perform the following generation steps: perform object segmentation processing on the food image to obtain a set of at least one image segmentation mask corresponding to at least one segmented object; perform segmentation mask fusion on the set of image segmentation masks corresponding to each segmented object in the at least one segmented object according to a preset semantic text library to generate an actual image segmentation mask; and bind the obtained at least one actual image segmentation mask with the corresponding semantic text to obtain at least one binary tuple pseudo label. The model training unit 202 is configured to train an initial food detection model according to a set of obtained binary tuple pseudo labels by a knowledge distillation technique to obtain a food detection model. The model deployment unit 203 is configured to deploy the food detection model to a video detection module in a restaurant monitoring terminal corresponding to a food production line, wherein the restaurant monitoring terminal further includes an image acquisition module. The food detection unit 204 is configured to, in response to receiving a real-time food image acquired by the image acquisition module, input the real-time food image into the food detection model deployed in the video detection module to obtain a food category.
[0121] It can be understood that the units described in the food detection device 200 correspond to the respective steps in the method described above. Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the food detection device 200 and the units included therein, which will not be described here again.
[0122] Reference is made below to Figure 3 which shows a structural schematic diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0123] As shown in Figure 3 , the electronic device 300 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 302 or loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0124] Generally, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 308 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 309. The communication devices 309 can allow the electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or fewer devices can alternatively be implemented or present. Figure 3 Each block shown in the flowcharts can represent a device or multiple devices as needed.
[0125] In particular, processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product including a computer program carried on a computer readable medium, the computer program containing program codes for executing the methods shown in the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network through the communication devices 309, or installed from the storage devices 308, or installed from the ROM 302. When the computer program is executed by the processing devices 301, the above-described functions defined in the methods of some embodiments of the present disclosure are performed.
[0126] Note that the computer readable medium in some embodiments of the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In some embodiments of the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used by an instruction execution system, apparatus or device, or that can be used by or in connection with an instruction execution system, apparatus or device. In some embodiments of the present disclosure, the computer readable signal medium can include a computer readable program code propagated on or through a carrier wave, in baseband or passed as a part of a carrier wave. Such propagated signals can take a wide variety of forms including, but not limited to, electro-magnetic signals, optical signals or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0127] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0128] The computer readable medium can be included in the electronic device, or exist separately from the electronic device. The computer readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: for each food image in a set of acquired food images, perform the following generation steps: perform object segmentation processing on the food image to obtain a set of at least one image segmentation mask corresponding to at least one segmented object; perform segmentation mask fusion on the set of image segmentation masks corresponding to each segmented object in the at least one segmented object according to a preset semantic text library to generate an actual image segmentation mask; bind the obtained at least one actual image segmentation mask to a corresponding semantic text to obtain at least one binary tuple pseudo label; train an initial food detection model according to the obtained set of binary tuple pseudo labels by a knowledge distillation technique to obtain a food detection model; and deploy the food detection model to a video detection module in a restaurant monitoring terminal corresponding to a food production line, wherein the restaurant monitoring terminal further includes an image acquisition module, and in response to receiving a real-time food image acquired by the image acquisition module, inputs the real-time food image to the food detection model deployed in the video detection module to obtain a food category.
[0129] Computer program code for carrying out operations of some embodiments of the disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0130] The flow and block diagrams in the drawings represent possible architectural, functional, and operational architectures of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block can represent a module, a segment, or a portion of code that comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.
[0131] The units described in some embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. The described units can also be arranged in a processor, for example, can be described as: a processor includes a generation unit, a model training unit, a model deployment unit and a food detection unit. Among them, the name of these units does not constitute a limitation to the unit itself in some cases, for example, the model training unit can also be described as: "a unit that obtains a food detection model by training an initial food detection model according to a set of obtained binary pseudo labels through a knowledge distillation technology".
[0132] The functions described above in the specification can be implemented, at least in part, by one or more hardware logic components. For example, and without limitation, non-limiting examples of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0133] The above description is merely some of the preferred embodiments of the present disclosure and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features are replaced with the technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions to form technical solutions.
Claims
1. A food detection method, comprising: For each food image in a set of acquired food images, the following generation steps are performed: performing object segmentation processing on the food image to obtain at least one image segmentation mask set corresponding to at least one segmented object; According to a preset semantic text library, the image segmentation mask set corresponding to each segmented object in the at least one segmented object is fused to generate an actual image segmentation mask; binding the obtained at least one actual image segmentation mask with the corresponding semantic text to obtain at least one binary tuple pseudo label; training an initial food detection model according to the obtained binary tuple pseudo label set through a knowledge distillation technique to obtain a food detection model; deploying the food detection model to a video detection module in a restaurant monitoring terminal corresponding to a food production line, wherein the restaurant monitoring terminal further comprises an image acquisition module; in response to receiving a real-time food image acquired by the image acquisition module, inputting the real-time food image into the food detection model deployed in the video detection module to obtain a food category.
2. The method of claim 1, wherein, The binding of the obtained at least one actual image segmentation mask with the corresponding semantic text to obtain at least one binary tuple pseudo label comprises: acquiring the semantic text corresponding to the actual image segmentation mask; binding the actual image segmentation mask and the semantic text to obtain a binary tuple pseudo label.
3. The method of claim 1, wherein, The object segmentation processing on the food image to obtain at least one image segmentation mask set corresponding to at least one segmented object comprises: extracting a feature map of the food image; constructing a feature pyramid through the feature map; sliding the feature map corresponding to the feature pyramid according to a preset sliding window to obtain a candidate box set; for each candidate box in the candidate box set, the following mask segmentation steps are performed: determining the object class in the corresponding original image within the candidate box to obtain at least one segmented object; segmenting the food image according to the at least one segmented object to obtain an initial image segmentation mask; for each segmented object, integrating the obtained initial image segmentation mask set to obtain an image segmentation mask set corresponding to the segmented object.
4. The method of claim 1, wherein, The image segmentation mask set corresponding to each segmented object in the at least one segmented object is fused according to a preset semantic text library to generate an actual image segmentation mask, comprising: extracting features from the original image region corresponding to each image segmentation mask in the image segmentation mask set to generate mask image feature information, obtaining a mask image feature information set; for each mask image feature information in the mask image feature information set, the following semantic matching steps are performed: determining the similarity information of the mask image feature information and each text feature information in the preset semantic text library to obtain a similarity information matrix; determining the main category label and weight value of the mask image feature information according to the similarity information matrix; weighting and fusing each image segmentation mask in the image segmentation mask set according to the obtained main category label set and the obtained weight value set to obtain an actual image segmentation mask.
5. The method of claim 1, wherein, The segmentation mask fusion on the image segmentation mask set corresponding to each of the at least one segmented object according to the preset semantic text library to generate an actual image segmentation mask comprises: performing overlapping region fusion on the image segmentation mask set corresponding to the segmented object to obtain an initial fusion mask; extracting geometric feature information of the initial fusion mask; generating a semantic query key according to the geometric feature information; retrieving a semantic value corresponding to the semantic query key from the preset semantic text library to obtain a category query vector; determining similarity information between a pixel position corresponding to the initial fusion mask and the category query vector to obtain a semantic attention map; performing confidence screening processing on each image segmentation mask in the image segmentation mask set to obtain a set of image segmentation masks to be fused; performing weighted fusion on each image segmentation mask in the set of image segmentation masks to be fused by using the semantic attention map to obtain an actual image segmentation mask.
6. A food detection device, comprising: a generation unit configured to, for each food image in a set of acquired food images, perform the following generation steps: performing object segmentation processing on the food image to obtain at least one image segmentation mask set corresponding to at least one segmented object; performing segmentation mask fusion on the image segmentation mask set corresponding to each of the at least one segmented object according to a preset semantic text library to generate an actual image segmentation mask; binding the obtained at least one actual image segmentation mask with a corresponding semantic text to obtain at least one binary tuple pseudo label; a model training unit configured to train an initial food detection model according to the obtained set of binary tuple pseudo labels by a knowledge distillation technique to obtain a food detection model; a model deployment unit configured to deploy the food detection model to a video detection module in a restaurant monitoring terminal corresponding to a food production line, wherein the restaurant monitoring terminal further comprises an image acquisition module; a food detection unit configured to, in response to receiving a real-time food image acquired by the image acquisition module, input the real-time food image into the food detection model deployed in the video detection module to obtain a food category.
7. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1-5.
8. A computer readable medium having stored thereon a computer program, wherein, The program is executed by the processor to implement the method of any one of claims 1-5. The program is executed by the processor to implement the method of any one of claims 1-5.