Incremental food detection model training method and food detection method

By adopting supervised comparison learning methods and class prototype perception constraints in the food detection model, the problem that food detection model is difficult to maintain the detection performance of old categories when facing new food categories is solved, and the model is quickly adapted and improved detection performance.

CN120107751APending Publication Date: 2025-06-06INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510221288.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When facing new food categories, it is difficult for existing food detection models to maintain detection performance for old categories. Due to large differences within food categories and small differences between categories, characteristics are confusing, making it difficult to achieve effective incremental learning.

Method used

The method based on supervised contrast learning is adopted, and the characteristic space of the food detection model is constrained through class prototype perception constraints and inter-class distance loss constraints, making the characteristics of the same class more compact and the characteristics of different classes more distinct, thereby realizing incremental learning.

Benefits of technology

Without adding a lot of time costs and computing resource overhead, the model is updated and expanded, allowing the model to quickly adapt to new food categories and improve detection performance and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107751A_ABST
    Figure CN120107751A_ABST
Patent Text Reader

Abstract

The invention provides a training method of an incremental food detection model. The training method comprises the following steps: initializing the incremental food detection model by utilizing a trained base class model; and training the incremental food detection model based on a supervised comparative learning method. Wherein the base class model is trained by using a first training set, the first training set comprises a first input image and a label corresponding to base class food, and the first input image comprises base class food data; the incremental food detection model is trained by using a second training set, the second training set comprises a second input image and a label corresponding to the new food, and the second input image comprises new food data. According to the method, the feature space is constrained based on a supervised comparative learning method, so that the features of the same class are more compact, the features of different classes are more distinct, and the detection performance and generalization ability of the model are improved. According to the invention, the model can quickly adapt to the requirements of new food categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a training method for an incremental food detection model and a food detection method. Background Art

[0002] The statements in this section are merely intended to provide background information related to the present invention to help understanding the present invention. Such background information does not necessarily constitute prior art.

[0003] Food is the basis for human life, and good eating habits can help people prevent various chronic diseases including diabetes and cardiovascular diseases. Food computing aims to process and analyze dietary data (such as food images) through artificial intelligence technology to promote the transformation and upgrading of the field of food science. Food detection is an important direction of food computing. It has important application value in automatic recording and settlement, dietary assessment and sustainable dietary monitoring, which helps people develop good personal dietary habits and thus provide protection for personal health. Due to the wide variety of food images in the real world, with the development of human society, the types of food are constantly updated and evolved, which poses a challenge to deploying food detection models in actual application scenarios. Suppose we want to deploy an intelligent settlement system in the cafeteria, and the canteen's menu needs to be updated from time to time. Then, when new dishes are added to the cafeteria, how to enable the system to accurately detect new dishes while maintaining the detection performance of old dishes is crucial. This requires the food detection model to have the ability of incremental learning (that is, the model continuously learns new categories of knowledge through training in each stage in turn, while avoiding forgetting old categories). In other words, we need to design an effective incremental detection model for food images.

[0004] There are several main development routes for existing incremental detection models: (1) Data playback-based methods. The contribution of this type of work is mainly to improve the construction of old class sample sets and propose loss functions for old class samples, using old class sample sets to alleviate the forgetting of target detectors on old classes. For example: Joseph et al. (Joseph KJ, Rajasegaran J, KhanS, et al. Incremental Object Detection via Meta-Learning[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 44(12):9209-9216.) proposed to construct a gradient correction layer to correct the gradient of model parameter updates through the classification and regression losses of old class samples. Liu et al. (Liu Y, Cong Y, Goswami D, et al. Augmented Box Replay: Overcoming Foreground Shift for Incremental Object Detection[C] / / Proceedingsof the IEEE / CVF International Conference on Computer Vision. 2023: 11367-11377.) made an innovation in the way of constructing sample sets, proposing to save only the image within the detection box instead of the entire image to avoid "new class-background" confusion. (2) Methods based on knowledge transfer. This type of work mainly contributes to the design of transfer loss functions, which avoids catastrophic forgetting of old classes by learning the knowledge of old models. For example: Shmelkov et al. (Shmelkov K, Schmid C, Alahari K. Incremental Learning of Object Detectors without Catastrophic Forgetting[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2017: 3400-3409.) designed a migration loss function and trained the model based on the classification results and detection box position differences between the new and old models on the old classes.Peng et al. (Peng C, Zhao K, Lovell B C. Faster ILOD: Incremental Learning for Object Detectors Based on FasterRCNN[J]. Pattern Recognition Letters, 2020, 140: 109-115.) added a migration loss based on the difference between backbone network features and region of interest (RoI) features on the basis of migrating the loss function and training the model, and designed the loss function based on the difference in response between the new and old models on the old classes. (3) Incremental detection work based on model extension. This type of work attempts to expand the model structure so that the model can preserve the features of the old classes; or measure the importance of model parameters to avoid large updates of important parameters related to the old classes, thereby suppressing the model from forgetting the old class objects. For example: Yang et al. (Yang B, Deng X, Shi H, et al. Continual Object Detection via Prototypical TaskCorrelation Guided Gating Mechanism[C] / / Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition. 2022: 9255-9264.) introduced a threshold mechanism in the detector, explicitly saving the information corresponding to a certain incremental step category into the parameters of the overall framework, and using the parameters to generate a feature mask of the corresponding category. Summary of the invention

[0005] In order to solve the above problems, the present invention proposes a training method for an incremental food detection model and a food detection method.

[0006] According to a first aspect, the present invention provides a training method for an incremental food detection model, wherein the incremental food detection model includes a neural network for generating candidate regions of an input image and a feature map of the input image based on an input image, and a region of interest head stage, wherein the region of interest head stage is used to classify targets and regress bounding boxes based on the candidate regions of the input image and the feature map of the input image; the training method includes: initializing the incremental food detection model using a trained base class model; training the incremental food detection model based on a supervised contrastive learning method; wherein the base class model is trained using a first training set, the first training set includes a first input image and labels corresponding to base class foods, and the first input image includes base class food data; in the step of training the incremental food detection model based on a supervised contrastive learning method, a second training set is used for training, the second training set includes a second input image and labels corresponding to new class foods, and the second input image includes new class food data.

[0007] Preferably, the incremental food detection model is trained based on a supervised contrastive learning method, including: obtaining the category of the target in the input image according to the candidate area of ​​the input image and the feature map of the input image; for each category, calculating the average value of the feature vector corresponding to the category, and determining the class prototype according to the average value; adjusting the category of the target in the input image according to the set intra-class distance loss constraint and the class prototype, and / or according to the set inter-class distance loss constraint and the class prototype.

[0008] Preferably, each time the training of the incremental food detection model reaches a set iteration round, the class prototype is updated in the following manner: a feature memory of the category is set for each category, the feature memory is used to store a feature queue corresponding to the category in each iteration round, and the feature queue includes a feature vector; the class prototype is updated according to the set momentum parameter, the value of the class prototype and the average value of the feature vector of the category in the current iteration round.

[0009] Preferably, the category of the target in the input image is adjusted according to the set intra-class distance loss constraint and the class prototype, including: for each new class target corresponding to the bounding box in the input image, determining the category of the target according to its corresponding label; determining the corresponding class prototype according to the category of the target; calculating the intra-class distance loss according to the feature vector of the target, the class prototypes corresponding to all categories, the mapping relationship of mapping the input image to the feature space and the set temperature coefficient; and adjusting the category of the target in the input image according to the intra-class distance loss.

[0010] Preferably, the category of the target in the input image is adjusted according to the set inter-class distance loss constraint and the class prototype, including: for each new class target corresponding to the bounding box in the input image, determining the category of the target according to its corresponding label; determining the corresponding class prototype according to the category of the target; calculating the second norm of the distance between the class prototype vector of each new class and the class prototype vectors of other classes, and taking the inverse of the second norm as the inter-class distance loss; adjusting the category of the target in the input image according to the inter-class distance loss.

[0011] Preferably, the neural network used to generate candidate regions of an input image and a feature map of the input image based on the input image includes a backbone network and a region proposal network; the backbone network is used to extract features from the input image and output the feature map of the input image; the region proposal network is used to generate candidate regions on the feature map, and the candidate regions are regions including the target.

[0012] Preferably, the backbone network includes one of VGG16, ResNet and ResNeXt.

[0013] According to the second aspect, the present invention proposes a food detection method, which includes: inputting a food image to be detected into a detection model, and the detection model outputs the category of the target in the food image and the bounding box coordinates of the target; wherein the detection model is an incremental food detection model trained by any of the methods described in the first aspect.

[0014] According to a third aspect, the present invention provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method described in any one of the first aspect and / or the second aspect are implemented.

[0015] The training method of the present invention can update and expand the model without increasing a lot of time cost and computing resource overhead, so that the model can quickly adapt to the needs of new food categories. The training method constrains the feature space based on the supervised contrastive learning method, making the features of the same category more compact and the features of different categories more distinct, thereby improving the detection performance and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic diagram of the structure of an incremental food detection model according to an embodiment of the present invention; Figure 2 is a flowchart of a training method for an incremental food detection model according to an embodiment of the present invention; Figure 3is a schematic diagram of a training process of an incremental detection model using a detector framework based on Faster RCNN according to an embodiment of the present invention; Figure 4 It is a schematic diagram of setting intra-class distance loss constraints and inter-class distance loss constraints to adjust the type of target and the bounding box coordinates of the target according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments herein are only used for illustration and are not intended to limit the present invention. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent to those of ordinary skill in the art that these specific details are not necessarily used to implement the present invention. In other examples, in order to avoid confusing the present invention, known procedures, materials or methods are not specifically described.

[0018] In the field of computer vision, object detection is a basic and crucial task, which aims to identify and locate objects of interest from images or videos. With the continuous development of technology, object detection technology plays an increasingly important role in many application scenarios, especially in the fields of food safety, quality control, and catering management. Food detection has become an important branch of object detection. However, compared with ordinary object detection scenarios, food detection faces more complex and unique challenges.

[0019] Ordinary object detection tasks usually target objects with rigid structures and relatively fixed shapes, such as vehicles and pedestrians. These objects have obvious inter-class differences and small intra-class differences, which makes it easier for the detection model to learn stable feature representations. However, in food detection, due to the lack of rigid structure of food itself and the variety of cooking methods, the same dish may present completely different visual effects under different production conditions. This characteristic of large intra-class differences makes traditional object detection methods perform poorly in food detection tasks. In addition, food detection also faces the problem of small inter-class differences. Since different ingredients may have similar visual features, foods made from these ingredients may also present similar appearances, such as chicken nuggets and potato nuggets are visually easy to confuse. This characteristic of small inter-class differences further increases the difficulty of food detection. The characteristics of small inter-class differences and large intra-class differences in food images easily cause confusion in feature space. In real life, the huge variety of food and diverse scenes (such as containers for food, the viewing angle of image acquisition equipment, and the lighting of the detection scene) also make it difficult to fully train food detection models. When new types of food appear, if all food data are retrained, it will greatly increase the time cost and computing resource overhead. If only the new types of food are used to fine-tune the original model, it is easy to cause catastrophic forgetting of the old types of data (that is, while learning to detect new categories, the detection performance of the old types of data is significantly reduced). Therefore, how to design an incremental detection model that can effectively deal with the problem of large intra-class differences and small inter-class differences in food detection has become a hot topic and difficulty in current research.

[0020] In response to the technical difficulties in food inspection, the inventors discovered through research that the problem of feature confusion in food inspection can be solved by extracting features of food images and then using supervised contrastive learning methods in each incremental training step to constrain the distribution of each current category in the feature space. This allows each category to form a compact cluster in the feature space, making features of the same category more compact and features of different categories more distinct.

[0021] 1. Model Structure The structure of the incremental food detection model in one embodiment of the present invention will be introduced below.

[0022] like Figure 1As shown, the incremental food detection model of the embodiment of the present invention adopts the framework of Faster R-CNN. Faster R-CNN is a two-stage target detection framework, including the Backbone Network, Region Proposal Network (RPN) and Region of Interest Head (ROI Head) stages. These three parts work together to complete the target detection task.

[0023] The backbone network is used to extract features from the input food image. In some embodiments, the backbone network can use a deep convolutional neural network such as a visual geometry group network VGG16, a residual network ResNet, or a residual network extended ResNeXt. The output of the backbone network is a feature map of the food image, which contains high-level semantic information of the food image.

[0024] The region proposal network (RPN) is used to generate candidate regions on the feature map and classify each candidate region (e.g., object or non-object) and regress bounding boxes (i.e., adjust the position of bounding boxes). These candidate regions are considered to be regions that may contain target objects. RPN predicts the best candidate region for each position by sliding a window on the feature map and using anchors. At each position in the feature map, RPN generates multiple anchors of different sizes and scales, which serve as the initial predictions of the candidate boxes. RPN contains two branches, one for classification (determining whether the anchor contains the target object) and the other for bounding box regression (adjusting the position and size of the anchor to better match the target object). The classification branch of RPN predicts whether each anchor is the foreground (containing the target object) or the background through the softmax function, while the regression branch of RPN predicts the offset of the anchor relative to the true target box. Since multiple anchors may predict the same target area, the non-maximum suppression (NMS) method can be used to screen out the best candidate regions, that is, while retaining the best predicted box, other predicted boxes that overlap too much can be removed. After the above steps, RPN outputs a series of screened and adjusted candidate regions, and inputs the candidate regions into the region of interest head (ROI Head) stage.

[0025] The ROI Head stage is used to further process the candidate regions and feature maps generated by the RPN to complete the classification of the target and the precise regression of the bounding box. The input of the ROI Head stage is the candidate regions and feature maps generated by the RPN stage, and the output is the category of the object and the corresponding bounding box coordinates. The ROI Head stage has two main functions: one is to classify each candidate region to determine which category they belong to; the other is to fine-tune the coordinates of each predicted candidate region, that is, to adjust the bounding box twice to locate the target more accurately. Due to the different sizes of the candidate regions, the ROI Head stage usually includes ROI Pooling or ROI Alignment operations, which adjust the candidate regions of different sizes to a uniform size so that they can be input into the subsequent network for feature extraction. The feature map after ROI Pooling or ROI Alignment will be sent to one or more fully connected layers, which are responsible for completing the tasks of classification and bounding box regression. During the training phase, the ROI Head calculates the classification loss and bounding box regression loss. The classification loss distinguishes the category of the target object, while the regression loss fine-tunes the position of the bounding box. During the test phase, the ROI Head calculates the probabilities of all candidate regions and adjusts the positions of the predicted candidate boxes. It then performs an NMS operation to remove overlapping bounding boxes and outputs the category and precise bounding box coordinates of each candidate region.

[0026] According to one embodiment of the present invention, the input image first extracts a feature map through the backbone network, and then the feature map is sent to the RPN to generate a candidate region. Finally, the candidate region and the feature map are Figure 1 The image is sent to the ROI Head stage for classification and bounding box regression. The backbone network provides basic feature extraction capabilities, the RPN is responsible for quickly generating candidate regions, and the ROI Head stage refines these candidate regions to ultimately achieve target detection. A key advantage of Faster R-CNN is that it can be trained end-to-end, that is, the backbone network, RPN, and ROI Head stages can be trained together and share the same loss function, which helps the model coordinate the performance of each part during training.

[0027] 2. Model Training like Figure 2As shown, according to one embodiment of the present invention, a training method for an incremental food detection model is proposed, comprising: step S201. Initializing the incremental food detection model using the trained base class model to achieve knowledge transfer. The base class model is trained using the first training set, the first input image of the first training set includes the base class food, and the first training set also includes the label corresponding to the base class food. The base class model refers to a model that has trained the image data of the previous stage. The base class food image refers to a basic or original food image. Step S202. The incremental food detection model is trained based on a supervised contrast learning method, the incremental food detection model is trained using the second training set, the second input image of the second training set includes the new class food, and the second training set also includes the label corresponding to the new class food. In some embodiments, the second input image can be based on the first input image and the new class food data is added (that is, an image can include both the instance of the new class food and the instance of the base class food), but in the process of training with the second training set, only the new class food instance added in this training stage and its corresponding label are used for training. Supervised contrast learning is a method that combines contrast learning technology and supervised learning. The key to supervised contrastive learning is to use category label information to define positive and negative samples. In this learning framework, samples with the same label are considered positive samples, while samples with different labels are considered negative samples. The goal of this method is to bring the feature representations of the same class of data closer in the feature space, while pushing the feature representations of different classes further apart. The main difference between supervised contrastive learning and unsupervised contrastive learning is that unsupervised contrastive learning usually relies on data augmentation to generate positive sample pairs (for example, different transformations of the same image), while supervised contrastive learning uses category label information to determine positive and negative sample pairs. This approach enables the model to utilize label information more effectively, thereby achieving better performance in classification tasks. New category food images refer to new category food images that are different from the categories in the base category food images.

[0028] Figure 3 FIG. 1 is a schematic diagram of a training process of an incremental detection model using a detector framework based on Faster RCNN according to an embodiment of the present invention. Figure 3As shown in Figure 1, the training process is divided into two stages: the first stage is the training stage of the base class model; the second stage is the incremental training stage of the incremental food detection model. In the training stage of the base class model, the base class model is trained using a training set containing only images of base class food. Through this process, a preliminary base class model can be obtained. Then, the incremental food detection model is initialized using the base class model, that is, the knowledge transfer of the base class model is used as the initialization of the incremental food detection model, and the incremental training of the new class model is prepared. In the incremental training stage, the food data of the new category is gradually added to the input image of the training set (only the food data of the new category and the corresponding labels are used for training, and the labels in the first stage training set will not be used for training in this stage), and training is performed based on the existing base class model. In this process, the proposed supervised contrastive learning method based on class prototype-aware constraints is used to constrain the feature space. Class prototype-aware constraints mean that in the process of updating model parameters, the class prototype is also updated with the update of the model. At the same time, using the class prototype as a reference point, each sample can be constrained to be distributed around the corresponding class prototype as much as possible. By continuously updating the model parameters and adjusting the weight of the loss function, an incremental food detection model is obtained. The incremental food detection model can detect the detection model of both base and new types of food at the same time.

[0029] In some embodiments, step S202. Training the incremental food detection model based on the supervised contrastive learning method further comprises: According to the candidate area of ​​the input image and the feature map of the input image, the category of the target in the input image is obtained. For each category, the average value of the feature vector corresponding to the category is calculated, and the class prototype is determined according to the average value. In representation learning, the class prototype is a central concept used to describe and represent categories. Specifically, the class prototype is a compact and idealized representation of a category, usually in the form of a certain central vector of the samples of that category. It is an induction or summary of the characteristics of the samples of that category, and can be used to characterize the overall characteristics of the category. The class prototype can be regarded as the clustering center of each class of samples. In some embodiments, when each training incremental food detection model reaches a set iteration round, the class prototype is updated in the following manner: a feature memory of the category is set for each category, and the feature memory is used to store the feature queue corresponding to the category of each iteration round, and the feature queue includes a feature vector; according to the set momentum parameter , the value of the class prototype And the average value of the feature vector of the category in the current iteration round Update the class prototype according to formula 1) .

[0030] 1) like Figure 4 As shown, in some embodiments, an intra-class distance loss constraint can be set, and the category of the target in the input image can be adjusted according to the set intra-class distance loss constraint and the class prototype. The intra-class distance loss constraint can be used as an attraction to cluster the features of the same class as much as possible around the class prototype vector of this class, so that the clustering clusters of each class are more concentrated within a certain range, and the features of the same class are constrained to be as far away as possible from the feature clustering clusters of other classes to prevent confusion. For each target corresponding to the bounding box in the input image, determine the category of the target. Determine the corresponding class prototype according to the category of the target. According to the feature vector of the target, the class prototypes corresponding to all categories, the mapping relationship of mapping the input image to the feature space, and the set temperature coefficient, calculate the intra-class distance loss according to Formula 2). . The categories of objects in the input image are adjusted according to the intra-class distance loss.

[0031] 2) in, , Represents all training samples, i.e., input food images, which can contain both new food image data and base food image data; Indicates Samples of new food examples; express Category of; Indicates category The corresponding class prototype vector; Indicates that except for category Other categories other than Indicates category The corresponding class prototype vector; Represents the mapping relationship from samples to feature space, that is, it can express the mapping function of the backbone network; Represents the temperature coefficient. The temperature parameter can adjust the smoothness of the output, enhance the difference between categories, and make the distribution sharper. In this process, the class prototype of the base class is saved, the category of the new class instance is determined by the label information of the new class, and the characteristics of the new class are constrained to be far away from the class prototype of the base class.

[0032] Specifically, for each new class sample feature extracted from the candidate region, the label information extracted in the ROI Head stage is used to determine the category of the target in the candidate region, and then the intra-class distance loss is calculated by formula 2). This loss consists of positive and negative comparison terms. With its class prototype vector Form a pair, the current feature With other class prototype vectors This loss calculates the similarity between the features of each new class sample and the class prototype vector, encouraging feature pairs with high similarity (i.e. positive pairs) to be closer and feature pairs with low similarity (i.e. negative pairs) to be farther away. In this way, features of the same class can form compact clusters in the feature space, effectively alleviating the problem caused by large intra-class differences.

[0033] Please refer to Figure 4 In some embodiments, in order to make the features of different classes more distinct in the feature space, an inter-class distance loss constraint is set, and the category of the target in the input image is adjusted according to the set inter-class distance loss constraint and the class prototype. The inter-class distance loss constraint can be regarded as a repulsive force, the purpose of which is to separate the clusters of each class from each other as much as possible, to constrain easily confused categories to stay away from each other, and to divide more obvious feature boundaries in the feature space. For each new class target corresponding to the bounding box in the input image, the category of the target is determined according to the new class label. The corresponding class prototype is determined according to the category of the target. The inter-class distance loss can be calculated according to Formula 3), that is, the class prototype vector of each new class is calculated. Class prototype vectors with other classes The second norm of the distance, and the inverse of the second norm is taken as the inter-class distance loss .

[0034] 3) in, represents the total number of all categories, Represents a set of categories. As can be seen from Equation 3), the inter-class distance loss decreases as the distance between each class prototype increases. This loss calculates the distance between the class prototype vector of each class and the class prototype vector of other classes, and encourages these distances to be as large as possible, so that the features of different classes can form a clear boundary in the feature space, thereby effectively alleviating the problem caused by small inter-class differences.

[0035] In some embodiments, the category of the object in the input image can be adjusted based on the intra-class distance loss and the inter-class distance loss in combination with the class prototype. In order to balance the weight between the intra-class distance loss and the inter-class distance loss, a hyperparameter needs to be introduced to adjust the relative importance of the two.

[0036] After the introduction of intra-class distance loss, the detection accuracy and recall rate of the model have been significantly improved. After the introduction of inter-class distance loss, the classification performance of the model has been significantly improved, and the ability to distinguish between easily confused categories has also been enhanced.

[0037] By introducing the supervised contrastive learning method, the model can learn the discriminative representations between different categories more specifically during the training process. After the introduction of supervised contrastive learning, the detection performance and generalization ability of the model have been significantly improved. At the same time, since supervised contrastive learning can make full use of label information to guide the learning process, it can also better adapt to the needs of different scenarios and data sets in practical applications.

[0038] According to one embodiment of the present invention, a food detection method is provided, comprising: inputting a food image to be detected into a detection model, and the detection model outputs the category of the target in the food image and the bounding box coordinates of the target. Among them, the detection model is an incremental food detection model trained according to the above-mentioned training method. By applying incremental detection technology to food detection tasks, it is possible to update and expand the model without increasing a lot of time cost and computing resource overhead. This incremental learning method enables the model to quickly adapt to the needs of new food categories and data sets, so that it has a wider application prospect in practical applications. After the introduction of incremental detection technology, the training cost of the model has been significantly reduced, and the detection performance for new categories has also been significantly improved.

[0039] The model trained by the training method proposed in the embodiment of the present invention has achieved significant improvement in detection accuracy, recall rate, classification performance, etc., and has good robustness and generalization ability in practical applications. The training method constrains the feature space by introducing intra-class distance loss and inter-class distance loss, making the features of the same class more compact and the features of different classes more distinct. At the same time, by introducing the supervised contrast learning method, the label information is fully utilized to guide the learning process, further improving the detection performance and generalization ability of the model.

[0040] Table 1. Table 1 shows the experimental setting of two-stage (40 classes in the first stage and 33 classes in the second stage) incremental detection on the food detection dataset UNIMIB2016, with and without using the inter-class distance loss. And intra-class distance loss The whole incremental training process is divided into two stages. The first stage (base class training) trains 40 classes. The model obtained after training these 40 classes is the base class model in the second stage (incremental training). It can be seen that no matter which supervised contrast learning loss function is removed, the mAP will decrease. The mAP (mean Average Precision) indicator is an important performance evaluation indicator in the detection field, which is used to measure the comprehensive performance of the target detection model. Among them, mAP1 represents the mAP indicator of the first stage training, mAP2 represents the mAP indicator of the second stage training, and mAPA represents the average mAP indicator of the above two training stages. The evaluation indicators in Table 1 prove that the two constraints on the loss of inter-class distance and intra-class distance play a role in the incremental learning process.

[0041] It should be noted that although the above describes the various steps in a specific order, it does not mean that the various steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.

[0042] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0043] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a protruding structure in a groove on which instructions are stored, and any suitable combination thereof.

[0044] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A training method for an incremental food detection model, wherein: The incremental food detection model includes a neural network for generating candidate regions of an input image and a feature map of the input image according to an input image, and a region of interest head stage, wherein the region of interest head stage is used to perform object classification and bounding box regression according to the candidate regions of the input image and the feature map of the input image; The training method comprises: Initializing the incremental food detection model using the trained base class model; Training the incremental food detection model based on a supervised contrastive learning method; The base class model is trained using a first training set, the first training set includes a first input image and labels corresponding to base class foods, and the first input image includes base class food data; in the step of training the incremental food detection model based on a supervised contrastive learning method, a second training set is used for training, the second training set includes a second input image and labels corresponding to new class foods, and the second input image includes new class food data.

2. The method according to claim 1, wherein: The incremental food detection model is trained based on a supervised contrastive learning method, including: Obtaining a category of an object in the input image according to the candidate region of the input image and the feature map of the input image; For each category, calculating the average value of the feature vectors corresponding to the category, and determining the class prototype according to the average value; The category of the object in the input image is adjusted according to the set intra-class distance loss constraint and the class prototype, and / or according to the set inter-class distance loss constraint and the class prototype.

3. The method according to claim 2, wherein: Each time the incremental food detection model is trained to reach a set iteration round, the class prototype is updated in the following manner: For each category, a feature memory of the category is set, the feature memory is used to store a feature queue corresponding to the category in each iteration round, the feature queue includes a feature vector; The class prototype is updated according to the set momentum parameter, the value of the class prototype and the average value of the feature vector of the category in the current iteration round.

4. The method according to claim 2, wherein: The category of the target in the input image is adjusted according to the set intra-class distance loss constraint and the class prototype, including: For each new class of target corresponding to the bounding box in the input image, determining the class of the target according to its corresponding label; Determining the corresponding class prototype according to the category of the target; Calculating the intra-class distance loss according to the feature vector of the target, the class prototypes corresponding to all classes, the mapping relationship of mapping the input image to the feature space, and the set temperature coefficient; The category of the object in the input image is adjusted according to the intra-class distance loss.

5. The method according to claim 2, wherein: The category of the target in the input image is adjusted according to the set inter-class distance loss constraint and the class prototype, including: For each new class of target corresponding to the bounding box in the input image, determining the class of the target according to its corresponding label; Determining the corresponding class prototype according to the category of the target; Calculate the bi-norm of the distance between the class prototype vector of each new class and the class prototype vectors of other classes, and take the inverse of the bi-norm as the inter-class distance loss; The category of the object in the input image is adjusted according to the inter-class distance loss.

6. The method according to claim 1, wherein: The neural network for generating a candidate region of an input image and a feature map of the input image according to the input image comprises a backbone network and a region proposal network; The backbone network is used to extract features from the input image and output a feature map of the input image; The region proposal network is used to generate a candidate region on the feature map, where the candidate region is a region including an object.

7. The method according to claim 6, wherein: The backbone network includes one of VGG16, ResNet and ResNeXt.

8. A food testing method, wherein: include: Inputting a food image to be detected into a detection model, the detection model outputting the category of an object in the food image and the bounding box coordinates of the object; Wherein, the detection model is an incremental food detection model trained according to the method according to any one of claims 1 to 7.

9. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.