Foreign object detection method and apparatus, non-transitory computer readable storage medium

The target detection model trained by self-supervised and supervised methods solves the problem of low efficiency in foreign object recognition in containers, and achieves efficient and accurate foreign object detection, which is suitable for terminal scenarios such as store refrigerators.

CN116630947BActive Publication Date: 2026-01-20BOE TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310220476.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2026-01-20
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficiently identifying and monitoring foreign objects in containers that do not belong to the container, and often rely on manual comparison, which is inefficient.

Method used

A self-supervised method is used to pre-train the teacher model, and a supervised method is used to fine-tune the teacher model. Knowledge distillation is then performed to obtain the object detection model. Feature vectors are extracted and processed to determine whether the object within the object detection box is a foreign object.

Benefits of technology

While ensuring accurate identification of product categories, it can efficiently detect foreign objects in the container, reduce the need for labeled data, improve identification accuracy, and is suitable for terminal scenarios with low computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630947B_ABST
    Figure CN116630947B_ABST
Patent Text Reader

Abstract

A foreign matter detection method and device, and a non-transitory computer readable storage medium, comprising: inputting a first image into a target detection model to obtain one or more target detection boxes; the target detection model is obtained by training as follows: pre-training a teacher model by a self-supervised method, and fine-tuning the teacher model by a supervised method; knowledge distillation is performed on a student model by using the teacher model to obtain the target detection model; a feature vector of each target detection box is extracted; and processing the extracted feature vector to determine whether an object in each target detection box is a foreign matter.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to, but are not limited to, the technical field of object detection, and in particular to a foreign object detection method and device, and a computer readable storage medium. BACKGROUND

[0002] Convolutional neural networks have strong learning ability and efficient feature expression ability, and play a huge advantage in computer vision tasks such as segmentation, detection, recognition, and tracking. In the field of intelligent retail, by using computer vision technology, through the camera installed on the refrigerator in the store, automatic vending machine, etc., the number and type of goods on the vending machine can be automatically identified, which can greatly save labor costs and improve customer shopping experience.

[0003] However, in some cases, some consumers may put items that do not belong to the vending machine, such as garbage or goods from other locations, on the vending machine. This situation is currently difficult to monitor. Merchants usually can only determine whether there are foreign objects on the vending machine by manual comparison. This method is tedious and inefficient. SUMMARY

[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0005] Embodiments of the present disclosure provide a foreign object detection method, comprising:

[0006] inputting a first image into an object detection model to obtain one or more target detection boxes; the object detection model is obtained by training through the following method: pre-training a teacher model through a self-supervised method, and fine-tuning the teacher model through a supervised method; using the teacher model to perform knowledge distillation on a student model to obtain the object detection model;

[0007] extracting a feature vector of each of the target detection boxes;

[0008] processing the extracted feature vectors to determine whether the objects in each of the target detection boxes are foreign objects.

[0009] Embodiments of the present disclosure also provide a foreign object detection device, comprising a memory; and a processor connected to the memory, the memory is used to store instructions, and the processor is configured to execute the steps of the foreign object detection method according to any one of the embodiments of the present disclosure based on the instructions stored in the memory.

[0010] Embodiments of the present disclosure also provide a non-transitory computer readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the foreign object detection method according to any one of the embodiments of the present disclosure.

[0011] Other aspects can become apparent from a review of the drawings and detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and serve to explain the principles of the present disclosure, and should not be considered limiting of the present disclosure. The shapes and sizes of the components in the drawings are for illustrative purposes only and are not meant to be limiting. The drawings are used to describe and explain the present disclosure.

[0013] Figure 1 A flowchart of a foreign matter detection method provided for an exemplary embodiment of the present disclosure;

[0014] Figure 2 An exemplary application scenario schematic diagram of an embodiment of the present disclosure is shown;

[0015] Figure 3 A structure schematic diagram of a target detection model provided for an exemplary embodiment of the present disclosure;

[0016] Figure 4 A knowledge distillation process schematic diagram provided for an exemplary embodiment of the present disclosure;

[0017] Figure 5 A commodity image provided for an exemplary embodiment of the present disclosure;

[0018] Figure 6 For Figure 5 A partial template image schematic diagram of the commodity image shown;

[0019] Figure 7 A structure schematic diagram of a foreign matter detection system provided for an exemplary embodiment of the present disclosure;

[0020] Figure 8 A structure schematic diagram of a foreign matter detection device provided for an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] In order to make the objects, technical solutions and advantages of the present disclosure clearer, the following will combine the drawings to make a detailed description of the embodiments of the present disclosure. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other in any manner without conflict.

[0022] Unless otherwise defined, technical terms and scientific terms used in the present disclosure shall have the meanings as understood by one of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects.

[0023] As shown in Figure 1 The present disclosure provides a foreign object detection method, comprising the following steps:

[0024] Step 101, inputting a first image into a target detection model to obtain one or more target detection boxes; the target detection model is obtained by training through the following method: pre-training a teacher model through a self-supervised method, fine-tuning the teacher model through a supervised method; performing knowledge distillation on a student model using the teacher model to obtain the target detection model;

[0025] Step 102, extracting a feature vector of each target detection box;

[0026] Step 103, processing the extracted feature vector to determine whether the object in each target detection box is a foreign object.

[0027] The foreign object detection method provided by the present disclosure can learn sufficient visual feature expression by pre-training a teacher model through a self-supervised method, fine-tuning the teacher model through a supervised method, performing knowledge distillation on a student model using the teacher model to obtain the target detection model, processing the feature vector of the target detection box output by the target detection model in the detection process, obtaining the foreign object detection result, ensuring correct identification of the product category, detecting the foreign object, and applying it to the terminal and other low-power scenarios.

[0028] The present disclosure has many application scenarios, which can be but are not limited to a store refrigerator, an automatic vending cabinet, etc. For example, for the refrigerator application scenario as shown in Figure 2 The foreign object detection method of the present disclosure can detect the products that do not belong to the refrigerator (for example, assuming that the foreign object in Figure 2 is a cold tea drink, and the foreign object detection result can be boxed with a rectangular dashed line, and other products can be boxed with a rectangular solid line). The product category in the refrigerator is limited, and the category is known, but the foreign object category is very large and unknown, and the present disclosure uses the idea of open set recognition to detect the foreign object.

[0029] In the embodiments of the present disclosure, for other non-foreign objects or goods, the objects or goods can be uniformly identified as objects or goods, or can be identified as multiple different categories, and the present disclosure does not limit this.

[0030] In the embodiments of the present disclosure, before the teacher model is used to perform knowledge distillation on the student model, the student model can also be pre-trained using a self-supervised method, thereby further improving the ability of the student model.

[0031] In Figure 2 In the application scenario shown, an angle sensor and an image acquisition device (for example, the image acquisition device can be a camera, which is directed towards the inside of the refrigerator) can be arranged on the refrigerator door, a target detection model and a foreign object detection device can be arranged on a terminal, and the foreign object detection device can acquire a rotation signal of the angle sensor; a rotation angle is calculated according to the rotation signal of the angle sensor, when the rotation angle is within a preset threshold range (i.e., after the refrigerator door is opened to a certain angle), the foreign object detection device instructs the image acquisition device to acquire an image, receives the image acquired by the image acquisition device, selects one or more images, inputs the selected image into the target detection model, obtains one or more target detection boxes, and extracts a feature vector of each target detection box; the extracted feature vector is processed to determine whether the object in each target detection box is a foreign object. Figure 2 In the embodiments of the present disclosure, the solid line box indicates that the object belongs to the goods inside the refrigerator, and the dashed line box indicates "other", which is a foreign object. The foreign object detection method of the present disclosure can detect foreign objects while ensuring correct identification of the classification of goods.

[0032] In the embodiments of the present disclosure, the target detection model is trained by the following method:

[0033] The teacher model is pre-trained by a self-supervised method;

[0034] The teacher model is fine-tuned by a supervised method;

[0035] The student model is knowledge distilled by the teacher model, and the target detection model is obtained.

[0036] In the embodiments of the present disclosure, the teacher model is pre-trained by a self-supervised learning method, which can well express visual features for downstream tasks without labeling, and can improve the ability of the pre-trained model. The data set used in the pre-training process is unlabeled image data, and the content of the unlabeled image data can be any object, i.e., it can not be goods in the final application scenario. Still taking Figure 2 For example, in the application scenario, when the teacher model is pre-trained by a self-supervised method, the content of the unlabeled image data used can be any object, and it does not have to be goods in the refrigerator.

[0037] In the commodity identification scene, since there are many similar commodities in the data set, the self-supervised method makes its own data augmentation as a positive sample and another image as a negative sample, so that the intra-class difference is reduced and the inter-class difference is increased, thereby learning the fine-grained features of the commodity and improving the identification accuracy.

[0038] In addition, the current collected data set has a large number of commodity images that can be cropped through detection, but there are not so many manpower to label the commodity categories. The present disclosure uses self-supervised learning to pre-train the teacher model, and then fine-tunes it with labeled data, which can greatly reduce the number of pictures that need to be labeled and greatly improve the accuracy of the prediction results.

[0039] In some example embodiments, the teacher model is pre-trained by a self-supervised method, comprising:

[0040] Each original image in the unlabeled image data is subjected to a plurality of transformations to obtain a first preset number of images, wherein the first preset number is the number of transformations;

[0041] The first preset number of images are respectively subjected to the teacher model before pre-training to obtain a first preset number of first prediction values corresponding to the teacher model;

[0042] The first preset number of first prediction values are compared and learned to further self-supervise the teacher model.

[0043] In the embodiment of the present disclosure, when the teacher model is pre-trained by a self-supervised method, the method of contrast learning is used to unsupervisedly learn the representation of vision. The purpose of contrast learning is to distinguish between same and non-same samples. The self-supervised contrast loss function adopts an infonce noise contrastive estimation (InfoNCE) loss function, and the expression is as follows:

[0044]

[0045] Wherein, q is the feature vector of the object to be searched, k + is the positive sample feature vector, k i is the negative sample feature vector, wherein q and k + are the same sample generated by data augmentation, and both belong to the same class. Then, the distance between q and the positive sample is narrowed and the distance between the negative samples is pushed away through contrast loss.

[0046] In some example embodiments, as Figure 3As shown, the target detection model includes a backbone, a fully connected layer and an output layer, wherein the backbone is used to extract features of an input image, the fully connected layer is used to predict the position of the target and the class score, and the output layer is used to output the classification label.

[0047] In some example embodiments, the teacher model includes a first backbone network, and the student model includes a second backbone network. The first backbone network can be a ReXNet network, and the second backbone network can be a MobileNeXt network. A SandGlass Bottleneck structure unique to the MobileNeXt network can greatly reduce the parameters of the student model. However, the present disclosure is not limited thereto, and the structures of the first backbone network and the second backbone network can be set as needed.

[0048] In the present disclosure, the fully connected layer can include one or more layers. For example, the fully connected layer can include two layers. However, the present disclosure is not limited thereto.

[0049] In some example embodiments, the knowledge distillation is feature distillation, and the total loss function of the student model is: Loss = Loss1 + lamda * Loss2, wherein lamda is a preset weight coefficient, Y is the true value, and cosface loss represents a cosine face recognition Cosface loss function; Loss2 = MSE (Feature_student, Feature_teacher), Feature_student is a feature vector of the student model, Feature_teacher is a feature vector of the teacher model, and MSE represents a mean square error loss function.

[0050] In the present disclosure, in order to improve the detection ability of the model, a large model (i.e., the teacher model) is pre-trained and fine-tuned first, and then a small model (i.e., the student model) is distilled. In the distillation process, only the features are distilled, so that the small model can fully learn the ability of the large model.

[0051] As Figure 4As shown, the present disclosure uses feature distillation, on the one hand, let the student model fit the soft label information output by the teacher model, so that the student model can learn some potential semantic information, and summarize the experience of the teacher model. When distilling, for the training sample X (the real label is Y), the feature vector Feature teacher of the teacher model is obtained after inputting the teacher model, and the feature vector Feature student of the student model is obtained after inputting the student model. The loss between Feature teacher and Feature student is calculated using MSEloss; on the other hand, the predicted value of the student model and the real hard label are used to calculate a cross-entropy loss, to understand the difference between the real data, and the classification recognition loss uses a cosface loss function. The two losses are added by a weight lamda to form the total loss.

[0052]

[0053] Loss2 = MSE (Feature student, Feature teacher)

[0054] The total training loss is Loss = Loss1 + lamda * Loss2.

[0055] In the feature distillation of the embodiments of the present disclosure, only one parameter needs to be adjusted: the weight lamda, and the temperature parameter T does not need to be adjusted. In the training process, adaptive sharpness aware minimization (ASAM) can be used for generalization. Although the general training method can converge to a good local optimum, the vicinity of the local optimum is usually very rough, that is, adding a small perturbation ωw+∈ to the weight ωw may lead to poor results. Therefore, the embodiments of the present disclosure use ASAM to adaptively adjust the weight perturbation range, take the model with the maximum loss, and make the model more robust.

[0056]

[0057] L S Loss (ω) = 1 |S| åx∈S L (x, ω) represents that the floating range of the weight ω is not more than ρ, and ∈ is the floating value.

[0058] In some example embodiments, the extracted feature vectors are processed to determine whether the object within each target detection frame is a foreign object, comprising:

[0059] According to the trained target detection model, the feature vector of each known category of article is extracted;

[0060] For each target detection frame, the following operations are performed:

[0061] The similarity of the extracted feature vector of the target detection frame and the feature vector of each known category of item is calculated, and whether the object in the target detection frame is a foreign object is determined according to the similarity result.

[0062] The embodiment of the present disclosure provides a method for determining whether the object in each target detection frame is a foreign object through similarity calculation. In the embodiment of the present disclosure, feature similarity is used to determine which commodity category or foreign object the object in the target detection frame belongs to. When using, the template image of M types of commodities needs to be prepared in advance, and M is the number of types of commodities to be identified. The more the types of commodities, the stronger the foreign object detection capability. When pre-training is performed by using a self-supervised learning method, any item data set can be used for pre-training, and when fine-tuning is performed, the commodity data set that can be found is used for fine-tuning. In some example embodiments, an N-type commodity data set is used for fine-tuning, and N is much larger than M.

[0063] Meaning of template image: we know in advance the M types of commodity images to be identified. Taking one of the commodity images as an example, suppose the commodity image is a Pepsi image as shown in Figure 5 When actually detecting, for each target detection frame, we obtain the image cropped from the target detection frame, extract the features, obtain Fquery, and compare Fquery with the features Fgallery of Pepsi. If the similarity of the two features is high, it means that the item detected by the target detection frame is Pepsi, that is, the Pepsi image as shown in Figure 5 is a template image of Pepsi. In actual use, images of Pepsi from different angles can be obtained as template images, as shown in Figure 6 , that is, the template image covers different angles as much as possible. The calculation formula for measuring feature similarity is:

[0064]

[0065] In some example embodiments, according to the similarity result, whether the object in the target detection frame is a foreign object is determined, including:

[0066] determining the maximum value in the calculated similarity;

[0067] comparing the determined maximum value with a preset similarity threshold value;

[0068] when the determined maximum value is greater than or equal to the preset similarity threshold value, determining that the object in the target detection frame is the known category of item corresponding to the determined maximum value;

[0069] when the determined maximum value is less than the preset similarity threshold value, determining that the object in the target detection frame is a foreign object.

[0070] Assuming that the i-th type of goods in the M type of goods has q i image as a template image, according to the trained target detection model, the features of each type of goods are extracted in advance, a total of features, when using, the features Fquery of the target detection box to be predicted are extracted, the cosine distances of the features Fquery and features are calculated respectively, and the closest cosine distance is found as the similarity calculation result, when the similarity calculation result is greater than or equal to the preset similarity threshold, it is determined that the object in the target detection box is the most similar feature corresponding to the type of goods, and when the similarity calculation result is less than the preset similarity threshold, it is determined that the object in the target detection box is a foreign object.

[0071] In some other exemplary embodiments, the extracted feature vectors (here, the extracted feature vectors are the softmax scores of each known category) are processed to determine whether the object in each target detection box is a foreign object, including:

[0072] According to the trained target detection model, determine the Weibull distribution model of each known category of goods;

[0073] For each target detection box, the following operations are performed:

[0074] Obtain the softmax score of the target detection box belonging to each known category; calculate the distance between the feature vector of the target detection box and the centroid of each known category, determine the probability of the target detection box belonging to each known category according to the calculated distance, and correct each softmax score using the calculated probability; calculate the probability of the object in the target detection box being a foreign object according to the corrected score.

[0075] The embodiments of the present disclosure also provide a method for determining whether the object in each target detection box is a foreign object through softmax score correction. Through the aforementioned self-supervised learning, learning with other goods category data can obtain a pre-trained model that can well express visual representation, and then the pre-trained teacher model is fine-tuned through a supervised method. Still taking the application scenario shown in the foregoing Figure 2 The application scenario is an example of placing an image acquisition device on the door of a refrigerator, and the image acquisition device acquires images inside the refrigerator to determine whether there are any items that do not belong to the category of items in the refrigerator. Assuming that the refrigerator stores M types of goods, M types of sample data can be actually collected for these categories, and when training the M type data, the model is fine-tuned to obtain a teacher model. The foreign object detection method of the present disclosure improves the ability of foreign object recognition without affecting the accuracy of normal recognition of goods.

[0076] In the similarity calculation mode, the feature vector extracted is the feature vector extracted from the layer before the softmax layer; in the softmax score correction mode, the feature vector extracted is the softmax score of each known category extracted from the softmax layer.

[0077] The general target detection model directly uses the result after the softmax as the final prediction, but the present disclosure needs to detect whether other goods not belonging to the M categories exist, and the embodiments of the present disclosure correct the softmax score by the following method to obtain the probability that the goods belong to the foreign matter.

[0078] Step 1, we have obtained the goods classification model through the aforementioned self-supervised learning and fine-tuning process. The M-class goods (P1, P2, Pi, … PM) data are collected to obtain the feature vector before the softmax. Taking the i-th class of goods as an example, for the goods in the i-th class, the feature F of the second-to-last layer of the goods is retained, and the total number of goods in the i-th class is s. Then the feature Fi set of the i-th class is {F1, F2, … Fs}, and the mean value of Fi is calculated as the centroid, denoted as mFi.

[0079]

[0080] The distance of each element in Fi to the centroid mF is calculated, denoted as Di, and the extreme value distribution in Di is fitted using the Weibull distribution, and the Weibull distribution formula is as follows:

[0081]

[0082] where x is a random variable, λ is a scale parameter, and k is a shape parameter, λ > 0, k > 0. The λ and k of each goods category are obtained by fitting calculation, and a Weibull distribution can be obtained.

[0083] Step 2, for the sample x input during inference, Fx is calculated, and after the softmax, the score of belonging to each class is calculated score = {score1, score2, …, score m}; the distance {d1, d2, di…dm} of Fx to each class centroid mF is calculated, and the maximum value appears in (-∞, d i i of the Weibull probability density function, and the maximum value appears in (-∞, di the probability p of the sample belonging to the i-th class i the probability p of the sample belonging to the i-th class i the probability p of the sample belonging to the i-th class i the probability p of the sample belonging to the i-th class i the probability p of the sample belonging to the i-th class i the probability p of the sample belonging to the i-th class

[0084] Step 3, correcting the score, the new score of each class after correction is recorded as new_score = {w1*score1, … w i *score i , w m *score m}, and the score probability of the foreign matter class is

[0085]

[0086] Compare all values in score foreignmatter and new_score, and take the class corresponding to the maximum score probability as the class corresponding to the object (foreign matter or commodity).

[0087] The two ways of determining whether the object in each target detection frame is foreign matter (similarity calculation and softmax score correction) are usually used separately. In some example embodiments, the two ways of determining whether the object in each target detection frame is foreign matter can be used jointly. For example, when the first image is obtained, if it is determined to be foreign matter by both of the above two ways, it is determined that the object in the target detection frame is foreign matter. If the results determined by the above two ways are different, i.e., it is determined to be foreign matter by one way and it is determined not to be foreign matter by the other way, several images can be obtained and detected again, or one of the two ways, such as the similarity calculation way, is used preferentially to determine whether the object in the target detection frame is foreign matter.

[0088] In some example embodiments, the method further comprises, before the method:

[0089] obtaining a rotation signal of an angle sensor;

[0090] calculating a rotation angle according to the rotation signal of the angle sensor;

[0091] when the calculated rotation angle is within a preset range, obtaining an image collected by an image collection device;

[0092] selecting one or more images of the images collected by the image collection device as the first image.

[0093] In some example embodiments, the method further comprises, before the method:

[0094] obtaining a rotation signal of an angle sensor;

[0095] calculating a rotation angle according to the rotation signal of the angle sensor;

[0096] when the calculated rotation angle is within a preset range, obtaining an image collected by an image collection device;

[0097] performing a rotation target detection on the image collected by the image collection device to obtain one or more rotation detection boxes, and taking the rotation detection boxes as a first image.

[0098] Still taking the application scenario of the foregoing Figure 2 , as an example, the foreign matter detection device obtains the angle sensor value to determine the opening angle of the refrigerator door. Since the foreign matter detection needs to see the overall appearance of the refrigerator interior, if the opening angle is too small, the collected image may be distorted; if the opening angle is too large, the body part of a person may be inserted into the refrigerator interior to block the goods, so the refrigerator interior image collected by the image collection device needs to be obtained when the opening angle of the refrigerator door is within a certain threshold range. The image collected by the image collection device is subjected to a rotation target detection to obtain one or more rotation detection boxes, and then a target detection model is used to perform a target detection on the image of each rotation detection box, and then the foreign matter detection and recognition are performed on the target detection boxes, and the final recognition result is sent to a Web end for display or an audible, visual, and electrical alarm is performed when the foreign matter is detected.

[0099] In the embodiments of the present disclosure, the product recognition accuracy can be improved by adding some strategies. For the same product, if the color and shape difference is not large, people are difficult to distinguish, and only the taste is different, then it can be identified as a category (correspondingly, in fine-tuning and making a template image, the data set and the template image can be classified into one category); for the same product, if the shape difference is not large, only the color is different, then it is also identified as two categories (correspondingly, in fine-tuning and making a template image, the data set and the template image can be classified into two categories).

[0100] As Figure 7 shown, the embodiments of the present disclosure also provide a foreign matter detection system, comprising a foreign matter detection device, an image collection device, and an angle sensor, wherein,

[0101] the image collection device is configured to collect an image;

[0102] the angle sensor is configured to detect a rotation signal;

[0103] The foreign matter detection device is configured to acquire a rotation signal of an angle sensor, calculate a rotation angle based on the rotation signal of the angle sensor, acquire an image captured by an image acquisition device when the calculated rotation angle is within a preset range, input the captured image into a target detection model to obtain one or more target detection boxes, extract a feature vector of each target detection box, and determine whether an object in each target detection box is a foreign matter by processing the extracted feature vector.

[0104] In some example embodiments, the foreign matter detection device is configured to acquire a rotation signal of an angle sensor, calculate a rotation angle based on the rotation signal of the angle sensor, acquire an image captured by an image acquisition device when the calculated rotation angle is within a preset range, perform rotation target detection on the image captured by the image acquisition device to obtain one or more rotation detection boxes, input the rotation detection boxes into a target detection model to obtain one or more target detection boxes, extract a feature vector of each target detection box, and determine whether an object in each target detection box is a foreign matter by processing the extracted feature vector.

[0105] The foreign matter detection device provided by the embodiments of the present disclosure further includes a memory and a processor connected to the memory, the memory is configured to store instructions, and the processor is configured to execute the steps of the foreign matter detection method based on the instructions stored in the memory.

[0106] As shown in FIG. 8, Figure 8 In one example, the foreign matter detection device can include a processor 810, a memory 820, and a bus system 830, wherein the processor 810 and the memory 820 are connected through the bus system 830, the memory 820 is configured to store instructions, and the processor 810 is configured to execute the instructions stored in the memory 820 to input a first image into a target detection model to obtain one or more target detection boxes, train the target detection model by pre-training a teacher model by a self-supervised method and fine-tuning the teacher model by a supervised method, perform knowledge distillation on a student model by using the teacher model to obtain the target detection model, extract a feature vector of each target detection box, and determine whether an object in each target detection box is a foreign matter by processing the extracted feature vector.

[0107] It should be understood that the processor 810 can be a central processing unit (CPU), and the processor 810 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), programmable logic devices (PLD), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0108] The memory 820 can include a read-only memory and a random access memory, and provide instructions and data for the processor 810. A part of the memory 820 can also include a non-volatile random access memory. For example, the memory 820 can also store device type information.

[0109] The bus system 830 can include a data bus in addition to a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, all the buses are marked as the bus system 830 in the Figure 8

[0110] In the implementation process, the processing performed by the processing device can be completed by the integrated logic circuit of the hardware in the processor 810 or the instructions in the form of software. That is, the method steps of the embodiments of the present disclosure can be embodied as hardware processor execution, or combined with hardware and software modules in the processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, or the like storage medium. The storage medium is located in the memory 820, and the processor 810 reads the information in the memory 820 and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0111] The embodiments of the present disclosure also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the foreign matter detection method according to any embodiment of the present disclosure.

[0112] In some possible implementation manners, various aspects of the foreign matter detection method provided by the present disclosure can also be implemented in the form of a program product, which includes program codes for causing a computer device to execute the steps in the foreign matter detection method according to various exemplary embodiments of the present disclosure described above in the specification, for example, the computer device can execute the foreign matter detection method described in the embodiments of the present disclosure.

[0113] ​The program product can employ any combination of one or more computer readable media. The computer readable media can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0114] Those of ordinary skill in the art will realize and understand that all or certain steps in the methods disclosed above and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Certain components or all components can be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer-readable media, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer. Furthermore, it is common and well understood by those of ordinary skill in the art that communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and can include any information delivery media.

[0115] Although the embodiments disclosed by the present disclosure are as above, the content described is only the embodiments adopted for the convenience of understanding the present disclosure, and is not intended to limit the present disclosure. Any person skilled in the art can make any modification and change in the form and details without departing from the spirit and scope of the present disclosure, but the patent protection scope of the present disclosure shall be subject to the scope defined by the appended claims.

Claims

1. A method for detecting foreign objects, characterized in that, include: The first image is input into the object detection model to obtain one or more object detection boxes; the object detection model is trained by the following method: a teacher model is pre-trained using a self-supervised method, and the teacher model is fine-tuned using a supervised method; the teacher model is used to perform knowledge distillation on the student model to obtain the object detection model; Extract the feature vector of each of the target detection boxes; The extracted feature vectors are processed to determine whether the object within each target detection box is a foreign object. This processing includes: determining the Weibull distribution model for each known category of item based on a trained target detection model; for each target detection box, performing the following operations: obtaining the softmax score of the target detection box belonging to each known category; calculating the distance from the feature vector of the target detection box to the centroid of each known category, determining the probability that the target detection box belongs to each known category based on the calculated distance, and correcting each softmax score using the calculated probability; and calculating the probability that the object within the target detection box is a foreign object based on the corrected score.

2. The method according to claim 1, characterized in that, The process of processing the extracted feature vectors to determine whether an object within each target detection box is a foreign object includes: Based on the trained object detection model, extract the feature vector for each known category of item; For each of the target detection boxes, perform the following operations: The similarity between the feature vector of the extracted target detection box and the feature vector of each known category of item is calculated, and the object within the target detection box is determined as a foreign object based on the similarity result.

3. The method according to claim 2, characterized in that, Determining whether the object within the target detection box is a foreign object based on the similarity result includes: Determine the maximum value among the calculated similarities; Compare the determined maximum value with the preset similarity threshold; When the determined maximum value is greater than or equal to the preset similarity threshold, the object within the target detection box is determined to be a known category item corresponding to the determined maximum value; When the determined maximum value is less than the preset similarity threshold, the object within the target detection box is determined to be a foreign object.

4. The method according to claim 1, characterized in that, The teacher model includes a first backbone network and a first feature extraction layer, and the student model includes a second backbone network and a second feature extraction layer. The first backbone network is a ReXNet network, and the second backbone network is a MobileNeXt network.

5. The method according to claim 1, characterized in that, The knowledge distillation is feature distillation, and the total loss function of the student model is: Loss = Loss1 + lamda * Loss2, where lamda is a preset weight coefficient, and Loss1 = cosfaceloss( Y), cosfaceloss represents the cosine loss function for face recognition. Y is the predicted value of the student model, and Y is the true value; Loss2 = MSE(Feature_student, Feature_teacher), where MSE represents the mean squared error loss function, Feature_student is the feature vector of the student model, and Feature_teacher is the feature vector of the teacher model.

6. The method according to claim 1, characterized in that, The method is preceded by: Acquire the rotation signal from the angle sensor; The rotation angle is calculated based on the rotation signal from the angle sensor; When the calculated rotation angle is within the preset range, the image acquired by the image acquisition device is obtained; One or more images acquired by the image acquisition device are selected as the first image.

7. The method according to claim 1, characterized in that, The method is preceded by: Acquire the rotation signal from the angle sensor; The rotation angle is calculated based on the rotation signal from the angle sensor; When the calculated rotation angle is within the preset range, the image acquired by the image acquisition device is obtained; The image acquired by the image acquisition device is subjected to rotation target detection to obtain one or more rotation detection boxes, and the rotation detection boxes are used as the first image.

8. A foreign object detection device, characterized in that, The method includes a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the foreign object detection method as described in any one of claims 1 to 7 based on the instructions stored in the memory.

9. A non-transient computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the foreign object detection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Agricultural pest image recognition method based on CNN few samples

    CN113177612A

  • Foreign matter detection method and device, electronic equipment and storage medium

    CN114724025A