A method and device for detecting dynamic foreign matter in the kitchen based on surveillance video
By adopting a detection method based on surveillance video in the dynamic foreign object detection in the kitchen, the combination of pixel difference value and the YOLOv7 model is used to solve the problems of low detection accuracy and high cost caused by the complexity of the kitchen environment, and efficient and accurate foreign object detection is achieved.
Patent Information
- Application Number
- CN202411818116.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-11
AI Technical Summary
In the dynamic foreign matter detection of kitchens, the prior art results in low detection accuracy and high cost due to complex environment, large lighting changes, high model error detection rate, and large computing resource requirements.
The dynamic foreign object detection method of the kitchen based on surveillance video is adopted, and the image set is extracted by preset frame extraction intervals, the pixel difference value of adjacent images is calculated, the image is denoised smoothly, the image changes are judged, and the change areas are input into the improved YOLOv7 model for detection.
It improves detection accuracy, reduces system computing power consumption, reduces error detection rate, reduces detection cost, and improves monitoring efficiency.
Smart Images

Figure CN119296047B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a method and device for detecting dynamic foreign matter in a kitchen based on surveillance video. Background Art
[0002] Today, the catering market is booming, and food safety is a top priority of people's livelihood issues. The standardization of the kitchen environment is an important factor affecting food safety. The appearance of animals such as cats, dogs, and mice in the kitchen environment will greatly affect the sanitary environment of the kitchen, thereby reducing the level of food safety.
[0003] At present, the supervision method mainly based on manual on-site inspection is difficult to ensure the hygiene of the back kitchen of the shop, which leads to a decline in food safety level. Therefore, building a dynamic foreign body detection system in the back kitchen based on target detection as a new food safety supervision method will help improve food safety problems and improve food safety levels.
[0004] As an application of target detection technology in the kitchen environment, dynamic foreign body detection in the kitchen mainly uses computer vision, deep learning, image processing and other technologies to detect targets such as cats, dogs, mice, cigarettes, etc. in the image that affect kitchen hygiene. Therefore, animal target detection technology plays a key role in building a dynamic foreign body detection system in the kitchen.
[0005] Existing animal target detection methods can be mainly divided into two categories, namely traditional image-based methods and deep learning-based methods. Traditional image-based methods rely on manually designed features and machine learning algorithms, which can achieve target detection in simple backgrounds, but have limited accuracy and generalization capabilities in complex scenes. In contrast, deep learning-based methods not only improve detection accuracy by automatically learning features, but also enhance the model's adaptability to complex scenes and diverse targets, thereby improving the overall generalization ability, but its accuracy still needs to be improved.
[0006] Moreover, although animal target detection technology is widely used in multiple scenarios such as smart breeding, smart agriculture, and wildlife protection; through real-time monitoring of animal behavior and health status, abnormal behavior can be discovered in time, so that corresponding preventive measures can be taken; through intelligent monitoring systems, the activities of wild animals in farmland can be monitored, and measures can be taken in time to protect crops and reduce economic losses; through monitoring of wild animal activities, scientific researchers and conservation workers can better understand the ecological habits and migration paths of wild animals, thereby protecting their habitats and preventing illegal hunting. Animal target detection technology has gradually become a research hotspot in many fields such as computer vision, agriculture, biology, and ecology. However, in the field of kitchen hygiene detection, the application of animal target detection technology has not been fully developed and utilized.
[0007] However, there are still the following difficulties in applying animal target detection technology to dynamic foreign body detection in the kitchen: First, the accuracy of small target detection is low. Dynamic foreign body detection in the kitchen is mainly applied to the kitchen monitoring screen. In the monitoring screen, rats and other animals that affect the sanitary environment of the kitchen are usually small targets. Since small targets account for a small proportion of the picture and provide little information, it is difficult for the model to learn enough feature details from them during the training process, which affects the performance of the detection model. Second, the model's false detection rate needs to be reduced. The kitchen monitoring images have problems such as complex background, large lighting changes, and motion blur, which increase the difficulty of target detection. For example, for the kitchen monitoring screen, the light may fluctuate greatly due to changes in natural light and lighting; there are many items in the kitchen, such as cooking utensils, ingredients, lockers, food residues, garbage, etc. These objects may be similar to animal features and easily cause interference; fast-moving animals in the kitchen environment, such as rats, may cause image blur and affect detection accuracy. Third, the demand for computing resources is high. Current deep learning-based target detection algorithms, such as YOLOv7, often require a lot of computing resources to ensure efficient reasoning. This type of model contains complex network structures and a large number of parameters, and has a high dependence on GPU computing power. In practical applications, especially in real-time monitoring systems, continuous high computing power requirements may lead to increased costs and correspondingly higher requirements for hardware facilities, which poses a challenge to some application scenarios with limited resources.
[0008] To sum up, in the field of dynamic foreign objects in the kitchen, since the detection target is small and moves fast, the kitchen environment can easily interfere with the collected images, resulting in a high incidence of false detection and missed detection and poor detection accuracy; and real-time monitoring requires a large amount of calculation, and has high requirements for computing resources and hardware, which can easily cause waste of resources and increase detection costs. Summary of the invention
[0009] To this end, the technical problem to be solved by the present invention is to overcome the problems of low detection accuracy and high cost in the prior art due to the particularity of the back kitchen environment and real-time monitoring.
[0010] In order to solve the above technical problems, the present invention provides a method for detecting dynamic foreign matter in a kitchen based on monitoring video, comprising:
[0011] Based on a preset frame extraction interval, multiple frames of images are extracted from the surveillance video of the kitchen to be inspected to form an image set to be inspected;
[0012] Calculate the pixel difference between the corresponding pixel points in the adjacent previous image and the next image in the image set to be detected, and obtain a difference image;
[0013] Smoothing and denoising the difference image to obtain an optimized difference image;
[0014] Calculate the sum of the pixel values of all pixels in the optimized difference image and compare it with the preset activation threshold:
[0015] If the sum of the pixel values of all pixels in the optimized difference image is not greater than the preset activation threshold, the next image in the current adjacent image is determined to be a static image, and the optimized difference image of the next group of adjacent images is determined;
[0016] If the sum of the pixel values of all pixels in the optimized difference image is greater than the preset activation threshold, the next image in the current adjacent image is determined to be a dynamic image, and the change area of the next image is obtained, including:
[0017] Initialize the row decision vector to a 1×N all-zero vector, and initialize the column decision vector to an M×1 all-zero vector, where M and N are the number of rows and columns of pixels in the optimized difference image, respectively;
[0018] Based on the pixel value at each pixel point in the optimized difference image and the preset change threshold, the row decision vector and the column decision vector are updated to obtain the target row decision vector and the target column decision vector;
[0019] Construct a decision matrix based on the target row decision vector and the target column decision vector;
[0020] Obtain the dot product result of the decision matrix and the subsequent image as the change area of the subsequent image;
[0021] The changed area of the subsequent image is input into the trained YOLOv7 model to obtain the dynamic foreign object category labels and corresponding confidence levels in multiple target detection frames in the subsequent image, and complete the detection of dynamic foreign objects in the kitchen at the moment of the subsequent image.
[0022] Preferably, after obtaining the dynamic foreign body category labels and corresponding confidences in multiple target detection frames in the next image, for each target detection frame, the following steps are performed:
[0023] Compare the confidence level of the target detection box with the preset confidence threshold:
[0024] If the confidence level corresponding to the target detection frame is not greater than the first confidence level threshold, it is determined that there is no dynamic foreign object in the kitchen at the time of the next image;
[0025] If the confidence level corresponding to the target detection frame is greater than the first confidence threshold, it is determined that there is a dynamic foreign object in the kitchen at the time of the next image, and based on the corresponding dynamic foreign object category label, it is determined whether the dynamic foreign object category is a category that requires an alarm. If it is a category that requires an alarm, the confidence level corresponding to the target detection frame is compared with the second confidence threshold:
[0026] If the confidence level corresponding to the target detection frame is greater than the second confidence threshold, an alarm is triggered;
[0027] If the confidence level corresponding to the target detection frame is not greater than the second confidence threshold, a single category detection is performed on the subsequent image;
[0028] The second confidence threshold is greater than the first confidence threshold.
[0029] Preferably, performing single-category detection on the latter image includes:
[0030] Input the latter image into the single-category detector to obtain the single-category confidence corresponding to the target detection box and compare it with the third confidence threshold:
[0031] If the single-category confidence corresponding to the target detection box is greater than the third confidence threshold, an alarm is triggered;
[0032] If the single-category confidence corresponding to the target detection box is not greater than the third confidence threshold, it is determined that there is a foreign object, but no alarm is required and the detection is terminated.
[0033] Preferably, the YOLOv7 model is an improved YOLOv7 model that introduces an efficient multi-scale attention mechanism module between the backbone network and the detection head, including:
[0034] Input the changed area of the image into the backbone network for feature extraction and output the feature tensor;
[0035] Input the feature tensor into the efficient multi-scale attention mechanism module and output a multi-scale feature tensor;
[0036] The multi-scale feature tensor is input into the detection head for decoding and classification to obtain the image detection result.
[0037] Preferably, the training process of the YOLOv7 model includes:
[0038] Obtain multiple kitchen images and perform label calibration to obtain the true label corresponding to each kitchen image and construct a training data set;
[0039] Train the YOLOv7 model based on the training data set until the improved loss function converges to obtain the improved YOLOv7 model;
[0040] Among them, the improved loss function , expressed as: ;
[0041] Represents the coordinate loss calculated using SIoU, expressed as: ; Indicates the overlap rate between the predicted box and the true box; Represents the distance loss, expressed as: , Represents the angle loss between the predicted box and the real box. hour, and Respectively Axis distance loss coordinate coefficient and Axis distance loss coordinate coefficient; Represents the shape loss, and the expression is: ,when hour, and Respectively represent the width coefficient and height coefficient of shape loss, Indicates the preset shape loss concern level;
[0042] Represents the confidence loss calculated using the BCEWithLogits Loss function, expressed as: ; and Represent the predicted confidence score and true label of the kitchen image, Indicates negative or positive examples, Sigmoid activation function ;
[0043] Represents the classification loss calculated using the BCEWithLogits Loss function, expressed as: , and Respectively represent The predicted scores and true labels of the categories, Indicates The category does not exist or exists, Sigmoid activation function , Indicates the total number of categories.
[0044] Preferably, the difference image is smoothed and denoised using morphological opening and closing operations to obtain an optimized difference image, including:
[0045] The difference image is first eroded and then expanded to obtain a denoised difference image.
[0046] The denoised difference image is first dilated and then eroded to obtain the optimized difference image.
[0047] Preferably, the updating of the row decision vector and the column decision vector based on the pixel value at each pixel point in the optimized difference image and the preset change threshold, and obtaining the target row decision vector and the target column decision vector, comprises:
[0048] Compare the pixel value dif[i][j] at each pixel point in the optimized difference image dif with the preset change threshold T2;
[0049] Update the corresponding elements in the row decision vector and column decision vector of the pixel points whose pixel values are greater than the preset change threshold to 1, which is expressed as:
[0050] If dif[i][j]>T2, then mark the row decision vector row[1][j]=1 and the column decision vector column[i][1]=1;
[0051] After each pixel in the optimized difference image is judged, a target row decision vector and a target column decision vector are obtained;
[0052] Among them, the value range of i is [1, M], the value range of j is [1, N], and M and N are the number of rows and columns of pixels in the optimized difference image respectively.
[0053] This embodiment provides a dynamic foreign body detection device in a kitchen based on surveillance video, comprising:
[0054] An image set building module is used to extract multiple frames of images from the surveillance video of the kitchen to be detected based on a preset frame extraction interval to form an image set to be detected;
[0055] The difference image acquisition module is used to calculate the pixel difference between the adjacent pixels of the previous image and the next image in the image set to be detected, and obtain the difference image;
[0056] The difference image optimization module is used to smooth and denoise the difference image to obtain an optimized difference image;
[0057] The static and dynamic discrimination module is used to calculate the sum of the pixel values of all pixels in the optimized difference image and compare it with the preset activation threshold: if the sum of the pixel values of all pixels in the optimized difference image is not greater than the preset activation threshold, the next image in the current adjacent image is determined to be a static image, and the difference image of the next group of adjacent images is judged; if the sum of the pixel values of all pixels in the optimized difference image is greater than the preset activation threshold, the next image in the current adjacent image is determined to be a dynamic image;
[0058] The change region acquisition module is used to initialize the row decision vector as a 1×N all-zero vector and the column decision vector as an M×1 all-zero vector, where M and N are the number of rows and columns of pixels in the optimized difference image, respectively; based on the pixel value at each pixel in the optimized difference image and a preset change threshold, the row decision vector and the column decision vector are updated to obtain the target row decision vector and the target column decision vector; based on the target row decision vector and the target column decision vector, a decision matrix is constructed; and the dot product result of the decision matrix and the subsequent image is obtained as the change region of the subsequent image;
[0059] The detection module is used to input the changed area of the subsequent image into the trained YOLOv7 model, obtain the dynamic foreign object category labels and corresponding confidence levels in multiple target detection frames in the subsequent image, and complete the detection of dynamic foreign objects in the kitchen at the moment of the subsequent image.
[0060] Preferably, a multi-category detection module is also included, which is used to:
[0061] Compare the confidence level of the target detection box with the preset confidence threshold:
[0062] If the confidence level corresponding to the target detection frame is not greater than the first confidence level threshold, it is determined that there is no dynamic foreign object in the kitchen at the time of the next image;
[0063] If the confidence level corresponding to the target detection frame is greater than the first confidence threshold, it is determined that there is a dynamic foreign object in the kitchen at the time of the next image, and based on the corresponding dynamic foreign object category label, it is determined whether the dynamic foreign object category is a category that requires an alarm. If it is a category that requires an alarm, the confidence level corresponding to the target detection frame is compared with the second confidence threshold:
[0064] If the confidence level corresponding to the target detection frame is greater than the second confidence threshold, an alarm is triggered;
[0065] If the confidence level corresponding to the target detection frame is not greater than the second confidence threshold, a single category detection is performed on the subsequent image;
[0066] The second confidence threshold is greater than the first confidence threshold.
[0067] Preferably, a single-category detection module is further included, which is used to:
[0068] Input the latter image into the single-category detector to obtain the single-category confidence corresponding to the target detection box and compare it with the third confidence threshold:
[0069] If the single-category confidence corresponding to the target detection box is greater than the third confidence threshold, an alarm is triggered;
[0070] If the single-category confidence corresponding to the target detection box is not greater than the third confidence threshold, it is determined that there is a foreign object, but no alarm is required and the detection is terminated.
[0071] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0072] The method for detecting dynamic foreign bodies in the kitchen based on surveillance video described in the present invention is that since the surveillance images of the kitchen are static in most cases when there are no staff, the image set to be detected is obtained by presetting the frame extraction interval, thereby reducing the number of detections and improving the monitoring efficiency; and the pixel difference between two adjacent images is used to determine whether the image currently detected has changed based on the previous image. The image processing method based on image subtraction shields the static area in the image, extracts the changed area, and then sends it to the YOLOv7 model for reasoning. The extraction of the changed area effectively avoids the problem of model misdetection caused by the complex background environment, while reducing the system computing power consumption and improving the detection accuracy. When extracting the changed area in the dynamic image, the present invention constructs a decision matrix that characterizes the change of pixel points through the optimized difference image between adjacent images, and uses the decision matrix to multiply the point of the latter image to obtain the changed area in the latter image, so as to achieve accurate extraction of the changed area, thereby improving the detection efficiency and accuracy.
[0073] After acquiring the target detection frame and the corresponding confidence in the subsequent image, the present invention further sets a first confidence threshold, a second confidence threshold and a third confidence threshold, combines multi-category detection with single-category detection, further distinguishes foreign objects in the target detection frame, and determines whether to issue an alarm immediately, so as to help further reduce the probability of false detection and improve detection accuracy.
[0074] The improved YLOLv7 model of the present invention introduces the SIoU loss function and the efficient multi-scale attention mechanism module to fully utilize the information of small and medium-sized targets in the extracted changing area, improve the representation ability of the extracted small target features, and thus improve the detection accuracy of the improved YLOLv7 model. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:
[0076] Figure 1 It is a flowchart of the steps of the method for dynamic foreign body detection in the kitchen based on monitoring video provided by the present invention;
[0077] Figure 2 It is a schematic diagram of the model structure of the improved YOLOv7 network;
[0078] Figure 3 It is a schematic diagram of the structure of the efficient multi-scale attention EMA module;
[0079] Figure 4 It is a flow chart of the steps of the change region extraction algorithm;
[0080] Figure 5This is a framework diagram of a dynamic foreign body detection system in the kitchen based on a combination of multi-category detectors and single-category detectors;
[0081] Figure 6 It is a comparison chart of the detection results of the present invention and the traditional method. DETAILED DESCRIPTION
[0082] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0083] Reference Figure 1 As shown, the flowchart of the method for dynamic foreign body detection in the kitchen based on monitoring video of the present invention comprises the following specific steps:
[0084] S101: extracting multiple frames of images from the surveillance video of the kitchen to be inspected based on a preset frame extraction interval to form an image set to be inspected;
[0085] S102: Calculate the pixel difference between the corresponding pixel points in the adjacent previous image and the next image in the image set to be detected, and obtain a difference image;
[0086] S103: Smoothing and denoising the difference image to obtain an optimized difference image;
[0087] S104: Calculate the sum of the pixel values of all pixels in the optimized difference image and compare it with the preset activation threshold:
[0088] S105: If the sum of the pixel values of all the pixels in the optimized difference image is not greater than the preset activation threshold, the next image in the current adjacent image is determined to be a static image, and the optimized difference image of the next group of adjacent images is determined;
[0089] S106: If the sum of the pixel values of all pixels in the optimized difference image is greater than the preset activation threshold, determining that the next image in the current adjacent image is a dynamic image, and obtaining a change region of the next image, including:
[0090] S106-1: Initialize the row decision vector to be a 1×N all-zero vector, and initialize the column decision vector to be an M×1 all-zero vector, where M and N are the number of rows and columns of pixels in the optimized difference image, respectively;
[0091] S106-2: based on the pixel value at each pixel point in the optimized difference image and the preset change threshold, updating the row decision vector and the column decision vector to obtain the target row decision vector and the target column decision vector;
[0092] S106-3: constructing a decision matrix based on the target row decision vector and the target column decision vector;
[0093] S106-4: Obtain the dot product result of the decision matrix and the next image as the change area of the next image;
[0094] S107: Input the changed area of the subsequent image into the trained YOLOv7 model, obtain the dynamic foreign object category labels and corresponding confidences in multiple target detection frames in the subsequent image, and complete the detection of dynamic foreign objects in the kitchen at the moment of the subsequent image.
[0095] Specifically, in step S103, this embodiment uses morphological opening and closing operations to smooth and denoise the difference image to obtain an optimized difference image, including:
[0096] S103-1: performing an erosion operation on the difference image and then performing an expansion operation to obtain a denoised difference image;
[0097] S103-2: Perform a dilation operation on the denoised difference image first, and then perform an erosion operation on the denoised difference image to obtain an optimized difference image.
[0098] Specifically, in step S106-2, the acquisition of the target row decision vector and the target column decision vector includes:
[0099] Compare the pixel value dif[i][j] at each pixel point in the optimized difference image dif with the preset change threshold T2;
[0100] Update the corresponding elements in the row decision vector and column decision vector of the pixel points whose pixel values are greater than the preset change threshold to 1, which is expressed as:
[0101] If dif[i][j]>T2, then mark the row decision vector row[1][j]=1 and the column decision vector column[i][1]=1;
[0102] After each pixel in the optimized difference image is judged, a target row decision vector and a target column decision vector are obtained;
[0103] Among them, the value range of i is [1, M], the value range of j is [1, N], and M and N are the number of rows and columns of pixels in the optimized difference image respectively.
[0104] In this embodiment, a decision matrix is calculated through the target row decision vector and the target column decision vector, and point multiplication is performed with the latter image to obtain the corresponding change area. The change area extraction based on image subtraction eliminates the influence of the image background, so as to effectively reduce the influence of the back kitchen environment on the model reasoning and reduce the model false detection rate, thereby greatly improving the detection accuracy of dynamic foreign objects.
[0105] Specifically, in step S107, after obtaining the dynamic foreign body category labels and corresponding confidences in multiple target detection frames in the next image, for each target detection frame, the following steps are performed:
[0106] Compare the confidence level of the target detection box with the preset confidence threshold:
[0107] If the confidence level corresponding to the target detection frame is not greater than the first confidence level threshold, it is determined that there is no dynamic foreign object in the kitchen at the time of the next image;
[0108] If the confidence level corresponding to the target detection frame is greater than the first confidence threshold, it is determined that there is a dynamic foreign object in the kitchen at the time of the next image, and based on the corresponding dynamic foreign object category label, it is determined whether the dynamic foreign object category is a category that requires an alarm. If it is a category that requires an alarm, the confidence level corresponding to the target detection frame is compared with the second confidence threshold:
[0109] If the confidence level corresponding to the target detection frame is greater than the second confidence threshold, an alarm is triggered;
[0110] If the confidence level corresponding to the target detection frame is not greater than the second confidence threshold, a single category detection is performed on the subsequent image;
[0111] The second confidence threshold is greater than the first confidence threshold.
[0112] Specifically, single-category detection is performed on the latter image, including:
[0113] Input the latter image into the single-category detector to obtain the single-category confidence corresponding to the target detection box and compare it with the third confidence threshold:
[0114] If the single-category confidence corresponding to the target detection box is greater than the third confidence threshold, an alarm is triggered;
[0115] If the single-category confidence corresponding to the target detection box is not greater than the third confidence threshold, it is determined that there is a foreign object, but no alarm is required and the detection is terminated.
[0116] After acquiring the target detection frame and the corresponding confidence in the latter image, the present invention further sets the first confidence threshold, the second confidence threshold and the third confidence threshold, combines multi-category detection with single-category detection, further distinguishes the foreign objects in the target detection frame, and determines whether to immediately issue an alarm, so as to help further reduce the probability of false detection and improve the detection accuracy. In this embodiment, common dynamic foreign object categories include cats, dogs and mice, and mice are usually set as the category that needs to be alarmed; when the dynamic foreign object category label is detected as a mouse, if the confidence is higher than the second confidence threshold, an alarm is directly issued, and if it is higher than the first confidence threshold but not higher than the second confidence threshold, a single-category detector is used to further detect the category that needs to be alarmed, so as to further determine the category of the dynamic foreign object in the current detection frame. Specifically, when the first confidence threshold, the second confidence threshold and the third confidence threshold are 0.6, 0.8 and 0.7 respectively, if the dynamic foreign object category label is rat, and its corresponding confidence threshold is greater than 0.6 and less than 0.8, at this time, it is determined that there may be a detection frame with a dynamic foreign object category label of rat in the image, but due to the possibility of false detection, the image is sent to a single category detector for detecting the rat category for re-detection. The obtained single category confidence is 0.82, which is greater than 0.7, then it is further determined that there is a dynamic foreign object with a dynamic foreign object category label of rat in the image, and an alarm is triggered.
[0117] Specifically, the embodiment of the present invention introduces an efficient multi-scale attention mechanism module EMA and SIoU loss function into the YOLOv7 model, proposes a change region extraction algorithm for the back kitchen environment, and proposes a system framework in combination with a multi-category detector and a single-category detector. Among them, the EMA module and the SIoU loss function are used to improve the accuracy of the algorithm's small target detection, and the change region extraction algorithm is used to reduce the model's misdetection problem and reduce the overhead of computing resources. Specifically, the YOLOv7 model is an improved YOLOv7 model that introduces an efficient multi-scale attention mechanism module between the backbone network and the detection head, including:
[0118] Input the changed area of the image into the backbone network for feature extraction and output the feature tensor;
[0119] Input the feature tensor into the efficient multi-scale attention mechanism module and output a multi-scale feature tensor;
[0120] The multi-scale feature tensor is input into the detection head for decoding and classification to obtain the image detection result.
[0121] Among them, the training process of the YOLOv7 model includes:
[0122] Obtain multiple kitchen images and perform label calibration to obtain the true label corresponding to each kitchen image and construct a training data set;
[0123] Train the YOLOv7 model based on the training data set until the improved loss function converges to obtain the improved YOLOv7 model;
[0124] Among them, the improved loss function , expressed as: ;
[0125] Represents the coordinate loss calculated using SIoU, expressed as: ; Indicates the overlap rate between the predicted box and the true box; Represents the distance loss, expressed as: , Represents the angle loss between the predicted box and the real box. hour, and Respectively Axis distance loss coordinate coefficient and Axis distance loss coordinate coefficient; Represents the shape loss, and the expression is: ,when hour, and Respectively represent the width coefficient and height coefficient of shape loss, Indicates the preset shape loss concern level;
[0126] Represents the confidence loss calculated using the BCEWithLogits Loss function, expressed as: ; and Represent the predicted confidence score and true label of the kitchen image, Indicates negative or positive examples, Sigmoid activation function ;
[0127] Represents the classification loss calculated using the BCEWithLogits Loss function, expressed as: , and Respectively represent The predicted scores and true labels of the categories, Indicates The category does not exist or exists, Sigmoid activation function , Indicates the total number of categories.
[0128] The improved YLOLv7 model of the present invention introduces the SIoU loss function and the efficient multi-scale attention mechanism module to fully utilize the information of small and medium-sized targets in the extracted changing area, improve the representation ability of the extracted small target features, and thus improve the detection accuracy of the improved YLOLv7 model.
[0129] Based on the above embodiments, the embodiment of the present invention sorts out the animal data in the public COCO dataset and combines it with the kitchen animal category data collected by itself into a complete Kitchen-Animal Dataset. Based on this dataset, the effectiveness of the kitchen dynamic foreign body detection method based on surveillance video provided by the embodiment of the present invention is verified. The specific steps include:
[0130] S201: Create data sets for experiments;
[0131] S201-1: Obtain animal data that may appear in the kitchen. In this embodiment, data of three categories, cat, dog, and mouse, are actually used. To make the experimental results more convincing, this embodiment writes a Python program to extract image data related to the experiment from the public COCO dataset, and a total of 8640 images are obtained.
[0132] S201-2: For the categories with a small number in the obtained data, this embodiment uses operations such as web crawling and video frame extraction to expand the data, and combines it with the data extracted from the COCO dataset to form the Kitchen-Animal Dataset, which includes 4298 pictures of the cat category, 4562 pictures of the dog category, and 4609 pictures of the mouse category.
[0133] In addition, to verify the effectiveness of the changed region extraction algorithm, this embodiment obtains 50 pairs of data through image segmentation and image synthesis to form a Frame-Extraction Dataset, which is used to simulate the video frame extraction operation.
[0134] S201-3: Use the labelimg tool to manually calibrate the image data in the Kitchen-Animal Dataset and Frame-Extraction Dataset, obtain the annotation files in txt format, and divide the complete Kitchen-Animal Dataset into training set, test set, and validation set in a ratio of 8:1:1.
[0135] S202: Improved YOLOv7 model;
[0136] S202-1: Replace CIoU loss function with SIoU loss function:
[0137] The loss function in the YOLOv7 model is: ;in, represents the coordinate loss, represents the confidence loss, Represents classification loss. Both confidence loss and classification loss are calculated using the BCEWithLogits Loss function, while coordinate loss is calculated using CIoU. The calculation formula is as follows:
[0138] ; ; ;
[0139] in, Represents the intersection-over-union ratio of the predicted box and the true box, represents the prediction box, represents the real frame, Represents the diagonal distance of the minimum closure area that can contain both the predicted box and the true box. is the balance parameter, Used to measure whether the aspect ratio is consistent. Represent the width and height of the real frame respectively, and Represent the width and height of the prediction box respectively.
[0140] It can be seen that when the aspect ratio of the predicted box is the same as that of the real box, When the value is 0, the aspect ratio penalty term does not work and the CIoU loss function cannot be expressed stably. Therefore, the present invention replaces the CIoU loss function of YOLOv7 with the SIoU loss function proposed by Zhora Gevorgyan. The SIoU loss function introduces a sorting mechanism and has position sensitivity, which improves the accuracy of bounding box regression and performs better when processing small objects of different shapes.
[0141] S202-2: Introduce an efficient multi-scale attention EMA module between the YOLOv7 backbone network and the neck to pass feature tensors; refer to Figure 2 As shown, it is a schematic diagram of the model structure of the improved YOLOv7 network;
[0142] Reference Figure 3 As shown in the figure, it is a schematic diagram of the structure of the efficient multi-scale attention EMA module. The EMA module first performs an input feature map Processing, where Indicates the number of channels, and Represent the height and width of the feature map respectively. This module converts the feature map The cross-channel dimension direction is divided into sub-feature groups, each sub-feature group . For each sub-feature group, EMA is processed using a parallel subnetwork. In the 1x1 branch, two 1D global average pooling operations are performed along the height and width directions respectively to encode channel information. Subsequently, the two pooled feature maps are concatenated along the height direction and further processed by 1x1 convolution, and their output is decomposed into two vectors and fit a 2D binomial distribution through two nonlinear Sigmoid functions. The channel attention maps of the two parallel paths are aggregated by multiplication. In the 3x3 branch, 3x3 convolutions are used to capture local cross-channel interactions, thereby expanding the feature space. In order to perform cross-space learning, 2D global average pooling is used to encode the global spatial information in the output of the 1x1 branch, and the output of the smallest branch will be converted to the corresponding dimensional shape before the joint activation mechanism of the channel features, that is, , the 2D global pooling operation can be described as:
[0143] ;
[0144] here It is Channels at position Next, the natural nonlinear function Softmax uses a 2D Gaussian map at the output of the 2D global average pooling to fit the linear transformation. The matrix dot product operation is performed on the output of the above parallel processing to obtain the first spatial attention map. Similarly, 2D global average pooling is used to encode the global spatial information in the 3x3 branch, and the 1x1 branch is directly converted to the corresponding dimensional shape before the joint activation mechanism of the channel features, that is, After that, the second spatial attention map is obtained by matrix dot product operation. Finally, the output feature map in each sub-feature group is aggregated by the two generated spatial attention weight values, and then processed by the Sigmoid function. The final output of EMA is similar to the input feature map are the same size.
[0145] S203: Train and improve the YOLOv7 model;
[0146] The improved YOLOv7 model was trained using the Kitchen-Animal Dataset to generate a dynamic foreign body detection model for the kitchen. The indicators that need to be considered in evaluating the model are precision P (Precision) and recall R (Recall). The area of the PR curve with P as the ordinate and R as the abscissa represents the average precision AP (Average Precision). The average value of all categories of AP is mAP (Mean Average Precision). The larger the mAP value, the higher the overall accuracy of the model. The formula is as follows:
[0147] ; ; ; ;
[0148] Among them, TP represents the number of positive predictions as positive, FP represents the number of negative predictions as positive, and FN represents the number of positive predictions as negative.
[0149] The GPU model used in the experiment is RXT 4090D (24GB), the deep learning framework is PyTorch 1.11.0, the programming language is Python 3.8.0, and the CUDA is 11.3. In order to reflect the impact of the introduction of SIoU loss function and EMA attention mechanism on model performance, each group of experiments uses the same dataset (Kitchen-Animal Dataset) and training parameters (YOLOv7 model default parameters). The total number of training rounds (epochs) in this experiment is 100, and the batch size (batchsize) of each round is 30. The experimental results are shown in Table 1: (Unit: %)
[0150] Table 1 Comparison of model detection results
[0151]
[0152] As can be seen from Table 1, compared with the original YOLOv7 model, the improved YOLOv7 model has improved the detection accuracy (P) by 2.9% and mAP0.5 by 1.2%, and the accuracy and mAP0.5 of each subcategory have improved.
[0153] S204: Acquire the changed region using a changed region extraction algorithm;
[0154] The current mainstream object detection model still cannot avoid false detection and missed detection, and the current research on this issue is mainly focused on improving the accuracy of the model. In order to reduce or avoid the problem of false detection, reducing the influence of irrelevant factors in the background environment on the model judgment is a key step. Therefore, in view of the particularity of the kitchen monitoring image, the present invention proposes a change area extraction algorithm to solve this problem.
[0155] The back kitchen monitoring screen has the characteristics of complex background and fixed scene, especially the back kitchen monitoring screen during non-working time periods such as nighttime and lunch break. Normally, there will be no large-scale pixel changes or local pixel mutations. Therefore, it is possible to judge whether there are illegal animals such as rats in the monitoring screen. The model only detects the monitoring screen when the change in the total pixel value of the monitoring screen is greater than the given threshold, which can greatly save computing resources. The appearance of illegal animals such as rats will only cause local pixel mutations. Therefore, the area where no pixel mutation occurs can be considered as an "irrelevant area" where no illegal animals appear, and it can be shielded, thereby reducing the probability of model misdetection. The area remaining after shielding the "irrelevant area" is the "changed area." Reference Figure 4 As shown, it is a flow chart of the steps of the change area extraction algorithm, and the specific steps include:
[0156] S204-1: extract two adjacent images img1 and img2 from the surveillance video according to a fixed frame extraction interval t;
[0157] S204-2: Setting activation threshold T1;
[0158] S204-3: Calculate pixel difference image of adjacent images: dif=img2-img1;
[0159] S204-4: Apply morphological opening and closing operations to the pixel difference image dif;
[0160] S204-5: Assume the total pixel value of the pixel difference image dif is S;
[0161] If S>T1, then assume that the number of rows and columns of pixels in the pixel difference image dif is M and N, initialize the row decision vector row to a 1×N all-zero vector, and initialize the column decision vector column to an M×1 all-zero vector; let the value range of i be [1, M], and the value range of j be [1, N];
[0162] S204-6: Setting threshold T2:
[0163] If img1[i][j]>T2, then mark row[1][j]=1, column[i][1]=1; until each pixel in the pixel difference image dif is judged, the target row decision vector and the target column decision vector are obtained;
[0164] S204-7: Based on the target row decision vector and the target column decision vector, construct a decision matrix decision=column*row;
[0165] S204-8: Based on the decision matrix, obtain the changed region of image img2 compared with image img1, expressed as: Interesting_region=img2*decision;
[0166] In order to verify the feasibility of the change region extraction method, the present invention simulates the images obtained by video frame extraction through artificial synthetic images, and organizes them into Frame-Extraction Dataset. The same model is allowed to perform inference on this dataset without using the change region extraction algorithm and with the change region extraction algorithm. The experimental results are shown in Table 2:
[0167] Table 2 Comparison of detection results based on change region extraction
[0168]
[0169] Based on Table 2, compared with not using the change region extraction algorithm, after using the change region extraction algorithm to preprocess the images sent to the model, the model detection accuracy (P) is improved by 1.5%, the recall rate (R) is improved by 27.0%, and the mAP0.5 is improved by 24.4%. The recall rate is greatly improved while the accuracy is also improved.
[0170] S205: Mixed use of multi-category detectors and single-category detectors: The present invention combines the characteristics of multi-category detectors and single-category detectors to propose a system framework to further reduce the model misdetection problem. Under the same system environment, data set, training parameters and model structure, the present invention conducted a comparative experiment of a multi-category detector with 3 categories, namely cat, dog and mouse, and a single-category detector with mouse as the category. The experimental results are shown in Table 3: (Unit: %)
[0171] Table 3 Comparison of detection accuracy between multi-category detector and single-category detector
[0172]
[0173] Based on Table 3, we can see that compared with the multi-category detector, the detection accuracy (P) of the single-category detector is 2.2% higher and the mAP0.5 is 1.6% higher, and the model performance is better.
[0174] Since the present invention is applied to the field of kitchen hygiene detection, when illegal animals that affect kitchen hygiene appear in the kitchen environment, the detection system should record them in time and send out an alarm at the right time. However, since the model may have misdetection, the alarm may be triggered by mistake, which will backfire. Therefore, the present invention proposes the following Figure 5The framework diagram of the back kitchen dynamic foreign object detection system based on the combination of multi-category detectors and single-category detectors is shown. By utilizing the characteristics of multi-category detectors that have multiple detection categories and single-category detectors that have high detection accuracy, an animal detection recording and alarm system for the back kitchen environment is constructed.
[0175] Reference Figure 6 As shown, it is a comparison diagram of the detection results of the present invention and the traditional method; in order to improve the detection accuracy of the model, the embodiment of the present invention introduces the SIoU loss function and the efficient multi-scale attention module (EMA module) into the target detection model YOLOv7, makes full use of the information of small targets, and extracts the features of small targets; in order to reduce the false detection rate of the model, the present invention proposes an image processing method based on image subtraction to extract the "changing area" in the image to be detected, and shields the content of the "static area" in the image before sending it to the model for reasoning, which can effectively avoid the problem of model false detection caused by complex background environment; in order to reduce the computing power consumption of the system, the present invention sets a suitable frame extraction interval, and only sends the monitoring image to the model for reasoning when the change amplitude of the monitoring picture is greater than a given threshold, which can effectively reduce the consumption of computing power resources by the system when the monitoring picture basically does not change during lunch break, night, etc.; in addition, the present invention combines the characteristics of multi-category detectors and single-category detectors to propose a system framework, which helps to further reduce the probability of false detection.
[0176] Based on the above embodiments, an embodiment of the present invention provides a dynamic foreign body detection device in a kitchen based on surveillance video, and the specific device may include:
[0177] An image set building module is used to extract multiple frames of images from the surveillance video of the kitchen to be detected based on a preset frame extraction interval to form an image set to be detected;
[0178] The difference image acquisition module is used to calculate the pixel difference between the adjacent pixels of the previous image and the next image in the image set to be detected, and obtain the difference image;
[0179] The difference image optimization module is used to smooth and denoise the difference image to obtain an optimized difference image;
[0180] The static and dynamic discrimination module is used to calculate the sum of the pixel values of all pixels in the optimized difference image and compare it with the preset activation threshold: if the sum of the pixel values of all pixels in the optimized difference image is not greater than the preset activation threshold, the next image in the current adjacent image is determined to be a static image, and the difference image of the next group of adjacent images is judged; if the sum of the pixel values of all pixels in the optimized difference image is greater than the preset activation threshold, the next image in the current adjacent image is determined to be a dynamic image;
[0181] The change region acquisition module is used to initialize the row decision vector as a 1×N all-zero vector and the column decision vector as an M×1 all-zero vector, where M and N are the number of rows and columns of pixels in the previous image, respectively; based on the pixel value at each pixel in the previous image and a preset change threshold, the row decision vector and the column decision vector are updated to obtain the target row decision vector and the target column decision vector; based on the target row decision vector and the target column decision vector, a decision matrix is constructed; and the dot product result of the decision matrix and the subsequent image is obtained as the change region of the subsequent image;
[0182] The detection module is used to input the changed area of the subsequent image into the trained YOLOv7 model, obtain the dynamic foreign object category labels and corresponding confidence levels in multiple target detection frames in the subsequent image, and complete the detection of dynamic foreign objects in the kitchen at the moment of the subsequent image.
[0183] The back kitchen dynamic foreign body detection device based on surveillance video of this embodiment is used to implement the aforementioned back kitchen dynamic foreign body detection method based on surveillance video. Therefore, the specific implementation of the back kitchen dynamic foreign body detection device based on surveillance video can be seen in the embodiment of the back kitchen dynamic foreign body detection method based on surveillance video in the previous text. For example, the image set construction module, the difference image acquisition module, and the difference image optimization module are respectively used to implement steps S101, S102 and S103 in the above-mentioned back kitchen dynamic foreign body detection method based on surveillance video; the static and dynamic discrimination module is used to implement steps S104, S105 and S106 in the above-mentioned back kitchen dynamic foreign body detection method based on surveillance video; the change area acquisition module is used to implement steps S106-1, S106-2, S106-3 and S106-4 in the above-mentioned back kitchen dynamic foreign body detection method based on surveillance video; the detection module is used to implement step S107 in the above-mentioned back kitchen dynamic foreign body detection method based on surveillance video. Therefore, its specific implementation can refer to the description of the corresponding embodiments of each part, which will not be repeated here.
[0184] Specifically, the back kitchen dynamic foreign body detection device based on surveillance video according to the embodiment of the present invention further includes a multi-category detection module for:
[0185] Compare the confidence level of the target detection box with the preset confidence threshold:
[0186] If the confidence level corresponding to the target detection frame is not greater than the first confidence level threshold, it is determined that there is no dynamic foreign object in the kitchen at the time of the next image;
[0187] If the confidence level corresponding to the target detection frame is greater than the first confidence threshold, it is determined that there is a dynamic foreign object in the kitchen at the time of the next image, and based on the corresponding dynamic foreign object category label, it is determined whether the dynamic foreign object category is a category that requires an alarm. If it is a category that requires an alarm, the confidence level corresponding to the target detection frame is compared with the second confidence threshold:
[0188] If the confidence level corresponding to the target detection frame is greater than the second confidence threshold, an alarm is triggered;
[0189] If the confidence level corresponding to the target detection frame is not greater than the second confidence threshold, a single category detection is performed on the subsequent image;
[0190] The second confidence threshold is greater than the first confidence threshold.
[0191] Specifically, the back kitchen dynamic foreign body detection device based on surveillance video according to the embodiment of the present invention further includes a single category detection module for:
[0192] Input the latter image into the single-category detector to obtain the single-category confidence corresponding to the target detection box and compare it with the third confidence threshold:
[0193] If the single-category confidence corresponding to the target detection box is greater than the third confidence threshold, an alarm is triggered;
[0194] If the single-category confidence corresponding to the target detection box is not greater than the third confidence threshold, it is determined that there is a foreign object, but no alarm is required and the detection is terminated.
[0195] The method for detecting dynamic foreign bodies in the kitchen based on surveillance video described in the present invention is that since the surveillance images of the kitchen are static in most cases when there are no staff, the image set to be detected is obtained by presetting the frame extraction interval, thereby reducing the number of detections and improving the monitoring efficiency; and the pixel difference between two adjacent images is used to determine whether the image currently detected has changed based on the previous image. The image processing method based on image subtraction shields the static area in the image, extracts the changed area, and then sends it to the YOLOv7 model for reasoning. The extraction of the changed area effectively avoids the problem of model misdetection caused by the complex background environment, while reducing the system computing power consumption and improving the detection accuracy. When extracting the changed area in the dynamic image, the present invention constructs a decision matrix that characterizes the change of pixel points through the optimized difference image between adjacent images, and uses the decision matrix to multiply the point of the latter image to obtain the changed area in the latter image, so as to achieve accurate extraction of the changed area, thereby improving the detection efficiency and accuracy. After acquiring the target detection frame and the corresponding confidence in the latter image, the present invention further sets the first confidence threshold, the second confidence threshold and the third confidence threshold, combines multi-category detection with single-category detection, further distinguishes the foreign matter in the target detection frame, and determines whether to immediately issue an alarm, so as to help further reduce the probability of false detection and improve the detection accuracy. The improved YLOLv7 model of the present invention introduces the SIoU loss function and the efficient multi-scale attention mechanism module to make full use of the information of the small and medium targets in the extracted change area, improve the characterization ability of the extracted small target features, and thus improve the detection accuracy of the improved YLOLv7 model.
[0196] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0197] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0198] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0199] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0200] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.
Claims
1. A method for detecting dynamic foreign matter in the kitchen based on monitoring video, characterized in that: include: Based on a preset frame extraction interval, multiple frames of images are extracted from the surveillance video of the kitchen to be inspected to form an image set to be inspected; Calculate the pixel difference between the corresponding pixel points in the adjacent previous image and the next image in the image set to be detected, and obtain a difference image; Smoothing and denoising the difference image to obtain an optimized difference image; Calculate the sum of the pixel values of all pixels in the optimized difference image and compare it with the preset activation threshold: If the sum of the pixel values of all pixels in the optimized difference image is not greater than the preset activation threshold, the next image in the current adjacent image is determined to be a static image, and the optimized difference image of the next group of adjacent images is determined; If the sum of the pixel values of all pixels in the optimized difference image is greater than the preset activation threshold, the next image in the current adjacent image is determined to be a dynamic image, and the change area of the next image is obtained, including: Initialize the row decision vector to a 1×N all-zero vector, and initialize the column decision vector to an M×1 all-zero vector, where M and N are the number of rows and columns of pixels in the optimized difference image, respectively; Based on the pixel value at each pixel point in the optimized difference image and the preset change threshold, the row decision vector and the column decision vector are updated to obtain the target row decision vector and the target column decision vector; Construct a decision matrix based on the target row decision vector and the target column decision vector; Obtain the dot product result of the decision matrix and the subsequent image as the change area of the subsequent image; The changed area of the subsequent image is input into the trained YOLOv7 model to obtain the dynamic foreign object category labels and corresponding confidence levels in multiple target detection frames in the subsequent image, and complete the detection of dynamic foreign objects in the kitchen at the moment of the subsequent image.
2. The method for detecting dynamic foreign matter in the kitchen based on monitoring video according to claim 1, characterized in that: After obtaining the dynamic foreign body category labels and corresponding confidences in multiple target detection frames in the next image, for each target detection frame, it includes: Compare the confidence level of the target detection box with the preset confidence threshold: If the confidence level corresponding to the target detection frame is not greater than the first confidence level threshold, it is determined that there is no dynamic foreign object in the kitchen at the time of the next image; If the confidence level corresponding to the target detection frame is greater than the first confidence threshold, it is determined that there is a dynamic foreign object in the kitchen at the time of the next image, and based on the corresponding dynamic foreign object category label, it is determined whether the dynamic foreign object category is a category that requires an alarm. If it is a category that requires an alarm, the confidence level corresponding to the target detection frame is compared with the second confidence threshold: If the confidence level corresponding to the target detection frame is greater than the second confidence level threshold, an alarm is triggered; If the confidence level corresponding to the target detection frame is not greater than the second confidence threshold, a single category detection is performed on the subsequent image; The second confidence threshold is greater than the first confidence threshold.
3. The method for detecting dynamic foreign matter in the kitchen based on monitoring video according to claim 2 is characterized in that: Perform single-category detection on the latter image, including: Input the latter image into the single-category detector to obtain the single-category confidence corresponding to the target detection box and compare it with the third confidence threshold: If the single-category confidence corresponding to the target detection box is greater than the third confidence threshold, an alarm is triggered; If the single-category confidence corresponding to the target detection box is not greater than the third confidence threshold, it is determined that there is a foreign object, but no alarm is required and the detection is terminated.
4. The method for dynamic foreign body detection in the kitchen based on monitoring video according to claim 1, characterized in that: The YOLOv7 model is an improved YOLOv7 model that introduces an efficient multi-scale attention mechanism module between the backbone network and the detection head, including: Input the changed area of the image into the backbone network for feature extraction and output the feature tensor; Input the feature tensor into the efficient multi-scale attention mechanism module and output a multi-scale feature tensor; The multi-scale feature tensor is input into the detection head for decoding and classification to obtain the image detection result.
5. The method for dynamic foreign body detection in the kitchen based on monitoring video according to claim 1, characterized in that: The training process of the YOLOv7 model includes: Obtain multiple kitchen images and perform label calibration to obtain the true label corresponding to each kitchen image and construct a training data set; Train the YOLOv7 model based on the training data set until the improved loss function converges to obtain the improved YOLOv7 model; Among them, the improved loss function , expressed as: ; Represents the coordinate loss calculated using SIoU, expressed as: ; Indicates the overlap rate between the predicted box and the true box; Represents the distance loss, expressed as: , Represents the angle loss between the predicted box and the real box. hour, and Respectively Axis distance loss coordinate coefficient and Axis distance loss coordinate coefficient; Represents the shape loss, and the expression is: ,when hour, and They represent the width coefficient and height coefficient of shape loss respectively, Indicates the preset shape loss concern level; Represents the confidence loss calculated using the BCEWithLogits Loss function, expressed as: ; and Represent the predicted confidence score and true label of the kitchen image, Indicates negative or positive examples, Sigmoid activation function ; Represents the classification loss calculated using the BCEWithLogits Loss function, expressed as: , and Respectively represent The predicted scores and true labels of the categories, Indicates The category does not exist or exists, Sigmoid activation function , Indicates the total number of categories.
6. The method for dynamic foreign body detection in the kitchen based on monitoring video according to claim 1, characterized in that: The difference image is smoothed and denoised using morphological opening and closing operations to obtain an optimized difference image, including: The difference image is first eroded and then expanded to obtain a denoised difference image. The denoised difference image is first dilated and then eroded to obtain the optimized difference image.
7. The method for detecting dynamic foreign matter in the kitchen based on monitoring video according to claim 1, characterized in that: The updating of the row decision vector and the column decision vector based on the pixel value at each pixel point in the optimized difference image and the preset change threshold, and obtaining the target row decision vector and the target column decision vector, comprises: Compare the pixel value dif[i][j] at each pixel point in the optimized difference image dif with the preset change threshold T2; Update the corresponding elements in the row decision vector and column decision vector of the pixel points whose pixel values are greater than the preset change threshold to 1, which is expressed as: If dif[i][j]>T2, then mark the row decision vector row[1][j]=1 and the column decision vector column[i][1]=1; After each pixel in the optimized difference image is judged, a target row decision vector and a target column decision vector are obtained; Among them, the value range of i is [1, M], the value range of j is [1, N], and M and N are the number of rows and columns of pixels in the optimized difference image respectively.
8. A dynamic foreign body detection device in the kitchen based on monitoring video, characterized in that: include: An image set building module is used to extract multiple frames of images from the surveillance video of the kitchen to be detected based on a preset frame extraction interval to form an image set to be detected; The difference image acquisition module is used to calculate the pixel difference between the adjacent pixels of the previous image and the next image in the image set to be detected, and obtain the difference image; The difference image optimization module is used to smooth and denoise the difference image to obtain an optimized difference image; The static and dynamic discrimination module is used to calculate the sum of the pixel values of all pixels in the optimized difference image and compare it with the preset activation threshold: if the sum of the pixel values of all pixels in the optimized difference image is not greater than the preset activation threshold, the next image in the current adjacent image is determined to be a static image, and the difference image of the next group of adjacent images is judged; If the sum of the pixel values of all pixels in the optimized difference image is greater than the preset activation threshold, the next image in the current adjacent image is determined to be a dynamic image; The change region acquisition module is used to initialize the row decision vector to be a 1×N all-zero vector and the column decision vector to be an M×1 all-zero vector, where M and N are the number of rows and columns of pixels in the optimized difference image respectively; Based on the pixel value at each pixel point in the optimized difference image and the preset change threshold, the row decision vector and the column decision vector are updated to obtain the target row decision vector and the target column decision vector; based on the target row decision vector and the target column decision vector, a decision matrix is constructed; Obtain the dot product result of the decision matrix and the subsequent image as the change area of the subsequent image; The detection module is used to input the changed area of the subsequent image into the trained YOLOv7 model, obtain the dynamic foreign object category labels and corresponding confidence levels in multiple target detection frames in the subsequent image, and complete the detection of dynamic foreign objects in the kitchen at the moment of the subsequent image.
9. The back kitchen dynamic foreign body detection device based on monitoring video according to claim 8, characterized in that: Also includes a multi-class detection module for: Compare the confidence level of the target detection box with the preset confidence threshold: If the confidence level corresponding to the target detection frame is not greater than the first confidence level threshold, it is determined that there is no dynamic foreign object in the kitchen at the time of the next image; If the confidence level corresponding to the target detection frame is greater than the first confidence threshold, it is determined that there is a dynamic foreign object in the kitchen at the time of the next image, and based on the corresponding dynamic foreign object category label, it is determined whether the dynamic foreign object category is a category that requires an alarm. If it is a category that requires an alarm, the confidence level corresponding to the target detection frame is compared with the second confidence threshold: If the confidence level corresponding to the target detection frame is greater than the second confidence level threshold, an alarm is triggered; If the confidence level corresponding to the target detection frame is not greater than the second confidence threshold, a single category detection is performed on the subsequent image; The second confidence threshold is greater than the first confidence threshold.
10. The back kitchen dynamic foreign body detection device based on monitoring video according to claim 9, characterized in that: Also includes a single-class detection module for: Input the latter image into the single-category detector to obtain the single-category confidence corresponding to the target detection box and compare it with the third confidence threshold: If the single-category confidence corresponding to the target detection box is greater than the third confidence threshold, an alarm is triggered; If the single-category confidence corresponding to the target detection box is not greater than the third confidence threshold, it is determined that there is a foreign object, but no alarm is required and the detection is terminated.
Citation Information
Patent Citations
Edge detection method, mobile detection method, pixel interpolating method and related apparatus
CN101420513A
Method for detecting picture displacement and related device
CN101527785A