A monitoring video quality assessment method, system and storage medium
By rating the surveillance videos in the area and targets, and using the weighted sum method to generate surveillance video scores, the problem of inability to effectively evaluate surveillance video quality in the prior art is solved, and efficient surveillance video quality evaluation and equipment adjustment are achieved.
Patent Information
- Application Number
- CN202210288485.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-23
AI Technical Summary
The existing surveillance system cannot effectively evaluate the quality of surveillance video, resulting in high cost and low efficiency of manual inspection, and the inability to accurately identify and adjust the problematic camera.
By scoring and weighting the monitoring video in both regional effectiveness and target effectiveness, using pre-trained scene segmentation model and object detection model, scene segmentation scores and object detection scores are generated, and weighted summed to obtain monitoring video scores, assisting manual detection and adjusting monitoring equipment.
It has achieved a comprehensive evaluation of the quality of surveillance video, reduced manual inspection costs, improved the detection efficiency and quality of the surveillance system, and can accurately identify and adjust unreasonable surveillance equipment layout.
Smart Images

Figure CN114708532B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video quality assessment, and particularly relates to a method, a system and a storage medium for monitoring video quality assessment. Background Art
[0002] In the currently operating monitoring systems, there are a large number of cameras that are not operating normally. Moreover, the operation and maintenance work of the monitoring systems still basically rely on manual detection and processing. Even if there are some fault monitoring mechanisms or means in the monitoring systems, they only simply judge functions such as whether the network is connected and whether the camera is online, and cannot truly identify the cameras with problems in the monitoring systems. On the other hand, with the rapid development of computer networks and monitoring networks today, the number of cameras has increased rapidly, and it has gradually become unrealistic to simply rely on manpower to judge the quality of monitoring cameras.
[0003] With the rapid development of artificial intelligence, it has become a trend to replace manual labor with machines to complete relatively fixed work. This can not only be used as an auxiliary means for manual labor, but also reduce labor costs. At present, most of the methods for video quality assessment are limited to the assessment of image quality, staying at the detection of aspects such as image blurring, abnormal colors, snowflake interference, and pan-tilt out of control, and no quality assessment method applicable to special videos such as monitoring cameras that can start from the monitoring content and conform to human subjective judgment has been proposed, which often shows limitations when analyzing monitoring videos. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, a system and a storage medium for monitoring video quality assessment. By scoring and weighting the video content in terms of regional effectiveness and target effectiveness to obtain an assessment score, the staff can determine the monitoring devices with unreasonable layout in the monitoring area according to the monitoring video score and make manual adjustments, reducing the labor cost in the detection process of the monitoring system.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] The first aspect of the present invention provides a method for monitoring video quality assessment, including:
[0007] Collect the monitoring video captured by the monitoring device, and extract the monitoring images of each frame from the monitoring video;
[0008] Input the monitoring images corresponding to the previous several frames into the anomaly detection module to determine whether the monitoring video has quality defects;
[0009] When it is determined that the surveillance video has quality defects, the evaluation is abandoned and a surveillance video permission warning is output; otherwise, the remaining surveillance images are input into the pre-trained scene segmentation model and the pre-trained object detection model according to the user settings to obtain the scene segmentation score and the object detection score;
[0010] The scene segmentation score and the target detection score are weightedly summed to obtain the surveillance video score, which is used to assist manual detection of surveillance equipment.
[0011] Preferably, the monitoring images corresponding to the first several frames are input into the abnormality detection module, and the method for determining whether the monitoring video has quality defects includes:
[0012] Grayscale the monitoring image and calculate the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold. When the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold is greater than 80% of the total number of pixels, a black screen or blue screen quality defect warning is output;
[0013] When the number of pixels whose grayscale values exceed the preset threshold in the monitoring image is less than 80% of the total number of pixels, the edge gradient G of the monitoring image is calculated by the Sobel algorithm; based on the edge gradient G, it is determined whether the monitoring image has clarity quality defects.
[0014] Preferably, the method for calculating the edge gradient G of the monitoring image by using the Sobel algorithm includes:
[0015] Perform planar convolution of the template matrix Sx and the template matrix Sy with the monitoring image to obtain the horizontal brightness difference Gx and the vertical brightness difference Gy;
[0016] The formulas of template matrix Sx and template matrix Sy are:
[0017]
[0018] The edge gradient G of the monitoring image is calculated based on the horizontal brightness difference Gx and the vertical brightness difference Gy. The formula is:
[0019]
[0020] Preferably, the monitoring image is input into a pre-trained scene segmentation model, and the method for obtaining the scene segmentation score includes:
[0021] Input the monitoring image into the scene segmentation model to obtain the prediction map; count the proportion of the target area and the scene area, and calculate the scene segmentation score S seg The expression formula is:
[0022] S seg =N target / (width×height)×100%
[0023] Among them, N target Indicates the number of pixels in the target area; width is the number of pixels in the width direction of the monitoring image; height is the number of pixels in the height direction of the monitoring image.
[0024] Preferably, the method of inputting the monitoring image into the scene segmentation model to obtain the prediction graph includes:
[0025] Extract monitoring features and shallow features from monitoring images through the backbone network of the scene segmentation model;
[0026] The monitoring features are sequentially input into the ASPP network structure for hole convolution calculation, and superimposed features are obtained by splicing; 1×1 convolution is performed on the superimposed features to obtain fused features;
[0027] The shallow features are refined by 1×1 convolution and then concatenated with the 4-fold upsampled fusion features in the channel direction. Then, the prediction graph is obtained by 3×3 convolution and 4-fold upsampling.
[0028] Preferably, the monitoring image is input into a pre-trained target detection model, and the method for obtaining the target detection score includes:
[0029] Input the monitoring image into the YOLO network to extract the target features, and classify the target features to obtain the confidence level;
[0030] The Euclidean distance between the center of the target feature and the midpoint of the lower edge of the image is calculated, and the product of the Euclidean distance and the confidence of the target feature is used as the target detection score.
[0031] Preferably, the calculation formula for the surveillance video score is:
[0032] S sum =α 1 ·S seg +(α 2 +α 3 )·S det
[0033] Among them, S sum Score surveillance video, S det Denotes the target detection score, α 1 Expressed as the weighting coefficient of scene segmentation score, α 2 and α 3 is the weighting coefficient of the target detection score. When the pixel area occupied by the target feature is greater than the historical maximum value, α 3 Equal to 0; when the pixel area occupied by the target feature is greater than the historical maximum value, α 3 Equal to the set value.
[0034] Preferably, the scene segmentation model and target detection model training process includes: collecting relevant training images of the monitoring site, annotating the training images through labelme, and constructing a training data set; training the scene segmentation model and target detection model through the training data set.
[0035] A second aspect of the present invention provides a monitoring video quality assessment system, comprising:
[0036] The acquisition module is used to acquire the surveillance video shot by the surveillance equipment and extract the surveillance image of each frame from the surveillance video;
[0037] The detection module is used to input the monitoring images corresponding to the previous frames into the abnormality detection module to determine whether there are quality defects in the monitoring video;
[0038] A scoring module is used to input the remaining surveillance images into a pre-trained scene segmentation model and a pre-trained target detection model according to user settings, to obtain a scene segmentation score and a target detection score;
[0039] The weighted summation module is used to perform weighted summation of the scene segmentation score and the target detection score to obtain the surveillance video score, and to assist manual detection of the surveillance equipment through the surveillance video score.
[0040] A third aspect of the present invention provides a computer-readable storage medium, characterized in that the computer-readable storage medium includes a stored program, wherein the program executes the monitoring video quality assessment method.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] (1) The present invention inputs the monitoring images corresponding to the first several frames into the anomaly detection module to determine whether the monitoring video has quality defects; when it is determined that the monitoring video has quality defects, the evaluation is abandoned and a monitoring video authority warning is output; monitoring videos with quality defects such as blue screen, black screen or poor clarity are preliminarily screened out to assist in manually determining the monitoring equipment with quality defects.
[0043] (2) The present invention inputs the remaining surveillance images into a pre-trained scene segmentation model and a pre-trained target detection model according to user settings, respectively, to obtain a scene segmentation score and a target detection score; the scene segmentation score and the target detection score are weightedly summed to obtain a surveillance video score, and the surveillance video score is used to assist manual detection of surveillance equipment; the scene segmentation score and the target detection score are weightedly summed to obtain a surveillance video score, and the staff determines the unreasonable deployment of surveillance equipment in the surveillance area according to the surveillance video score and makes manual adjustments, thereby improving the surveillance quality and reducing the labor intensity of the staff during the detection process. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a flow chart of a monitoring video quality assessment method provided by an embodiment of the present invention;
[0045] Figure 2 is a structural diagram of a scene segmentation model provided by an embodiment of the present invention;
[0046] Figure 3 It is a structural diagram of the target detection model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0048] Embodiment 1
[0049] like Figures 1 to 3 As shown, a monitoring video quality assessment method includes:
[0050] Collect the surveillance video shot by the surveillance equipment, decode the surveillance video using ffmpeg, convert the decoded YUV format data into RGB format data, and extract the surveillance image of each frame from the surveillance video;
[0051] The monitoring images corresponding to the first several frames are input into the anomaly detection module, and the method for determining whether the monitoring video has quality defects includes:
[0052] The monitoring image is converted into a grayscale image through opencv, and the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold is calculated; when the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold is greater than 80% of the total number of pixels, a black screen or blue screen quality defect warning is output;
[0053] When the number of pixels whose grayscale values exceed the preset threshold in the monitoring image is less than 80% of the total number of pixels, the edge gradient G of the monitoring image is calculated by the Sobel algorithm; based on the edge gradient G, it is determined whether the monitoring image has clarity quality defects.
[0054] The method of calculating the edge gradient G of the monitoring image by the Sobel algorithm includes:
[0055] Perform planar convolution of the template matrix Sx and the template matrix Sy with the monitoring image to obtain the horizontal brightness difference Gx and the vertical brightness difference Gy;
[0056] The formulas of template matrix Sx and template matrix Sy are:
[0057]
[0058] The edge gradient G of the monitoring image is calculated based on the horizontal brightness difference Gx and the vertical brightness difference Gy. The formula is:
[0059]
[0060] When it is determined that the surveillance video has quality defects, the evaluation is abandoned and a surveillance video permission warning is output; otherwise, the remaining surveillance images are input into the pre-trained scene segmentation model and the pre-trained object detection model according to the user settings to obtain the scene segmentation score and the object detection score;
[0061] The scene segmentation model and target detection model training process includes: collecting relevant training images of the monitoring site, annotating the training images through labelme, and constructing a training data set; and training the scene segmentation model and the target detection model through the training data set.
[0062] The surveillance image is input into the pre-trained scene segmentation model, and the method of obtaining the scene segmentation score includes:
[0063] Perform deeplabv3+ operations on surveillance images through the backbone network of the scene segmentation model to extract surveillance features and shallow features;
[0064] like Figure 2 As shown in the figure, the monitoring features are sequentially input into the atrous spatial pyramid pooling (ASPP) network structure for hole convolution calculation, and the superimposed features are obtained by splicing. The ASPP network structure can capture more scale information and avoid the loss of details caused by pooling when the full convolution network extracts features. The encoder-decoder structure in the ASPP network structure can better protect the edge information in the original image and realize pixel-level semantic segmentation; the superimposed features are convolved by 1×1 to obtain the fused features;
[0065] The shallow features are refined by 1×1 convolution and then concatenated with the 4-fold upsampled fusion features in the channel direction. Then, the prediction graph is obtained by 3×3 convolution and 4-fold upsampling.
[0066] Calculate the ratio of target area to scene area in the prediction image and calculate the scene segmentation score S seg The expression formula is:
[0067] S seg =N target / (width×height)×100%
[0068] Among them, N targetIndicates the number of pixels in the target area; width is the number of pixels in the width direction of the monitoring image; height is the number of pixels in the height direction of the monitoring image; the target area mainly includes roads, bridges, buildings, and grasslands where human and vehicle targets can move. The proportion of the target area in the scene is used to measure whether the angle and position of the camera are reasonable.
[0069] The method of inputting the surveillance image into the pre-trained object detection model and obtaining the object detection score includes:
[0070] The monitoring image is input into the YOLO network to extract target features, which mainly include vehicles (including passenger cars, buses, trucks and electric vehicles), pedestrians and license plates; the target features are classified to obtain confidence; the YOLO network in this embodiment is YOLOv3, and its backbone network is as follows Figure 3 Darknet-53, which consists of 52 convolutional layers and 1 fully connected layer, borrows the residual structure of ResNet, avoids the degradation of deep networks through residual, and avoids using pooling layers. Instead, it uses stride to obtain a larger receptive field, which makes it have no significant decrease in accuracy compared to ResNet-101, but the calculation speed has been significantly improved. At the same time, Yolov3 uses feature maps of multiple scales for prediction in the backbone network, taking into account the detection accuracy of large, medium and small targets.
[0071] The Euclidean distance between the center of the target feature and the midpoint of the lower edge of the image is calculated, and the product of the Euclidean distance and the confidence of the target feature is used as the target detection score.
[0072] The scene segmentation score and the target detection score are weightedly summed to obtain the surveillance video score, which is used to assist manual detection of surveillance equipment.
[0073] The calculation formula of the surveillance video score is:
[0074] S sum =α 1 ·S seg +(α 2 +α 3 )·S det
[0075] Among them, S sum Score surveillance video. det Denotes the target detection score, α 1 Expressed as the weighting coefficient of scene segmentation score, α 2 and α 3 is the weighting coefficient of the target detection score. When the pixel area occupied by the target feature is greater than the historical maximum value, α 3Equal to 0; when the pixel area occupied by the target feature is greater than the historical maximum value, α 3 Equal to the set value.
[0076] During the detection process, for cameras within a certain local area network, the front-end page is used to realize the selectivity of camera evaluation. According to the evaluation time and passing score of a single camera set by the user, the cameras selected by the user are polled and evaluated. After the evaluation, an alarm is issued for cameras with scores lower than the passing score. At the same time, the camera video during the evaluation process is provided for the staff to review, and the staff will confirm whether there is a problem with the camera.
[0077] Embodiment 2
[0078] A second aspect of the present invention provides a surveillance video quality assessment system, which can be applied to the surveillance video quality assessment method described in Embodiment 1, including:
[0079] The acquisition module is used to acquire the surveillance video shot by the surveillance equipment and extract the surveillance image of each frame from the surveillance video;
[0080] The detection module is used to input the monitoring images corresponding to the previous frames into the abnormality detection module to determine whether there are quality defects in the monitoring video;
[0081] A scoring module is used to input the remaining surveillance images into a pre-trained scene segmentation model and a pre-trained target detection model according to user settings, to obtain a scene segmentation score and a target detection score;
[0082] The weighted summation module is used to perform weighted summation of the scene segmentation score and the target detection score to obtain the surveillance video score, and to assist manual detection of the surveillance equipment through the surveillance video score.
[0083] Embodiment 3
[0084] A third aspect of the present invention provides a computer-readable storage medium, characterized in that the computer-readable storage medium includes a stored program, and the program executes the monitoring video quality assessment method described in Example 1.
[0085] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0086] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0087] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0089] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A surveillance video quality assessment method, It is characterized in that include: Collect surveillance videos shot by surveillance equipment and extract surveillance images of each frame from the surveillance videos; The surveillance images corresponding to the first several frames are input into the anomaly detection module to determine whether the surveillance video has quality defects. The process includes: Grayscale the monitoring image and calculate the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold. When the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold is greater than 80% of the total number of pixels, a black screen or blue screen quality defect warning is output. When the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold is less than 80% of the total number of pixels, the edge gradient G of the monitoring image is calculated using the Sobel algorithm. The process includes: Perform planar convolution of the template matrix Sx and the template matrix Sy with the monitoring image to obtain the horizontal brightness difference Gx and the vertical brightness difference Gy; The formulas of template matrix Sx and template matrix Sy are: ; The edge gradient G of the monitoring image is calculated based on the horizontal brightness difference Gx and the vertical brightness difference Gy. The formula is: ; Judging whether the monitoring image has clarity quality defects according to the edge gradient G; When it is determined that the surveillance video has quality defects, the evaluation is abandoned and a surveillance video permission warning is output; otherwise, the remaining surveillance images are input into the pre-trained scene segmentation model and the pre-trained object detection model according to the user settings to obtain the scene segmentation score and the object detection score; The scene segmentation score and the target detection score are weightedly summed to obtain the surveillance video score, which is used to assist manual detection of surveillance equipment.
2. The surveillance video quality assessment method according to claim 1, It is characterized in that The surveillance image is input into the pre-trained scene segmentation model, and the method of obtaining the scene segmentation score includes: Input the monitoring image into the scene segmentation model to obtain the prediction map; count the proportion of the target area and the scene area, and calculate the scene segmentation score S seg The expression formula is: S seg =N target / (width×height)×100%; Among them, N target Indicates the number of pixels in the target area; width is the number of pixels in the width direction of the monitoring image; height is the number of pixels in the height direction of the monitoring image.
3. The monitoring video quality assessment method according to claim 2, It is characterized in that The method of inputting the monitoring image into the scene segmentation model to obtain the prediction map includes: Extract monitoring features and shallow features from monitoring images through the backbone network of the scene segmentation model; The monitoring features are sequentially input into the ASPP network structure for hole convolution calculation, and superimposed features are obtained by splicing; 1×1 convolution is performed on the superimposed features to obtain fused features; The shallow features are refined by 1×1 convolution and then concatenated with the 4-fold upsampled fusion features in the channel direction. Then, the prediction graph is obtained by 3×3 convolution and 4-fold upsampling.
4. The surveillance video quality assessment method according to claim 3, It is characterized in that The method of inputting the surveillance image into the pre-trained object detection model and obtaining the object detection score includes: Input the monitoring image into the YOLO network to extract the target features, and classify the target features to obtain the confidence level; The Euclidean distance between the center of the target feature and the midpoint of the lower edge of the image is calculated, and the product of the Euclidean distance and the confidence of the target feature is used as the target detection score.
5. The monitoring video quality assessment method according to claim 4, It is characterized in that The calculation formula of the surveillance video score is: S sum = a 1 ·S seg + (a) 2 +a 3 )·S det ; Among them, S sum Score surveillance video. det Denotes the target detection score, α 1 Expressed as the weighting coefficient of scene segmentation score, α 2 and α 3 is the weighting coefficient of the target detection score. When the pixel area occupied by the target feature is greater than the historical maximum value, α 3 Equal to 0; when the pixel area occupied by the target feature is greater than the historical maximum value, α 3 Equal to the set value.
6. The monitoring video quality assessment method according to claim 1, It is characterized in that The scene segmentation model and target detection model training process includes: collecting relevant training images of the monitoring site, annotating the training images through labelme, and constructing a training data set; and training the scene segmentation model and the target detection model through the training data set.
7. A surveillance video quality assessment system, It is characterized in that include: The acquisition module is used to acquire the surveillance video shot by the surveillance equipment and extract the surveillance image of each frame from the surveillance video; The detection module is used to input the monitoring images corresponding to the previous frames into the abnormality detection module to determine whether there are quality defects in the monitoring video; A scoring module is used to input the remaining surveillance images into a pre-trained scene segmentation model and a pre-trained target detection model according to user settings, to obtain a scene segmentation score and a target detection score; A weighted summation module is used to perform a weighted summation of the scene segmentation score and the target detection score to obtain a surveillance video score, and to assist manual detection of surveillance equipment through the surveillance video score; The detection module inputs the monitoring images corresponding to the previous several frames into the abnormality detection module to determine whether the monitoring video has quality defects. The process includes: Grayscale the monitoring image and calculate the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold. When the number of pixels in the monitoring image whose grayscale value exceeds the preset threshold is greater than 80% of the total number of pixels, a black screen or blue screen quality defect warning is output. When the number of pixels in the surveillance image whose grayscale value exceeds the preset threshold is less than 80% of the total number of pixels, the edge gradient G of the surveillance image is calculated using the Sobel algorithm. The process includes: Perform planar convolution of the template matrix Sx and the template matrix Sy with the monitoring image to obtain the horizontal brightness difference Gx and the vertical brightness difference Gy; The formulas of template matrix Sx and template matrix Sy are: ; The edge gradient G of the monitoring image is calculated based on the horizontal brightness difference Gx and the vertical brightness difference Gy. The formula is: ; The edge gradient G is used to determine whether the monitored image has clarity quality defects.
8. A computer-readable storage medium, It is characterized in that The computer-readable storage medium includes a stored program, wherein the program executes the monitoring video quality assessment method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Abnormal behavior supervision method and device based on action recognition and storage medium
CN113052029A
Parking detection method and device based on monitoring video
WO2018130016A1