Video quality evaluation method and device for monitoring video, medium and program product

By performing defect detection and multimodal feature extraction of surveillance video frames, combined with traditional image processing and multimodal large model, the accuracy of surveillance video quality evaluation in complex scenarios is solved, and more efficient video quality evaluation is achieved.

CN120374593APending Publication Date: 2025-07-25ROPEOK TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510540538.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing surveillance video quality evaluation methods are insufficient in complex scenarios, especially under low light, rain and fog, and lack effective utilization of video semantic information.

Method used

By detecting defects on surveillance video frames, using multimodal feature extraction and matching technology, combining traditional image processing algorithms and multimodal large models, video quality evaluation results are generated.

Benefits of technology

The adaptability of the video quality evaluation method to complex situations is improved, and the accuracy and reliability of the evaluation results are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374593A_ABST
    Figure CN120374593A_ABST
Patent Text Reader

Abstract

The invention provides a video quality evaluation method and device for a monitoring video, a medium and a program product. The method comprises the following steps: performing video frame extraction according to an obtained monitoring video stream; defect detection is carried out on the extracted video frames, and video frames without quality defects are determined; performing multi-modal feature extraction on the video frame without the quality defect to obtain a corresponding image feature; matching the image features with preset features in a feature library, wherein the preset features are used for describing pictures with various quality defects and normal display; and generating a corresponding video quality evaluation result according to the matching result. According to the technical scheme provided by the embodiment of the invention, the adaptability to complex conditions can be improved, and the accuracy of a video quality evaluation result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and particularly to a method, device, medium, and program product for video quality assessment of surveillance videos. Background Art

[0002] In the field of video surveillance, video quality assessment is crucial for application scenarios such as security surveillance, intelligent transportation, and industrial production. In current technical solutions, video quality assessment technologies mainly include methods based on traditional image processing and methods based on deep learning. Methods based on traditional image processing usually evaluate by analyzing single indicators such as image sharpness, contrast, and noise. Although the computational efficiency is relatively high, the accuracy significantly decreases when dealing with complex scenarios (such as low light, rainy or foggy weather, etc.). While methods based on deep learning have improved in terms of accuracy, they perform poorly when dealing with complex situations such as small targets and occlusions, and most methods only start from visual features, lacking effective utilization of video semantic information and making it difficult to accurately determine whether the video content meets the actual scenario requirements.

[0003] Thus, how to improve the adaptability to complex situations and ensure the accuracy of video quality assessment results has become a technical problem to be urgently solved. Summary of the Invention

[0004] The present disclosure provides a method, device, medium, and program product for video quality assessment of surveillance videos.

[0005] According to one aspect of the present disclosure, a method for video quality assessment of surveillance videos includes:

[0006] Performing video frame extraction according to the acquired surveillance video stream;

[0007] Performing defect detection on the extracted video frames to determine video frames without quality defects;

[0008] Performing multi-modal feature extraction on the video frames without quality defects to obtain corresponding image features;

[0009] Matching the image features with each preset feature in the feature library, where the preset features are used to describe various quality defects and normal display pictures;

[0010] Generating a corresponding video quality assessment result according to the matching result.

[0011] A method for evaluating the video quality of a surveillance video according to at least one embodiment of the present disclosure extracts video frames based on the acquired surveillance video stream, detects defects in the extracted video frames, determines the video frames without quality defects, then extracts multi-modal features from the video frames without quality defects to obtain corresponding image features, and then matches the image features with each preset feature in the feature library. The preset feature is used to describe various quality defects and normal display pictures, and then according to the matching result, a corresponding video quality evaluation result is generated.

[0012] In this way, first detect obvious quality defects in the extracted video frames. If there are no obvious quality defects, then extract multi-modal features from the video frames, match the extracted image features with the preset features in the feature library, so as to perform further video quality detection, and then generate corresponding video quality evaluation results according to the matching results. This can improve the adaptability of the video quality evaluation method to complex situations and ensure the accuracy of the video quality evaluation results.

[0013] In some embodiments of the present disclosure, generating a corresponding video quality evaluation result according to the matching result includes:

[0014] Generating a corresponding video quality evaluation result according to the defect detection result and the matching result.

[0015] According to another aspect of the present disclosure, there is provided a device for evaluating the video quality of a surveillance video, including:

[0016] An extraction module for extracting video frames according to the acquired surveillance video stream;

[0017] A detection module for detecting defects in the extracted video frames and determining the video frames without quality defects;

[0018] An extraction module for extracting multi-modal features from the video frames without quality defects to obtain corresponding image features;

[0019] A matching module for matching the image features with each preset feature in the feature library, and the preset feature is used to describe various quality defects and normal display pictures;

[0020] A processing module for generating a corresponding video quality evaluation result according to the matching result.

[0021] According to another aspect of the present disclosure, there is provided an electronic device, including: a memory storing execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the method for evaluating the video quality of a surveillance video according to any one of the embodiments of the present disclosure.

[0022] According to another aspect of the present disclosure, there is provided a readable storage medium storing execution instructions, which are used to implement the video quality assessment method for surveillance videos according to any embodiment of the present disclosure when executed by a processor.

[0023] According to still another aspect of the present disclosure, there is provided a computer program product including a computer program, which implements the video quality assessment method for surveillance videos according to any embodiment of the present disclosure when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.

[0025] Figure 1 FIG. shows a schematic flowchart of a video quality assessment method for surveillance videos according to an embodiment of the present disclosure;

[0026] Figure 2 FIG. shows a schematic flowchart of a video quality assessment method for surveillance videos according to another embodiment of the present disclosure;

[0027] Figure 3 is a schematic block diagram of a video quality assessment device for surveillance videos according to an embodiment of the present disclosure;

[0028] Figure 4 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The present disclosure will be further described in detail below with reference to the drawings and examples. It can be understood that the specific examples described herein are only for explaining the relevant content and do not limit the present disclosure. Additionally, it should be noted that for the sake of description, only parts related to the present disclosure are shown in the drawings.

[0030] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the drawings and embodiments.

[0031] In the field of surveillance videos, video quality is crucial for many applications such as security surveillance, intelligent transportation, and industrial production. Currently, there are various existing surveillance video quality assessment methods, but each has its own deficiencies.

[0032] Some traditional methods rely on manual inspection, which is inefficient and subjective and difficult to meet the requirements of large-scale surveillance videos. For example, in the security surveillance system of a large city, it is almost an impossible task to manually check the video quality one by one.

[0033] Automatic evaluation methods based on a single technology also have limitations. Some methods based on traditional image processing algorithms, such as evaluating video quality only relying on image clarity metrics, have a significant drop in accuracy when facing complex scenes (such as fog, light changes, etc.). In addition, although deep learning-based methods have made some progress, they still face some problems. For example, some models perform poorly when dealing with small targets or occlusion situations in videos, and model training usually requires a large amount of labeled data, with a high acquisition cost. At the same time, most existing methods only start from the visual features of videos and lack effective utilization of semantic information, making it difficult to accurately judge whether the video content meets the actual scene requirements.

[0034] Therefore, the present disclosure proposes a video quality evaluation method for surveillance videos. It should be noted that the video quality evaluation method for surveillance videos provided by the embodiments of the present disclosure can be applied to terminal devices or servers. Among them, the terminal device can include, but is not limited to, one or more of a smart phone, a tablet computer, a portable computer, and a desktop computer; the server can be a physical server or a cloud server.

[0035] It should be noted that in addition to the field of surveillance videos, the video quality evaluation method provided by the embodiments of the present disclosure can also be applied to other video-related fields, such as video conferencing, virtual reality (VR) / augmented reality (AR) fields, etc. The present disclosure makes no special limitation on this.

[0036] The following takes the application of this method to a server as an example for illustration. Figure 1 The flowchart of the video quality evaluation method for surveillance videos according to an embodiment of the present disclosure is shown.

[0037] As Figure 1 shown, the video quality evaluation method for surveillance videos at least includes steps S110 to S150, which are introduced in detail as follows:

[0038] In step S110, video frames are extracted according to the acquired surveillance video stream.

[0039] Among them, the surveillance video stream can be a continuous video data stream obtained in real time by a camera or other video acquisition devices, which is usually transmitted in the form of a frame sequence.

[0040] In this embodiment, the server can be communicatively connected to one or more video capture devices, so as to receive in real time the monitoring video stream transmitted by the video capture devices. The monitoring video stream is captured by the video capture devices for a specified monitoring area.

[0041] After receiving the monitoring video stream, the server can extract video frames from the monitoring video stream at a predetermined time interval. The predetermined time interval can be pre-determined by those skilled in the art according to prior experience. For example, the predetermined time interval can be 1 second, 5 seconds, etc., that is, extract one frame every 1 second or extract one frame every 5 seconds, etc. It should be noted that the above numbers are only exemplary examples, and those skilled in the art can dynamically adjust according to the scenario requirements to balance real-time performance and computational resource consumption, and no special limitation is made thereto.

[0042] In one example, due to the relatively low real-time requirement for video quality detection, extraction can be performed at intervals of a certain time. For example, the predetermined time interval can be 10 seconds, 15 seconds, etc.

[0043] In one embodiment, the server can access the monitoring video source (i.e., the video capture device) through RTSP (Real-Time Streaming Protocol) or ONVIF (Open Network Video Interface), and use tools such as FFmpeg or GStreamer to decode the video stream to obtain the original frame sequence, so as to ensure compatibility with devices of different manufacturers and achieve efficient parsing and format unification of the video stream.

[0044] In addition, when setting the predetermined time interval, it can be set to be fixed or dynamic. For example, when the predetermined time interval is a dynamic value, the predetermined time interval can be automatically adjusted according to the complexity of the video content. The higher the complexity, the shorter the predetermined time interval, etc.

[0045] In one example, after extracting the video frames, the server can perform size normalization (such as scaling to a unified resolution, etc.) and format standardization on the extracted video frames, and associate metadata such as timestamps and camera numbers, so as to facilitate subsequent defect detection and result traceability.

[0046] In step S120, defect detection is performed on the extracted video frames to determine the video frames without quality defects.

[0047] Among them, defect detection can be a process of identifying quality defects (such as blurring, color cast, frozen frames, etc.) existing in video frames through algorithms. That is to say, a preset image processing algorithm can be used to perform defect detection on the extracted video frames to determine whether there are obvious quality defects in each of the video frames, and then determine the video frames without quality defects. It should be noted that at this time, the video frames without quality defects do not necessarily mean that they are video frames with normal display, but it is possible that the quality defects they have cannot be recognized by traditional image processing algorithms.

[0048] The preset image processing algorithm can be a traditional image analysis method, which can include but is not limited to one or more of signal loss detection, black and white image recognition, abnormal exposure analysis, image color cast judgment, frozen frame detection, and image blurring evaluation.

[0049] In this embodiment, the server can use a preset image processing algorithm to perform defect detection on the extracted video frames, and then determine whether there are obvious quality defects (i.e., the several quality defects listed above) in the video frames.

[0050] Specifically, when performing signal loss detection and black and white image recognition, the server can use a random sampling method to randomly select a series of pixel points in the video frame (for example, select 100 pixel points covering different regions of the video frame) for detection, and judge whether their pixel values conform to normal rules. If the pixel values are abnormal, such as a large number of pixel values being 0 or exceeding the normal range, it can be judged that there are signal loss or black and white image problems.

[0051] In an example, if more than 80% of the selected pixel points meet one of the following conditions, it can be determined that there are obvious quality defects:

[0052] A. All RGB channel values are 0 (pure black) or 255 (pure white);

[0053] B. The RGB three-channel values are equal and the difference is less than 5.

[0054] When performing abnormal exposure analysis and image color cast judgment, the server can use a histogram calculation method to analyze the histograms of the R, G, and B channels of the video frame. If the brightness difference between the channels is too large or does not conform to the normal ratio, it can be judged that there are problems of abnormal exposure or image color cast in the video frame.

[0055] When performing frozen frame detection, the server can compare the pixels of two consecutive video frames and simultaneously intercept the time stamp position information of the video frame. If the pixels of two consecutive video frames hardly change (i.e., less than a certain range) and the time shown by the time stamp has a certain shift, it can be judged that this video frame is a frozen frame.

[0056] When performing image blurring recognition, the server can use the edge gradient detection method to calculate the image sharpness. In one example, the server can use the Laplace operator to determine the image sharpness of a video frame according to the following formula:

[0057]

[0058] where f(x,y) is the pixel value of the video frame at the coordinate (x,y), G(x,y) is the gradient magnitude of this pixel point, and by setting a threshold T, when G(x,y) < T, it is determined as a video frame with low sharpness.

[0059] In one embodiment, based on the above image processing algorithm, the server can also combine OCR technology for auxiliary judgment. Specifically, the server can identify the time information in the video frame through OCR technology to determine whether defects such as black and white pictures are normal. For example, when the identified time information is during the night monitoring period, it can be considered that the black and white picture in the video frame is in a normal state, thus avoiding misjudgment.

[0060] In this way, through multi-level detection, the server can quickly and accurately determine video frames with obvious quality defects, so as to reduce the computational burden of subsequent processing.

[0061] Please continue to refer to Figure 1 , in step S130, multi-modal feature extraction is performed on video frames without quality defects to obtain corresponding image features.

[0062] Among them, video frames without quality defects can be video frames that are determined to have no obvious quality defects (such as signal loss, black and white images, abnormal exposure, etc.) after preliminary detection by the foregoing image processing algorithm. These video frames are relatively normal in appearance, but there may still be some potential problems that are not easily detected by traditional image processing algorithms and need to be further identified.

[0063] Multi-modal feature extraction can refer to the process of extracting multiple types of features (such as visual features of images, text features related to semantics, etc.) from the same video frame. By fusing the feature information of different modalities, the characteristics of the video frame can be more comprehensively and accurately characterized, providing a richer information basis for subsequent video quality assessment.

[0064] In this embodiment, those skilled in the art can pre-construct and train a multimodal large model, which can be a large-scale artificial intelligence model capable of simultaneously processing and understanding multiple types of data (such as text, images, audio, video, etc.). It can integrate the relevance of different modal information and achieve cross-modal semantic understanding and feature extraction through deep learning techniques. After determining that there are no obvious quality defects in the video frames, the server can input the video frames without obvious quality defects into the multimodal large model, and the multimodal large model can then perform multimodal feature extraction based on the input to obtain image features corresponding to the input video frames.

[0065] It should be noted that only a small amount of labeled data is required for pre-training the multimodal large model to adapt to new tasks (for example, only dozens of samples are required for new defect types), thereby reducing the difficulty of obtaining training data and ensuring the training effect of the model.

[0066] In one embodiment, the multimodal large model can be the Chinese-CLIP model pre-trained based on Chinese text-image data, which can simultaneously understand image content and Chinese semantic descriptions. Its core architecture is ViT-14L-336px (Vision Transformer, 14-layer network, input image resolution 336×336). The multimodal large model can map the input video frames into high-dimensional vectors, which can contain the semantic information and visual features of the video frames for subsequent matching.

[0067] It should be noted that those skilled in the art can also adopt other existing multimodal large models for feature extraction according to actual implementation needs, and the present disclosure does not make special limitations on this.

[0068] In step S140, the image features are matched with each preset feature in the feature library, and the preset features are used to describe various quality defect situations and normal display pictures.

[0069] Among them, the feature library can be a structured database pre-constructed by those skilled in the art, which can contain several preset features. These preset features can be used to describe various quality defect situations and normal display pictures. That is to say, a preset feature can correspond to a picture with a certain type of quality defect or a picture with normal display. In one example, those skilled in the art can pre-construct the feature library and store it using the vector database ChromaDB.

[0070] In this embodiment, after obtaining the image features corresponding to the video frames without obvious quality defects, the server can match them with each preset feature in the feature library respectively, so as to obtain the similarity between the image features and each preset feature. It should be understood that if the preset feature corresponds to a picture with a certain type of quality defect, the higher the similarity between it and the image features, the more likely it is that the video frame has the same quality defect; if the preset feature corresponds to a normally displayed picture, the higher the similarity between it and the image features, the higher the possibility that the video frame has no quality defect.

[0071] In one embodiment, when calculating the similarity between the image features and the preset features, the cosine similarity between the two can be calculated to determine the degree of similarity between them. It should be understood that the cosine similarity can be an index used to measure the similarity of the directions of two vectors, and its value range is [-1, 1]. The closer the value is to 1, the more similar the features are.

[0072] Please continue to refer to Figure 1 , in step S150, according to the matching result, a corresponding video quality evaluation result is generated.

[0073] In this embodiment, the server can generate a video quality evaluation result corresponding to the monitoring video stream according to the matching result between the subsequent image features and the preset features.

[0074] In one embodiment, those skilled in the art can pre-construct video quality evaluation rules based on possible matching results, so as to finally evaluate the video quality of the monitoring video stream based on the actual matching results and output the corresponding video quality evaluation result. For example, a basic score can be set for each monitoring video stream, and corresponding scores can be set for different types of quality defects. When the similarity between the image features and the preset features corresponding to a certain type of quality defect is greater than a certain threshold, the corresponding score is deducted from the basic score. In this way, the finally obtained video quality score is used as the video quality evaluation result of the monitoring video stream.

[0075] In one example, the video quality evaluation result can be output in the form of a report. The video quality evaluation report can include, but is not limited to, information such as the overall video quality score and the specific types of quality defects existing. That is to say, the video quality evaluation report can include the total quality evaluation score, the defect list (such as quality defect type, location, confidence level, recommended measures, etc.), visual annotation (such as marking the defect area in the video frame, a red box indicates a serious quality defect, a yellow box indicates a general quality defect, etc.).

[0076] Thus, based on Figure 1In the illustrated embodiment, first, based on an image processing algorithm, the extracted video frames are detected for obvious quality defects. If there are no obvious quality defects, then a multi-modal large model is combined for feature extraction. The extracted image features are matched with preset features in a pre-constructed feature library, so as to perform further video quality detection. Then, according to the defect detection result and / or the matching result, a corresponding video quality evaluation result is generated. Thus, the adaptability of the video quality evaluation method to complex situations can be improved, and further the accuracy of the video quality evaluation result can be ensured.

[0077] In one embodiment, since some video frames may be detected to have obvious quality defects during the aforementioned defect detection. Then the server can also generate a corresponding video quality evaluation result based on the defect detection result and the matching result. That is to say, those skilled in the art can pre-construct a video quality evaluation rule that combines the matching result and the defect detection result, so as to finally evaluate the video quality of the monitored video stream based on the detection result of the image processing algorithm and the matching result processed by the multi-modal large model, and output a corresponding video quality evaluation result. For example, the matching result and the defect detection result can be subjected to a weighted sum operation to obtain a corresponding video quality score as the video quality evaluation result, etc.

[0078] In some embodiments of the present disclosure, matching the image features with each preset feature in the feature library includes: for each quality defect, calculating the similarity between the image features and the preset features corresponding to the normal display picture and the picture with the quality defect respectively; performing normalization processing according to the calculated similarity to determine the confidence corresponding to the quality defect, and the confidence is used to represent the possibility that the video frame has the quality defect; if there are multiple preset features corresponding to the same quality defect, then take the confidence with the largest value among the values corresponding to the multiple preset features as the confidence corresponding to the quality defect.

[0079] In this embodiment, when matching the image features with the preset features, the server can detect each type of quality defect one by one. Specifically, for each quality defect, the server can first determine the preset feature corresponding to the quality defect, and then match the determined preset feature with the image features one by one, that is, calculate the similarity between the two.

[0080] It should be noted that due to the random differences in feature vectors, to ensure the effectiveness of the subsequent video quality evaluation result, the server can calculate the similarity between the image features and the preset features corresponding to the pictures with the detected quality defects and the preset features corresponding to the pictures with normal display respectively. Then, normalization processing is performed according to the two calculated similarity values to determine the confidence corresponding to the detected quality defect, and this confidence can be used to represent the possibility that this video frame has a quality defect.

[0081] For example, the image feature corresponding to the video frame is f i , and the preset feature corresponding to the quality defect of noise interference is f t , and the preset feature corresponding to the normally displayed picture is f n , the server first calculates the cosine similarity between f i and f n to obtain sim in , then calculates the cosine similarity between f i and f t to obtain sim it , then, after performing softmax normalization on sim in and sim it , softmax in and softmax it are obtained respectively. At this time, softmax it can be regarded as the confidence of the video frame for noise interference, that is, the probability that there is noise interference in the video frame.

[0082] In one example, the cosine similarity between the image feature and the preset feature can be calculated according to the following formula:

[0083]

[0084] where A and B are two feature vectors (such as the aforementioned f i and f n , f i and f t ).

[0085] In addition, the softmax normalization of the two cosine similarities can be performed according to the following formula:

[0086]

[0087] where z is the input vector (that is, the vector composed of the similarities calculated from the image feature and at least one preset feature), j represents the element position in the vector, K is the vector dimension, and here z are sim in and sim it .

[0088] After the above processing, the server can determine the confidence of the video frame relative to each quality defect. It should be noted that when there are multiple preset features corresponding to a certain quality defect, the server can use the maximum confidence value among the preset features corresponding to the quality defect as the confidence corresponding to the quality defect.

[0089] In some embodiments of the present disclosure, after determining the confidence level corresponding to the quality defect, the method further includes: if the value of the confidence level is within a preset threshold range, marking the corresponding video frame to trigger a review; and adding the video frame triggering the review to a feature library.

[0090] Among them, the preset threshold range can be a threshold range preset by those skilled in the art to determine whether a review is required. That is, this preset threshold range is the "uncertainty range" for the system to judge quality defects and requires a review. For example, if the preset threshold range is 60% - 70%, then when the confidence level corresponding to a certain quality defect is within this range, it indicates that a review is needed.

[0091] In this embodiment, after determining the confidence level of a video frame corresponding to a certain type of quality defect, the server can compare this confidence level with the values in the preset threshold range to determine whether the value of this confidence level is within the preset threshold range. If so, mark the corresponding video frame to trigger a review. In one example, a marker field (such as the quality defect and its corresponding confidence level, etc.) can be added to the metadata of the video frame, and then the video frame to be reviewed and its associated data (such as time stamp, camera ID, original image, etc.) are stored in a temporary database, and the manual review platform is notified through a message queue.

[0092] Then, the manual review platform can display the video frame to be reviewed and the system's judgment basis (such as similarity comparison graph, confidence level, etc.) in the interface. The reviewer then reviews the video frame to be reviewed according to the review annotation and confirms the review result. For example, after the review, if the video frame actually has the corresponding quality defect, the reviewer can mark the specific quality defect type or severity level, etc. If the video frame does not have this quality defect, the reviewer can reject the system's judgment and mark the reason for rejection (such as "misjudgment caused by light interference", etc.).

[0093] After the manual review, the server can also add the video frame triggering the review to the feature library, that is, it can extract features from this video frame through a multi-modal large model to add the extracted feature vectors to the feature library for direct comparison in subsequent algorithms, thereby continuously updating the detection ability of the system.

[0094] In this way, through the setting of the preset threshold range, corresponding fuzzy judgments can be intercepted to ensure the reliability of the final video quality assessment report. It should be noted that other larger or more highly trained models can also be used to automatically review the video frames to improve the review efficiency, and no special limitation is made in this regard.

[0095] In some embodiments of the present disclosure, the preset features include preset text features and preset image features;

[0096] The method further includes: constructing Chinese prompt words to describe various types of quality defects and normal display pictures, and extracting features of the Chinese prompt words to obtain corresponding preset text features; performing multi-modal feature extraction on pictures with various types of quality defects and normal display pictures collected in advance to obtain corresponding preset image features.

[0097] In this embodiment, the preset features may include preset text features and preset image features, both of which can be used to judge image quality defects. Specifically, when generating preset text features, several corresponding Chinese prompt words can be constructed for each quality defect or normal picture to describe various types of quality defects and normal display pictures. The Chinese prompt words can be as follows:

[0098] Noise interference: "Due to too high sensitivity, a large number of noise points appear".

[0099] Low contrast: "Lens contamination, ghosting, flare, light spot, smoke, local blurring and distortion"

[0100] Strip interference: "Signal interference, stripe noise, ripple noise".

[0101] Image occlusion: "Abnormal situation caused by large-area occlusion of the picture, usually manifested as a large number of leaves, branches or other occlusions in the picture".

[0102] Image color cast: "White balance error, color distortion, picture biased towards green",

[0103] "White balance error, color distortion, picture biased towards purple",

[0104] "White balance error, color distortion, picture biased towards orange",

[0105] "White balance error, color distortion, picture biased towards blue".

[0106] Normal picture: "Normal picture, clear and real, color balanced, brightness appropriate".

[0107] Then, use a multi-modal large model (such as the aforementioned Chinese-CLIP multi-modal model) to extract text features from each Chinese prompt word, so as to obtain the corresponding preset text features and store them in the feature library.

[0108] For the preset image features, pictures with various types of quality defects or normal display can be collected in advance, which can be obtained by collecting in the actual monitoring scenario. Then use a multi-modal large model to perform multi-modal feature extraction on each picture image, so as to obtain the corresponding preset image features and store them in the feature library.

[0109] It should be noted that although the preset features include preset text features and preset image features, due to the modality differences between text and images, during the aforementioned matching process, the preset text features and preset image features in the feature library need to be compared separately and cannot be used interchangeably.

[0110] In some embodiments of the present disclosure, based on a hierarchical matching strategy, different levels of feature libraries are assigned for matching the image features according to the complexity of the quality defects detected currently.

[0111] Among them, the hierarchical matching strategy can be a mechanism that dynamically selects different levels of feature libraries for matching according to the complexity of quality defects (such as single defects, combined defects, or subdivided defects).

[0112] In this embodiment, those skilled in the art can preset the hierarchical matching strategy according to prior experience. When constructing the feature library, it can be divided into a general feature library and a professional feature library. Among them, the general feature library can be used to store the preset features of common quality defects (such as blurring, frozen frames), so as to quickly perform a preliminary screening with a simple index structure. The professional feature library can be used to store the preset features of subdivided defects (such as 4 categories of color cast and 3 levels of noise), adopt a high-precision index, and support fine-grained matching.

[0113] In this way, according to the complexity of quality defects, different levels of feature libraries are used for matching and judgment. For images without special circumstances, the general feature library is used for quick detection; while for complex or subtle defects, a more subdivided professional feature library is used for in-depth detection. For example, for the color cast defect of an image, due to the differences in defect subdivision, the prompt words for various color casts will affect each other, resulting in a decline in detection ability. Therefore, it is split into 4 prompt words. If it is found during the preliminary detection or the detection by the general feature library that the image has a color cast, it is necessary to perform detection and comparison with these 4 prompt words respectively to clarify the specific defect points and defect confidence levels, otherwise this detection content can be ignored.

[0114] In some embodiments of the present disclosure, the method further includes: based on a business prompt word feature library, performing business analysis on the extracted content of the video frame to obtain a corresponding business analysis result.

[0115] Among them, the business prompt word feature library can be a database that stores Chinese description texts related to specific business scenarios and their corresponding feature vectors (such as "person breaking in", "fire smoke", "vehicle illegally parked"), and is used for business event recognition.

[0116] In this embodiment, those skilled in the art can define key event types according to business requirements (such as security, traffic management, etc.) and design corresponding Chinese prompt words. For example, for the security category, the corresponding Chinese prompt words can be "illegal entry of personnel", "package left behind", "climbing over the fence", etc.; for the traffic category, the corresponding Chinese prompt words can be "illegal lane change of vehicle", "pedestrian crossing the road", "traffic congestion", etc.

[0117] Then, use the multi-modal large model to convert the corresponding Chinese prompt words into corresponding feature vectors and store them in the business prompt word feature library.

[0118] In the actual business analysis process, the server can convert the extracted video frames into corresponding feature vectors through the multi-modal large model and compare them with each feature vector in the business prompt word feature library to obtain the corresponding business analysis results. The business analysis results can include but are not limited to event type, location, timestamp, and confidence score, etc.

[0119] In this way, the dual functions of quality defect detection and business monitoring can be realized, highlighting the high generalization ability of the multi-modal large model.

[0120] Based on the technical solution of the above embodiment, the following introduces a specific application scenario of the embodiment of the present application:

[0121] Figure 2 The flowchart of the video quality assessment method for surveillance videos according to another embodiment of the present disclosure is shown.

[0122] As Figure 2 shown, this method uses traditional image processing algorithms and multi-modal semantic information processing technology to achieve refined judgment and efficient evaluation of the quality of surveillance videos through a multi-stage detection and analysis process.

[0123] Specifically, first, perform video frame extraction processing on the obtained surveillance video stream. Subsequently, use traditional image processing algorithms to perform preliminary detection on the extracted video frames, identify and mark video frames with obvious defects, such as signal loss, black and white images, abnormal exposure, image color cast, frozen frames, and blurred images, etc.

[0124] For video frames without serious problems, they are further input into the open-source Chinese-CLIP multi-modal model for image feature extraction. In the multi-modal algorithm multi-level detection stage, the system constructs a feature library and stores it using a vector database, and compares the extracted image features with the text features and image features in the feature library.

[0125] By calculating the similarity and performing softmax processing, the system determines whether a specific defect type exists in the video frame and its confidence level. If the detection result meets a certain defect definition, the detection results of the traditional algorithm and the multi-modal model are further combined for comprehensive evaluation, and finally a quality evaluation report containing information such as the overall video quality score and specific defect types is output.

[0126] During the detection process, for samples with medium confidence, the system marks them as "to be reviewed" samples and feeds them back to professionals for manual review. The samples that pass the review will be incorporated into the image feature library to continuously update and optimize the detection ability of the system.

[0127] Based on any of the above embodiments, the present disclosure also provides a video quality evaluation device for surveillance videos.

[0128] Figure 3 It is a structural schematic block diagram of a video quality evaluation device for surveillance videos according to an embodiment of the present disclosure.

[0129] As Figure 3 shown, the video quality evaluation device for surveillance videos includes:

[0130] An extraction module for extracting video frames according to the acquired surveillance video stream;

[0131] A detection module for detecting defects in the extracted video frames to determine video frames without quality defects;

[0132] An extraction module for extracting multi-modal features from video frames without quality defects to obtain corresponding image features;

[0133] A matching module for matching the image features with each preset feature in the feature library, where the preset features are used to describe various quality defects and normal display pictures;

[0134] A processing module for generating corresponding video quality evaluation results according to the matching results.

[0135] The above video quality evaluation device for surveillance videos can be in the form of computer software, and each module of the above video quality evaluation device for surveillance videos can be implemented by computer software modules.

[0136] In some embodiments of the present disclosure, matching the image features with each preset feature in the feature library includes:

[0137] For each quality defect, calculating the similarity between the image features and the preset features corresponding to normal display pictures and pictures with the quality defect respectively;

[0138] Normalize the calculated similarity to determine the confidence corresponding to the quality defect, where the confidence is used to characterize the possibility that the video frame has the quality defect;

[0139] If there are multiple preset features corresponding to the same quality defect, take the confidence with the largest value among the values corresponding to the multiple preset features as the confidence corresponding to the quality defect.

[0140] In some embodiments of the present disclosure, after determining the confidence corresponding to the quality defect, the processing module is further configured to:

[0141] If the value of the confidence is within a preset threshold range, mark the corresponding video frame to trigger a review;

[0142] Add the video frame that triggers the review to the feature library.

[0143] In some embodiments of the present disclosure, the preset features include preset text features and preset image features;

[0144] The processing module is further configured to:

[0145] Construct Chinese prompt words to describe the pictures with various quality defects and normal displays, and extract features from the Chinese prompt words to obtain corresponding preset text features;

[0146] Perform multi-modal feature extraction on the pictures with various quality defects and normal displays collected in advance to obtain corresponding preset image features.

[0147] In some embodiments of the present disclosure, the matching module is further configured to:

[0148] Based on a hierarchical matching strategy, assign different levels of feature libraries for matching to the image features according to the complexity of the quality defect detected currently.

[0149] In some embodiments of the present disclosure, the processing module is further configured to:

[0150] Based on the business prompt word feature library, perform business analysis on the extracted picture content of the video frame to obtain a corresponding business analysis result.

[0151] Figure 4 The structure diagram of the computer system of the electronic device suitable for implementing the embodiments of the present application is shown.

[0152] It should be noted that Figure 4 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.

[0153] Such asFigure 4 As shown, the computer system includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 402 or the program loaded from the storage section 408 into the Random Access Memory (RAM) 403, such as executing the method described in the above embodiments. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An Input / Output (I / O) interface 405 is also connected to the bus 404.

[0154] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read from it can be installed into the storage section 408 as needed.

[0155] Specifically, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the Central Processing Unit (CPU) 401, various functions defined in the system of the present application are executed.

[0156] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0158] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0159] As another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the methods described in the above embodiments.

[0160] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0161] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented in software or in the form of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the methods according to the embodiments of this application.

[0162] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of this application. This application is intended to cover any variations, uses, or adaptations of this application, which follow the general principles of this application and include known common general knowledge or conventional technical means in the technical field not disclosed in this application.

[0163] It should be understood that this application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.

Claims

1. A method for evaluating the video quality of surveillance videos, characterized in that, Including: Performing video frame extraction based on the acquired surveillance video stream; Performing defect detection on the extracted video frames to determine the video frames without quality defects; Performing multi-modal feature extraction on the video frames without quality defects to obtain corresponding image features; Matching the image features with each preset feature in the feature library, where the preset features are used to describe various quality defects and normal display pictures; Generating a corresponding video quality evaluation result according to the matching result.

2. The method according to claim 1, wherein Matching the image features with each preset feature in the feature library includes: For each quality defect, calculating the similarity between the image features and the preset features corresponding to the normal display picture and the picture with the quality defect respectively; Performing normalization processing according to the calculated similarity to determine the confidence corresponding to the quality defect, where the confidence is used to represent the possibility that the video frame has the quality defect; If there are multiple preset features corresponding to the same quality defect, taking the confidence with the largest value corresponding to the multiple preset features as the confidence corresponding to the quality defect.

3. The method according to claim 2, wherein After determining the confidence corresponding to the quality defect, the method further includes: If the value of the confidence is within the preset threshold range, marking the corresponding video frame to trigger a review; Adding the video frames triggering the review to the feature library.

4. The method according to claim 2, wherein The preset features include preset text features and preset image features; The method further includes: Constructing Chinese prompt words to describe various quality defects and normal display pictures, and performing feature extraction on the Chinese prompt words to obtain corresponding preset text features; Performing multi-modal feature extraction on the pre-collected pictures with various quality defects and normal display pictures to obtain corresponding preset image features.

5. The method according to any one of claims 1-4, characterized in that, Based on a hierarchical matching strategy, allocating different levels of feature libraries for the image features to match according to the complexity of the currently detected quality defect.

6. The method according to any one of claims 1-4, characterized in that, The method further includes: Based on the business prompt word feature library, performing business analysis on the picture content of the extracted video frames to obtain a corresponding business analysis result.

7. A video quality assessment device for monitoring videos, characterized in that, Including: An extraction module for performing video frame extraction based on the acquired surveillance video stream; A detection module for performing defect detection on the extracted video frames to determine the video frames without quality defects; An extraction module for performing multi-modal feature extraction on the video frames without quality defects to obtain corresponding image features; A matching module for matching the image features with each preset feature in the feature library, where the preset features are used to describe various quality defects and normal display pictures; A processing module for generating a corresponding video quality evaluation result according to the matching result.

8. An electronic device, characterized in that, Including: A memory that stores execution instructions; And A processor that executes the execution instructions stored in the memory, so that the processor executes the video quality evaluation method for surveillance videos according to any one of claims 1 to 6.

9. A readable storage medium, characterized in that, The executable instructions are stored in the readable storage medium, and when executed by the processor, are used to implement the video quality evaluation method for surveillance videos according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for evaluating the video quality of a monitored video according to any one of claims 1 to 6.