Monitoring video quality evaluation method, device and equipment and computer readable medium
By framing and spatio-temporal sampling of the surveillance video, generating image feature information, and using neural network models for quality evaluation, the evaluation inaccuracy problem caused by multiple distortions of the surveillance video is solved, and higher evaluation accuracy is achieved.
Patent Information
- Application Number
- CN202510304693.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-04
AI Technical Summary
In the evaluation of monitoring video quality in the prior art, manual features or statistical features cannot effectively handle the various distortions generated by monitoring video during acquisition and transmission, resulting in poor evaluation accuracy.
By performing frame processing and spatiotemporal sampling of the monitoring video, a collection of image feature information is generated, and a pre-trained monitoring image quality evaluation model is used for evaluation, considering the distortion of the video at different times and spaces.
It improves the accuracy of surveillance video quality evaluation and can more comprehensively handle various distortion types such as real distortion, synthetic distortion and enhanced distortion.
Smart Images

Figure CN120259856A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to a method, apparatus, device, and computer-readable medium for evaluating the quality of surveillance videos. Background Art
[0002] Image quality evaluation technology is a method for evaluating the image quality of images or videos, aiming to quantify the visual effects of images and provide a more objective evaluation criterion for the quality of the image quality. Currently, when evaluating the quality of a video, the commonly adopted method is to perform reference-free image quality evaluation on each frame image of the video based on manual features or simple statistical characteristics, so as to determine the video quality.
[0003] However, when performing image quality evaluation in the above manner, there are often the following technical problems: Manual features or statistical characteristics are applicable to natural videos with fewer types of distortions. However, surveillance videos will generate different types of distortions (for example, real distortions, synthetic distortions, and enhancement distortions) during the processes of acquisition, transmission, and use, resulting in poor accuracy in evaluating the quality of surveillance videos.
[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0005] This content part of the present disclosure is used to briefly introduce the concepts, which will be described in detail in the following detailed implementation part. This content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure propose a method, apparatus, device, and computer-readable medium for evaluating the quality of surveillance videos to solve one or more of the technical problems mentioned in the above background art section.
[0007] In a first aspect, some embodiments of the present disclosure provide a method for evaluating the quality of surveillance videos. The method includes: obtaining a surveillance video as a video to be evaluated; performing frame splitting on the video to be evaluated to obtain a set of target video frames; performing spatio-temporal sampling on the set of target video frames to obtain a set of regional images; generating a set of image feature information according to the set of regional images and the set of target video frames, where the image feature information in the set of image feature information corresponds to the target video frames in the set of target video frames; generating surveillance video quality evaluation information corresponding to the video to be evaluated according to the set of image feature information and a pre-trained surveillance image quality evaluation model.
[0008] Optionally, perform frame splitting on the to-be-evaluated video to obtain a target video frame set, and perform spatio-temporal sampling on the to-be-evaluated video frame sequence to obtain a set of regional images, including: parsing the to-be-evaluated video to obtain a to-be-evaluated video frame sequence; extracting to-be-evaluated video frames as target video frames from the to-be-evaluated video frame sequence at at least one time interval to obtain a target video frame set.
[0009] Optionally, perform spatio-temporal sampling on the target video frame set to obtain a set of regional images, including: for each target video frame in the target video frame set, perform the following steps: perform salient region detection on the target video frame to generate salient region detection information, where the salient region detection information includes at least one salient region position information; perform cropping processing on the target video frame according to the at least one salient region position information to obtain at least one regional image; determine the obtained regional images as the set of regional images.
[0010] Optionally, perform cropping processing on the target video frame according to the at least one salient region position information to obtain at least one regional image, including: for each salient region position information in the at least one salient region position information, perform the following cropping steps: in response to determining that the size of the image salient region represented by the salient region position information is less than or equal to the preset input size, crop the target video frame according to the salient region position information to obtain a regional image; in response to determining that the size of the image salient region represented by the salient region position information is greater than the preset input size, generate respective regional sub-images according to the salient region position information and the preset input size; select a regional sub-image that satisfies the minimum coverage condition from the generated respective regional sub-images as the regional image.
[0011] Optionally, generate a set of image feature information according to the set of regional images and the target video frame set, including: for each target video frame in the target video frame set, perform global feature extraction on the target video frame to generate image global feature information; for each regional image in the set of regional images, perform regional feature extraction processing on the regional image to generate image regional feature information; for each generated image global feature information, perform feature fusion processing on the image global feature information and the corresponding respective image regional feature information to generate image feature information, to obtain a set of image feature information
[0012] Optionally, global feature extraction is performed on the above-mentioned target video frame to generate image global feature information, including: in response to determining that the image size of the above-mentioned target video frame is less than or equal to the network input size, performing feature extraction processing on the above-mentioned target video frame to generate image global feature information; in response to determining that the image size of the above-mentioned target video frame is greater than the network input size, performing the following feature extraction steps: extracting sub-images from the above-mentioned target video frame according to the network input size to generate each video frame sub-image; performing feature extraction processing on each video frame sub-image to generate each video frame sub-image feature; performing splicing processing on the above-mentioned each video frame sub-image feature to obtain image global feature information.
[0013] In a second aspect, some embodiments of the present disclosure provide a monitoring video quality evaluation device, the device includes: an acquisition unit configured to acquire a monitoring video as a video to be evaluated; a frame division unit configured to perform frame division processing on the above-mentioned video to be evaluated to obtain a set of target video frames; a spatio-temporal sampling unit configured to perform spatio-temporal sampling on the above-mentioned set of target video frames to obtain a set of regional images; a first generation unit configured to generate a set of image feature information according to the above-mentioned set of regional images and the above-mentioned set of target video frames, wherein the image feature information in the above-mentioned set of image feature information corresponds to the target video frames in the above-mentioned set of target video frames; a second generation unit configured to generate monitoring video quality evaluation information corresponding to the above-mentioned video to be evaluated according to the above-mentioned set of image feature information and a pre-trained monitoring image quality evaluation model.
[0014] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above first aspect.
[0015] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium on which a computer program is stored, wherein when the program is executed by a processor, the method described in any implementation manner of the above first aspect is implemented.
[0016] The above - mentioned various embodiments of the present disclosure have the following beneficial effects: The monitoring video quality evaluation method according to some embodiments of the present disclosure provides a video quality evaluation method for monitoring videos, which can take into account different types of distortions generated during the acquisition, transmission, and use of monitoring videos, thereby improving the accuracy of quality evaluation. Specifically, the reason for the inaccurate quality evaluation of monitoring videos is that manual features or statistical characteristics are applicable to natural videos with fewer types of distortions. However, different types of distortions (e.g., real distortion, synthetic distortion, and enhancement distortion) will occur during the acquisition and transmission of monitoring videos, resulting in poor accuracy of the quality assessment of monitoring videos. Based on this, the monitoring video quality evaluation method according to some embodiments of the present disclosure first obtains a monitoring video as the video to be evaluated. Then, the above - mentioned video to be evaluated is frame - processed to obtain a set of target video frames. Thus, the video can be converted into an image sequence through frame - processing and some images can be extracted as evaluation images, so as to evaluate the video quality in units of images. After that, spatio - temporal sampling is performed on the above - mentioned set of target video frames to obtain a set of regional images. Thus, by sampling in both the time and space dimensions, regions with representativeness, importance, and in line with human visual attention can be selected for analysis, so as to take into account the differences in image quality at different times and different positions of the video to be evaluated. Next, an image feature information set is generated based on the above - mentioned set of regional images and the above - mentioned set of target video frames. Among them, the image feature information in the above - mentioned image feature information set corresponds to the target video frames in the above - mentioned set of target video frames. Thus, compared with manual features, feature extraction based on a neural network model can extract image pixel features more comprehensively, so that the generated image feature information can cover different types of distortion situations. Finally, based on the above - mentioned image feature information set and a pre - trained monitoring image quality evaluation model, monitoring video quality evaluation information corresponding to the above - mentioned video to be evaluated is generated. Thus, through each image feature information representing the image quality at different time nodes and space nodes, different types of distortions generated during the acquisition and transmission of monitoring videos can be taken into account, thereby improving the accuracy of quality evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In combination with the accompanying drawings and with reference to the following specific embodiments, the above - mentioned and other features, advantages, and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0018] Figure 1 is a flowchart according to some embodiments of the monitoring video quality evaluation method of the present disclosure;
[0019] Figure 2It is a flowchart of some other embodiments of the monitoring video quality evaluation method according to the present disclosure;
[0020] Figure 3 It is a schematic structural diagram of some embodiments of the monitoring video quality evaluation device according to the present disclosure;
[0021] Figure 4 It is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners
[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0023] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0024] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0025] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0027] Embodiments of the present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0028] Figure 1 Flow 100 of some embodiments of the monitoring video quality evaluation method according to the present disclosure is shown. The monitoring video quality evaluation method includes the following steps:
[0029] Step 101, obtain a monitoring video as a video to be evaluated.
[0030] In some embodiments, the execution subject (such as a computing device) of the monitoring video quality evaluation method may obtain a monitoring video as the video to be evaluated. The above execution subject may be a server or a service end. The above monitoring video may be a video collected by a monitoring device and processed by compression. The above monitoring device may be a monitoring camera.
[0031] Step 102: Perform frame splitting on the video to be evaluated to obtain a set of target video frames.
[0032] In some embodiments, the above execution subject may perform frame splitting on the above video to be evaluated to obtain a set of target video frames. Among them, the target video frames in the above set of target video frames may be video frames extracted from the above video to be evaluated for image quality evaluation.
[0033] In some optional implementation manners of some embodiments, the above execution subject may perform frame splitting on the above video to be evaluated through the following steps to obtain a set of target video frames:
[0034] The first step: Perform parsing processing on the above video to be evaluated to obtain a sequence of video frames to be evaluated. In practice, the above execution subject may perform frame-by-frame parsing processing on the above video to be evaluated through relevant library functions (such as the cv2.VideoCapture() function in the OpenCV library in Python) or a video encoder (such as the Ffmpeg video encoder) to obtain a sequence of video frames to be evaluated.
[0035] The second step: Extract video frames to be evaluated as target video frames from the above sequence of video frames to be evaluated according to at least one time interval to obtain a set of target video frames. Among them, the above time interval may be a preset time length for intermittently extracting video frames from the above sequence of video frames to be evaluated. For example, the time intervals are 1s and 5s, indicating that a video frame to be evaluated is extracted from the above sequence of video frames to be evaluated every 1 second and 5 seconds as a target video frame. The duration of the above monitoring video may be 10s. Then the video times corresponding to the target video frames extracted through the above two time intervals are the 1st second, the 2nd second, the 5th second, the 6th second, and the 10th second. In practice, since the similarity between adjacent video frames is relatively high and there is a large amount of redundant information, by adopting the method of taking frames at intervals for the sequence of video frames to be evaluated, the redundant information and complexity in the video image quality evaluation process can be reduced.
[0036] Step 103: Perform spatio-temporal sampling on the set of target video frames to obtain a set of regional images.
[0037] In some embodiments, the above-mentioned execution entity may perform spatio-temporal sampling on the above-mentioned target video frame set to obtain a set of regional images. Among them, the regional images in the above-mentioned set of regional images may be images of significant regions cropped from the corresponding target video frames.
[0038] In some optional implementation manners of some embodiments, the above-mentioned execution entity may perform spatio-temporal sampling on the above-mentioned target video frame set through the following steps to obtain a set of regional images:
[0039] First, for each target video frame in the above-mentioned target video frame set, perform the following steps:
[0040] In the first step, perform significant region detection on the above-mentioned target video frame to generate significant region detection information. Among them, the above-mentioned significant region detection information includes at least one significant region position information. The above-mentioned significant region detection may be to detect the regions with objects or people and high-frequency information (such as image edges) in the above-mentioned target video frame through an object detection algorithm (for example, an object detection algorithm based on the YOLO series model) or an edge detection algorithm (such as the Sobel algorithm or the Canny algorithm). The above-mentioned significant region position information may be a set of image coordinates representing the range where the detected object or person is located in the above-mentioned target video frame.
[0041] In the second step, perform cropping processing on the above-mentioned target video frame according to the above-mentioned at least one significant region position information to obtain at least one regional image.
[0042] Then, determine the obtained regional images as the set of regional images.
[0043] In some optional implementation manners of some embodiments, the above-mentioned execution entity may perform cropping processing on the above-mentioned target video frame according to the above-mentioned at least one significant region position information through the following steps to obtain at least one regional image, including:
[0044] First, for each significant region position information in the above-mentioned at least one significant region position information, perform the following cropping steps:
[0045] First step, in response to determining that the size of the image significant region characterized by the above-mentioned significant region position information is less than or equal to the preset input size, crop the above-mentioned target video frame according to the above-mentioned significant region position information to obtain a region image. Wherein, the above-mentioned preset input size may be the image input size of the feature extraction network. As an example, the above-mentioned preset input size may be 224×224. In practice, in response to determining that the size of the image significant region characterized by the above-mentioned significant region position information is less than or equal to the preset input size, the above-mentioned execution entity may directly crop a video frame sub-image from the above-mentioned target video frame as the region image according to the size of the significant region characterized by the above-mentioned significant region position information.
[0046] Second step, in response to determining that the size of the image significant region characterized by the above-mentioned significant region position information is greater than the above-mentioned preset input size, generate each region sub-image according to the above-mentioned significant region position information and the above-mentioned preset input size. In practice, in response to determining that the size of the image significant region characterized by the above-mentioned significant region position information is greater than the above-mentioned preset input size, the above-mentioned execution entity may, within the image significant region characterized by the above-mentioned significant region position information, use the above-mentioned preset input size as the window size and crop each region sub-image through a sliding window.
[0047] Third step, select a region sub-image that meets the minimum coverage condition from each of the generated region sub-images as the region image. Wherein, the above-mentioned minimum coverage condition may be that the proportion of non-significant region pixels in the region image is the smallest. In practice, the above-mentioned execution entity may select a region sub-image that meets the minimum coverage condition from each of the generated region sub-images as the region image.
[0048] Step 104, generate a set of image feature information according to the set of region images and the set of target video frames.
[0049] In some embodiments, the above-mentioned execution entity may generate a set of image feature information according to the above-mentioned set of region images and the above-mentioned set of target video frames. Wherein, the image feature information in the above-mentioned set of image feature information corresponds to the target video frames in the above-mentioned set of target video frames. The image feature information in the above-mentioned set of image feature information may be the image feature vector of the corresponding target video frame.
[0050] In practice, first, the above-mentioned execution entity can use a feature extraction network to extract features from each regional image in the regional image set to obtain respective image region features. Then, the above-mentioned execution entity can use the feature extraction network to extract features from each target video frame in the target video frame set to obtain respective image global features. Finally, the above-mentioned execution entity can perform feature fusion on each obtained image global feature and the corresponding respective image region features (i.e., the image region features corresponding to the same target video) through a feature fusion network model to generate image feature information and obtain an image feature information set. Among them, the above-mentioned feature extraction network can be a neural network model for extracting image features. The above-mentioned feature extraction network can be, but is not limited to, a convolutional neural network or a residual neural network. The above-mentioned feature fusion network model can be a neural network model that takes two image feature vectors as inputs and outputs a fused image feature vector. As an example, the above-mentioned feature fusion network model can be a Transformer neural network or a derivative network model based on the Transformer network architecture or other improved neural network architectures based on the attention mechanism.
[0051] Step 105: Generate monitoring video quality evaluation information corresponding to the video to be evaluated according to the image feature information set and a pre-trained monitoring image quality evaluation model.
[0052] In some embodiments, the above-mentioned execution entity can generate monitoring video quality evaluation information corresponding to the video to be evaluated according to the above-mentioned image feature information set and a pre-trained monitoring image quality evaluation model. In practice, the above-mentioned execution entity can splice the image feature information in the above-mentioned image feature information set and input it into the above-mentioned monitoring image quality evaluation model to obtain a monitoring video quality score as the monitoring video quality evaluation information. Among them, the above-mentioned monitoring image quality evaluation model can be a prediction model for predicting the quality of monitoring video frame images or predicting the quality of monitoring video images. The above-mentioned monitoring image quality evaluation model can be a neural network model that takes image feature information as an input and outputs a monitoring video quality score. The above-mentioned monitoring video quality score can be a numerical value output by the fully connected layer of the model, representing the corresponding monitoring image quality or monitoring video quality. As an example, the above-mentioned monitoring image quality evaluation model can be a Transformer neural network model or a derivative network model based on the Transformer network architecture
[0053] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The monitoring video quality evaluation method according to some embodiments of the present disclosure provides a video quality evaluation method for monitoring videos, which can take into account different types of distortions generated during the acquisition, transmission, and use of monitoring videos, thereby improving the accuracy of quality evaluation. Specifically, the reasons for inaccurate monitoring video quality evaluation are as follows: Manual features or statistical characteristics are applicable to natural videos with fewer types of distortions. However, different types of distortions (such as real distortion, synthetic distortion, and enhancement distortion) will be generated during the acquisition, transmission, and use of monitoring videos, resulting in poor accuracy of the quality assessment of monitoring videos. Based on this, the monitoring video quality evaluation method according to some embodiments of the present disclosure first obtains a monitoring video as a video to be evaluated. Then, the above-mentioned video to be evaluated is frame-processed to obtain a set of target video frames. Thus, the video can be converted into an image sequence through frame processing and some images can be extracted as evaluation images, so as to evaluate the video quality in units of images. After that, spatio-temporal sampling is performed on the above-mentioned set of target video frames to obtain a set of regional images. After that, spatio-temporal sampling is performed on the above-mentioned sequence of video frames to be evaluated to obtain a set of regional images. Thus, by sampling in both the time and space dimensions, regions with representativeness, importance, and in line with human visual attention can be selected for analysis, so as to take into account the differences in image quality at different times and different positions of the video to be evaluated. Then, according to the above-mentioned set of regional images and the above-mentioned set of target video frames, a set of image feature information is generated. Among them, the image feature information in the above-mentioned set of image feature information corresponds to the target video frames in the above-mentioned set of target video frames. Thus, compared with manual features, the feature extraction based on the neural network model can extract image pixel features more comprehensively, so that the generated image feature information can cover different types of distortion situations. Finally, according to the above-mentioned set of image feature information and a pre-trained monitoring image quality evaluation model, monitoring video quality evaluation information corresponding to the above-mentioned video to be evaluated is generated. Thus, through each piece of image feature information representing the image quality at different time nodes and space nodes, different types of distortions generated during the acquisition and transmission of the monitoring video can be taken into account, thereby improving the accuracy of quality evaluation.
[0054] Further referring to Figure 2 , which shows the flow 200 of some other embodiments of the monitoring video quality evaluation method. The flow 200 of this monitoring video quality evaluation method includes the following steps:
[0055] Step 201, obtain a monitoring video as a video to be evaluated.
[0056] In some embodiments, the above-mentioned execution subject may obtain a monitoring video as a video to be evaluated.
[0057] Step 202: Perform frame splitting on the video to be evaluated to obtain a set of target video frames.
[0058] In some embodiments, the above-mentioned execution entity may perform frame splitting on the above-mentioned video to be evaluated to obtain a sequence of video frames to be evaluated.
[0059] Step 203: Perform spatio-temporal sampling on the set of target video frames to obtain a set of regional images.
[0060] In some embodiments, the above-mentioned execution entity may perform spatio-temporal sampling on the sequence of video frames to be evaluated to obtain a set of regional images.
[0061] In some embodiments, the specific implementation manners and the technical effects brought by the above steps 201-203 may refer to Figure 1 Steps 101-103 in the corresponding embodiments, which will not be elaborated here.
[0062] Step 204: For each target video frame in the set of target video frames, perform global feature extraction on the target video frame to generate image global feature information.
[0063] In some embodiments, for each target video frame in the set of target video frames, the above-mentioned execution entity may perform global feature extraction on the target video frame to generate image global feature information. Among them, the above-mentioned image global feature information may be an image global feature vector.
[0064] In some optional implementation manners of some embodiments, the above-mentioned execution entity may perform global feature extraction on the above-mentioned target video frame through the following steps to generate image global feature information:
[0065] First step, in response to determining that the image size of the above-mentioned target video frame is less than or equal to the network input size, perform feature extraction processing on the above-mentioned target video frame to generate image global feature information. In practice, in response to determining that the image size of the above-mentioned target video frame is less than or equal to the network input size, the above-mentioned execution entity may directly perform feature extraction processing on the above-mentioned target video frame through a feature extraction network to generate image global feature information (i.e., an image global feature vector). Among them, the above-mentioned network input size may be the maximum image input size of the feature extraction network. The above-mentioned feature extraction network may be a neural network model for extracting image features. As an example, the above-mentioned feature extraction network may be, but is not limited to: a convolutional neural network or a residual neural network.
[0066] Second step, in response to determining that the image size of the above-mentioned target video frame is greater than the network input size, perform the following feature extraction steps:
[0067] The first sub-step is to extract sub-images from the above-mentioned target video frames according to the network input size to generate sub-images of each video frame. In practice, the above-mentioned execution entity can use the above-mentioned network input size as the window size and extract sub-images from the above-mentioned target video frames through a non-overlapping sliding window to generate sub-images of each video frame.
[0068] The second sub-step is to perform feature extraction processing on each sub-image of the video frame to generate feature information of each sub-image of the video frame. In practice, the above-mentioned execution entity can perform feature extraction processing on each sub-image of the video frame through the above-mentioned feature extraction network to generate feature information of each sub-image of the video frame.
[0069] The third sub-step is to perform splicing processing on the above-mentioned feature information of each sub-image of the video frame to obtain global image feature information. In practice, the above-mentioned execution entity can perform sequential splicing processing on the above-mentioned feature information of each sub-image of the video frame according to the position order of each sub-image of the video frame in the above-mentioned target video frame to obtain global image feature information.
[0070] Step 205: For each region image in the region image set, perform region feature extraction processing on the region image to generate image region feature information.
[0071] In some embodiments, for each region image in the region image set, the above-mentioned execution entity can perform region feature extraction processing on the region image through the above-mentioned feature extraction network to generate image region feature information (i.e., image region feature vector).
[0072] Step 206: For each generated global image feature information, perform feature fusion processing on the global image feature information and the corresponding respective image region feature information to generate image feature information, obtaining a set of image feature information.
[0073] In some embodiments, for each generated global image feature information, the above-mentioned execution entity can perform feature fusion processing on the global image feature information and the corresponding respective image region feature information (i.e., image region feature information corresponding to the same target video frame) to generate image feature information, obtaining a set of image feature information. In practice, the above-mentioned execution entity can perform feature fusion processing on each obtained global image feature information and the corresponding respective image region features (i.e., image region features corresponding to the same target video) through a feature fusion network model to generate image feature information, obtaining a set of image feature information. The above-mentioned feature fusion network model can be a neural network model that takes two image feature vectors as inputs and outputs a fused image feature vector. As an example, the above-mentioned feature fusion network model can be a Transformer neural network or a derivative network model based on the Transformer network architecture or other improved neural network architectures based on the attention mechanism.
[0074] Step 207: Generate monitoring video quality evaluation information corresponding to the video to be evaluated according to the image feature information set and the pre-trained monitoring image quality evaluation model.
[0075] In some embodiments, the above-mentioned execution entity can generate monitoring video quality evaluation information corresponding to the video to be evaluated according to the image feature information set and the pre-trained monitoring image quality evaluation model. In practice, the above-mentioned execution entity can splice the image feature information in the above-mentioned image feature information set and input it into the above-mentioned monitoring image quality evaluation model to obtain a monitoring video quality score as the monitoring video quality evaluation information.
[0076] Further referring to Figure 3 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a monitoring video quality evaluation device. These device embodiments correspond to Figure 1 the method embodiments shown, and the monitoring video quality evaluation device can be specifically applied to various electronic devices.
[0077] As shown in Figure 2 , some embodiments of the monitoring video quality evaluation device 200 include: an acquisition unit 301, a frame division unit 302, a spatio-temporal sampling unit 303, a first generation unit 304, and a second generation unit 305. Among them, the acquisition unit 301 is configured to acquire a monitoring video as the video to be evaluated; the frame division unit 302 is configured to perform frame division processing on the above-mentioned video to be evaluated to obtain a target video frame set; the spatio-temporal sampling unit 303 is configured to perform spatio-temporal sampling on the above-mentioned target video frame set to obtain a regional image set; the first generation unit 304 is configured to generate an image feature information set according to the above-mentioned regional image set and the above-mentioned target video frame set, where the image feature information in the above-mentioned image feature information set corresponds to the target video frames in the target video frame set; the second generation unit 305 is configured to generate monitoring video quality evaluation information corresponding to the above-mentioned video to be evaluated according to the above-mentioned image feature information set and the pre-trained monitoring image quality evaluation model.
[0078] It can be understood that the units described in the monitoring video quality evaluation device 300 correspond to the respective steps in the method described with reference to Figure 1 . Therefore, the operations, features, and beneficial effects described above for the method also apply to the web page generation device 300 and the units included therein, and will not be repeated here.
[0079] Next, referring to Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing some embodiments of the present disclosure. Figure 4The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0080] As Figure 4 shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to the programs stored in the read-only memory 402 or the programs loaded from the storage device 408 into the random access memory 403. In the random access memory 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the read-only memory 402, and the random access memory 403 are connected to each other through a bus 404. The input / output interface 405 is also connected to the bus 404.
[0081] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 can allow the electronic device 400 to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 3 the electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively. Figure 3 Each block shown in
[0082] particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the read-only memory 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are executed.
[0083] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0084] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0085] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain a surveillance video as a video to be evaluated; perform frame splitting on the above video to be evaluated to obtain a set of target video frames; perform spatio-temporal sampling on the above set of target video frames to obtain a set of regional images; generate a set of image feature information according to the above set of regional images and the above set of target video frames, wherein the image feature information in the above set of image feature information corresponds to the target video frames in the above set of target video frames; generate surveillance video quality evaluation information corresponding to the above video to be evaluated according to the above set of image feature information and a pre-trained surveillance image quality evaluation model.
[0086] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0088] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a framing unit, a spatio-temporal sampling unit, a first generation unit, and a second generation unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring a surveillance video as a video to be evaluated".
[0089] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0090] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for evaluating the quality of surveillance videos, comprising: Obtaining a surveillance video as a video to be evaluated; Performing frame division processing on the video to be evaluated to obtain a set of target video frames; Performing spatio-temporal sampling on the set of target video frames to obtain a set of regional images; Generating a set of image feature information according to the set of regional images and the set of target video frames, wherein the image feature information in the set of image feature information corresponds to the target video frames in the set of target video frames; Generating surveillance video quality evaluation information corresponding to the video to be evaluated according to the set of image feature information and a pre-trained surveillance image quality evaluation model.
2. The method according to claim 1, wherein, The performing frame division processing on the video to be evaluated to obtain a set of target video frames includes: Performing parsing processing on the video to be evaluated to obtain a sequence of video frames to be evaluated; Extracting video frames to be evaluated as target video frames from the sequence of video frames to be evaluated at at least one time interval to obtain a set of target video frames.
3. The method according to claim 2, wherein The performing spatio-temporal sampling on the set of target video frames to obtain a set of regional images includes: For each target video frame in the set of target video frames, perform the following steps: Performing significant region detection on the target video frame to generate significant region detection information, wherein the significant region detection information includes at least one significant region position information; Performing cropping processing on the target video frame according to the at least one significant region position information to obtain at least one regional image; Determining the obtained regional images as the set of regional images.
4. The method according to claim 3, wherein, The performing cropping processing on the target video frame according to the at least one significant region position information to obtain at least one regional image includes: For each significant region position information in the at least one significant region position information, perform the following cropping steps: In response to determining that the size of the significant region of the image represented by the significant region position information is less than or equal to a preset input size, cropping the target video frame according to the significant region position information to obtain a regional image; In response to determining that the size of the significant region of the image represented by the significant region position information is greater than the preset input size, generating respective regional sub-images according to the significant region position information and the preset input size; Selecting a regional sub-image that satisfies the minimum coverage condition from the generated respective regional sub-images as the regional image.
5. The method according to claim 1, wherein The generating a set of image feature information according to the set of regional images and the set of target video frames includes: For each target video frame in the set of target video frames, performing global feature extraction on the target video frame to generate image global feature information; For each regional image in the set of regional images, performing regional feature extraction processing on the regional image to generate image regional feature information; For each generated image global feature information, performing feature fusion processing on the image global feature information and the corresponding respective image regional feature information to generate image feature information, obtaining a set of image feature information.
6. The method according to claim 5, wherein The performing global feature extraction on the target video frame to generate image global feature information includes: In response to determining that the image size of the target video frame is less than or equal to the network input size, perform feature extraction processing on the target video frame to generate image global feature information; In response to determining that the image size of the target video frame is greater than the network input size, perform the following feature extraction steps: According to the network input size, perform sub-image extraction on the target video frame to generate respective video frame sub-images; Perform feature extraction processing on each video frame sub-image to generate respective video frame sub-image features; Perform splicing processing on the respective video frame sub-image features to obtain image global feature information.
7. A monitoring video quality evaluation device, comprising: An acquisition unit configured to acquire a monitoring video as a video to be evaluated; A frame division unit configured to perform frame division processing on the video to be evaluated to obtain a target video frame set; A spatio-temporal sampling unit configured to perform spatio-temporal sampling on the target video frame set to obtain a regional image set; A first generation unit configured to generate an image feature information set according to the regional image set and the target video frame set, wherein the image feature information in the image feature information set corresponds to the target video frames in the target video frame set; A second generation unit configured to generate monitoring video quality evaluation information corresponding to the video to be evaluated according to the image feature information set and a pre-trained monitoring image quality evaluation model.
8. An electronic device, comprising: One or more processors; A storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 6.