Method and apparatus for detecting quality of video stream, electronic device and storage medium

By fusing quality features and performing time-domain-frequency domain transformation on the current and historical frames of the video stream, and combining attention mechanisms and long short-term memory networks, the problem of low accuracy in video stream quality detection is solved, achieving efficient and accurate detection of underground monitoring video streams.

CN120612279BActive Publication Date: 2026-02-06CHINA COAL RES INST +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510495078.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2026-02-06
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Existing technologies for video stream quality detection have low accuracy, especially in complex image scenarios where they are prone to missed or false positives, and they also consume a lot of computational resources.

Method used

By performing quality detection on the current frame image of the target video stream, fusing the target quality features of historical frame images and the initial quality features of the current frame image, performing time-domain-frequency domain transformation, and combining attention mechanism and long short-term memory network, the detection accuracy is improved.

Benefits of technology

It achieves comprehensive and dynamic video stream quality inspection, improves the accuracy of inspection, is applicable to underground monitoring video stream scenarios, and reduces the consumption of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612279B_ABST
    Figure CN120612279B_ABST
Patent Text Reader

Abstract

The application provides a quality detection method and device of a video stream, electronic equipment and a storage medium. The method comprises the following steps: performing quality detection on a current frame image of a target video stream to obtain initial quality features of the current frame image; fusing target quality features of historical frame images and the initial quality features of the current frame image to obtain first time domain features; performing time domain-frequency domain conversion on the first time domain features to obtain first frequency domain features; obtaining target quality features of the current frame image based on the first frequency domain features; and obtaining a target quality detection result of the current frame image based on the target quality features of the current frame image. Thus, the target quality features of the historical frame images and the initial quality features of the current frame image are comprehensively considered to obtain the first time domain features, then the time domain-frequency domain conversion is performed to obtain the first frequency domain features, and the current frame image is detected based on the first frequency domain features, so that the accuracy of the video stream quality detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly relates to a quality detection method and device of a video stream, an electronic device and a storage medium. BACKGROUND

[0002] At present, in application scenarios such as security monitoring, video conference, remote education and film and television production, quality detection of a video stream is mostly needed, such as detection of whether a picture blur or picture occlusion occurs in the video stream. However, in the related art, the accuracy of the video stream quality detection is low. SUMMARY

[0003] The present application aims to at least partly solve one of the technical problems in the related art.

[0004] To this end, a first object of the present application is to provide a quality detection method of a video stream.

[0005] A second object of the present application is to provide a quality detection device of a video stream.

[0006] A third object of the present application is to provide an electronic device.

[0007] A fourth object of the present application is to provide a computer-readable storage medium.

[0008] A fifth object of the present application is to provide a computer program product.

[0009] To achieve the above objects, a quality detection method of a video stream is provided in an embodiment of the first aspect of the present application, comprising: performing quality detection on a current frame image of a target video stream to obtain an initial quality feature of the current frame image; fusing a target quality feature of a historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time-domain feature; performing time-domain-frequency domain conversion on the first time-domain feature to obtain a first frequency-domain feature; obtaining a target quality feature of the current frame image based on the first frequency-domain feature; and obtaining a target quality detection result of the current frame image based on the target quality feature of the current frame image.

[0010] To achieve the above object, the second aspect of the present application proposes a quality detection device of a video stream, comprising: a first detection module, configured to detect a quality of a current frame image of a target video stream to obtain an initial quality feature of the current frame image; a fusion module, configured to fuse a target quality feature of a historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time domain feature; a conversion module, configured to perform time domain-frequency domain conversion on the first time domain feature to obtain a first frequency domain feature; a second detection module, configured to obtain a target quality feature of the current frame image based on the first frequency domain feature; and a third detection module, configured to obtain a target quality detection result of the current frame image based on the target quality feature of the current frame image.

[0011] To achieve the above object, the third aspect of the present application proposes an electronic device, comprising: a processor, and a memory connected with the processor in communication; the memory stores computer execution instructions; and the processor executes the computer execution instructions stored in the memory to implement the quality detection method of the video stream as described in the first aspect of the present application.

[0012] To achieve the above object, the fourth aspect of the present application proposes a computer readable storage medium, wherein the computer readable storage medium stores computer execution instructions; and the computer execution instructions are executed by a processor to implement the quality detection method of the video stream as described in the first aspect of the present application.

[0013] To achieve the above object, the fifth aspect of the present application proposes a computer program product, comprising a computer program; and the computer program is executed by a processor to implement the quality detection method of the video stream as described in the first aspect of the present application.

[0014] The method, device, electronic device and storage medium provided by the application can comprehensively consider the target quality features of the historical frame images and the initial quality features of the current frame image to obtain the first time domain features, then perform time domain-frequency domain conversion to obtain the first frequency domain features, and consider the first frequency domain features to perform quality detection on the current frame image, thereby improving the accuracy of video stream quality detection compared with the related art which only performs quality detection on a frame image, and realizing comprehensive and dynamic quality detection of the video stream, and the method is suitable for the quality detection of downhole monitoring video stream.

[0015] Additional aspects and advantages of the application will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following and / or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0016] The above and / or additional aspects and advantages of the application will become apparent and be readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:

[0017] Figure 1 A flowchart of a video stream quality detection method provided by an embodiment of the application;

[0018] Figure 2 A flowchart of another video stream quality detection method provided by an embodiment of the application;

[0019] Figure 3 A flowchart of another video stream quality detection method provided by an embodiment of the application;

[0020] Figure 4 A flowchart of another video stream quality detection method provided by an embodiment of the application;

[0021] Figure 5 A schematic diagram of a video stream quality detection method provided by an embodiment of the application;

[0022] Figure 6 A schematic diagram of a time domain-frequency domain detection module provided by an embodiment of the application;

[0023] Figure 7A schematic diagram of a frequency domain attention unit provided by an embodiment of the present application.

[0024] Figure 8 A structural schematic diagram of a video stream quality detection device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, in which the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.

[0026] With the advancement of industrialization and intelligentization, the detection of working conditions through a monitoring system and the supervision of personnel have become more and more popular, which can save human and material resources. Since a large number of monitoring cameras are deployed in different corners, the deployment environment of many cameras is very poor, which is easy to cause different picture problems, such as picture blur, picture color cast, picture occlusion, and strong light direct radiation. When the corresponding problems occur, it will affect the judgment of the observation content by artificial or intelligent algorithm, so timely troubleshooting is needed. However, under the condition of observing more video cameras at the same time, the efficiency of human judgment of picture abnormality problems is relatively low. How to accurately judge through algorithm needs to design the algorithm from the root cause of the problem and the structural problem of the problem picture.

[0027] The traditional solution to picture abnormality problems is to judge by artificial vision, which is subjective and difficult to ensure consistency and accuracy. Even if a computer-aided judgment scheme is adopted, most of them use relatively simple algorithms and indicators. When facing complex images, it often shows insufficient performance and has problems such as difficult accurate identification. When light and other factors change in the picture and external factors interfere, problems such as missed judgment and misjudgment occur. With the gradual maturity of artificial intelligence methods, the solution to this problem can be transformed into an image classification problem. This method needs to create a corresponding data set and label the data set, and use a model for training. The main difficulty of this method is data set collection, and how to ensure that the collected data set is representative is a big problem.

[0028] How to design an algorithm scheme suitable for the current scene to accurately judge and analyze video data and real-time prompt abnormalities is a very meaningful topic in the field of image processing.

[0029] In addition, the process is equivalent to pre-processing the video, which consumes some computing resources. In some scenarios where intelligent algorithms are required, if the picture itself is abnormal, it will hinder the further implementation of functions, causing abnormality or failure in subsequent judgment. Therefore, under the premise of ensuring that it does not affect the application of subsequent intelligent algorithms, solving the preprocessing process with reasonable computing resources is also one of the core problems of the current intelligent analysis system.

[0030] To solve the above problems, the embodiment of the present application provides a video stream quality detection method. The method comprises the following steps: performing quality detection on a current frame image of a target video stream to obtain initial quality features of the current frame image; fusing target quality features of historical frame images of the target video stream and the initial quality features of the current frame image to obtain first time domain features; performing time domain-frequency domain conversion on the first time domain features to obtain first frequency domain features; obtaining target quality features of the current frame image based on the first frequency domain features; and obtaining a target quality detection result of the current frame image based on the target quality features of the current frame image. Thus, the target quality features of the historical frame images and the initial quality features of the current frame image are comprehensively considered to obtain the first time domain features, and then the time domain-frequency domain conversion is performed to obtain the first frequency domain features. The quality of the current frame image is detected based on the first frequency domain features. Compared with the related art which only detects the quality of a frame image based on the frame image, the accuracy of video stream quality detection is improved, and comprehensive and dynamic quality detection of the video stream can be realized, which is suitable for the quality detection of the downhole monitoring video stream.

[0031] The present application focuses on the causes of the problems and designs algorithms based on the structural features of the problem pictures to screen the problems. Most of the current algorithms are designed to solve a single problem and are deployed for a specific problem scenario. The present application solves several common picture problems and designs an algorithm processing flow to detect and warn abnormal pictures as much as possible while occupying few system resources. The present application can maintain good detection effect in various scenarios and can detect various common abnormalities in the mine, such as strong light direct radiation, fog obstruction, dirty lens, and insufficient light.

[0032] The video stream quality detection method, device, electronic equipment and storage medium provided by the embodiment of the present application are described below with reference to the accompanying drawings.

[0033] Figure 1 A flowchart of a video stream quality detection method provided by the embodiment of the present application is shown in the figure.

[0034] As Figure 1 shown, the method comprises the following steps:

[0035] S101, performing quality detection on a current frame image of a target video stream to obtain initial quality features of the current frame image.

[0036] It should be noted that the target video stream is not limited too much, for example, taking the underground monitoring scene as an example, the target video stream can include the video stream collected by the underground camera.

[0037] The initial quality feature refers to the initial quality feature of the image. The quality feature is not limited too much, for example, it can include whether the jth quality problem occurs, the abnormality degree of the occurrence of the jth quality problem, the score in each quality dimension, etc. The categories of quality problems include blur, occlusion, noise, underexposure, overexposure (hereinafter referred to as overexposure), overdarkness, overbrightness, distortion, artifact, etc. The quality dimensions include sharpness, contrast, brightness, color cast, exposure, etc. j is a positive integer not greater than M, and M is the number of quality problem categories.

[0038] The abnormality degree of different types of quality problems can be represented by different indicators.

[0039] For example, lens dirt can cause the image to have an occlusion type quality problem, in which case the abnormality degree of the occlusion type quality problem can be represented by a dirty occlusion area percentage, which refers to the proportion of the lens dirt area in the image to the total area of the image multiplied by 100%.

[0040] For example, fog occlusion can cause the image to have an occlusion type quality problem, in which case the abnormality degree of the occlusion type quality problem can be represented by an average picture light transmittance, which refers to the average value of the light transmittance of multiple pixel points in the image.

[0041] For example, strong light direct radiation can cause the image to have an overexposure type quality problem, in which case the abnormality degree of the overexposure type quality problem can be represented by a strong light area percentage, which refers to the proportion of the strong light area in the image to the total area of the image multiplied by 100%.

[0042] For example, insufficient illumination can cause the image to have an overdarkness type quality problem, in which case the abnormality degree of the overdarkness type quality problem can be represented by an image darkness.

[0043] The quality detection of the current frame image can use any image quality detection method in related technologies, which is not limited too much here.

[0044] Optionally, the quality detection of the current frame image of the target video stream obtains the initial quality feature of the current frame image, including feature extraction of the current frame image to obtain the image feature of the current frame image, and obtaining the initial quality feature of the current frame image based on the image feature of the current frame image.

[0045] Optionally, the quality detection model comprises an image classification module, and the current frame image is input to the image classification module, and initial quality features of the current frame image are output by the image classification module. The image classification module can be implemented by using any image classification network in the related art, and will not be limited here.

[0046] For example, as shown in FIG. 10, when monitoring frame t=t0-1, that is, the current frame is the (t0-1)th frame, the (t0-1)th frame image is input to the image classification module, and the image abnormal type code and the image abnormal degree information of the (t0-1)th frame image are output by the image classification module as initial quality features of the (t0-1)th frame image. Figure 5

[0047] When monitoring frame t=t0, that is, the current frame is the t0th frame, the t0th frame image is input to the image classification module, and the image abnormal type code and the image abnormal degree information of the t0th frame image are output by the image classification module as initial quality features of the t0th frame image.

[0048] The image abnormal type code is used to indicate the quality problem category of the image, and the image abnormal degree information is used to indicate the abnormal degree of the image in the jth quality problem category. The image abnormal type code output by the image classification module is the initial abnormal type code of the image, and the image abnormal degree information output by the image classification module is the initial abnormal degree information of the image.

[0049] For example, the image classification module comprises a down-sampling layer, a convolution feature extraction layer, a full connection layer and a Softmax layer, the current frame image is down-sampled by the down-sampling layer to obtain a down-sampled image, the down-sampled image is extracted by the convolution feature extraction layer to obtain a plurality of initial feature maps, the plurality of initial feature maps are integrated by the full connection layer to obtain a target feature map, and the initial quality features of the current frame image are obtained based on the target feature map by the Softmax layer.

[0050] Firstly, the down-sampling layer effectively reduces the calculation complexity and improves the processing efficiency by reducing the resolution of the image. Down-sampling not only reduces the size of the image, but also removes some unnecessary details and retains the most important structural information in the image. For example, the down-sampling layer can comprise an AVG (Average Pooling) layer, and using average pooling can effectively reduce the calculation amount and enhance the invariance of the model to translation and deformation.

[0051] For example, the processing process of down-sampling the current frame image by the down-sampling layer to obtain a down-sampled image is as follows:

[0052] A2=avg_pool(A1,poolsize=2)

[0053] ​wherein, A1 is the current frame image, A2 is the down-sampled image, avg_pool(·) is an average pooling function, and poolsize = 2 represents that the size of the pooling window is 2*2. If then At this time, B is the batch size of the current frame image, i.e., the number of images processed at a time, C is the number of channels of the current frame image, H is the height of the current frame image, and W is the width of the current frame image.

[0054] In addition, the convolutional down-sampling method is also applied to the down-sampling process of the image. The stride adjustment of the convolution kernel enables the down-sampling operation to not only reduce the resolution of the image but also extract the key features of the image, thereby ensuring the effect of subsequent feature extraction.

[0055] Next, the role of the convolution feature extraction layer is to extract different levels of features of the down-sampled image through a multi-layer convolutional neural network. In this process, the convolution operation can capture local information such as edges, textures, and colors in the image through sliding convolution of the down-sampled image by multiple convolution kernels.

[0056] For example, the process of obtaining multiple initial feature maps through multiple feature extractions of the down-sampled image by the convolution feature extraction layer is as follows:

[0057] A3 = Conv2d(A2, stride = 2, C out = 2xC in )

[0058] wherein, A3 is the initial feature map, Conv2d(·) is a two-dimensional convolution function, stride = 2 represents that the convolution kernel moves 2 pixels at a time, i.e., the step length of the convolution kernel is 2, and C out = 2xC in represents that the output channel number is 2 times the input channel number. If then At this time, B is the batch size of the down-sampled image, i.e., the number of images processed at a time, C is the number of channels of the down-sampled image, H is the height of the down-sampled image, and W is the width of the down-sampled image.

[0059] With the deepening of the convolution layer, the network can gradually extract from simple low-level features to more complex high-level features. By increasing the feature channels of the convolution layer, the network can extract multiple features in parallel, so that each layer can focus on different visual information. The size and step length of each convolution kernel directly affect the effect of feature extraction, especially multi-scale feature extraction, which enables the network to understand image content at different scales, thereby improving the recognition ability of diversified images.

[0060] The features extracted by the convolutional layer need to be further integrated and classified, which is done by the fully connected layer. The fully connected layer flattens the multi-dimensional feature map output by the convolutional layer and maps the local features to the global feature space through full connection operation. Through the feature aggregation in this stage, the network can recognize the overall pattern in the image and perform classification. The fully connected layer can obtain the whole image feature information by processing the features in different positions in parallel.

[0061] For example, the processing process of obtaining the target feature map by integrating multiple initial feature maps through the fully connected layer is as follows:

[0062] A4 = fully_connect(reshape(A3))

[0063] Where A4 is the target feature map, reshape(·) is a reshaping function, such as flattening A3 into a two-dimensional tensor, and if then the reshape operation reshapes the At this time, each initial feature map is flattened into a one-dimensional vector with a length of CxHxW. At this time, B is the batch size of the initial feature map, i.e. the number of images processed at a time, C is the channel number of the initial feature map, H is the height of the initial feature map, and W is the width of the initial feature map. fully_connect(·) is a full connection function.

[0064] The quality problem category of the current frame image is evaluated by the Softmax layer, which obtains the probability of the current frame image under each quality problem category according to the target feature map output by the fully connected layer, to determine the quality problem category of the current frame image. The abnormality degree of the jth quality problem of the current frame image is directly output after being processed by the Sigmoid function. The Sigmoid function is a nonlinear activation function.

[0065] For example, the first quality feature and the second quality feature of the current image frame are obtained by segmenting the target feature map through the Softmax layer. The first quality feature of the current image frame is used to indicate the quality problem category of the current image frame, and the second quality feature of the current image frame is used to indicate the abnormality degree of the jth quality problem category of the current image frame.

[0066] The probability of the current frame image under each quality problem category is obtained by converting the first quality feature through the Softmax layer, as the third quality feature of the current image frame, and the second quality feature is nonlinearly transformed to obtain the fourth quality feature of the current image frame. The third quality feature of the current image frame and the fourth quality feature of the current image frame are spliced to obtain the initial quality feature of the current image frame.

[0067] It should be noted that the information segmentation manner is not limited too much, for example, a split function can be used to implement the split function is an information segmentation function.

[0068] For example, based on the target feature map through the Softmax layer, the process of obtaining the initial quality feature of the current frame image is as follows:

[0069] A5, A6 = split(A4)

[0070] Y = cat[softmax(A5), sigma(A6)]

[0071] Wherein, A5 is the first quality feature of the current image frame, A6 is the second quality feature of the current image frame, split(Y) is the split function, Y is the initial quality feature of the current image frame, softmax(·) is the Softmax function, sigma(·) is the Sigmoid function, and cat(·) is the concatenation function.

[0072] S102, the target quality feature of the historical frame image of the target video stream and the initial quality feature of the current frame image are fused to obtain the first time domain feature.

[0073] It should be noted that the historical frame image is not limited too much, for example, it can include the previous frame image of the current frame image, and the continuous multiple frame images before the current frame image. The target quality feature refers to the final quality feature of the image, for example, taking the historical frame image as an example, the target quality feature of the historical frame image is used to obtain the target quality detection result of the historical frame image, and the target quality detection result refers to the final quality detection result of the image. The quality detection result is not limited too much, for example, it can include the quality problem category of the image, the abnormal degree of the quality problem category of the image, etc.

[0074] The first time domain feature is used to indicate the change of the quality feature of the video stream in the time domain, for example, to indicate the change of the quality problem information of the video stream in the time domain.

[0075] In this application, the target quality feature of the historical frame image and the initial quality feature of the current frame image are fused to obtain the first time domain feature, so as to detect the quality of the current frame image, that is, the time sequence information such as the target quality feature of the historical frame image and the initial quality feature of the current frame image is comprehensively considered to detect the quality of the current frame image. For example, the final quality problem category and the abnormal degree of the current frame image are obtained by comprehensively considering the final quality problem category and the abnormal degree of the historical frame image and the initial quality problem category and the abnormal degree of the current frame image. Compared with the related art which only detects the quality of a frame image according to the frame image, the accuracy of video stream quality detection is improved.

[0076] In addition, the application can effectively process the time sequence information of the video stream, and can realize comprehensive and dynamic quality detection of the video stream, and has a significant advantage in a scene requiring continuous monitoring and detection.

[0077] It should be noted that the information fusion manner is not limited too much.

[0078] Optionally, the target quality features of the historical frame images of the target video stream and the initial quality features of the current frame image are fused to obtain the first time domain feature, including splicing the target quality features of the historical frame images and the initial quality features of the current frame image to obtain the first time domain feature.

[0079] Optionally, the target quality features of the historical frame images of the target video stream and the initial quality features of the current frame image are fused to obtain the first time domain feature, including splicing the target quality features of the historical frame images and the initial quality features of the current frame image to obtain the first time domain feature.

[0080] Optionally, the target quality features of the historical frame images of the target video stream and the initial quality features of the current frame image are fused to obtain the first time domain feature, including splicing the target quality features of the historical frame images and the initial quality features of the current frame image to obtain the first time domain feature.

[0081] S103, time domain-frequency domain conversion is performed on the first time domain feature to obtain the first frequency domain feature.

[0082] It should be noted that the time domain-frequency domain conversion refers to converting the time domain information into the frequency domain information. The first frequency domain feature is used to indicate the change of the quality feature of the video stream in the frequency domain, such as being used to indicate the change of the quality problem information of the video stream in the frequency domain. In the application, the first frequency domain feature of the video stream can be considered to perform quality detection on the current frame image, thereby improving the accuracy of the quality detection of the video stream.

[0083] The time domain-frequency domain conversion manner is not limited too much, such as DFT (Discrete Fourier Transform), FFT (Fast Fourier Transform), etc.

[0084] For example, the process of performing time domain-frequency domain conversion on the first time domain feature to obtain the first frequency domain feature is as follows:

[0085]

[0086] wherein x k is the kth frequency component of the first frequency domain feature x, X n is the nth sample point of the first time domain feature X, and N is the number of sample points of the first time domain feature X, and also the number of frequency components of the first frequency domain feature x. It should be noted that the number of frequency components of the first to third frequency domain features is consistent.

[0087] S104, obtaining a target quality feature of the current frame image based on the first frequency domain feature.

[0088] Optionally, obtaining the target quality feature of the current frame image based on the first frequency domain feature comprises obtaining the target quality feature of the current frame image based on the first time domain feature and the first frequency domain feature. Thus, the first time domain feature and the first frequency domain feature of the video stream can be comprehensively considered for quality detection of the current frame image, which improves the accuracy of video stream quality detection compared with related art in which quality detection of a frame image is performed only according to the frame image.

[0089] S105, obtaining a target quality detection result of the current frame image based on the target quality feature of the current frame image.

[0090] Optionally, obtaining the target quality detection result of the current frame image based on the target quality feature of the current frame image comprises extracting a probability of the current frame image under each quality problem category from the target quality feature of the current frame image, and taking the quality problem category with the maximum probability as the quality problem category of the current frame image.

[0091] Optionally, obtaining the target quality detection result of the current frame image based on the target quality feature of the current frame image comprises extracting an abnormality degree of the quality problem category of the current frame image from the target quality feature of the current frame image.

[0092] Optionally, as shown in Figure 5 , the quality detection model further comprises a time-frequency domain detection module, and steps S102-S105 are performed by the time-frequency domain detection module. Alternatively, as shown in Figure 5 , the quality detection model further comprises an abnormality screening module, and steps S102-S104 are performed by the time-frequency domain detection module, and step S105 is performed by the abnormality screening module.

[0093] For example, as shown in Figure 5As shown, when monitoring frame t=t0, that is, the current frame is the t0th frame, the time-frequency domain detection module outputs the final abnormal type code and the final abnormal degree information of the t0th frame image as the target quality feature of the t0th frame image.

[0094] The final abnormal type code output by the time-frequency domain detection module is the final abnormal type code of the image, and the final abnormal degree information output by the time-frequency domain detection module is the final abnormal degree information of the image.

[0095] When monitoring frame t=t0-1, the processing process of the time-frequency domain detection module can refer to the related content of monitoring frame t=t0, which will not be described here.

[0096] In summary, according to the video stream quality detection method provided in the embodiments of the present application, the initial quality feature of the current frame image of the target video stream is obtained by performing quality detection on the current frame image, the first time domain feature is obtained by fusing the target quality feature of the historical frame image of the target video stream and the initial quality feature of the current frame image, the first frequency domain feature is obtained by performing time-frequency domain conversion on the first time domain feature, the target quality feature of the current frame image is obtained based on the first frequency domain feature, and the target quality detection result of the current frame image is obtained based on the target quality feature of the current frame image. Therefore, the target quality feature of the historical frame image and the initial quality feature of the current frame image can be comprehensively considered to obtain the first time domain feature, and then the time-frequency domain conversion is performed to obtain the first frequency domain feature, and the current frame image is detected based on the first frequency domain feature. Compared with the related art which only detects the quality of a frame image based on the frame image, the accuracy of video stream quality detection is improved, and comprehensive and dynamic quality detection of the video stream can be realized, which is suitable for the quality detection of downhole monitoring video stream.

[0097] In the above embodiments, regarding the step S104 of obtaining the target quality feature of the current frame image based on the first frequency domain feature, the following can be combined for further understanding. Figure 2 Figure 2 Another flowchart of a video stream quality detection method provided by the embodiments of the present application is shown in FIG. 6. As shown in FIG. 6, the method can include the following steps: Figure 2

[0098] S201, performing quality detection on the current frame image of the target video stream to obtain the initial quality feature of the current frame image.

[0099] ​​S202, fuse the target quality feature of the historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time domain feature.

[0100] S203, perform time domain-frequency domain conversion on the first time domain feature to obtain a first frequency domain feature.

[0101] The related content of steps S201-S203 can be referred to the above embodiments, which will not be repeated here.

[0102] S204, process the first frequency domain feature based on an attention mechanism to obtain a second frequency domain feature.

[0103] It should be noted that the attention mechanism is a technology for enhancing (also called reinforcing) specific information in the input information (such as the information most critical to the task), and weakening the remaining information in the input information, so as to automatically focus on the specific information in the input information and ignore the remaining information in the input information during processing.

[0104] The attention mechanism is not limited too much, such as soft attention, hard attention, self-attention, local attention, hierarchical attention, etc.

[0105] In this application, the first frequency domain feature is processed based on the attention mechanism to obtain the second frequency domain feature, such as enhancing the quality detection related features in the first frequency domain feature, which helps to improve the accuracy of video stream quality detection. For example, enhancing the frequency components with significant quality problems in the first frequency domain feature helps to identify the areas with significant quality problems in the frequency domain of the image, thereby improving the accuracy of video stream quality detection.

[0106] Optionally, processing the first frequency domain feature based on the attention mechanism to obtain the second frequency domain feature includes segmenting the first frequency domain feature to obtain a third frequency domain feature and initial weights of N frequency components of the third frequency domain feature, N is an integer greater than 1, normalizing the initial weight of the kth frequency component to obtain the target weight of the kth frequency component, k is a positive integer not greater than N, and weighting and summing the N frequency components based on the target weights of the N frequency components to obtain the second frequency domain feature, so as to enhance the first frequency domain feature by attention to obtain the second frequency domain feature.

[0107] It should be noted that the normalization processing manner is not limited too much, such as a Sigmoid function can be used. Based on the target weights of the N frequency components, the N frequency components are weighted and summed to obtain the second frequency domain feature, including obtaining the product of the target weight of the kth frequency component and the kth frequency component as the kth result, and summing the N results to obtain the second frequency domain feature.

[0108] For example, the process of processing the first frequency domain feature based on the attention mechanism to obtain the second frequency domain feature is as follows:

[0109] x q ,x v =split(x)

[0110] x′=σ(x q )⊙x v

[0111] Wherein, x q is an initial weight vector composed of the initial weights of the N frequency components of the third frequency domain feature, x v is the third frequency domain feature, x′ is the second frequency domain feature, σ(x q ) is a target weight vector composed of the target weights of the N frequency components of the third frequency domain feature, and is a Hadamard product operator symbol.

[0112] S205, the second frequency domain feature is converted from frequency domain to time domain to obtain the second time domain feature.

[0113] It should be noted that the frequency domain-time domain conversion refers to converting frequency domain information into time domain information. The frequency domain-time domain conversion manner is not limited too much, such as it can include IDFT (Inverse Discrete Fourier Transform), IFFT (Inverse Fast Fourier Transform) and the like.

[0114] For example, the process of converting the second frequency domain feature from frequency domain to time domain to obtain the second time domain feature is as follows:

[0115]

[0116] Wherein, x k ′ is the kth frequency component of the second frequency domain feature, and x n ′ is the nth sample point of the second time domain feature.

[0117] For example, as Figure 6As shown, the time-frequency domain detection module includes a frequency domain attention unit, and steps S203-S205 are performed through the frequency domain attention unit.

[0118] As shown, Figure 7 As shown, the frequency domain attention unit includes an FFT subunit, a full connection layer subunit, a multiplication operation subunit, a SiLU (Sigmoid Linear Unit) subunit, a layer normalization subunit, a GELU (Gaussian Error Linear Unit) subunit, and an IFFT subunit. It should be noted that, Figure 6 7 In the symbol is used to represent the multiplication operation subunit, and the symbol is used to represent the SiLU subunit.

[0119] The first time domain feature is time-frequency domain converted through the FFT subunit to obtain a first frequency domain feature, the first frequency domain feature is feature extracted through the full connection layer subunit to obtain a fourth frequency domain feature, the first frequency domain feature is updated as the fourth frequency domain feature, and the first frequency domain feature is segmented to obtain a third frequency domain feature and an initial weight vector.

[0120] The initial weight vector is normalized through the SiLU subunit to obtain a target weight vector, and the target weight vector and the third frequency domain feature are Hadamard product operated through the multiplication unit to obtain a second frequency domain feature.

[0121] The second frequency domain feature is layer normalized through the layer normalization subunit to obtain a fifth frequency domain feature, the fifth frequency domain feature is nonlinearly transformed through the GELU subunit to obtain a sixth frequency domain feature, the second frequency domain feature is updated as the sixth frequency domain feature, and the second frequency domain feature is frequency-time domain converted through the IFFT subunit to obtain a second time domain feature.

[0122] S206, based on the second time domain feature, a target quality feature of the current frame image is obtained.

[0123] Optionally, based on the second time domain feature, a target quality feature of the current frame image is obtained, including based on the second time domain feature and the second frequency domain feature, a target quality feature of the current frame image is obtained. Thus, the second time domain feature and the second frequency domain feature of the video stream can be comprehensively considered for quality detection of the current frame image, which improves the accuracy of video stream quality detection compared to related art in which only a certain frame image is used for quality detection of the frame image.

[0124] ​Optionally, based on the second time domain feature, the target quality feature of the current frame image is obtained, including obtaining a cell state of the current frame based on the cell state of the historical frame and the second time domain feature, and obtaining the target quality feature of the current frame image based on the cell state of the current frame and the second time domain feature. Thus, the LSTM (Long Short-Term Memory) can be used to process the second time domain feature, which can effectively capture and retain the time domain feature critical to quality detection, and help improve the accuracy of video stream quality detection.

[0125] For example, the time-frequency domain detection module further includes an LSTM, and step S206 is performed by the LSTM.

[0126] It can be understood that the LSTM can effectively capture and retain the time domain feature critical to quality detection through its unique gating mechanism, i.e., the forget gate, the input gate, and the output gate. The forget gate is responsible for deciding which information should be forgotten at each time step, the input gate controls the introduction of new information, and the output gate determines which information will be passed to the next layer as output. The cell state is also called the memory vector.

[0127] Hereinafter, the frequency domain attention unit is referred to as Attn, the (t-1)th frame is referred to as a historical frame, and the tth frame is referred to as a current frame. The forget gate formula is as follows:

[0128] i t =σ(Attn[cat(h t-1 ,x t )])

[0129] Where h t-1 is the hidden state of the (t-1)th frame, which is also the target quality feature of the (t-1)th frame image, x t is the initial quality feature of the tth frame image, i t is the output of the forget gate at the tth frame, and Attn(·) is a function corresponding to the frequency domain attention unit.

[0130] The input gate formula is as follows:

[0131]

[0132] Where tanh(·) is a tanh function, is the output of the input gate at the tth frame.

[0133] The update formula of the cell state C is as follows:

[0134]

[0135] Where C t-1 is the cell state of the (t-1)th frame, and C tLet t represent the cell state in frame t.

[0136] The output formula of LSTM is:

[0137] h t =σ(Attn[cat(h t-1 ,x t )])⊙tanh(C t )

[0138] Among them, h t Let be the hidden state of frame t, and also the target quality feature of the image in frame t.

[0139] Through these gating mechanisms, LSTM can effectively suppress noise during the transmission of temporal information while preserving temporal features crucial for quality detection. Finally, the time-frequency domain detection module performs quality detection on the current image frame based on the LSTM-processed temporal information, combined with frequency domain features and temporal dependencies.

[0140] For example, such as Figure 6 As shown, at monitoring frame t = t0, i.e., when the current frame is frame t0, the result obtained by the time-frequency domain detection module in frame (t0-1) is used as historical information, i.e., h t-1 C t-1 As historical information, h is detected by the time-frequency domain detection module. t-1 The image anomaly type code and image anomaly degree information of the t0th frame are superimposed to obtain the first temporal feature.

[0141] The first time-domain feature is input into the frequency-domain attention unit, which outputs the second time-domain feature. The second time-domain feature is then processed by the time-domain-frequency-domain detection module to obtain h. t C t , will h t As the module output, this is the output of the time-domain-frequency domain detection module at frame t0. It should be noted that the second time-domain feature is processed by the time-domain-frequency domain detection module to obtain h. t C t The processing procedure can be found in the above embodiments, and will not be repeated here.

[0142] When monitoring frame t = t0 + 1, that is, when the current frame is the (t0 + 1)th frame, the result obtained by the time-frequency domain detection module in the t0th frame is used as historical information, that is, h t C t As this is historical information, it should be noted that the processing procedure of the time-frequency domain detection module at monitoring frame t = t0 + 1 can be found in the relevant content at monitoring frame t = t0, and will not be repeated here.

[0143] Figure 6Middle symbol For representing the add operation, the symbol For representing the tanh function operation, the tanh function is a tanh nonlinear activation function (also called hyperbolic tangent function).

[0144] S207, based on the target quality feature of the current frame image, obtaining a target quality detection result of the current frame image.

[0145] The related content of step S207 can be referred to the above embodiments, which will not be described here.

[0146] In summary, according to the quality detection method of the video stream provided in the embodiments of the present application, the first frequency domain feature is processed based on the attention mechanism to obtain the second frequency domain feature, the second frequency domain feature is subjected to frequency domain-time domain conversion to obtain the second time domain feature, and the target quality feature of the current frame image is obtained based on the second time domain feature. Therefore, the first frequency domain feature can be processed based on the attention mechanism to obtain the second frequency domain feature, and then subjected to frequency domain-time domain conversion to obtain the second time domain feature, and the current frame image is subjected to quality detection considering the second time domain feature, which helps to identify the area of the image having significant quality problems in the frequency domain, thereby improving the accuracy of the quality detection of the video stream.

[0147] In the above embodiments, regarding step S104, based on the first frequency domain feature, obtaining the target quality feature of the current frame image, the above-mentioned Figure 3 further understanding can be combined. Figure 3 Another flowchart of the quality detection method of the video stream provided in the embodiments of the present application is shown in FIG. 6. As shown in FIG. 6, the method can include the following steps: Figure 3

[0148] S301, performing quality detection on the current frame image of the target video stream to obtain an initial quality feature of the current frame image.

[0149] S302, fusing the target quality feature of the historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time domain feature.

[0150] S303, performing time domain-frequency domain conversion on the first time domain feature to obtain a first frequency domain feature.

[0151] The related content of steps S301-S303 can be referred to the above embodiments, which will not be described here.

[0152] S304, based on the first frequency domain feature, obtaining an intermediate quality feature of the current frame image.

[0153] ​It should be noted that the intermediate quality feature refers to the quality feature of the image obtained based on the first frequency domain feature in the time period after the initial quality feature is obtained and before the target quality feature is obtained, that is, the multiple quality features of the image are obtained in the order of the initial quality feature, the intermediate quality feature, and the target quality feature.

[0154] The related content of step S304 can refer to the related content of steps S104, S204-S206, which will not be repeated here.

[0155] S305, the target quality features of the plurality of historical frame images are aggregated to obtain an aggregated quality detection result of the target video stream.

[0156] It should be noted that the information aggregation manner and the aggregated quality detection result are not limited too much, for example, the aggregated quality detection result includes the frequency, duration, and average abnormality of the jth quality problem of the target video stream, the quality problem category with the maximum frequency of the target video stream, the quality problem category with the maximum duration of the target video stream, and the quality problem category with the maximum abnormality of the target video stream.

[0157] The frequency of the jth quality problem of the video stream refers to the frame number of the image of the jth quality problem in the video stream, the duration of the jth quality problem of the video stream refers to the product of the frame number of the jth quality problem in the video stream and a frame duration, and the average abnormality of the jth quality problem of the video stream refers to the average of the abnormality of the jth quality problem of the multiple frames in the video stream.

[0158] Optionally, the aggregated quality detection result includes at least one of the frequency, duration, and average abnormality of the jth quality problem of the target video stream.

[0159] S306, based on the aggregated quality detection result and the intermediate quality feature of the current frame image, the target quality feature of the current frame image is obtained.

[0160] In this application, the aggregated quality detection result and the intermediate quality feature of the current frame image are comprehensively considered for quality detection of the current frame image, that is, the quality detection result aggregated by the plurality of historical frame images is used to assist the quality detection of the current frame image, compared with the related art which only depends on a frame image for quality detection of the frame image, the accuracy of the video stream quality detection is improved.

[0161] Optionally, the target quality feature of the current frame image is obtained based on the aggregated quality detection result and the intermediate quality feature of the current frame image, including: in response to at least one of the following conditions being met: the target video stream meets a condition that a frequency of occurrence of the jth quality problem of the target video stream is less than or equal to a first set threshold value, and a condition that a duration of occurrence of the jth quality problem of the target video stream is less than or equal to a second set threshold value, and the intermediate quality feature of the current frame image indicates that the current frame image has the jth quality problem, the intermediate quality feature of the current frame image is corrected to obtain a quality feature indicating that the current frame image does not have the jth quality problem as the target quality feature of the current frame image.

[0162] It can be understood that, in the embodiment, the frequency and / or the duration of occurrence of the jth quality problem of the target video stream is small, and the intermediate quality feature of the current frame image indicates that the current frame image has the jth quality problem, so the current frame image is likely to be misjudged as having the jth quality problem. Therefore, the intermediate quality feature of the current frame image can be corrected so that the final quality feature of the current frame image indicates that the current frame image does not have the jth quality problem, which can avoid the situation that a frame image is misjudged as having a certain quality problem, and improve the accuracy of video stream quality detection and reduce the false positive rate of video stream quality detection.

[0163] Optionally, the target quality feature of the current frame image is obtained based on the aggregated quality detection result and the intermediate quality feature of the current frame image, including: in response to at least one of the following conditions being met: the target video stream meets a condition that a frequency of occurrence of the jth quality problem of the target video stream is greater than or equal to a third set threshold value, and a condition that a duration of occurrence of the jth quality problem of the target video stream is greater than or equal to a fourth set threshold value, and the intermediate quality feature of the current frame image indicates that the current frame image does not have the jth quality problem, the intermediate quality feature of the current frame image is corrected to obtain a quality feature indicating that the current frame image has the jth quality problem as the target quality feature of the current frame image.

[0164] It can be understood that, in the embodiment, the frequency and / or the duration of occurrence of the jth quality problem of the target video stream is large, and the intermediate quality feature of the current frame image indicates that the current frame image does not have the jth quality problem, so the current frame image is likely to miss the jth quality problem. Therefore, the intermediate quality feature of the current frame image can be corrected so that the final quality feature of the current frame image indicates that the current frame image has the jth quality problem, which can avoid the situation that a frame image misses a certain quality problem, and improve the accuracy of video stream quality detection.

[0165] Optionally, the method further includes obtaining an average abnormality degree of the jth quality problem of the target video stream as the abnormality degree of the jth quality problem of the current frame image.

[0166] It should be noted that the first to fourth thresholds are not limited too much, for example, the first set threshold is less than the third set threshold, and the second set threshold is less than the fourth set threshold, for example, the first set threshold and the second set threshold are both zero.

[0167] S307, obtaining a target quality detection result of the current frame image based on the target quality feature of the current frame image.

[0168] The related content of step S307 can be referred to the above-mentioned embodiments, which will not be repeated here.

[0169] In summary, according to the quality detection method of the video stream provided in the embodiments of the present application, the intermediate quality feature of the current frame image is obtained based on the first frequency domain feature, the target quality features of the plurality of historical image frames are aggregated to obtain the aggregated quality detection result of the target video stream, and the target quality feature of the current frame image is obtained based on the aggregated quality detection result and the intermediate quality feature of the current frame image. Therefore, the aggregated quality detection result and the intermediate quality feature of the current frame image can be comprehensively considered for quality detection of the current frame image, and the accuracy of the video stream quality detection is improved.

[0170] The embodiments of the present application provide another quality detection method of a video stream.

[0171] As shown in Figure 4 , the method can include the following steps:

[0172] S401, performing quality detection on a current frame image of a target video stream to obtain an initial quality feature of the current frame image.

[0173] S402, fusing a target quality feature of a historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time domain feature.

[0174] S403, performing time domain-frequency domain conversion on the first time domain feature to obtain a first frequency domain feature.

[0175] S404, obtaining a target quality feature of the current frame image based on the first frequency domain feature.

[0176] S405, obtaining a target quality detection result of the current frame image based on the target quality feature of the current frame image.

[0177] The related content of steps S401-S405 can be referred to the above-mentioned embodiments, which will not be repeated here.

[0178] S406, determining that the current frame image has a jth quality problem based on the target quality detection result of the current frame image.

[0179] S407, in response to the j-th quality problem being the target quality problem, generate target alarm information to indicate that the current frame image has the j-th quality problem.

[0180] It should be noted that there are no strict restrictions on the target quality issues, which can be customized by the user. For example, target quality issues that have a significant impact on the quality of the video stream can include blurring, occlusion, and excessive darkness.

[0181] In this application, when a quality problem of type j occurs in the current frame image, and the quality problem of type j is a target quality problem, a target alarm message is generated to indicate that the current frame image has a quality problem of type j. That is, when the quality problem category of the current frame image is a set category, a target alarm message is generated to inform the user in a timely manner that a quality problem has occurred in the current frame image, so that the user can deal with the quality problem of the set category in a timely manner. For example, the user can deal with the quality problem that has a significant impact on the quality of the video stream in a timely manner, thereby improving the monitoring quality in the monitoring scenario based on the video stream.

[0182] S408, in response to a non-target quality problem of the j-th type of quality problem, and the abnormality of the j-th type of quality problem in the current frame image is greater than or equal to the fifth preset threshold, generate target alarm information.

[0183] It should be noted that no excessive restrictions are placed on the fifth threshold setting.

[0184] In this application, when a quality problem of type j occurs in the current frame image, and the quality problem of type j is not a target type quality problem, and the degree of abnormality of the quality problem of type j in the current frame image is relatively large, target alarm information is generated to indicate that the current frame image has a quality problem of type j. That is, when the quality problem category of the current frame image is not a set category, and the degree of abnormality of the quality problem category of the current frame image is relatively large, target alarm information is generated to promptly inform the user that the current frame image has a quality problem, so that the user can deal with quality problems with a large degree of abnormality in a timely manner. In other words, the user can deal with quality problems that have a significant impact on the quality of the video stream in a timely manner, thereby improving the monitoring quality in the monitoring scenario based on the video stream.

[0185] S409, in response to the fact that the j-th quality problem is not a target quality problem, and the degree of abnormality of the j-th quality problem in the current frame image is less than the fifth set threshold, the target alarm information is refused to be generated.

[0186] In this application, if a quality problem of type j occurs in the current frame image, and the quality problem of type j is not a target quality problem, and the degree of abnormality of the quality problem of type j in the current frame image is small, the target alarm information is refused to be generated. That is, if the quality problem category in the current frame image is not a set category, and the degree of abnormality of the quality problem category in the current frame image is small, the target alarm information is refused to be generated. This allows users to focus more on dealing with quality problems of the set category and / or with a large degree of abnormality. In other words, users can deal with quality problems that have a significant impact on the quality of the video stream in a timely manner, thereby improving the monitoring quality in video stream-based monitoring scenarios.

[0187] Optionally, in response to a non-target quality problem of type j, and after the degree of abnormality of type j quality problem in the current frame image is less than a fifth preset threshold, non-alarm information is also generated.

[0188] For example, alarms will not be triggered for situations such as water mist, short-term strong light exposure, or small-area dirt generated by mining machinery under normal operating conditions, thereby reducing false alarms and allowing relevant personnel to focus on dealing with serious anomalies that affect monitoring.

[0189] For example, such as Figure 5 As shown, when the monitoring frame is t=t0, that is, when the current frame is the t0th frame, the final anomaly type code and final anomaly degree information of the t0th frame image are input to the anomaly filtering module. The anomaly filtering module outputs the final output at t=t0, such as the anomaly filtering module executing step S406, and executing any one of steps S407-S409.

[0190] At monitoring frame t = t0-1, the final anomaly type code and final anomaly degree information of the (t0-1)th frame image are input to the anomaly filtering module, and the anomaly filtering module outputs the final output at time t = t0-1.

[0191] The final output at any given time is either an alarm message or a non-alarm message.

[0192] In summary, the video stream quality detection method according to the embodiments of this application can take into account the type of quality problem appearing in the current frame image and the degree of abnormality of the type of quality problem appearing in the current frame image, and determine whether to generate target alarm information, so that users can deal with quality problems that have a significant impact on the quality of the video stream in a timely manner, thereby improving the monitoring quality in monitoring scenarios based on video streams.

[0193] To achieve the above embodiments, this application also proposes a video stream quality detection device.

[0194] Figure 8 This is a schematic diagram of the structure of a video stream quality detection device provided in an embodiment of this application.

[0195] As Figure 8 shown in the figure, the video stream quality detection device 100 includes a first detection module 110, a fusion module 120, a conversion module 130, a second detection module 140, and a third detection module 150.

[0196] The first detection module 110 is configured to perform quality detection on a current frame image of a target video stream to obtain an initial quality feature of the current frame image.

[0197] The fusion module 120 is configured to fuse a target quality feature of a historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time-domain feature.

[0198] The conversion module 130 is configured to perform time-domain-frequency domain conversion on the first time-domain feature to obtain a first frequency-domain feature.

[0199] The second detection module 140 is configured to obtain a target quality feature of the current frame image based on the first frequency-domain feature.

[0200] The third detection module 150 is configured to obtain a target quality detection result of the current frame image based on the target quality feature of the current frame image.

[0201] Further, in a possible implementation manner of the embodiment of the present application, the second detection module 140 is further configured to: process the first frequency-domain feature based on an attention mechanism to obtain a second frequency-domain feature; perform frequency-domain-time-domain conversion on the second frequency-domain feature to obtain a second time-domain feature; and obtain the target quality feature of the current frame image based on the second time-domain feature.

[0202] Further, in a possible implementation manner of the embodiment of the present application, the second detection module 140 is further configured to: segment the first frequency-domain feature to obtain a third frequency-domain feature and initial weights of N frequency components of the third frequency-domain feature, N being an integer greater than 1; perform normalization processing on the initial weight of the kth frequency component to obtain a target weight of the kth frequency component, k being a positive integer not greater than N; and perform weighted summation on the N frequency components based on the target weights of the N frequency components to obtain the second frequency-domain feature.

[0203] Further, in a possible implementation manner of the embodiment of the present application, the second detection module 140 is further configured to: obtain a cell state of a current frame based on a cell state of a historical frame and the second time-domain feature; and obtain the target quality feature of the current frame image based on the cell state of the current frame and the second time-domain feature.

[0204] Further, in a possible implementation form of the embodiment of the application, the second detection module 140 is further configured to: obtain an intermediate quality feature of the current frame image based on the first frequency domain feature; aggregate the target quality features of the plurality of historical frame images to obtain an aggregated quality detection result of the target video stream; and obtain the target quality feature of the current frame image based on the aggregated quality detection result and the intermediate quality feature of the current frame image.

[0205] Further, in a possible implementation form of the embodiment of the application, the aggregated quality detection result comprises at least one of a frequency of occurrence of the jth quality problem of the target video stream and a duration of occurrence of the jth quality problem of the target video stream.

[0206] Further, in a possible implementation form of the embodiment of the application, the second detection module 140 is further configured to: in response to at least one of the target video stream satisfying a condition that a frequency of occurrence of the jth quality problem of the target video stream is less than or equal to a first set threshold and a duration of occurrence of the jth quality problem of the target video stream is less than or equal to a second set threshold, and the intermediate quality feature of the current frame image indicating that the current frame image has the jth quality problem, correct the intermediate quality feature of the current frame image to obtain a quality feature indicating that the current frame image does not have the jth quality problem as the target quality feature of the current frame image.

[0207] Further, in a possible implementation form of the embodiment of the application, the second detection module 140 is further configured to: in response to at least one of the target video stream satisfying a condition that a frequency of occurrence of the jth quality problem of the target video stream is greater than or equal to a third set threshold and a duration of occurrence of the jth quality problem of the target video stream is greater than or equal to a fourth set threshold, and the intermediate quality feature of the current frame image indicating that the current frame image does not have the jth quality problem, correct the intermediate quality feature of the current frame image to obtain a quality feature indicating that the current frame image has the jth quality problem as the target quality feature of the current frame image.

[0208] Further, in a possible implementation of the embodiment of the present application, the third detection module 150 is further configured to: determine that the current frame image has a jth quality problem based on the target quality detection result of the current frame image; in response to the jth quality problem being a target quality problem, generate target alarm information indicating that the current frame image has the jth quality problem; or, in response to the jth quality problem not being the target quality problem and the abnormal degree of the jth quality problem of the current frame image being greater than or equal to a fifth preset threshold, generate the target alarm information; or, in response to the jth quality problem not being the target quality problem and the abnormal degree of the jth quality problem of the current frame image being less than the fifth preset threshold, refuse to generate the target alarm information.

[0209] It should be noted that the foregoing explanation of the embodiment of the video stream quality detection method is also applicable to the embodiment of the video stream quality detection device, which will not be described here.

[0210] In summary, the video stream quality detection device of the embodiment of the present application performs quality detection on a current frame image of a target video stream to obtain an initial quality feature of the current frame image, fuses a target quality feature of a historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time-domain feature, performs time-domain-frequency conversion on the first time-domain feature to obtain a first frequency-domain feature, obtains a target quality feature of the current frame image based on the first frequency-domain feature, and obtains a target quality detection result of the current frame image based on the target quality feature of the current frame image. In this way, the target quality feature of the historical frame image and the initial quality feature of the current frame image are comprehensively considered to obtain the first time-domain feature, and then the time-domain-frequency conversion is performed to obtain the first frequency-domain feature, and the current frame image is detected based on the first frequency-domain feature, which improves the accuracy of video stream quality detection compared with the related art in which only a certain frame image is detected for quality detection of the frame image, and can realize comprehensive and dynamic quality detection of the video stream, and is suitable for the quality detection of the downhole monitoring video stream.

[0211] To implement the above-mentioned embodiments, the present application further provides an electronic device, comprising: a processor and a memory connected with the processor in communication; the memory stores computer execution instructions; and the processor executes the computer execution instructions stored in the memory to implement the video stream quality detection method provided by the foregoing embodiments.

[0212] To implement the above-mentioned embodiments, the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by a processor to implement the video stream quality detection method provided by the foregoing embodiments.

[0213] To achieve the above-mentioned embodiments, the application further provides a computer program product comprising a computer program which, when executed by a processor, implements the quality detection method of the video stream provided by the foregoing embodiments.

[0214] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the present application comply with relevant laws and regulations and do not violate public order and good customs.

[0215] It should be noted that the personal information from the user should be collected for legal and reasonable purposes, and should not be shared or sold outside these legal uses. In addition, such collection / sharing should be carried out after the user's informed consent is received, including but not limited to informing the user to read the user agreement / user notice before the user uses the function, and signing the agreement / authorization including authorization of relevant user information. In addition, any necessary steps should be taken to safeguard and protect access to such personal information data and ensure that other people with access to personal information data comply with their privacy policies and processes.

[0216] The present application is expected to provide embodiments in which the user can selectively prevent the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk is minimized by limiting data collection and deleting data. In addition, such personal information is de-identified, if applicable, to protect the privacy of the user.

[0217] In the foregoing embodiment description, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples, without contradiction.

[0218] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as controlling or implying relative importance or implicitly indicating the number of technical features controlled. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise specifically limited.

[0219] Any processes or methods described in the flowcharts or otherwise described herein can be understood as representing modules, segments, or portions of code that include one or more executable instructions for implementing specific logical functions or steps, and the various embodiments of the application can include additional or fewer steps performing the same or equivalent functions as those shown or discussed, in a different order, or in combination with other steps, as would be understood by one of ordinary skill in the art.

[0220] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a list of executable instructions for implementing the logic function, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can specifically include the following, which are non-exhaustive listings: electrical connections (electrical apparatus), portable computer disks (magnetic apparatus), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber devices, and portable compact disk read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium upon which the program can be printed, as the program can be electronically captured, for example, via the optical scanning of the paper or other medium, followed by the electronic conversion of the optically scanned program into a form that can be edited, compiled, or interpreted or otherwise processed into an electronically usable form, and then stored in the computer memory.

[0221] It should be understood that portions of the application can be implemented in hardware, software, firmware, or combinations thereof. In the above embodiments, the various steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. As such, if implemented in hardware, the various steps or methods can be implemented in any one or combination of the following technologies, which are all well known in the art: discrete logic circuitry having logic gates for implementing logic functions upon data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and so forth.

[0222] Those skilled in the art of the present technology can understand that all or part of the steps carried out by the above-mentioned embodiment method can be completed by programs instructing related hardware, and the programs can be stored in a computer readable storage medium. When the program is executed, it includes one of the steps of the method embodiment or a combination thereof.

[0223] In addition, each functional unit in each embodiment of the present application can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0224] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above-mentioned embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-mentioned embodiments within the scope of the present application.

Claims

1. A method for quality detection of a video stream, characterized in that, The method comprises the following steps: performing quality detection on a current frame image of a target video stream to obtain an initial quality feature of the current frame image; fusing a target quality feature of a historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time-domain feature; performing time-domain-frequency domain conversion on the first time-domain feature to obtain a first frequency-domain feature; obtaining a target quality feature of the current frame image based on the first frequency-domain feature; obtaining a target quality detection result of the current frame image based on the target quality feature of the current frame image; wherein the step of obtaining the target quality feature of the current frame image based on the first frequency-domain feature comprises the following steps: processing the first frequency-domain feature based on an attention mechanism to obtain a second frequency-domain feature; performing frequency-domain-time-domain conversion on the second frequency-domain feature to obtain a second time-domain feature; obtaining the target quality feature of the current frame image based on the second time-domain feature; the step of processing the first frequency-domain feature based on an attention mechanism to obtain a second frequency-domain feature comprises the following steps: segmenting the first frequency domain feature to obtain a third frequency domain feature, and initial weights of respective frequency components of the third frequency domain feature, N is an integer greater than 1, N is an integer greater than 1, For the first k The initial weights of the i-th frequency components are normalized to obtain the i-th frequency component. k The target weight of each frequency component k Not greater than N Positive integers; Based on the above N The target weights of each frequency component are given by the following: N The second frequency domain feature is obtained by weighted summation of the frequency components.

2. The method of claim 1, wherein, the step of obtaining the target quality feature of the current frame image based on the second time-domain feature comprises the following steps: obtaining a cell state of a current frame based on a cell state of a historical frame and the second time-domain feature; obtaining the target quality feature of the current frame image based on the cell state of the current frame and the second time-domain feature.

3. The method according to any one of claims 1-2, characterized in that, the step of obtaining the target quality feature of the current frame image based on the first frequency-domain feature comprises the following steps: obtaining an intermediate quality feature of the current frame image based on the first frequency-domain feature; aggregating target quality features of a plurality of historical frame images to obtain an aggregated quality detection result of the target video stream; obtaining the target quality feature of the current frame image based on the aggregated quality detection result and the intermediate quality feature of the current frame image.

4. The method of claim 3, wherein, The aggregation quality detection result includes at least one of a frequency and a duration of the target video stream appearing in the first j quality problem.

5. The method of claim 4, wherein, the step of obtaining the target quality feature of the current frame image based on the aggregated quality detection result and the intermediate quality feature of the current frame image comprises the following steps: In response to the target video stream satisfying the occurrence of the target video stream at the first j The frequency of a certain type of quality problem is less than or equal to a first set threshold, and the target video stream appears at the first... j The duration of the quality problem is less than or equal to at least one of the conditions in the second set threshold, and the intermediate quality feature of the current frame image indicates that the current frame image has a first quality problem. j Quality issues; correcting the intermediate quality features of the current frame image to obtain quality features indicating that the current frame image does not have the first j quality problem as the target quality features of the current frame image.

6. The method of claim 4, wherein, the step of obtaining the target quality feature of the current frame image based on the aggregated quality detection result and the intermediate quality feature of the current frame image comprises the following steps: In response to the target video stream satisfying the occurrence of the target video stream at the first j The frequency of a certain type of quality problem is greater than or equal to a third preset threshold, and the target video stream appears at the first... j The duration of the quality problem is greater than or equal to at least one of the fourth set thresholds, and the intermediate quality features of the current frame image indicate that the current frame image does not exhibit the first quality problem. j Quality issues; correcting the intermediate quality feature of the current frame image to obtain a quality feature indicating that the current frame image has the quality problem of the first type as a target quality feature of the current frame image. j correcting the intermediate quality feature of the current frame image to obtain a quality feature indicating that the current frame image has the quality problem of the first type as a target quality feature of the current frame image.

7. The method of any one of claims 1-2, wherein, after the step of obtaining the target quality detection result of the current frame image based on the target quality feature of the current frame image, the method further comprises the following steps: determine that the current frame image occurs the first quality problem based on the target quality detection result of the current frame image j quality problem; in response to the first j target class quality problem, generate target alarm information for indicating that the current frame image appears the first j target class quality problem; or, In response to the first j The target class quality problem is not the class quality problem, and the abnormal degree of the current frame image appearing the first j class quality problem is greater than or equal to a fifth set threshold, the target alarm information is generated; or, In response to the first j The quality problem is not the target quality problem, and the current frame image shows the first occurrence of the quality problem. j If the degree of abnormality of a quality problem is less than the fifth preset threshold, the target alarm information will not be generated.

8. A quality detection apparatus of a video stream, characterized by comprising: The method comprises the following steps: a first detection module is configured to perform quality detection on a current frame image of a target video stream to obtain an initial quality feature of the current frame image; a fusion module is configured to fuse a target quality feature of a historical frame image of the target video stream and the initial quality feature of the current frame image to obtain a first time-domain feature; a conversion module is configured to perform time-domain-frequency domain conversion on the first time-domain feature to obtain a first frequency-domain feature; a second detection module is configured to obtain a target quality feature of the current frame image based on the first frequency-domain feature; a third detection module is configured to obtain a target quality detection result of the current frame image based on the target quality feature of the current frame image. The second detection module is further configured to: process the first frequency domain feature based on an attention mechanism to obtain a second frequency domain feature; perform frequency domain-time domain conversion on the second frequency domain feature to obtain a second time domain feature; and obtain a target quality feature of the current frame image based on the second time domain feature. The second detection module is further configured to: segment the first frequency domain feature to obtain a third frequency domain feature, and the third frequency domain feature... N The initial weights of each frequency component. N It is an integer greater than 1; for the th k The initial weights of the i-th frequency components are normalized to obtain the i-th frequency component. k The target weight of each frequency component k Not greater than N Positive integers; based on the N The target weights of each frequency component are given by the following: N The second frequency domain feature is obtained by weighted summation of the frequency components.

9. An electronic device, comprising: The method comprises: a processor, and a memory connected to the processor in communication; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are executed by the processor to implement the method of any one of claims 1-7.

11. A computer program product, characterised in that, The computer program is executed by the processor to implement the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Image-based fingerprint quality evaluation method and device and electronic equipment

    CN111179265A

  • Non-reference ultra-high-definition video quality objective evaluation method based on deep reinforcement learning

    CN114915777A

  • Industrial time series data anomaly detection method and system based on time-frequency domain feature enhancement

    CN119167243A