Automatic video quality guarantee method and device based on articulated naturality web

By introducing an automated video quality assurance method based on video networking into the video audit system, using deep learning and pre-trained neural networks for video content and quality detection, the existing system's shortcomings in processing speed, accuracy and comprehensiveness are solved, and more efficient, accurate and comprehensive video auditing and quality detection are achieved.

CN120201239APending Publication Date: 2025-06-24CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510385138.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing video auditing system has shortcomings in processing speed and accuracy, making it difficult to comprehensively and quickly review a large number of video content, and the video quality detection method is mainly limited to the quantification of technical parameters and lacks a comprehensive evaluation of user experience.

Method used

An automated video quality assurance method based on visual network is adopted, keyframes and audio clips are extracted from video data through deep learning algorithms, pre-trained deep neural network model is used for content audit, and video playback quality is detected by simulating the user's viewing process, and video quality reports are generated. Finally, the video violation level is determined based on the preset violation threshold and the playback and release are stopped.

Benefits of technology

It improves the efficiency, accuracy and comprehensiveness of video audits, and can more effectively detect illegal content and playback quality problems in videos, ensuring the quality and safety of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201239A_ABST
    Figure CN120201239A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic video quality guarantee method and device based on an articulated naturality web, and the method comprises the steps: obtaining original video data from various audio sources, and extracting key frames and audio clips from the original video data based on a deep learning algorithm; and content auditing is carried out on the key frame and the audio clip through a pre-trained deep neural network model to judge whether illegal content exists in the key frame and the audio clip, and if the illegal content exists, an alarm is triggered and fed back to the platform management terminal. The method comprises the following steps: playing original video data by simulating a user watching process to detect whether the original video data has an abnormal phenomenon in the playing process, and recording video parameters associated with the abnormal phenomenon to obtain a video quality report. And judging the violation level of the original video data based on a preset violation threshold in combination with the video quality report and the violation content, and stopping playing and publishing of the current video when it is detected that serious violation content or a playing quality problem exists in the original video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of the Internet and video content review, and in particular, to an automated video quality assurance method and device based on a visual Internet. Background Art

[0002] With the explosive growth of Internet video content, the traditional manual review method can no longer meet the need for quickly and accurately reviewing a large amount of video content. The video automatic frame extraction review system automates this process by extracting key frames from the video and using image recognition technology. At the same time, video quality detection is also a key link to ensure video playback quality, which usually includes dimensions such as clarity evaluation and color saturation analysis.

[0003] Currently, existing video review systems mainly rely on manual review or simple rule-based automated tools. For example, the frame extraction method based on a fixed interval: This method usually extracts frames from the video at a preset time interval. Although this method is easy to implement, it is difficult to ensure that the extracted frames can represent the core content of the original video. Especially when the video content changes rapidly, this method may result in the loss of key information. The frame extraction method based on key point detection: This method attempts to select key frames by identifying important events or scene transition points in the video. However, this method often requires complex algorithm support and may need customized adjustments for different video types (such as sports events, news reports, etc.), increasing the complexity and deployment difficulty of the system. Content review based on deep learning: Using deep learning models for content review has become a trend, and this system can automatically learn and identify illegal content in the video. However, most of the existing technologies focus on single-aspect review and ignore other types of prohibited content. Video quality detection: Current video quality detection systems mostly focus on the quantification of technical parameters, such as resolution, frame rate, etc., but involve less subjective quality evaluation (such as visual comfort, viewing experience, etc.) that is crucial for the user experience.

[0004] In summary, there are certain deficiencies in the existing technologies in terms of processing speed and accuracy. For example, extracting frames only based on a fixed frame interval is simple and easy to implement, but it may miss important information segments, resulting in an incomplete review. In addition, most of the existing video quality detection methods are limited to basic quality parameter measurements, such as resolution, bit rate, etc., and lack the comprehensive evaluation ability of the overall video quality. Summary of the Invention

[0005] Based on this, it is necessary to provide an automated video quality assurance method and device based on a visual Internet with high video review efficiency, accuracy, and comprehensiveness for the above technical problems.

[0006] The present invention provides an automated video quality assurance method based on the vision Internet, and the method includes: Obtain the original video data from various audio sources, and extract key frames and audio segments from the original video data based on a deep learning algorithm; Conduct content review on the key frames and audio segments through a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, trigger an alarm and feedback it to the platform management terminal; Play the original video data by simulating the user viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and record the video parameters associated with the abnormal phenomena to obtain a video quality report; Based on a preset violation threshold, combine the video quality report and the illegal content to determine the violation level of the original video data. When serious illegal content or playback quality problems are detected in the original video data, stop the playback and release of the current video; Wherein, the abnormal phenomena include video stuttering and playback delay, and the video parameters include video buffering time and video loading speed.

[0007] In one embodiment, the obtaining the original video data from various audio sources and extracting key frames and audio segments from the original video data based on a deep learning algorithm includes: Based on the deep learning algorithm, dynamically extract the key frames from the original video data through fixed interval extraction combined with scene detection based on MediaPipe; Extract the speech content from the speech data of the original video data through automatic speech recognition technology, and transcribe the speech content to obtain the audio segments.

[0008] In one embodiment, the obtaining the original video data from various audio sources and extracting key frames and audio segments from the original video data based on a deep learning algorithm further includes: Call a lightweight SIFT topic to extract key points from the original video data as local features, and extract deep texture representations in the original video data as global features through a pre-trained ResNet-18; Use Faster R-CNN to detect the entity objects in the original video data, and conduct scene understanding on the original video data through a ViT-Hybrid model to classify the scenes in the original video data to obtain semantic features; Based on the local features, global features, entity objects, and semantic features, the original video data is annotated using the Kaldi toolbox, and the frame extraction interval of the key frames is dynamically adjusted according to the annotation results, and the key frames are output.

[0009] In one embodiment, the content of the key frames and audio segments is audited by a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, an alarm is triggered and fed back to the platform management terminal, including: Call a pre-trained deep neural network model to analyze the text information, image data, and voice data in the key frames and audio segments to obtain an analysis result; Configure different audit thresholds for the text information, image data, and voice data according to the business scenario, and when the analysis result exceeds the corresponding audit threshold, it is determined that there is illegal content in the text information and / or image data and / or voice data.

[0010] In one embodiment, the content of the key frames and audio segments is audited by a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, an alarm is triggered and fed back to the platform management terminal, and further includes: When there is illegal content in the text information and / or image data and / or voice data, an alarm is triggered and an execution instruction is sent to the platform management terminal through the API. The platform management terminal takes blocking or traffic limiting measures in response to the execution instruction.

[0011] In one embodiment, the original video data is played by simulating the user viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and the video parameters associated with the abnormal phenomena are recorded to obtain a video quality report, including: Calculate the buffer time based on the first frame rendering time and the total duration of the video playback buffer time, and calculate the playback delay based on the difference between the actual playback timestamp and the theoretical playback timestamp; When the proportion of frames with a continuous frame playback time difference exceeding a set first threshold, calculate the frame drop rate of the original video data, and calculate the video loading speed based on the playback completion time of the original video data; Among them, the buffer time, playback delay, frame drop rate, and video loading speed are all monitoring indicators for determining whether there are any abnormal phenomena in the original video data.

[0012] In one embodiment, determining the violation level of the original video data based on a preset violation threshold in combination with the video quality report and the violation content, and when serious violation content or playback quality problems are detected in the original video data, stopping the playback and release of the current video, including: Determining the preset violation threshold based on the service type of the original video data, and determining the violation level of the original video data with violation content according to the violation threshold, where the violation level is preset based on the monitoring indicators, including minor violations, moderate violations, and serious violations; When it is detected that the violation level of the violation content in the original video data is a serious violation, a troubleshooting instruction is sent to the platform management terminal, and the platform management terminal responds to the troubleshooting instruction to troubleshoot the upstream and downstream links of the original video data to determine whether there are abnormalities in the video source of the original video data.

[0013] The present invention also provides an automated video quality assurance device based on the vision network. The device includes: An intelligent frame extraction module, configured to obtain the original video data from various audio sources, and extract key frames and audio segments from the original video data based on a deep learning algorithm; A content review module, configured to review the content of the key frames and audio segments through a pre-trained deep neural network model to determine whether there is any violation content in the key frames and audio segments. If there is violation content, an alarm is triggered and fed back to the platform management terminal; A playback quality detection module, configured to play the original video data by simulating the user's viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and record the video parameters associated with the abnormal phenomena to obtain a video quality report; A decision execution module, configured to determine the violation level of the original video data based on a preset violation threshold in combination with the video quality report and the violation content, and when serious violation content or playback quality problems are detected in the original video data, stop the playback and release of the current video; Wherein, the abnormal phenomena include video stuttering and playback delay, and the video parameters include video buffering time and video loading speed.

[0014] The present invention also provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the automated video quality assurance method based on the vision network as described in any one of the above.

[0015] The present invention also provides a computer storage medium storing a computer program, which when executed by a processor implements the method for automated video quality assurance based on the vision network as described in any one of the above.

[0016] The above method and apparatus for automated video quality assurance based on the vision network obtain original video data from various audio sources, and extract key frames and audio segments from the original video data based on a deep learning algorithm. Then, a pre-trained deep neural network model is used to conduct content review on the key frames and audio segments to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, an alarm is triggered and fed back to the platform management terminal. After that, the original video data is played by simulating the user viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and the video parameters associated with the abnormal phenomena are recorded to obtain a video quality report. Finally, based on a preset violation threshold, combined with the video quality report and the illegal content, the violation level of the original video data is determined, and when serious illegal content or playback quality problems are detected in the original video data, the playback and release of the current video are stopped. This method integrates the extraction of key frames by a deep learning algorithm and the content review and video quality detection driven by a deep neural network, and incorporates the simulation of the user viewing process during the video review and detection, which not only improves the efficiency and accuracy of video review and detection, but also has better comprehensiveness in review. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a flowchart of the method for automated video quality assurance based on the vision network provided by the present invention; Figure 2 It is a schematic diagram of the overall video review process of the method for automated video quality assurance based on the vision network in a specific embodiment provided by the present invention; Figure 3 It is a schematic diagram of the key frame extraction process of the method for automated video quality assurance based on the vision network in a specific embodiment provided by the present invention; Figure 4 It is a schematic diagram of the multi-dimensional content review process of the method for automated video quality assurance based on the vision network in a specific embodiment provided by the present invention; Figure 5Schematic diagram of the video playback quality assessment process of the automated video quality assurance method based on the Visual Internet provided by the present invention; Figure 6 Schematic diagram of the structure of the automated video quality assurance device based on the Visual Internet provided by the present invention; Figure 7 Internal structure diagram of the electronic device provided by the present invention. Detailed implementation manners

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] The following combines Figures 1 to 7 to describe the automated video quality assurance method and device based on the Visual Internet of the present invention.

[0021] As Figure 1 shown, in one embodiment, an automated video quality assurance method based on the Visual Internet includes the following steps: Step S110: Obtain the original video data from various audio sources, and extract key frames and audio segments from the original video data based on a deep learning algorithm.

[0022] Specifically, the server obtains the original video data from various audio sources, and extracts key frames and audio segments from the original video data through a deep learning algorithm.

[0023] In some embodiments, for the automated video quality assurance method based on the Visual Internet provided by the present invention, step S110 specifically includes the following steps: Step S111: Dynamically extract key frames from the original video data based on a deep learning algorithm by extracting at fixed intervals in combination with scene detection based on MediaPipe.

[0024] Step S112: Extract the speech content from the speech data of the original video data through automatic speech recognition technology, and transcribe the speech content to obtain an audio segment.

[0025] In some embodiments, for the automated video quality assurance method based on the Visual Internet provided by the present invention, step S110 specifically further includes the following steps: Step S113: Call the lightweight SIFT algorithm to extract key points from the original video data as local features, and use the pre-trained ResNet-18 to extract the deep texture representation in the original video data as global features.

[0026] Step S114: Use Faster R-CNN to detect entity objects in the original video data, and perform scene understanding on the original video data through the ViT-Hybrid model to classify the scenes in the original video data and obtain semantic features.

[0027] Step S115: Based on the local features, global features, entity objects, and semantic features, use the Kaldi toolbox to annotate the original video data, and dynamically adjust the frame extraction interval of key frames according to the annotation results, and output key frames.

[0028] Step S120: Use the pre-trained deep neural network model to conduct content review on the key frames and audio segments to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, trigger an alarm and feedback it to the platform management terminal.

[0029] Specifically, the server uses the pre-trained deep neural network model to conduct content review on the key frames and audio segments extracted in step S110, and determines whether there is any illegal content that does not meet the review criteria in the key frames and audio segments according to the preset review criteria in the system. If it is determined that there is illegal content, trigger the corresponding alarm and feedback it to the platform management terminal.

[0030] In some embodiments, for the automated video quality assurance method based on the visual network provided by the present invention, step S120 specifically includes the following steps: Step S121: Call the pre-trained deep neural network model to analyze the text information, image data, and voice data in the key frames and audio segments to obtain the analysis results.

[0031] Step S122: Configure different review thresholds for the text information, image data, and voice data according to the business scenario, and determine that there is illegal content in the text information and / or image data and / or voice data when the analysis results exceed the corresponding review thresholds.

[0032] In some embodiments, for the automated video quality assurance method based on the visual network provided by the present invention, step S120 specifically further includes the following steps: Step S123: When there is illegal content in the text information and / or image data and / or voice data, trigger an alarm and send an execution instruction to the platform management terminal through the API. The platform management terminal takes blocking or traffic limiting measures in response to the execution instruction.

[0033] Step S130: Play the original video data by simulating the user's viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and record the video parameters associated with the abnormal phenomena to obtain a video quality report.

[0034] Among them, the abnormal phenomena include video stuttering and playback delay, and the video parameters include video buffering time and video loading speed.

[0035] Specifically, the server regularly plays the original video data by simulating the user's viewing process to detect whether there are abnormal phenomena such as video stuttering and playback delay during the playback of the original video data, and records the video parameters of video buffering time and video loading speed associated with the detected abnormal phenomena. By integrating these, the corresponding video quality report can be obtained.

[0036] In some embodiments, for the automated video quality assurance method based on the video networking provided by the present invention, step S130 specifically includes the following steps: Step S131: Calculate the buffering time based on the first frame rendering time and the total duration of the video playback buffering time, and calculate the playback delay based on the difference between the actual playback timestamp and the theoretical playback timestamp.

[0037] Step S132: When the proportion of frames with a continuous frame playback time difference exceeding a set first threshold, calculate the stuttering rate of the original video data, and calculate the video loading speed based on the playback completion time of the original video data.

[0038] Among them, the buffering time, playback delay, stuttering rate, and video loading speed are all monitoring indicators for determining whether there are abnormal phenomena in the original video data.

[0039] Step S140: Based on a preset violation threshold, combine the video quality report and the violation content to determine the violation level of the original video data, and stop the playback and release of the current video when serious violation content or playback quality problems are detected in the original video data.

[0040] Specifically, the server determines the violation level of the current video data based on the preset violation threshold in its system, combined with the video quality report and the violation content obtained in step S130, and stops the playback and release of the current video when serious violation content or playback quality problems are detected in the current video data.

[0041] In some embodiments, for the automated video quality assurance method based on the video networking provided by the present invention, step S140 specifically includes the following steps: Step S141: Determine a preset violation threshold based on the service type of the original video data, and determine the violation level of the original video data with violation content according to the violation threshold. The violation level is preset based on monitoring indicators and includes minor violations, moderate violations, and serious violations.

[0042] Step S142: When it is detected that the violation level of the violation content in the original video data is a serious violation, send a troubleshooting instruction to the platform management terminal. The platform management terminal responds to the troubleshooting instruction to troubleshoot the upstream and downstream links of the original video data to determine whether there is an abnormality in the video source of the original video data.

[0043] The above-mentioned automated video quality assurance method based on the Visual Internet of Things obtains the original video data from various audio sources, and extracts key frames and audio segments from the original video data based on the deep learning algorithm. Then, the pre-trained deep neural network model is used to conduct content review on the key frames and audio segments to determine whether there is violation content in the key frames and audio segments. If there is violation content, an alarm is triggered and feedback is sent to the platform management terminal. After that, the original video data is played by simulating the user viewing process to detect whether there are abnormal phenomena during the playback of the original video data, and the video parameters associated with the abnormal phenomena are recorded to obtain a video quality report. Finally, the violation level of the original video data is determined based on the preset violation threshold in combination with the video quality report and the violation content. When serious violation content or playback quality problems are detected in the original video data, the playback and release of the current video are stopped. This method integrates the extraction of key frames by the deep learning algorithm and the content review and video quality detection driven by the deep neural network, and incorporates the simulation of the user viewing process during the video review and detection, which not only improves the efficiency and accuracy of video review and detection, but also has better comprehensiveness in review.

[0044] Combined with Figures 2 to 5 As shown, in a specific embodiment, the automated video quality assurance method based on the Visual Internet of Things provided by the present invention is mainly implemented by a content receiving module, an intelligent key frame extraction module, a content review module, a playback quality detection module, and a decision execution module.

[0045] In this embodiment, the content receiving module is responsible for receiving the original data stream of the live broadcast or short video, that is, docking various video sources through a network interface and supporting multiple protocols such as RTMP and HLS. The intelligent key frame extraction module is used to dynamically extract key frames and audio sources according to the content of the video.

[0046] Specifically, the intelligent key-frame extraction module is based on deep learning algorithms and realizes dynamic key-frame extraction by adopting fixed-interval extraction + scene detection based on MediaPipe. For the voice part, ASR (Automatic Speech Recognition) technology is used to extract and transcribe the voice content. Task scheduling is triggered, supporting dynamic switching between FFmpeg and OpenCV, and GPU hardware acceleration is supported.

[0047] See Figure 3 As shown, during the process of extracting key frames, the input video stream extracts scene features through the feature extraction layer (texture features, semantic features). Among them, texture feature extraction includes: local features (extracting key points with lightweight SIFT) and global features (extracting deep texture representations using pre-trained ResNet-18). Semantic feature extraction includes: object recognition (detecting main objects using Faster R-CNN) and scene classification (using the ViT-Hybrid model for scene understanding). Then, a classifier is used to judge the scene transition points. Data augmentation is achieved through manual annotation using the Kaldi toolbox to prepare training data. Finally, the key-frame extraction interval is dynamically adjusted according to the classification results, and the extracted key frames are output. The key-frame extraction strategy can be based on interval adjustment of confidence and context-aware sampling.

[0048] In this embodiment, the content review module is used to review the extracted key frames and voice for illegal content.

[0049] Specifically, the content review module uses a pre-trained deep neural network model to analyze image and text information, and configures different thresholds for text, pictures, and voice according to business scenarios. If the review result is higher than the specified threshold, it is determined that there is illegal content. If illegal content is found, an alarm is triggered and the platform management end is notified through the API to take corresponding measures such as banning and traffic limiting. It is also possible to configure the review method for different review sources through configuration policies: machine review, human review, machine review + human review, and review closed.

[0050] Among them, the illegal content and its thresholds are shown in Table 1 as follows: Table 1

[0051] See Figure 4 As shown, multi-dimensional content review is adopted during the video content review process, including: Image review branch: ResNet and Faster R-CNN are used to extract and identify illegal content (such as violence and sensitive information).

[0052] Voice review branch: Voice compliance analysis is realized through BERT-based.

[0053] Review result integration: When political content is detected, the image review weight value is automatically increased to 0.9; in lecture videos, the voice review weight is increased to 0.9 to achieve priority review of corresponding content. When there are differences in the image and voice review results, LSTM context analysis is triggered to synthesize the image and voice review results and generate a final review opinion.

[0054] This multi-dimensional video content review method can demonstrate the workflow of multi-modal review and the relationships among them, ensuring the comprehensiveness of video review.

[0055] In this embodiment, the playback quality detection module is used to simulate the user viewing experience and detect quality problems during video playback.

[0056] Specifically, the playback quality detection module monitors abnormal phenomena such as freezing and latency by regularly playing the video stream, and records relevant parameters such as buffering time and loading speed. Technical means such as FFmpeg stream analysis, end-to-end playback testing with Selenium, and gstreamer real-time stream detection are used for indicator monitoring. For example, the buffering time is calculated based on the total duration of the first frame rendering time + subsequent buffering time; the playback latency is calculated based on the difference between the actual playback timestamp and the theoretical playback timestamp; the freezing rate is calculated based on the proportion of frames with a consecutive frame playback time difference greater than 100 ms; the loading speed is calculated based on the completion time recorded in the video metadata. Additionally, for live broadcasts, attention is paid to their low latency, freezing rate, and packet loss rate; for on-demand videos, attention is paid to the buffering completion rate and average loading time.

[0057] Among them, the monitoring indicators of playback quality, their calculation methods, and thresholds are shown in Table 2 as follows: Table 2

[0058] See Figure 5 As shown, in the review process of video playback quality, the following evaluation mechanism is adopted: Simulate the user viewing environment: Build an immersive simulated viewing environment, simulate multi-network environments such as 3G, 4G, and 5G based on container technology, and simulate a high-load environment on a cloud server. Adopt a multi-dimensional quality index acquisition system to collect basic performance indicators (theoretical frame rate, actual rendering frame rate, first frame latency, and consecutive frame latency), deep content features (picture clarity PSNR, dynamic range HDR, and noise level), and audio quality indicators (maximum peak value, root mean square volume, and signal-to-noise ratio).

[0059] Intelligent quality evaluation: Automatically learn the normal range and perform anomaly detection through a dynamic threshold adjustment algorithm, and dynamically calculate the threshold based on historical data; extract and process image features, extract and process audio features, and perform cross-modal fusion evaluation through a multi-modal fusion evaluation algorithm.

[0060] Generate quality reports and provide timely feedback: Generate visual reports, including buffer time distribution, frame rate trend charts, PSNR heat maps, etc.; Provide automated feedback, including dynamically reducing the video bitrate, switching between H.264|H.265 encoding, enabling AAC audio encoding, etc.

[0061] In this embodiment, the decision execution module is used to make decisions based on the review results and the playback quality report.

[0062] Specifically, based on the violation thresholds set by the business, the decision execution module determines the violation level. When serious violation content or playback quality problems are detected, it automatically stops the current live broadcast or prevents the release of short videos. At the same time, it triggers an alarm to notify the relevant responsible terminals to coordinate the investigation of the quality of the upstream and downstream links and whether there are abnormalities in the video source.

[0063] The above-mentioned automated video quality assurance method based on the visual network dynamically adjusts the frame extraction density according to the video content complexity (such as the frequency of scene changes and the amplitude of actions), which can ensure the comprehensive coverage of key frames. The linkage between the frame extraction strategy and the review results can automatically increase the frame extraction frequency to enhance the review accuracy when potential violation content is detected. It evaluates the technical quality of the video (such as resolution, bitrate, and freezing rate) and the content quality (such as subtitle integrity and picture jitter), and uses a lightweight model (such as MobileNet) to achieve edge-side deployment, which can support real-time quality detection. In addition, an associated database of "low-quality videos - potential violation risks" is established, and the review priority is optimized. When low-quality videos (such as blurring and noise) are detected, it can automatically trigger in-depth reviews to identify hidden violation content. Finally, this method can reverse-optimize the frame extraction strategy and the review model according to the review results and the quality detection report, and improve the overall performance of the video review system through continuous learning (such as online updating of model parameters).

[0064] The automated video quality assurance device based on the visual network provided by the present invention will be described below. The automated video quality assurance device based on the visual network described below can be correspondingly referred to the automated video quality assurance method described above.

[0065] As Figure 6 shown, in one embodiment, an automated video quality assurance device based on the visual network includes an intelligent frame extraction module 610, a content review module 620, a playback quality detection module 630, and a decision execution module 640.

[0066] The intelligent frame extraction module 610 is used to obtain the original video data from various audio sources and extract key frames and audio segments from the original video data based on deep learning algorithms.

[0067] The content review module 620 is used to perform content review on key frames and audio segments through a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, an alarm is triggered and feedback is sent to the platform management terminal.

[0068] The playback quality detection module 630 is used to play the original video data by simulating the user's viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and record the video parameters associated with the abnormal phenomena to obtain a video quality report.

[0069] The decision execution module 640 is used to determine the violation level of the original video data based on a preset violation threshold in combination with the video quality report and the illegal content, and stop the playback and release of the current video when serious illegal content or playback quality problems are detected in the original video data.

[0070] Among them, the abnormal phenomena include video stuttering and playback delay, and the video parameters include video buffering time and video loading speed.

[0071] In this embodiment, for the automated video quality assurance device based on the video networking provided by the present invention, the intelligent frame extraction module 610 is specifically used for: Based on the deep learning algorithm, key frames are dynamically extracted from the original video data by fixed-interval extraction combined with scene detection based on MediaPipe.

[0072] The voice content is extracted from the voice data of the original video data through automatic speech recognition technology, and the voice content is transcribed to obtain an audio segment.

[0073] In this embodiment, for the automated video quality assurance device based on the video networking provided by the present invention, the intelligent frame extraction module 610 is specifically further used for: Call the lightweight SIFT topic to extract key points from the original video data as local features, and extract the deep texture representation in the original video data as global features through the pre-trained ResNet-18.

[0074] Use Faster R-CNN to detect the entity objects in the original video data, and perform scene understanding on the original video data through the ViT-Hybrid model to classify the scenes in the original video data to obtain semantic features.

[0075] Based on the local features, global features, entity objects, and semantic features, the original video data is labeled through the Kaldi toolbox, and the frame extraction interval of the key frames is dynamically adjusted according to the labeling results, and the key frames are output.

[0076] In this embodiment, for the automated video quality assurance device based on the Visual Internet provided by the present invention, the content review module 620 is specifically configured to: Call a pre-trained deep neural network model to analyze the text information, image data, and voice data in the key frames and audio segments, and obtain an analysis result.

[0077] Configure different review thresholds for the text information, image data, and voice data according to the business scenario, and when the analysis result exceeds the corresponding review threshold, determine that there is illegal content in the text information and / or image data and / or voice data.

[0078] In this embodiment, for the automated video quality assurance device based on the Visual Internet provided by the present invention, the content review module 620 is further specifically configured to: When there is illegal content in the text information and / or image data and / or voice data, trigger an alarm and send an execution instruction to the platform management terminal through the API, and the platform management terminal takes blocking or traffic limiting measures in response to the execution instruction.

[0079] In this embodiment, for the automated video quality assurance device based on the Visual Internet provided by the present invention, the playback quality detection module 630 is specifically configured to: Calculate the buffer time according to the first frame rendering time and the total duration of the video playback buffer time, and calculate the playback delay based on the difference between the actual playback timestamp and the theoretical playback timestamp.

[0080] When the proportion of frames with a continuous frame playback time difference exceeding a set first threshold, calculate the stuttering rate of the original video data, and calculate the video loading speed based on the playback completion time of the original video data.

[0081] Among them, the buffer time, playback delay, stuttering rate, and video loading speed are all monitoring indicators for determining whether there are abnormal phenomena in the original video data.

[0082] In this embodiment, for the automated video quality assurance device based on the Visual Internet provided by the present invention, the decision execution module 640 is specifically configured to: Determine a preset violation threshold based on the business type of the original video data, and determine the violation level of the original video data with illegal content according to the violation threshold. The violation level is preset based on the monitoring indicators, including minor violations, moderate violations, and serious violations.

[0083] When it is detected that the violation level of the illegal content in the original video data is a serious violation, send a troubleshooting instruction to the platform management terminal, and the platform management terminal troubleshoots the upstream and downstream links of the original video data in response to the troubleshooting instruction to determine whether there is an abnormality in the video source of the original video data.

[0084] Figure 7 Illustrates a schematic diagram of the physical structure of an electronic device, which can be a smart terminal, and its internal structure diagram can be as Figure 7 shown. The electronic device includes a processor, an internal memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an automated video quality assurance method based on the visual network, and the method includes: Obtain original video data from various audio sources, and extract key frames and audio segments from the original video data based on a deep learning algorithm; Perform content review on the key frames and audio segments through a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, trigger an alarm and feedback it to the platform management terminal; Play the original video data by simulating the user viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and record the video parameters associated with the abnormal phenomena to obtain a video quality report; Determine the violation level of the original video data based on a preset violation threshold in combination with the video quality report and the illegal content, and stop the playback and release of the current video when serious illegal content or playback quality problems are detected in the original video data; Among them, the abnormal phenomena include video stuttering and playback delay, and the video parameters include video buffer time and video loading speed.

[0085] Those skilled in the art can understand that Figure 7 the structure shown in

[0086] is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the electronic device to which the solution of the present invention is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout. Obtain original video data from various audio sources, and extract key frames and audio segments from the original video data based on a deep learning algorithm; Conduct content review on key frames and audio segments through a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, trigger an alarm and feedback it to the platform management terminal; Play the original video data by simulating the user viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and record the video parameters associated with the abnormal phenomena to obtain a video quality report; Based on a preset violation threshold, combined with the video quality report and the illegal content, determine the violation level of the original video data. When serious illegal content or playback quality problems are detected in the original video data, stop the playback and release of the current video; Among them, the abnormal phenomena include video stuttering and playback delay, and the video parameters include video buffering time and video loading speed.

[0087] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. When the processor of the electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, an automated video quality assurance method based on the video networking is implemented. The method includes: Obtain the original video data from various audio sources, and extract key frames and audio segments from the original video data based on the deep learning algorithm; Conduct content review on key frames and audio segments through a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio segments. If there is illegal content, trigger an alarm and feedback it to the platform management terminal; Play the original video data by simulating the user viewing process to detect whether there are any abnormal phenomena during the playback of the original video data, and record the video parameters associated with the abnormal phenomena to obtain a video quality report; Based on a preset violation threshold, combined with the video quality report and the illegal content, determine the violation level of the original video data. When serious illegal content or playback quality problems are detected in the original video data, stop the playback and release of the current video; Among them, the abnormal phenomena include video stuttering and playback delay, and the video parameters include video buffering time and video loading speed.

[0088] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory.

[0089] By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0090] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0091] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. An automated video quality assurance method based on visual networking, characterized in that: The method comprises: Obtaining raw video data from various audio sources, and extracting key frames and audio clips from the raw video data based on a deep learning algorithm; Conduct content review on the key frames and audio clips through a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio clips. If there is any illegal content, an alarm is triggered and fed back to the platform management terminal; Playing the original video data by simulating a user's viewing process to detect whether the original video data has an abnormality during the playback process, and recording video parameters associated with the abnormality to obtain a video quality report; Determine the violation level of the original video data based on a preset violation threshold combined with the video quality report and the violation content, and stop playing and publishing the current video when serious violation content or playback quality problems are detected in the original video data; The abnormal phenomena include video freeze and playback delay, and the video parameters include video buffering time and video loading speed.

2. The automated video quality assurance method based on visual networking according to claim 1, characterized in that: The obtaining of raw video data from various audio sources and extracting key frames and audio clips from the raw video data based on a deep learning algorithm includes: Based on the deep learning algorithm, dynamically extract the key frame from the original video data by fixed interval extraction combined with scene detection based on MediaPipe; The voice content is extracted from the voice data of the original video data by using automatic voice recognition technology, and the voice content is transcribed to obtain the audio clip.

3. The automated video quality assurance method based on visual networking according to claim 2 is characterized in that: The method of obtaining raw video data from various audio sources and extracting key frames and audio clips from the raw video data based on a deep learning algorithm further includes: Calling lightweight SIFT to extract key points from the original video data as local features, and extracting deep texture representations in the original video data as global features through pre-trained ResNet-18; Using Faster R-CNN to detect physical objects in the original video data, and using the ViT-Hybrid model to perform scene understanding on the original video data to classify the scenes in the original video data and obtain semantic features; Based on the local features, global features, entity objects and semantic features, the original video data is annotated by the Kaldi toolbox, and the frame extraction interval of the key frame is dynamically adjusted according to the annotation results to output the key frame.

4. The automated video quality assurance method based on visual networking according to claim 1, characterized in that: The key frames and audio clips are audited by the pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio clips. If there is any illegal content, an alarm is triggered and fed back to the platform management terminal, including: Calling a pre-trained deep neural network model to analyze the text information, image data, and voice data in the key frames and audio clips to obtain analysis results; Different audit thresholds are configured for text information, image data and voice data according to business scenarios, and when the analysis result exceeds the corresponding audit threshold, it is determined that there is illegal content in the text information and / or image data and / or voice data.

5. The automated video quality assurance method based on visual networking according to claim 4 is characterized in that: The method further includes: performing content review on the key frames and audio clips through the pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio clips. If there is any illegal content, an alarm is triggered and fed back to the platform management terminal, and further includes: When there is illegal content in the text information and / or image data and / or voice data, an alarm is triggered and an execution instruction is sent to the platform management terminal through the API. The platform management terminal takes blocking or flow limiting measures in response to the execution instruction.

6. The automated video quality assurance method based on visual networking according to claim 5, characterized in that: Playing the original video data by simulating a user viewing process to detect whether the original video data has an abnormality during the playback process, and recording video parameters associated with the abnormality to obtain a video quality report, includes: The buffering time is calculated based on the first frame rendering time and the total duration of the video playback buffering time, and the playback delay is calculated based on the difference between the actual playback timestamp and the theoretical playback timestamp; When the difference in the playback time of consecutive frames exceeds the frame ratio of a set first threshold, calculating the jam rate of the original video data, and calculating the video loading speed based on the playback completion time of the original video data; Among them, the buffering time, playback delay, freeze rate and video loading speed are all monitoring indicators for determining whether there is any abnormality in the original video data.

7. The automated video quality assurance method based on visual networking according to claim 6, characterized in that: The determining the violation level of the original video data based on the preset violation threshold combined with the video quality report and the violation content, and stopping the playback and release of the current video when serious violation content or playback quality problems are detected in the original video data, includes: Determining the preset violation threshold based on the service type of the original video data, and determining the violation level of the original video data with illegal content according to the violation threshold, wherein the violation level is preset based on the monitoring indicator and includes a minor violation, a moderate violation, and a serious violation; When it is detected that the violation level of the illegal content in the original video data is a serious violation, a check instruction is sent to the platform management terminal, and the platform management terminal checks the upstream and downstream links of the original video data in response to the check instruction to determine whether there is any abnormality in the video source of the original video data.

8. An automated video quality assurance device based on visual networking, characterized in that: The device comprises: An intelligent frame extraction module, used to obtain raw video data from various audio sources, and extract key frames and audio clips from the raw video data based on a deep learning algorithm; A content review module is used to review the key frames and audio clips through a pre-trained deep neural network model to determine whether there is any illegal content in the key frames and audio clips. If there is any illegal content, an alarm is triggered and fed back to the platform management terminal; A playback quality detection module, used to play the original video data by simulating a user viewing process to detect whether the original video data has abnormal phenomena during the playback process, and record video parameters associated with the abnormal phenomena to obtain a video quality report; A decision execution module, configured to determine the violation level of the original video data based on a preset violation threshold combined with the video quality report and the violation content, and to stop the playback and release of the current video when serious violation content or playback quality problems are detected in the original video data; The abnormal phenomena include video freeze and playback delay, and the video parameters include video buffering time and video loading speed.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the automated video quality assurance method based on visual networking described in any one of claims 1 to 7 are implemented.

10. A computer storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the automated video quality assurance method based on visual networking described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Video key frame quality enhancement method and device, equipment and storage medium

    CN121305446A