Communication video quality real-time automatic detection method and system based on AI technology

By using an AI-based real-time automatic detection method for communication video quality, detection vectors of the source and target videos are generated and merged to calculate video quality. This solves the problems of limited video detection functionality and the influence of the source video in existing technologies, and achieves multi-dimensional and accurate quality detection.

CN121056683BActive Publication Date: 2026-02-13HANGZHOU JINSHUO INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511575190.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-13
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

Existing technologies lack comprehensive video service quality detection functions, and fail to consider the impact of source video on quality detection, resulting in inaccurate detection results.

Method used

A real-time automatic detection method for communication video quality based on AI technology is adopted. By setting a video quality classification set, source video and target video are generated, images are cropped separately, detection vectors are constructed, and the data are merged and calculated to eliminate the influence of source video on quality detection.

Benefits of technology

It enables multi-dimensional detection of various video quality conditions in video communication services, improves the accuracy of quality detection, and eliminates the influence of source video on the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056683B_ABST
    Figure CN121056683B_ABST
Patent Text Reader

Abstract

The application discloses a communication video quality real-time automatic detection method and system based on AI technology, and the method comprises the following steps: setting a video quality classification set, and generating a source video for each classification set; establishing a video communication service, selecting the video quality classification set to play the source video at a sending end, and generating a target video corresponding to the classification set at a receiving end; respectively performing picture interception on the source video and the target video to obtain a video picture set; generating a label representing the quality of the video picture, and constructing a first detection vector of the source video and the target video; performing two classification on the first detection vector of the source video and the target video based on a relative quality detection target, and constructing a second detection vector of the source video and the target video; merging and calculating the second detection vector of the source video and the target video to obtain a video quality value, and determining a video quality condition. The application evaluates the video quality from multiple dimensions, and discards the influence of the source video on the quality detection, so that the target video quality detection is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, in particular to a communication video quality real-time automatic detection method and system based on AI technology. BACKGROUND

[0002] With the continuous promotion and extensive application of video communication services such as video call and video conference, quality problems such as freezing and video mosaic have become a problem that cannot be ignored in video service application. To solve the quality problem of video service, it is necessary to first detect whether these problems exist in a specific communication service, and then determine the cause. Before the emergence of multi-modal large models, it is very difficult to automatically analyze and detect quality problems in video services using software. Current video detection methods have single functions and can only perform single quality analysis, such as video freezing, and lack methods for multi-aspect video quality detection. Moreover, video detection methods often only analyze target videos, without considering the influence of source videos, resulting in inaccurate quality detection results.

[0003] The application with the application number 202411559364.3 and the name of video freezing detection method, system, electronic device and readable storage medium discloses real-time acquisition of a video, i.e., a target video, analysis of the time difference between adjacent frame images of the video to determine whether the video freezes. This patent only discloses detection of video freezing, has a single function, lacks multi-aspect video quality detection, and only analyzes adjacent frame images of the target video in video freezing detection, without considering the influence of source videos, resulting in inaccurate detection results. SUMMARY

[0004] The present application mainly solves the problem of difficulty in automatically detecting video service quality using software or machines in the prior art, and provides a communication video quality real-time automatic detection method and system based on AI technology.

[0005] The present application solves the problem of single function and lack of multi-aspect quality detection in video service quality detection in the prior art, and provides a communication video quality real-time automatic detection method and system based on AI technology.

[0006] The present application solves the problem of inaccurate quality detection results caused by only using target videos for analysis without considering the influence of source videos on quality detection in the prior art, and provides a communication video quality real-time automatic detection method and system based on AI technology.

[0007] The above technical problems of the present application are mainly solved by the following technical solution: a communication video quality real-time automatic detection method based on AI technology, comprising the following steps:

[0008] S1. Set a video quality classification set, and generate a source video for each classification set;

[0009] S2. Establish a video communication service, select the video quality classification set to play the source video at a sending end, and generate a target video corresponding to the classification set at a receiving end;

[0010] S3. Capture pictures from the source video and the target video respectively to obtain a video picture set;

[0011] S4. Generate a label representing the quality of the video pictures according to a quality detection target, and construct a first detection vector of the source video and the target video;

[0012] S5. Perform binary classification on the first detection vectors of the source video and the target video based on a relative quality detection target, and construct a second detection vector of the source video and the target video;

[0013] S6. Merge and calculate the second detection vectors of the source video and the target video to determine the quality of the video.

[0014] The application realizes detection of various quality conditions of a video communication service, and evaluates the quality of the video from multiple dimensions. In the detection process, the source video and the target video detection vectors are merged and calculated, the influence of the source video on the quality is discarded, and the quality detection of the target video is more accurate.

[0015] As a preferred scheme, according to the evaluation dimension of the video quality, a video quality classification set corresponding to the evaluation dimension is set, a video generating pattern information of the relative evaluation dimension is generated by shooting, software or a multi-modal large model, and the video is set as the source video of the video quality classification set.

[0016] The evaluation dimension of the video quality in the application includes whether the video is stuck, delayed, whether the video is flickering, and mosaic. According to the evaluation dimension, a corresponding video quality classification set is generated, such as a video stuck quality classification set generated according to the evaluation dimension of whether the video is stuck, delayed, and other transmission delays. According to the evaluation dimension of whether the video is flickering and mosaic, a video flickering quality classification set is generated. And an attribute is established for each video quality classification set, including the corresponding first frame video image or last frame video image, video generation prompt content, video quality detection target prompt word, video information extraction vector, etc. A video with relative evaluation dimension setting pattern information is generated for the video quality classification set, such as a video stuck quality classification set. Relative to whether it is stuck, a video that does not contain stuck specific pattern information is generated. This video is the source video in the video communication service, which is set to the corresponding video quality classification set. The video played in the video communication service includes a video captured in real time or a generated video. Specifically, a video captured in real time through a camera or a video with corresponding pattern information generated by a multi-modal large model according to generated information.

[0017] As a preferred scheme, a video with relative evaluation dimension setting pattern information is generated by a multi-modal large model as a source video of video quality classification, specifically including:

[0018] According to the video generation prompt content, a first frame image or / and a last frame image of the video is generated;

[0019] Based on the multi-modal large model video generation requirement, the video generation prompt content is converted into a video generation prompt word;

[0020] The multi-modal large model generates a video according to the video generation prompt word, receives the video and sets it as the source video of the video quality classification set.

[0021] According to the video generation prompt content of the video quality classification set, the format required by the video is generated, and the first frame image or / and the last frame image of the video is generated. The video generation prompt content, for example, generates the first frame image img0, and places the digital stopwatch in the first frame image. The preferred video generation prompt content includes the required first frame or last frame image, video length and frame rate, and the requirement of the classification quality (such as not containing flickering and mosaic). The minimum scale of the digital stopwatch is determined according to the video frame rate, for example, for 25FPS, the minimum scale of the digital stopwatch should be below 40ms, and the minimum scale of the digital stopwatch can be set to 10ms.

[0022] According to the multi-modal large model video generation requirement, the video generation prompt content is converted into a video generation prompt word, for example, the video generation prompt content is to generate a source video with a set length and frame rate according to the first frame image, and each frame image contains a natural running digital stopwatch. If the digital stopwatch runs naturally in each frame of the video, the obtained video generation prompt word is:

[0023] prompt = "Please generate a video of {{sec}} seconds based on the first frame image,

[0024] the video frame rate is {{fps}} FPS,

[0025] a digital stopwatch with natural elapsed time must be included,

[0026] the minimum scale of the digital stopwatch is 10 ms;

[0027] """

[0028] Send the first frame image and video generation prompt to the multi-modal large model, where the video length in seconds and the video frame rate in FPS are sent as parameters together with the video generation prompt to the multi-modal large model. The multi-modal large model generates the corresponding video, and the system receives the generated video and sets it as the source video of the corresponding video quality classification set.

[0029] When establishing a video communication service, select a video quality classification set, which includes real-time shooting videos and generated videos. The source video of the classification set is played at the service sending end, and the video is received and recorded at the service receiving end as the target video of the video quality classification set.

[0030] As a preferred solution, video pictures are obtained by picture interception of the source video and the target video, specifically including:

[0031] According to the frame rate, the source video and the target video are intercepted to generate the source video picture set and the target video picture set.

[0032] This solution uses existing software tools to perform screen capture processing on the source video and the target video. The screen capture method is to uniformly capture the number of pictures per second of video as needed, for example, for a frame rate of 25 FPS, the number of pictures per second is not less than 25 pictures, at this time, the screen capture time interval T = 40 ms, and the source video picture set and the target video picture set are obtained.

[0033] As a preferred solution, according to the quality detection target, a label representing the quality of the video picture is generated, and a first detection vector of the source video and the target video is constructed, specifically including:

[0034] The video picture set generates a time vector according to the timestamp sequence of the intercepted pictures;

[0035] For each time component, analyze the features related to the video picture quality detection target, and set the label representing the quality of the video picture by judging whether the features meet the quality detection target;

[0036] According to the obtained label, a first detection vector of the source video and the target video is constructed.

[0037] This scheme employs a multimodal large model to analyze the source and target video image sets. Based on the classification quality to be detected, it returns a vector composed of features representing the quality of each image in the video image set. Specifically, a digital stopwatch is set for each video image according to the frame rate. The digital stopwatch represents the timestamp when the image was captured. Based on this timestamp, the video image set is converted into an n-dimensional time vector representing time. According to the classification quality corresponding to the video quality set, features reflecting that classification quality are obtained. For example, for the video stuttering quality set, the corresponding quality detection target is stuttering, and the feature reflecting whether there is stuttering is the time difference ratio between adjacent frame images; for the video screen tearing quality set, the corresponding quality detection target is screen tearing, and the feature reflecting whether there is screen tearing is the screen tearing pattern in the video image. Feature analysis is used to determine whether the features meet the quality detection target. Specifically, time difference analysis uses a threshold to determine whether there is stuttering, and screen tearing pattern analysis determines whether there is screen tearing. The determined quality is labeled, for example, with "true" and "false" to indicate whether there is stuttering or screen tearing. Finally, the first detection vector for the source and target videos is constructed using the labels corresponding to each time component.

[0038] As a preferred approach, the video quality classification set includes a video stuttering quality classification set, with stuttering being the quality detection target;

[0039] The time component of each time vector corresponding to the video stuttering quality classification set is compared with the time component of the previous time moment to obtain the time difference score. The self-difference vectors of the source video and the target video are constructed based on the time difference score.

[0040] Set a first stuttering threshold, calculate the time difference ratio of each component of the source video self-difference vector, compare the time difference ratio of each component with the first stuttering threshold to determine whether there is stuttering, match a label for each component according to the judgment result, and establish the first detection vector of the source video according to the stuttering labels of each component.

[0041] Set a second stuttering threshold, calculate the time difference ratio of each component of the target video's self-difference vector, compare the time difference ratio of each component with the second stuttering threshold to determine whether there is stuttering, match a label for each component based on the determination result, and establish the first detection vector of the target video based on the stuttering labels of each component.

[0042] The video quality classification set includes a video freezing quality classification set and a video screen flashing quality classification set, and the video communication service is established by using the video freezing quality classification set, the quality detection target is determined to be freezing according to the video freezing quality classification set, and the characteristics reflecting freezing are calculated. The time vector corresponding to the source video is taken, and each component is subtracted from the previous time component from the first component of the time vector to the high component minus the low component. Specifically, each component is subtracted from the previous time component to obtain the time difference value, and the source video self-difference vector is constructed according to the time difference value. The target video self-difference vector is constructed in the same way. The source video and target video self-difference vectors are characteristics related to the freezing quality detection target. The freezing of the video is judged according to the deviation of the self-difference vector from the actual time of watching the video, and the first threshold value and the second threshold value are set for the source video and the target video respectively to judge the freezing in different ranges. As another scheme, the same threshold value can also be set for the source video and the target video. Through comparison with the threshold value, it is determined whether the time component is frozen, and the corresponding label is matched, the label including freezing and non-freezing, and the label is preferably represented by true and false, or can be represented in other forms. Finally, the first detection vector of the source video and the target video is established according to the label.

[0043] As a preferred scheme, the video quality classification set includes a video screen flashing quality classification set, and the quality detection target is screen flashing;

[0044] The video image corresponding to each time component of the time vector is subjected to screen flashing pattern detection to judge whether it is screen flashing, and a label is matched for each time component according to the judgment result, and the first detection vector of the source video and the target video is established according to the screen flashing label of each time component.

[0045] The video communication service is established by using the video screen flashing quality classification set, the quality detection target is determined to be screen flashing according to the video screen flashing quality classification set, and the characteristics reflecting screen flashing, i.e. screen flashing patterns, are detected. The method of the present application adopts a multi-modal large model, and the video picture set corresponding to the time vector of the source video and the target video is respectively sent to the multi-modal large model, and the analysis result of whether each video picture is screen flashing is returned. A label is matched for each time component according to the analysis result, and the label is screen flashing and non-screen flashing, and the label is preferably represented by true and false, or can be represented in other forms. Finally, the first detection vector of the source video and the target video is established according to the label.

[0046] As a preferred scheme, the first detection vector of the source video is subjected to binary classification with the non-quality detection target as the judgment condition, and the second detection vector of the source video is constructed according to the classified value;

[0047] The first detection vector of the target video is subjected to binary classification with the quality detection target as the judgment condition, and the second detection vector of the target video is constructed according to the classified value.

[0048] The non-quality detection target in the scheme represents the opposite target of the quality detection target, such as using a video stall quality classification set, and the quality detection target is stall, and the non-quality detection target represents non-stall. The result of the binary classification is represented by 1 and 0. The non-quality detection target is used as the binary classification judgment subject condition, and 1 is used to represent the case where the non-quality detection target is judged, otherwise 0 is used. Similarly, the quality detection target is used as the binary classification subject judgment condition, and 1 is used to represent the case where the quality detection target is judged, otherwise 0 is used. In the scheme, the component of 1 in the first detection vector of the source video indicates that the video at the screenshot time of the source video is not stall or not screen, and participates in subsequent calculation, and the component of 0 in the first detection vector of the source video indicates that the video at the screenshot time of the source video is stall or screen, and is discarded in subsequent operation. The second detection vector of the target video is opposite to the source video, and the component of 1 in the second detection vector of the target video indicates that the video at the screenshot time of the target video is stall or screen, and is calculated in order to count the stall or screen of the target video. The component of 0 in the second detection vector of the target video indicates that the video at the screenshot time of the target video is not stall or not screen, and is removed from the statistics of the video at the screenshot time in the subsequent calculation.

[0049] As a preferred scheme, the dot product operation is performed on the second detection vector of the source video and the target video to obtain a video quality value.

[0050] A quality threshold is set, and the video quality value is compared with the quality threshold to determine the video quality.

[0051] According to the setting of the second detection vector of the source video and the target video, the dot product operation is performed on the two, the source video meets the non-stall or non-screen and participates in the subsequent calculation, and whether the screenshot time is stall or screen is determined by the target video. If the source video does not meet the condition, that is, there is stall or screen, the dot product operation is 0, regardless of the target video, this screenshot does not contribute to the calculation, and the influence of the source video can be considered to be discarded. The scheme adopts the dot product operation to discard the influence of the source video on the quality, and the quality detection of the target video is more accurate.

[0052] An AI technology communication video quality real-time automatic detection system, comprising:

[0053] A video classification module generates a source video for each classification set according to the evaluation dimension to set a video quality classification set.

[0054] A service sending module selects a video quality classification set to play a source video.

[0055] A service receiving module receives a source video and records to form a target video.

[0056] The quality detection module obtains a video picture set by picture interception on the source video and the target video, generates a label representing the quality of the video picture according to a quality detection target, constructs a first detection vector of the source video and the target video, performs binary classification on the first detection vectors of the source video and the target video based on the relative quality detection target, constructs a second detection vector of the source video and the target video, merges and calculates the second detection vectors of the source video and the target video to obtain a video quality value, and determines the video quality condition.

[0057] Therefore, the present application has the following advantages:

[0058] The video communication service video quality in various conditions is detected, and the video quality is evaluated from multiple dimensions.

[0059] In the detection process, the source video and the target video detection vectors are merged and calculated, the influence of the source video on the quality detection is discarded, and the quality detection of the target video is more accurate. BRIEF DESCRIPTION OF DRAWINGS

[0060] Figure 1 is a flowchart of the method of the present application. DETAILED DESCRIPTION

[0061] The technical solutions of the present application will be further specifically described below by examples in combination with the drawings.

[0062] Example 1:

[0063] The embodiment is a communication video quality real-time automatic detection method based on AI technology, which is used for detecting video communication services such as video telephone ViLTE or video conference, and automatically detecting whether there is a video quality problem such as video lag, delay, screen flicker, mosaic, etc. in ViLTE or video conference in real time. Taking ViLTE video telephone application as an example, the source video is played at the calling end of the ViLTE video telephone, and the video is received at the called end of the ViLTE video telephone and recorded to the target video. Through AI technology, the information of the source video and the target video is automatically compared and analyzed in real time, and whether there is a video quality problem in the ViLTE video telephone is detected in real time: first, the calling party of the ViLTE video telephone plays the source video, and each frame of the source video contains a timestamp pattern. The timestamp is generally a digital stopwatch which displays dynamically according to the natural time of each frame instead of being fixed. The called party of the ViLTE video telephone records the target video, and the timestamp pattern of the source video is transmitted through the ViLTE network and recorded to the target video at the called side. Secondly, the source video and the target video are respectively screened at a certain time interval to obtain a set of screen images, and the timestamp of each screen of the set of screen images of the source video and the target video is automatically detected by using the AI multi-modal large model video understanding function. The timestamp of each screen is generated into a source video time vector and a target video time vector, respectively. By matching and calculating the source video time vector and the target video time vector, whether there is a video quality problem in the video communication service is calculated.

[0064] The embodiment method specifically includes the following steps:

[0065] S1. Set a video quality classification set, and generate a source video for each classification set.

[0066] According to the evaluation dimension of the video quality, set the video quality classification set corresponding to the evaluation dimension, generate the video with the pattern information set according to the relative evaluation dimension by shooting or multi-modal large model, and set the video as the source video of the video quality classification set.

[0067] In the embodiment, the evaluation dimension of the video quality includes whether the video is lagging, delayed, flickering, mosaic, etc. According to the evaluation dimension, the corresponding video quality classification set is generated, such as generating a video lag quality classification set according to the evaluation dimension of whether the video is lagging, delayed, etc. transmission delay, and generating a video flicker quality classification set according to the evaluation dimension of whether the video is flickering, mosaic, etc. And establish the attribute of each video quality classification set, which includes the corresponding first frame video image or last frame video image, video generation prompt content, video quality detection target prompt word, video information extraction vector, etc.

[0068] The video containing the set pattern information is generated for the relative evaluation dimension setting pattern information of the video quality classification set, such as the video stutter quality classification set, and the video containing no stutter is generated for the relative stutter, the video is the source video in the video communication service, and is set to the corresponding video quality classification set. The video played in the video communication service includes a video captured in real time or a generated video, specifically a video captured in real time through a camera or a video generated by a multi-modal large model according to generated information to generate corresponding pattern information.

[0069] Preferably, the source video containing the set pattern information is generated by the AI large model, and the step of generating the source video includes:

[0070] S11. Generate a first frame image or / and a last frame image of the video according to the video generation prompt content;

[0071] S12. Convert the video generation prompt content into a video generation prompt word based on the multi-modal large model video generation requirement;

[0072] S13. The multi-modal large model generates a video according to the video generation prompt word, receives the video and sets it as the source video of the video quality classification set.

[0073] According to the video generation prompt content of the video quality classification set, the format corresponding to the video requirement, the first frame image or / and the last frame image of the video is generated, and the video generation prompt content is, for example, to generate a first frame image img0, and a digital stopwatch is placed in the first frame image. Preferably, the video generation prompt content includes the required first frame or last frame image, video duration and frame rate, and the requirement corresponding to the classification quality (such as no containing screen, mosaic). The minimum scale of the digital stopwatch is determined according to the video frame rate, for example, for 25FPS, the minimum scale of the digital stopwatch should be below 40ms, and the minimum scale of the digital stopwatch can be set to 10ms.

[0074] According to the multi-modal large model video generation requirement, the video generation prompt content is converted into a video generation prompt word, for example, the video generation prompt content is to generate a source video with a set duration and frame rate according to the first frame image, and each frame image contains a natural running digital stopwatch. If the digital stopwatch runs naturally in each frame of the video, the obtained video generation prompt word is:

[0075] prompt="""Please generate a video of {{sec}} seconds according to the first frame image,

[0076] the video frame rate is {{fps}}FPS,

[0077] a natural running digital stopwatch must be included,

[0078] the minimum scale of the digital stopwatch is 10ms;

[0079] """

[0080] The first frame image and the video generation prompt word are sent to the multi-modal large model, wherein the video length in seconds and the video frame rate in FPS are sent to the multi-modal large model as parameters together with the video generation prompt word. The multi-modal large model generates a corresponding video, and the system receives the generated video and sets it as the source video of the corresponding video quality classification set.

[0081] S2. Establish a video communication service, select a video quality classification set to play a source video at the sending end and generate a target video of the corresponding classification set at the receiving end.

[0082] When establishing a video communication service, one of the video quality classification sets is selected, which includes a real-time shot video and a generated video. The source video of the classification set is played at the service sending end, and the video is received and recorded at the service receiving end as the target video of the video quality classification set.

[0083] Taking ViLTE video phone as an example, after establishing a video call, the source video is played at the calling end and the target video is recorded at the called end.

[0084] S3. Picture interception is performed on the source video and the target video respectively to obtain a video picture set.

[0085] Picture interception is performed on the source video and the target video respectively according to the frame rate to generate a source video picture set and a target video picture set.

[0086] In this embodiment, the existing software tool is used to perform screen capture processing on the source video and the target video. The screen capture mode is to uniformly capture the pictures of each second of video according to the number of pictures to be captured, for example, for a frame rate of 25 FPS, the number of pictures captured per second is not less than 25 pictures, at this time, the screen capture time interval T = 40 ms, and the source video picture set and the target video picture set are obtained, which are represented as follows:

[0087] 1 second source video picture set = {img1, img2, …, img25};

[0088] 1 second target video picture set = {img'1, img'2, …, img'25}.

[0089] S4. According to the quality detection target, a label representing the quality of the video picture is generated, and a first detection vector of the source video and the target video is constructed.

[0090] S41. The video picture set generates a time vector according to the timestamp sequence of the captured pictures;

[0091] S42. For each time component, analyze the features related to the video picture quality detection target, and set a label representing the quality of the video picture by judging whether the features meet the quality detection target.

[0092] S43. Construct the first detection vector of the source video and the target video according to the obtained label.

[0093] In this embodiment, a multimodal large model is used to analyze the picture set of the source video and the target video, and a vector composed of features representing the quality of each picture in the video picture set is returned according to the classification quality to be detected. Specifically, a digital stopwatch is set for each video picture according to the frame rate, and the digital stopwatch represents the timestamp of the intercepted picture. According to the timestamp, the video picture set is converted into an n-dimensional time vector representing time. For example, for a 25FPS video, the screen is captured once every T=40ms, and the digital stopwatch of a certain screen capture is denoted as P(n*T), where n is the screen capture order. The timestamp sequence of the screen capture pictures of the source video is obtained:

[0094]

[0095] The timestamp sequence of the screen capture pictures of the target video is:

[0096]

[0097] According to the timestamp sequence of the intercepted pictures, a time vector is generated, including generating an n-dimensional source video time vector:

[0098] {P(1*T), P(2*T),…, P(n*T)};

[0099] The n-dimensional target video time vector is:

[0100] {P' (1*T), P' (2*T),…, P' (n*T)}.

[0101] According to the classification quality corresponding to the video quality classification set, the feature reflecting the classification quality is obtained, such as the video stutter quality classification set, the corresponding quality detection target is stutter, and the feature reflecting whether it is stutter is the adjacent frame image time difference value; the video screen flashing quality classification set, the corresponding quality detection target is screen flashing, and the feature reflecting whether it is screen flashing is the screen flashing pattern of the video image. Whether the feature meets the quality detection target is judged by analyzing the feature, that is, whether the stutter is judged by analyzing the time difference value through the threshold, and whether the screen flashing is judged by analyzing the screen flashing pattern; the quality after judgment is marked with a label, for example, marked with true and false labels, indicating whether it is stutter or screen flashing, and finally the first detection vector of the source video and the target video is constructed according to the label corresponding to each time component.

[0102] S5. Based on the relative quality detection target, the first detection vector of the source video and the target video is binary classified to construct the second detection vector of the source video and the target video.

[0103] S51. Binary classification is performed on the first detection vector of the source video with the non-quality detection target as the judgment condition, and the second detection vector of the source video is constructed according to the classified value;

[0104] S52. Binary classification is performed on the first detection vector of the target video with the quality detection target as the judgment condition, and the second detection vector of the target video is constructed according to the classified value.

[0105] In this embodiment, the non-quality detection target represents the opposite target of the quality detection target. For example, if the quality detection target is stutter, the non-quality detection target represents non-stutter. The result of binary classification is represented by 1 and 0. If the non-quality detection target is the main condition for binary classification, 1 is used to represent the case where the non-quality detection target is judged, otherwise 0 is used. Similarly, if the quality detection target is the main condition for binary classification, 1 is used to represent the case where the quality detection target is judged, otherwise 0 is used. In this scheme, the component 1 in the first detection vector of the source video represents that the video at this screenshot time of the source video is non-stutter or non-flashing, which participates in subsequent calculation. The component 0 in the first detection vector of the source video represents that the video at this screenshot time of the source video is stutter or flashing, which is discarded in subsequent operation. The second detection vector of the target video is opposite to the source video. The component 1 in the second detection vector of the target video represents that the video at this screenshot time of the target video is stutter or flashing, which is calculated in subsequent calculation in order to count the stutter or flashing of the target video. The component 0 in the second detection vector of the target video represents that the video at this screenshot time of the target video is non-stutter or non-flashing, which is removed from the statistics of the video at this screenshot time in subsequent calculation. According to the value obtained after binary classification, the second detection vectors of the source video and the target video are constructed.

[0106] S6. The second detection vectors of the source video and the target video are combined to calculate the video quality ratio, and the video quality condition is determined.

[0107] S61. The second detection vectors of the source video and the target video are subjected to dot product operation to obtain a video quality value;

[0108] S62. A quality threshold is set, and the video quality value is compared with the quality threshold to determine the video quality condition.

[0109] According to the setting of the second detection vectors of the source video and the target video, the dot product operation is performed on the two, the source video meets the non-stutter or non-flashing condition and participates in subsequent calculation, then whether the screenshot time is stutter or flashing is determined by the target video. If the source video does not meet the condition, that is, there is stutter or flashing, the dot product operation is 0, regardless of the target video, this screenshot does not contribute to the calculation, and the influence of the source video can be considered to be discarded. In this scheme, the dot product operation discards the influence of the source video on the quality, and the quality detection of the target video is more accurate.

[0110] Embodiment 2:

[0111] The embodiment of the application is a communication video quality real-time automatic detection method of AI technology. The embodiment detects the video freezing quality, and the specific steps are as follows:

[0112] S1. Set a video quality classification set, and generate a source video for each classification set.

[0113] According to the evaluation dimension of the video quality, set a video quality classification set corresponding to the evaluation dimension, generate a video with pattern information relative to the evaluation dimension by camera or multi-modal large model, and set the video as the source video of the video quality classification set.

[0114] In the embodiment, the video freezing quality classification set is generated according to whether the motion pattern in the communication service has freezing and other transmission delays as the evaluation dimension, and the frame rate is 25 FPS, that is, 25 frames per second.

[0115] Generate a video with pattern information relative to the freezing evaluation dimension by camera or multi-modal large model, and set the video as the source video of the video freezing quality classification set.

[0116] S2. Establish a video communication service, select a video quality classification set to play a source video at a sending end, and generate a target video corresponding to the classification set at a receiving end.

[0117] When the video communication service is established, the video freezing quality classification set is selected, the source video of the video freezing quality classification set is played at the service sending end, and the video is received and recorded at the service receiving end as the target video of the video freezing quality classification set.

[0118] S3. Capture pictures of the source video and the target video respectively to obtain a video picture set.

[0119] Capture pictures of the source video and the target video respectively according to the frame rate to generate a source video picture set and a target video picture set.

[0120] For a frame rate of 25 FPS, the number of pictures captured per second is not less than 25 pictures, at this time, the screenshot time interval T is 40 ms, and the source video picture set and the target video picture set are obtained.

[0121] S4. Generate a label representing the quality of the video picture according to the quality detection target, and construct a first detection vector of the source video and the target video.

[0122] S41. The video picture set generates a time vector according to the timestamp sequence of the captured pictures.

[0123] For a 25 FPS video, the digital stopwatch is recorded as P(n*T), n is the screenshot order, and an n-dimensional source video time vector is generated:

[0124] {P'(1*T), P'(2*T),…, P'(n*T)}.

[0125] n-dimensional target video time vector:

[0126] {P'(1*T), P'(2*T),…, P'(n*T)}.

[0127] S42. For each time component, analyze the features related to the quality detection target of the video picture, and set a label representing the quality of the video picture by judging whether the features meet the quality detection target.

[0128] According to the obtained label, construct the first detection vector of the source video and the target video.

[0129] According to the quality classification set of the video freezing, determine that the quality detection target is freezing, calculate the feature reflecting freezing, and the feature is the difference value of the time difference of adjacent frames.

[0130] S421. The difference between each time component of the time vector corresponding to the video freezing quality classification set and the previous time component is obtained, and the self-difference vector of the source video and the target video is constructed according to the time difference value.

[0131] Specifically, the time vector corresponding to the source video is taken, starting from the first component of the time vector, and the high component is subtracted from the low component, that is, each component is subtracted from the previous time component to obtain the time difference value, and the self-difference vector of the source video is constructed according to the time difference value.

[0132] For example, for a 25FPS video, taking a screenshot every T=40ms, the unit time is 1 second, and the source video time vector is:

[0133] { P(t+1T),P(t+2T),…, P(t+25T)};

[0134] Where t is the starting time, T is the screenshot interval, and the previous time component of the first component P(t+1T) of the source video time vector is P(t), that is, the component of the time vector of the last screenshot in the previous unit time.

[0135] The self-difference vector ΔP(t,T) of the source video is calculated and obtained:

[0136] { P(t+1T)- P(t),P(t+2T) - P(t+1T),…, P(t+25T)- P(t+24T)};

[0137] Similarly, for the target video time vector:

[0138] { P'(t+1T),P'(t+2T),…, P'(t+25T)};

[0139] where t is the starting time, T is the screenshot interval, and the previous time component of the first component P'(t+1T) of the target video time vector is P'(t), i.e., the component of the time vector of the last screenshot of the last unit time.

[0140] The target video difference vector ΔP'(t,T) is calculated as follows:

[0141] {P'(t+1T)-P'(t),P'(t+2T)-P'(t+1T),…,P'(t+25T)-P'(t+24T)}.

[0142] S422. Set the first stutter threshold, calculate the time difference ratio of each component of the source video difference vector, compare the time difference ratio of each component with the first stutter threshold to determine whether there is stutter, match a label to each component according to the determination result, and establish a first detection vector of the source video according to the stutter labels of each component.

[0143] The first stutter threshold μ is set, and the value of μ is 105%, indicating that the deviation between the clock of the moving object in the source video and the clock of the person watching the video is not higher than 5%, which is determined as non-stutter, otherwise it is determined as stutter. The labels true and false are set to represent stutter and non-stutter respectively, and the corresponding label is set according to the determination of whether there is stutter. The time difference ratio is the ratio of the difference vector to the screenshot time interval.

[0144] The n components of the first detection vector of the source video are represented as:

[0145] If the time difference ratio of the n components of the source video time vector is greater than μ, then the n components of the first detection vector of the source video are true;

[0146] If the time difference ratio of the n components of the source video time vector is less than μ, then the n components of the first detection vector of the source video are false.

[0147] The specific representation is as follows:

[0148] The n components of the first detection vector of the source video are represented as: .

[0149] The first detection vector of the source video is established according to the stutter labels of each component:

[0150] R(t,T)=(r1,r2,…,rn).

[0151] For example, for a 25FPS video, the source video first detection vector is (false,false,false,…,false) if the screenshots are taken every T=40ms and the unit time is 1 second.

[0152] S423. Set the second stalling threshold, calculate the time difference ratio of each component of the target video difference vector, compare the time difference ratio of each component with the second stalling threshold to determine whether it is stalling, match the label for each component according to the determination result, and establish the first detection vector of the target video according to the stalling label of each component.

[0153] Set the first stalling threshold λ, and the value of λ is 120%, which means that when the deviation between the moving object in the source video and the viewer's clock is higher than 1.2 times, it is determined to be stalling, otherwise it is determined to be non-stalling. Set the labels true and false to represent stalling and non-stalling respectively, and set the corresponding label according to whether it is stalling,

[0154] The n components of the first detection vector of the target video are represented as:

[0155] If the time difference ratio of the n components of the target video time vector is greater than λ, then the n components of the first detection vector of the target video are true;

[0156] If the time difference ratio of the n components of the target video time vector is less than λ, then the n components of the first detection vector of the target video are false.

[0157] The specific representation is as follows:

[0158] The n components of the first detection vector of the target video are represented as: .

[0159] According to the stalling label of each component, the first detection vector of the target video is established:

[0160] R' (t,T)=(r'1,r'2,…,r'n).

[0161] For example, for a 25FPS video, take a screenshot every T=40ms, the unit time is 1 second, and the first detection vector of the target video is: (false,true,false, …, false).

[0162] S5. Based on the relative quality detection target, the first detection vector of the source video and the target video is binary classified, and the second detection vector of the source video and the target video is constructed.

[0163] S51. Take the non-quality detection target as the judgment condition, and perform binary classification on the first detection vector of the source video, and construct the second detection vector of the source video according to the classified value;

[0164] Specifically, take the non-stalling as the judgment condition, perform binary classification on the first detection vector of the source video, and the n components of the second detection vector of the source video are represented as:

[0165] When the n-th component of the first detection vector of the source video is true, the n-th component of the second detection vector of the source video is 0.

[0166] When the n-th component of the first detection vector of the source video is false, the n-th component of the second detection vector of the source video is 1.

[0167] The specific representation is as follows:

[0168] The n-th component of the second detection vector of the source video ,

[0169] The second detection vector of the source video is constructed as follows:

[0170] Q = (q1, q2, …, qn),

[0171] For example: (1, 1, 1, …, 1).

[0172] S52. The first detection vector of the target video is classified according to the quality detection target as the judgment condition, and the second detection vector of the target video is constructed according to the classified value.

[0173] The first detection vector of the target video is classified according to the choppiness as the judgment condition, and the n-th component of the second detection vector of the target video is represented as follows:

[0174] When the n-th component of the first detection vector of the target video is true, the n-th component of the second detection vector of the target video is 1.

[0175] When the n-th component of the first detection vector of the target video is false, the n-th component of the second detection vector of the target video is 0.

[0176] The specific representation is as follows:

[0177] The n-th component of the second detection vector of the target video ,

[0178] The second detection vector of the target video is constructed as follows:

[0179] Q' = (q'1, q'2, …, q'n),

[0180] For example: (0, 1, 0, …, 0).

[0181] S6. The second detection vectors of the source video and the target video are combined to calculate the video quality ratio, and the video quality situation is determined.

[0182] S61. The second detection vectors of the source video and the target video are subjected to dot product operation to obtain the video quality value.

[0183] The second detection vector Q of the source video and the second detection vector Q' of the target video are subjected to dot product operation:

[0184] Q·Q'=q1*q'1+q2*q'2+q3*q'3+…+qn*q'n

[0185] Each arbitrary component qi, q'i of each vector has only two values of 1 and 0, if qi is 1, it means that the source video at the screenshot time meets the requirements, qi*q'i=1*q'i, that is, whether the target video determines whether the screenshot time is stuck, if there is stuck q'i=1, then qi*q'i=1*q'i=q'i=1, otherwise qi*q'i=0.

[0186] If qi is 0, it means that the source video at the screenshot time does not meet the requirements, that is, stuck, regardless of the target video, this screenshot has no contribution to the stuck calculation, that is, the influence of the source video is abandoned.

[0187] The dot product operation value is the video quality value, which represents the number of screenshots of the target video with stuck after abandoning the influence of the original video.

[0188] S62. Set the quality threshold, compare the video quality value with the quality threshold, and judge the video quality.

[0189] The vector dimension n represents the total number of screenshots in a unit time, and the video quality ratio in a unit time is: Q·Q' / n.

[0190] Set the ratio of the target video to appear stuck delay image after abandoning the source video factor to be greater than the set proportion, which means that the video service has stuck, which is represented as:

[0191] Q·Q' / n≥a

[0192] Convert the formula to:

[0193] Q·Q'≥an

[0194] Set the quality threshold an, wherein a is set to 20%, and the cumulative number of screenshots of the video clock and the observer clock with stuck at the screenshot time accounts for more than 20% of the total number of screenshots in a unit time:

[0195] Q·Q'≥20%n

[0196] It is represented that the ratio of the target video to appear stuck delay image after abandoning the source video factor is greater than the quality threshold, which means that the video service has stuck.

[0197] Embodiment 3:

[0198] The communication video quality real-time automatic detection method of the AI technology in this embodiment detects the video screen quality, and the specific steps are as follows:

[0199] S1. Set a video quality classification set, and generate source videos for each classification set.

[0200] In this embodiment, whether the motion pattern in the communication service has a screen flicker or mosaic is used as an evaluation dimension to generate a video flicker quality classification set. A video containing pattern information in the relative flicker evaluation dimension is generated by camera or multi-modal large model, and the video is set as the source video of the video flicker quality classification set.

[0201] In this embodiment, the source video containing the set pattern information is preferably generated by an AI large model. The steps of generating the source video include:

[0202] According to the video generation prompt content of the video quality classification set, the first frame image img0 and the last frame image img of the video are generated according to the required format of the video, and the images are required to be clear and do not contain flicker or mosaic.

[0203] According to the requirements of the multi-modal large model video generation, the video generation prompt content is converted into a video generation prompt word, which is expressed as follows:

[0204] prompt = """Please generate a video of {{sec}} seconds according to the first frame image and the last frame image,

[0205] the video frame rate is {{fps}} FPS,

[0206] the images must be cleaned and the objects in the images must move naturally and smoothly,

[0207] and there must be no mosaic or flicker;

[0208] """

[0209] The first frame image, the last frame image, and the video generation prompt word are sent to the multi-modal large model, wherein the video length in seconds and the video frame rate in FPS are sent to the multi-modal large model as parameters together with the video generation prompt word. The multi-modal large model generates a corresponding video, and the system receives the generated video and sets it as the source video of the video flicker quality classification set.

[0210] S2. Establish a video communication service, select a video quality classification set to play a source video at the sending end, and generate a target video corresponding to the classification set at the receiving end.

[0211] When establishing a video communication service, the video flicker quality classification set is selected, the source video of the video flicker quality classification set is played at the service sending end, and the video is received and recorded at the service receiving end as the target video of the video flicker quality classification set.

[0212] S3. Pictures are captured from the source video and the target video respectively to obtain a video picture set.

[0213] According to the frame rate, the source video and the target video are respectively picture intercepted to generate a source video picture set and a target video picture set.

[0214] For a frame rate of 25 FPS, the number of pictures intercepted per second is not less than 25 pictures, at this time, the screenshot time interval T = 40 ms, and the source video picture set and the target video picture set are obtained.

[0215] S4. According to the quality detection target, a label representing the quality of the video picture is generated, and a first detection vector of the source video and the target video is constructed.

[0216] S41. The video picture set generates a time vector according to the timestamp sequence of the intercepted pictures.

[0217] For a 25 FPS video, the screen is captured once every T = 40 ms, and the digital stopwatch is recorded as P(n*T), n is the screenshot order, and an n-dimensional source video time vector is generated:

[0218] {P(1*T), P(2*T),…, P(n*T)};

[0219] n-dimensional target video time vector:

[0220] {P' (1*T), P' (2*T),…, P' (n*T)}.

[0221] S42. For each time component, analyze the features related to the quality detection target of the video picture, and set the label representing the quality of the video picture by judging whether the features meet the quality detection target. According to the obtained label, a first detection vector of the source video and the target video is constructed.

[0222] According to the video screen quality classification set, the quality detection target is determined to be screen and mosaic, and the feature reflecting the stall is calculated, which is the screen pattern and the mosaic pattern.

[0223] The current step specifically includes:

[0224] The video image corresponding to each time component of the time vector is detected for screen pattern, and it is judged whether it is a screen.

[0225] According to the judgment result, each time component is matched with a label, and a first detection vector of the source video and the target video is respectively established according to the screen label of each time component.

[0226] In this embodiment, the video image screen and mosaic pattern are recognized by a multi-modal large model. According to the video understanding requirement of the multi-modal large model, a video detection prompt word is generated, which is used to identify whether each picture in the video picture set contains a screen or a mosaic pattern, for example:

[0227] prompt = "Please check whether each image contains a screen pattern or a mosaic pattern, and the result is true (contains a screen or mosaic pattern), false indicates,

[0228] Output in json format;

[0229] """

[0230] Send the video detection prompt word to the multi-modal large model together with the source video picture set and the target video picture set, receive the analysis result of each picture returned by the large model, that is, the label, and construct the source video first detection vector and the target video first detection vector according to the label according to the correspondence between the video picture set and the time vector.

[0231] Source video first detection vector: R(t, T) = (r1, r2, …, rn).

[0232] For example, for a 25FPS video, take a screenshot every T=40ms, and the unit time is 1 second, and the source video first detection vector is: ( false, false, false,…, false ).

[0233] Target video first detection vector: R'(t, T) = (r'1, r'2, …, r'n).

[0234] For example, for a 25FPS video, take a screenshot every T=40ms, and the unit time is 1 second, and the target video first detection vector is: ( false, true, false,…, false ).

[0235] S5. Based on the relative quality detection target, the source video and the target video first detection vector are classified, and the source video and the target video second detection vector are constructed.

[0236] S51. Take the non-quality detection target as the judgment condition, and classify the source video first detection vector, and construct the source video second detection vector according to the classified value;

[0237] Specifically, take the non-screen or mosaic as the judgment condition, classify the source video first detection vector, and represent as follows:

[0238] Source video second detection vector n component ,

[0239] Construct the source video second detection vector:

[0240] Q = (q1, q2, …, qn),

[0241] For example: ( 1, 1, 1,…, 1 ).

[0242] S52. Take the quality detection target as the judgment condition, and perform binary classification on the target video first detection vector. According to the classified value, the target video second detection vector is constructed.

[0243] Take the screen or mosaic as the judgment condition, and perform binary classification on the target video first detection vector, which is represented as follows:

[0244] Target video second detection vector n component ,

[0245] Construct the target video second detection vector:

[0246] Q'=(q'1,q'2,…,q'n),

[0247] For example: (0, 1, 0, …, 0).

[0248] S6. Merge the source video and the target video second detection vector to calculate the video quality ratio and determine the video quality.

[0249] S61. Perform dot product operation on the source video and the target video second detection vector to obtain the video quality value.

[0250] Perform dot product operation on the source video second detection vector Q and the target video second detection vector Q':

[0251] Q·Q'=q1*q'1+q2*q'2+q3*q'3+…+qn*q'n

[0252] Each component qi, q'i of the vector has only two values of 1 and 0. If qi is 1, it means that the source video meets the requirements at that screenshot time, qi*q'i=1*q'i, that is, whether the screen is blank or mosaic is determined by the target video. If there is a stall q'i=1, then qi*q'i=1*q'i=q'i=1, otherwise qi*q'i=0.

[0253] If qi is 0, it means that the source video does not meet the requirements at that screenshot time, that is, there is a blank or mosaic, regardless of the target video. This screenshot has no contribution to the blank or mosaic calculation, that is, the influence of the source video is discarded.

[0254] The dot product operation value is the video quality value, which represents the number of screenshots with blank or mosaic in the target video after discarding the influence of the original video.

[0255] S62. Set the quality threshold, compare the video quality value with the quality threshold, and judge the video quality.

[0256] If the quality threshold is set to 1, then

[0257] Q·Q'≥1

[0258] It is indicated that the target video has a large number of screen flicker and mosaic images greater than the quality threshold after discarding the source video factors.

[0259] Embodiment 4:

[0260] The embodiment is a communication video quality real-time automatic detection method based on AI technology, which detects various classification qualities of videos. The embodiment includes steps S1-S6, which are the same as steps S1-S6 in Embodiment 1. The embodiment also includes step S7, which is specifically:

[0261] S7. Generate multiple quality subsets according to different evaluation standards of video communication services. A matrix is generated by using multiple target video second detection vectors at the same time, with the same screenshot duration and the same dimension, which comprehensively represents the quality of video communication services.

[0262] Embodiment 5:

[0263] The embodiment is an AI technology-based communication video quality real-time automatic detection system, which is used to implement the methods in Embodiments 1-4. The system includes:

[0264] A video classification module sets a video quality classification set according to the evaluation dimensions and generates a source video for each classification set.

[0265] A service sending module selects a video quality classification set to play a source video.

[0266] A service receiving module receives a source video and records it to form a target video.

[0267] A quality detection module performs picture interception on the source video and the target video to obtain a video picture set, generates a label representing the quality of the video picture according to a quality detection target, constructs a source video and a target video first detection vector, performs binary classification on the source video and the target video first detection vector based on a relative quality detection target, constructs a source video and a target video second detection vector, combines and calculates the source video and the target video second detection vector to determine the quality of the video.

[0268] The specific embodiments described herein are merely illustrative of the spirit of the present application. Those skilled in the art of the present application can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, without deviating from the spirit of the present application or exceeding the scope defined by the appended claims.

Claims

1. An AI technology-based real-time automatic detection method for communication video quality, characterized in that, Comprise: Set a video quality classification set, generate source video for each classification set; Establish a video communication service, select a video quality classification set to play source video at the sending end, and generate target video corresponding to the classification set at the receiving end; According to the frame rate, the source video and the target video are respectively subjected to picture interception to obtain source video and target video picture sets; Analyze the features related to the quality detection target of the video pictures, determine whether the features meet the label representing the quality of the video pictures generated by the quality detection target, and construct the first detection vector of the source video and the target video according to the obtained label; Take the non-quality detection target as the judgment condition, and perform binary classification on the first detection vector of the source video, and represent it as 1 in the case of non-quality detection target, otherwise as 0, and construct the second detection vector of the source video according to the classified value; Take the quality detection target as the judgment condition, and perform binary classification on the first detection vector of the target video, and represent it as 1 in the case of quality detection target, otherwise as 0, and construct the second detection vector of the target video according to the classified value; The dot product operation is performed on the second detection vectors of the source video and the target video to obtain the video quality value, and the video quality condition is determined.

2. The communication video quality real-time automatic detection method based on AI technology according to claim 1, wherein: According to the evaluation dimension of video quality, set the video quality classification set corresponding to the evaluation dimension, generate the video with pattern information relative to the evaluation dimension through camera, software or multi-modal large model, and set the video as the source video of the video quality classification set. 3.The AI technology-based real-time automatic communication video quality detection method of claim 2, wherein, The multi-modal large model generates a video with pattern information relative to the evaluation dimension as the source video of the video quality classification, specifically including: Generate the first frame image or / and the last frame image of the video according to the video generation prompt content; Based on the multi-modal large model video generation requirement, convert the video generation prompt content into video generation prompt words; The multi-modal large model generates a video according to the video generation prompt words, receives the video and sets it as the source video of the video quality classification set. 4.The AI technology-based real-time automatic communication video quality detection method of claim 1, wherein: According to the label representing the quality of the video pictures generated by the quality detection target, construct the first detection vector of the source video and the target video, specifically including: The video picture set generates a time vector according to the timestamp sequence of the intercepted pictures; For each time component, analyze the features related to the quality detection target of the video pictures, and set the label representing the quality of the video pictures by judging whether the features meet the quality detection target; According to the obtained label, construct the first detection vector of the source video and the target video.

5. The communication video quality real-time automatic detection method based on AI technology according to claim 4, wherein: The video quality classification set includes a video stutter quality classification set, and the quality detection target is stutter; Obtain the time difference value by subtracting each time component of the time vector corresponding to the video stutter quality classification set from the previous time component, and construct the self-difference vector of the source video and the target video according to the time difference value; Setting a first freezing threshold, calculating the time difference ratio of each component of the source video autocorrelation difference vector, comparing the time difference ratio of each component with the first freezing threshold to determine whether it is frozen, matching the label for each component according to the judgment result, and establishing a source video first detection vector according to the freezing labels of each component; Setting a second freezing threshold, calculating the time difference ratio of each component of the target video autocorrelation difference vector, comparing the time difference ratio of each component with the second freezing threshold to determine whether it is frozen, matching the label for each component according to the judgment result, and establishing a target video first detection vector according to the freezing labels of each component.

6. The AI technology-based communication video quality real-time automatic detection method according to claim 4, characterized in that: The video quality classification set includes a video screen quality classification set, and the quality detection target is a screen; Performing screen pattern detection on the video image corresponding to each time component of the time vector, determining whether it is a screen, matching the label for each time component according to the judgment result, and respectively establishing a source video and a target video first detection vector according to the screen labels of each time component.

7. The AI technology-based communication video quality real-time automatic detection method according to claim 1, characterized in that: Performing dot product operation on the source video and target video second detection vectors to obtain a video quality value; Setting a quality threshold, comparing the video quality value with the quality threshold, and determining the video quality condition.

8. An AI technology communication video quality real-time automatic detection system, implementing the method of any one of claims 1-7, characterized in that, It includes: A video classification module that generates a source video for each classification set according to the evaluation dimension and sets a video quality classification set; A service sending module that selects a video quality classification set to play a source video; A service receiving module that receives a source video and records it to form a target video; A quality detection module that cuts the source and target videos to obtain a video picture set, generates labels for the video pictures, constructs a source video and a target video first detection vector, performs two classification on the first detection vector, constructs a source video and a target video second detection vector, combines the source video and target video second detection vectors to calculate a video quality value, and determines the video quality condition.

Citation Information

Patent Citations

  • Video lag detection method and system, electronic equipment and readable storage medium

    CN119364108A

  • Video transmission quality evaluation method based on multi-task deep learning

    CN110913207A

  • Video playing quality detection method and device

    CN112511818A