Intelligent inspection method and equipment for video terminal equipment, and storage medium

By initiating inspection calls to video terminal devices, collecting and separating audio, mainstream video, and auxiliary video track data, performing anomaly detection, and generating structured reports, the problem of time-consuming and labor-intensive manual inspections in existing technologies is solved. This achieves automated, comprehensive, and quantifiable intelligent inspections of video terminal devices, ensuring accurate assessment of audio and video quality and business continuity.

CN121547561APending Publication Date: 2026-02-17SHENZHEN XINGWANG XINTONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511829115.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing video terminal equipment inspection methods rely on manual inspection, which is time-consuming and labor-intensive, and cannot comprehensively and accurately assess audio and video quality. As a result, audio and video quality problems are not detected in a timely manner, affecting the continuity of video conferencing and security monitoring.

Method used

By initiating inspection call commands to video terminal devices, detection data of audio track, mainstream video track and auxiliary video track are collected, separated into audio stream, mainstream video stream and auxiliary video stream, anomaly detection is performed, and a structured inspection report is generated in combination with network quality detection results.

Benefits of technology

It enables automated, comprehensive, and quantifiable intelligent inspection of video terminal devices, accurately identifies audio and video quality anomalies, ensures timely detection and repair of potential faults in critical scenarios, and guarantees business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547561A_ABST
    Figure CN121547561A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent inspection method and device for video terminal equipment and a storage medium, and relates to the technical field of intelligent inspection of equipment, and the method comprises the steps: collecting detection data including an audio track, a main stream video track and an auxiliary stream video track; extracting an audio stream, a main stream video stream and an auxiliary stream video stream based on the detection data; performing audio anomaly detection on the audio stream to obtain an audio detection result; performing main stream video anomaly detection on the main stream video stream to obtain a main stream video detection result, and performing auxiliary stream video anomaly detection on the auxiliary stream video stream to obtain an auxiliary stream video detection result; and obtaining a network quality detection result of the target video terminal equipment, and generating an inspection detection report of the target video terminal equipment, so that the technical problem that the audio and video quality of the video terminal equipment in the operation process cannot be automatically and comprehensively evaluated is solved, and automatic full-coverage intelligent inspection of the target video terminal equipment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent inspection of equipment, and in particular to an intelligent inspection method of a video terminal device, equipment and a storage medium. BACKGROUND

[0002] In the current audio and video communication technology is widely used in video conferencing, distance education, security monitoring and other fields, the stable operation of various video terminal devices (such as conference terminal, camera, large screen, etc.) is crucial. However, the main way of routine maintenance and state checking of the video terminal device in the prior art is to rely on manual inspection by operation and maintenance personnel. Specifically, the operation and maintenance personnel need to initiate actual calls or on-site viewing to each video terminal device, and evaluate whether the device can be normally started, whether the picture is displayed, and whether there is sound and other states through subjective judgment. The existing device intelligent inspection relies on manual testing of audio and video functions of each device to determine whether they are normal, which not only consumes time and effort, is prone to detection omissions, but also cannot systematically evaluate the quality status of the audio and video tracks of the device, so as to cause that the actual running state of the video terminal device cannot be comprehensively and accurately mastered, and the continuity of video conference, security monitoring and other business scenarios is affected due to the audio and video quality problems not being discovered in advance. SUMMARY

[0003] The main purpose of the present application is to provide an intelligent inspection method, equipment and storage medium for a video terminal device, aiming to solve the technical problem that the existing inspection method cannot automatically and comprehensively evaluate the audio and video quality of the video terminal device during operation.

[0004] To achieve the above-mentioned purpose, the present application provides an intelligent inspection method for a video terminal device, applied to an intelligent inspection system for a video terminal device, which comprises: initiating an inspection call instruction to a target video terminal device, wherein the target video terminal device collects detection data containing an audio track, a main stream video track and an auxiliary stream video track in response to the inspection call instruction, and sends the detection data to the intelligent inspection system for the video terminal device, the main stream video track is a video picture collected by a camera of the target video terminal device, and the auxiliary stream video track is a screen sharing video picture of the target video terminal device; extracting an audio stream, a main stream video stream and an auxiliary stream video stream based on the detection data; performing audio anomaly detection on the audio stream to obtain an audio detection result; performing main stream video anomaly detection on the main stream video stream to obtain a main stream video detection result, and performing auxiliary stream video anomaly detection on the auxiliary stream video stream to obtain an auxiliary stream video detection result; obtaining a network quality detection result of the target video terminal device, and generating an inspection detection report of the target video terminal device according to the network quality detection result, the audio detection result, the main stream video detection result and the auxiliary stream video detection result.

[0005] In addition, to achieve the above object, the present application further provides a video terminal device intelligent inspection method and device. The video terminal device intelligent inspection device comprises: a video collection module configured to initiate an inspection call instruction to a target video terminal device, wherein the target video terminal device collects detection data comprising an audio track, a main stream video track and an auxiliary stream video track in response to the inspection call instruction, and sends the detection data to a video terminal device intelligent inspection system, the main stream video track being a video picture collected by a camera of the target video terminal device, and the auxiliary stream video track being a screen sharing video picture of the target video terminal device; a multi-track data separation module configured to extract an audio stream, a main stream video stream and an auxiliary stream video stream based on the detection data; an audio detection module configured to perform audio anomaly detection on the audio stream to obtain an audio detection result; a video detection module configured to perform main stream video anomaly detection on the main stream video stream to obtain a main stream video detection result, and perform auxiliary stream video anomaly detection on the auxiliary stream video stream to obtain an auxiliary stream video detection result; a comprehensive evaluation module configured to obtain a network quality detection result of the target video terminal device, and generate an inspection detection report of the target video terminal device according to the network quality detection result, the audio detection result, the main stream video detection result and the auxiliary stream video detection result.

[0006] In addition, to achieve the above object, the present application further provides a video terminal device intelligent inspection device. The device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the video terminal device intelligent inspection method as described above.

[0007] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program. When the computer program is executed by a processor, the steps of the video terminal device intelligent inspection method as described above are implemented.

[0008] In addition, to achieve the above object, the present application further provides a computer program product, which comprises a computer program. When the computer program is executed by a processor, the steps of the video terminal device intelligent inspection method as described above are implemented.

[0009] The one or more technical solutions provided in the application have at least the following technical effects: a patrol call instruction is initiated to a target video terminal device, the target video terminal device collects complete audio and video data containing an audio track, a main stream video track and an auxiliary stream video track, i.e., detection data, in response to the patrol call instruction, and sends the detection data to an intelligent patrol system of the video terminal device, the main stream video track is a video picture collected by a camera of the target video terminal device, and the auxiliary stream video track is a screen sharing video picture of the target video terminal device. Based on the detection data, an audio stream, a main stream video stream and an auxiliary stream video stream are extracted to provide a basis for multi-dimensional anomaly detection. Audio anomaly detection is performed on the audio stream to obtain an audio detection result to evaluate an audio processing related function, main stream video anomaly detection is performed on the main stream video stream to obtain a main stream video detection result to evaluate a camera and a physical environment, and auxiliary stream video anomaly detection is performed on the auxiliary stream video stream to obtain an auxiliary stream video detection result to evaluate a screen sharing function and content display quality, so as to accurately identify quality anomalies of each track. A network quality detection result of the target video terminal device is obtained, and the network quality detection result, the audio detection result, the main stream video detection result and the auxiliary stream video detection result are integrated to generate a structured patrol detection report of the target video terminal device, thereby solving a technical problem that an existing patrol method cannot automatically and comprehensively evaluate audio and video quality of a video terminal device in a running process, and realizing automatic, full coverage and quantifiable intelligent patrol of the target video terminal device, so as to help ensure that potential faults can be found and repaired before a key scene such as a conference or security, and business continuity is ensured. The intelligent patrol method of the video terminal device provided in the application initiates a patrol call actively, so that the target video terminal device outputs complete detection data containing audio, a main stream (camera picture) and an auxiliary stream (screen sharing) in a real business state, then separates the detection data to obtain independent audio stream, main stream / auxiliary stream video stream, and performs anomaly detection on each track, obtains detection results of the audio, main stream / auxiliary stream video track, and simultaneously combines a network quality detection result to generate a structured patrol detection report, realizes a closed loop from automatic collection to multi-dimensional detection and comprehensive evaluation, solves a technical problem that an existing patrol method cannot automatically and comprehensively evaluate audio and video quality of a video terminal device in a running process, and realizes automatic, full coverage and quantifiable intelligent patrol of the target video terminal device. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the application and serve to explain the principles of the application.

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating an embodiment of the intelligent inspection method for video terminal equipment in this application. Figure 2 This is a schematic diagram of the module structure of the intelligent inspection device for video terminal equipment according to an embodiment of this application; Figure 3 This is a schematic diagram of the hardware operating environment involved in the intelligent inspection method of the video terminal device in this application embodiment. Detailed Implementation

[0013] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0014] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0015] The main solution of this application embodiment is as follows: An inspection call command is initiated to the target video terminal device. The target video terminal device responds to the inspection call command by collecting detection data including audio track, main video track, and auxiliary video track, and sends the detection data to the intelligent inspection system of the video terminal device. The main video track is the video image captured by the camera of the target video terminal device, and the auxiliary video track is the screen-shared video image of the target video terminal device. Based on the detection data, audio stream, main video stream, and auxiliary video stream are extracted. Audio anomaly detection is performed on the audio stream to obtain audio detection results. Main video anomaly detection is performed on the main video stream to obtain main video detection results, and auxiliary video anomaly detection is performed on the auxiliary video stream to obtain auxiliary video detection results. The network quality detection results of the target video terminal device are obtained. Based on the network quality detection results, audio detection results, main video detection results, and auxiliary video detection results, an inspection report for the target video terminal device is generated.

[0016] In this embodiment, for ease of description, the intelligent inspection system of the video terminal device will be used as the execution subject for the following description.

[0017] Currently, the primary method for routine maintenance and status checks of video terminal equipment relies on manual inspections by maintenance personnel. Specifically, these personnel need to call or physically inspect each video terminal individually, subjectively assessing whether the equipment can be turned on normally, whether the image is displayed, and whether there is sound. Existing intelligent equipment inspection methods, which rely on manual testing of audio and video functions on each device, are not only time-consuming and labor-intensive, prone to omissions, but also lack a systematic assessment of the quality of the audio and video tracks. This results in an inability to comprehensively and accurately grasp the actual operating status of the video terminal equipment, making it easy for audio and video quality issues to go undetected in advance, impacting the continuity of business scenarios such as video conferencing and security monitoring.

[0018] This application provides a solution that proactively initiates inspection calls to enable target video terminal devices to output complete detection data, including audio, mainstream (camera feed), and secondary (screen sharing) streams, under real-world business conditions. The detection data is then separated into independent audio and mainstream / secondary video streams, and targeted anomaly detection is performed on each track to obtain detection results for the audio and mainstream / secondary video tracks. Simultaneously, network quality detection results are combined to generate a structured inspection report, achieving a closed loop from automatic data collection to multi-dimensional detection and comprehensive evaluation. This solves the technical problem that existing inspection methods cannot automatically and comprehensively evaluate the audio and video quality of video terminal devices during operation, enabling automated, comprehensive, and quantifiable intelligent inspection of target video terminal devices.

[0019] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an intelligent inspection device for video terminal devices capable of performing the above functions. The following description uses an intelligent inspection system for video terminal devices as an example to illustrate this embodiment and the subsequent embodiments.

[0020] Based on this, embodiments of this application provide an intelligent inspection method for video terminal devices, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the intelligent inspection method for video terminal equipment according to this application.

[0021] In this embodiment, the intelligent inspection method for video terminal equipment includes steps 101-105: Step 101: Initiate an inspection call command to the target video terminal device. The target video terminal device responds to the inspection call command by collecting detection data including audio track, main video track and auxiliary video track, and sends the detection data to the intelligent inspection system of the video terminal device. The main video track is the video image captured by the camera of the target video terminal device, and the auxiliary video track is the screen-shared video image of the target video terminal device. Step 102: Extract the audio stream, main video stream, and auxiliary video stream based on the detection data.

[0022] Specifically, the inspection call command is a standardized communication request sent by the intelligent inspection system (such as an operation and maintenance platform) of the video terminal equipment to the target video terminal equipment. It can be based on the Session Initiation Protocol (SIP) or Hypertext Transfer Protocol (HTTP) to trigger the terminal to initiate the audio and video acquisition process, replacing manual operation to trigger the detection. The target video terminal equipment is the video terminal equipment to be inspected, which can be a video conferencing terminal, surveillance camera, large-screen display device, or other devices with audio and video acquisition / output capabilities. The audio track is a dedicated data stream track in the detection data that carries audio signals and can store audio information such as voice and ambient sound collected by the terminal. The mainstream video track is the video data stream track in the detection data that carries the "camera-captured images." The image source is the camera of the target video terminal device (such as images of participants in a meeting). The auxiliary video track is the video data stream track in the detection data that carries the "screen-shared images." The image source is the desktop / document sharing function of the target video terminal device (such as shared PPTs or desktop content during a meeting). The detection data is a composite audio and video file generated by the target video terminal device after responding to the inspection command. It includes the audio track, mainstream video track, and auxiliary video track. The detection data is the raw data carrier for subsequent detection. The audio stream, mainstream video stream, and auxiliary video stream are independent data streams obtained after separating the detection data, corresponding to the audio track, mainstream video track, and auxiliary video track of the detection data, respectively.

[0023] In some embodiments, firstly, the intelligent inspection system of the video terminal device sends a standardized inspection call instruction to the target video terminal device (such as a video conferencing terminal, an integrated collaboration terminal, etc.). This inspection call instruction can be transmitted via SIP or HTTP protocols. The inspection call instruction contains audio and video acquisition parameter configuration information, such as video acquisition resolution, frame rate, audio acquisition sampling rate, bit rate, and audio and video acquisition duration. After receiving the inspection call instruction, the target video terminal device automatically enters a preset self-test mode, activates its local camera to capture real-time environmental images as the main video track, simultaneously activates its microphone to capture audio signals to form an audio track, and automatically enables screen sharing according to a built-in strategy, loading a preset standard test content, such as a PPT file containing specific text, color blocks, and graphics, to generate an auxiliary video track. Subsequently, the target video terminal device encapsulates the data from the aforementioned audio track, main video track, and auxiliary video track into complete test data and sends it back to the intelligent inspection system of the video terminal device. After receiving the detection data, the intelligent inspection system of the video terminal device can call the FFmpeg-based audio-video separation algorithm to separate the detection data. Specifically, by parsing the track metadata in the detection data, it extracts independent audio streams, main video streams (corresponding to the camera view, which can be labeled as the main video stream), and auxiliary video streams (corresponding to screen-shared content, which can be labeled as auxiliary video streams or presentation streams). The audio stream can be in Waveform Audio File (WAV) format, while the main and auxiliary video streams can use the H.264 video encoding standard. Simultaneously, during the separation process, it ensures that the audio stream, main video stream, and auxiliary video stream are time-aligned and format-consistent. This provides a clearly structured and non-interfering input data source for subsequent parallel and specialized anomaly detection targeting audio quality, the authenticity and clarity of the main image, and the readability of the auxiliary content. This provides an independent and standardized data source for subsequent anomaly detection for each track.

[0024] Step 103: Perform audio anomaly detection on the audio stream to obtain the audio detection results.

[0025] Specifically, the audio detection result is a structured result output after audio anomaly detection, which may include audio anomaly type labels, audio quality scores, audio dimension scores for each preset audio dimension, and detection conclusions of the audio detection results.

[0026] In some embodiments, multi-dimensional anomaly detection is performed on the separated audio stream. First, the audio stream can be preprocessed, for example, by using an adaptive Wiener filter for noise reduction. Next, acoustic features (peak amplitude, real-time decibel level, noise power, signal-to-noise ratio, etc.) and semantic features (the degree of matching between the audio stream and standard text) of the preprocessed audio stream are extracted. Subsequently, various acoustic and semantic features are compared with preset thresholds (silence threshold, noise threshold, semantic matching degree threshold, etc.) to identify audio anomalies such as silence, feedback, noise, and incomplete speech, generate an audio quality score, and obtain structured audio detection results.

[0027] Step 104: Perform mainstream video anomaly detection on the mainstream video stream to obtain the mainstream video detection result, and perform auxiliary video anomaly detection on the auxiliary video stream to obtain the auxiliary video detection result.

[0028] Specifically, the mainstream video stream is an independent video data stream separated from the detection data, formed by the images captured by the camera of the target video terminal device. The auxiliary video stream is an independent video data stream separated from the detection data, formed by the screen-shared images of the target video terminal device. The mainstream video detection result is a structured result output after anomaly detection of the mainstream video stream, which may include mainstream video anomaly type labels, mainstream video quality scores, static image quality scores, image anomaly detection results for each mainstream keyframe, number of static anomaly frames, dynamic smoothness scores, total number of frame skips, total number of stutters, and the detection conclusion of the mainstream video detection result. The auxiliary video detection result is a structured result output after anomaly detection of the auxiliary video stream, which may include video frame quality scores for each auxiliary stream keyframe, auxiliary stream video anomaly type labels, auxiliary stream video quality scores, and the detection conclusion of the auxiliary stream video detection result.

[0029] In some embodiments, the intelligent inspection system of the video terminal device extracts key frames from the mainstream video stream, uses a large visual model to evaluate sharpness and brightness, and combines the Structural Similarity Index Measure (SSIM) algorithm to detect stuttering or freezing. For the auxiliary video stream, it can perform optical character recognition (OCR) recognition, compare the extracted text with a preset standard reference template, and analyze the text clarity, color deviation, and text rendering of the text area to detect missing content, blurring, or color distortion. Finally, it generates detection results for the mainstream video and auxiliary video stream respectively, realizing accurate evaluation of the quality of different video output scenarios of the target video terminal device.

[0030] Step 105: Obtain the network quality test results of the target video terminal device, and generate an inspection report of the target video terminal device based on the network quality test results, audio test results, mainstream video test results, and auxiliary stream video test results.

[0031] Specifically, the network quality test results are quantitative data collected on the network transmission status of the target video terminal device, including key parameters such as network latency, packet loss rate, network jitter, and available bandwidth. These results serve as the core basis for evaluating the stability of audio and video transmission. The inspection report is a comprehensive diagnostic output generated by integrating the above multi-source information (i.e., network quality test results, audio test results, mainstream video test results, and auxiliary stream video test results), used to characterize the overall health status of the target video terminal device.

[0032] Optionally, during audio and video anomaly detection, the network quality detection results of the target video terminal device are acquired simultaneously. These network quality detection results can be obtained in various ways, such as: real-time monitoring of four parameters—network latency, packet loss rate, network jitter, and available bandwidth—within a preset time period (e.g., 1 minute) based on Transmission Control Protocol (TCP) / User Datagram Protocol (UDP). Additionally, the network terminal's own network interface operation data is received and normalized to form a standardized network quality detection result. Subsequently, the network quality detection result is comprehensively evaluated along with the previously obtained audio detection results, mainstream video detection results, and auxiliary stream video detection results. Based on a preset fusion rule base or weighted scoring model (e.g., Comprehensive Health Score = 0.25 × Audio Quality Score + 0.3 × Mainstream Video Quality Score + 0.2 × Auxiliary Stream Quality Score + 0.25 × Network Quality Score), the comprehensive health score of the target video terminal device is calculated. Simultaneously, anomaly root cause correlation analysis is performed. Based on predefined "phenomenon-root cause" mapping rules, multi-dimensional cross-judgments are made. For example, when video stuttering and high network packet loss are detected simultaneously, the root cause is associated with network jitter; when silence is detected and audio energy is zero, the root cause is associated with microphone malfunction or the mute switch being on. Finally, the system generates a structured inspection report. This report includes the target video terminal device's device identifier, scores for each dimension's sub-items, a list of anomaly types, a comprehensive health score, and possible root cause suggestions. It is output in JavaScript Object Notation (JSON) or Hypertext Markup Language (HTML) format for automatic alarm, statistical analysis, or work order dispatch by the intelligent inspection system for video terminal devices, thereby achieving a comprehensive, interpretable, and automated assessment of the video terminal device's operating status.

[0033] Based on the intelligent inspection method for video terminal devices provided in this application, an inspection call command is initiated to the target video terminal device. The target video terminal device responds to the inspection call command by collecting complete audio and video data, i.e., detection data, including the audio track, the main video track, and the auxiliary video track, and sends the detection data to the intelligent inspection system of the video terminal device. The main video track is the video image captured by the camera of the target video terminal device, and the auxiliary video track is the screen-shared video image of the target video terminal device. The detection data is separated and processed to obtain the audio stream, the main video stream, and the auxiliary video stream, providing a basis for multi-dimensional anomaly detection. Audio anomaly detection is performed on the audio stream to obtain audio detection results, which are used to evaluate audio processing-related functions. Mainstream video anomaly detection is performed on the main video stream to obtain main video detection results, which are used to evaluate the camera and physical environment. Auxiliary video anomaly detection is performed on the auxiliary video stream to obtain auxiliary video detection results, which are used to evaluate the screen-shared function and content display quality, accurately identifying quality anomalies in each track. This system acquires network quality test results for target video terminal devices and integrates these results with audio, mainstream video, and auxiliary video streams to generate a structured inspection report. This addresses the technical challenge of existing inspection methods failing to automatically and comprehensively assess the audio and video quality of video terminal devices during operation. It enables automated, comprehensive, and quantifiable intelligent inspection of target video terminal devices, helping to identify and repair potential faults in critical scenarios such as conferencing and security, thus ensuring business continuity.

[0034] In some embodiments, the step of performing audio anomaly detection on an audio stream to obtain audio detection results includes: Calculate the peak amplitude sequence, decibel value sequence, and audio energy value sequence of the audio stream in the time domain; Perform a short-time Fourier transform on the audio stream to obtain the audio energy spectrum corresponding to the audio stream, and determine the energy proportion of the human voice frequency band based on the audio energy spectrum; The signal-to-noise ratio of the audio stream is calculated based on the audio energy spectrum and audio energy value sequence. The audio stream is converted into corresponding audio recognition text, and the character-level text matching degree between the audio recognition text and the preset standard text in the inspection task is calculated. The audio detection results corresponding to the audio stream are determined based on the decibel value sequence, peak amplitude sequence, energy proportion of human voice frequency band, signal-to-noise ratio, and character-level text matching degree.

[0035] Specifically, the audio energy spectrum is the output of the short-time Fourier transform, presenting the energy distribution of audio across different frequency bands and time frames in matrix form, serving as the core data carrier for frequency domain feature extraction. The human voice frequency band energy percentage is the percentage of the total energy in the core human voice frequency band (e.g., 300-3400Hz) out of the total energy of the entire frequency band in the audio energy spectrum, used to assess the clarity and effectiveness of the human voice signal. The signal-to-noise ratio (SNR) is the ratio of effective audio signal energy to noise signal energy (unit: dB). A higher SNR value indicates less noise interference in the audio signal; SNR is a core indicator for evaluating audio quality. The peak amplitude sequence is time-series data composed of the maximum amplitude of each audio frame; the decibel value sequence is volume time-series data calculated based on the mean squared amplitude of each frame; and the audio energy value sequence is energy time-series data composed of the sum of the squared amplitudes of each audio signal frame.

[0036] As an example, the raw audio stream separated from the detection data undergoes noise reduction and signal normalization sequentially. Noise reduction can employ a deep learning-based speech enhancement model, such as a recurrent neural network noise suppression (RNNoise) algorithm or a deep complex convolutional recurrent network (DCCRN), to suppress ambient background noise and obtain a denoised audio stream. The amplitude of the denoised audio stream is then uniformly scaled to a standard range (e.g., [-1,1]) to eliminate gain differences between different devices, resulting in a preprocessed audio stream. Subsequently, the preprocessed audio stream is framed, and the peak amplitude, decibel value, and root mean square audio energy of each frame are calculated, forming a temporal sequence of peak amplitude, decibel value, and audio energy value. Simultaneously, a Short-Time Fourier Transform (STFT) is performed on the preprocessed audio stream to obtain the audio energy spectrum in the time-frequency domain. Based on this audio energy spectrum, the energy of the main human voice frequency band in the range of 300 Hz to 3400 Hz is integrated and divided by the total energy of the entire frequency band to obtain the energy proportion of the human voice frequency band. The energy proportion of the human voice frequency band is used to determine whether there is valid speech content. Further, combining the audio energy value sequence and the audio energy spectrum, speech activity detection (VAD) is used to divide speech segments and non-speech segments, and the signal power and noise power are estimated respectively, thereby calculating the signal-to-noise ratio (SNR). In addition, the preprocessed audio stream is input into a pre-trained Automatic Speech Recognition (ASR) model, such as Whisper or PaddleSpeech, to generate corresponding audio recognition text, which is then compared character-level with the preset standard text in the inspection task (such as "the audio of this inspection test") to calculate the text matching degree (e.g., accuracy = number of correct characters / total number of characters). Finally, by integrating multiple dimensions of features, including decibel value sequence (assessing whether the volume is too low), peak amplitude sequence (detecting clipping or silence), human voice frequency band energy ratio (judging sound quality abnormalities), signal-to-noise ratio (measuring clarity), and character-level text matching degree (verifying speech intelligibility), and based on preset audio anomaly detection rules (e.g., if text matching degree <70% and signal-to-noise ratio <15 dB, it is judged as "speech distortion"; if peak amplitude is consistently close to 0 and decibel value <20 dB, it is judged as "silence"), a structured audio detection result is generated, which includes audio quality score and audio anomaly type label (such as silence, clipping, feedback, sound quality abnormality, noise, or speech distortion). This provides a reliable basis for subsequent comprehensive inspection reports.

[0037] In some embodiments, the step of determining the audio detection result corresponding to the audio stream based on the decibel value sequence, peak amplitude sequence, energy proportion of human voice frequency band, signal-to-noise ratio, and character-level text matching degree includes: According to the preset audio anomaly detection rules, the decibel value sequence, peak amplitude sequence, human voice frequency band energy ratio, signal-to-noise ratio and character-level text matching degree are mapped by rules to generate audio dimension scores of the audio stream in each preset audio dimension; The audio stream is compared with the audio dimension score of each preset audio dimension to determine whether the audio stream has at least one audio abnormality among mute, clipping, howling, abnormal sound quality, noise or speech distortion. Obtain the first weighting coefficient corresponding to the score of each audio dimension, and perform a weighted summation of each first weighting coefficient and the score of each audio dimension to obtain the initial audio quality score; If at least one audio anomaly exists, the first preset score is determined as the audio quality score, and the audio anomaly is used as the audio anomaly type label. If no audio anomalies are found, the initial audio quality score is set to the audio quality score, and the audio anomaly type label is set to empty. If the audio quality score is less than the audio pass threshold, the audio detection result is determined to be unqualified; if the audio quality score is greater than or equal to the audio pass threshold, the audio detection result is determined to be qualified, and the audio pass threshold is greater than the first preset score. The audio detection result is defined by the audio anomaly type label, audio quality score, audio dimension score of each preset audio dimension, and the detection conclusion of the audio detection result.

[0038] Specifically, the preset audio anomaly detection rules are pre-defined judgment logics for detecting audio anomalies. These rules map audio features (decibels, peak amplitude, vocal frequency band energy percentage, signal-to-noise ratio, and character-level text matching degree) to audio dimension scores for each preset audio dimension. The preset audio dimensions are core dimensions defined based on audio quality assessment requirements, such as volume stability, signal strength, anti-interference capability, effective signal percentage, and semantic accuracy. The audio dimension score is a quantitative score (range 0-100 points) converted from each audio feature according to the preset audio anomaly detection rules, used to characterize the quality level of the corresponding preset audio dimension. The dimension threshold is the critical value for judging whether a preset audio dimension is abnormal; for example, noise anomalies exist when the signal-to-noise ratio is <15dB. Audio anomalies are quality problems in the preset audio dimensions that do not meet the dimension threshold requirements, including silence, clipping (peak amplitude exceeding the standard), feedback (low signal-to-noise ratio), abnormal sound quality (vocal frequency band percentage / signal-to-noise ratio not meeting the standard), noise (low signal-to-noise ratio), and speech distortion (low character-level text matching degree). Audio anomalies correspond to preset audio dimensions. For example, silence corresponds to the volume stability dimension, clipping to the signal strength dimension, howling and noise to the signal strength dimension, sound quality anomalies to the effective signal ratio dimension, and speech distortion to the semantic accuracy dimension. The first weighting coefficient represents the importance weight of each preset audio dimension to the overall quality of the audio stream (e.g., 30% for semantic accuracy, 25% for anti-interference capability, 20% for effective signal ratio, 15% for volume stability, and 10% for signal strength). The first preset score is a lower quality score (e.g., 50, 40, 30, etc.) forcibly set when any preset audio dimension anomaly exists. The audio pass threshold is the threshold for judging whether the overall audio stream is "passable" (e.g., 70 points), which is higher than the first preset score. The audio anomaly type label is an identifier that records the specific type of audio anomaly (such as "noise + speech distortion" or "silence"). The audio anomaly type label is empty when there is no anomaly.

[0039] As an example, the intelligent inspection system for video terminal devices has pre-configured preset audio anomaly detection rules. These rules establish a one-to-one mapping relationship between various audio features (decibel value sequence, peak amplitude sequence, human voice frequency band energy proportion, signal-to-noise ratio, character-level text matching degree) and preset audio dimensions (volume stability dimension, signal strength dimension, effective signal proportion dimension, anti-interference capability dimension, semantic accuracy dimension). This allows each audio feature to be converted into an audio dimension score of 0-100. For example, the volume stability dimension score is generated based on the mean of the decibel value sequence (e.g., mean 120-200dB corresponds to 80-100 points, mean <80dB or >250dB corresponds to 0-30 points, and other intervals correspond to 30-80 points); the signal strength dimension score is generated based on the proportion of frames exceeding 0.9V in the peak amplitude sequence; and the effective signal proportion score is generated based on the human voice frequency band energy proportion (e.g., ≥60% corresponds to 80-100 points). 40%–60% corresponds to 50–80 points, <40% corresponds to 0–50 points; audio dimension scores based on signal-to-noise ratio mapping are generated for anti-interference capability (e.g., ≥25dB ​​corresponds to 80–100 points, 15–25dB corresponds to 50–80 points, <15dB corresponds to 0–50 points); audio dimension scores based on character-level text matching degree mapping are generated for semantic accuracy (e.g., ≥70% corresponds to 80–100 points, 50%–70% corresponds to 50–80 points, <50% corresponds to 0–50 points). Subsequently, the threshold values ​​corresponding to each preset audio dimension are extracted (60 points for volume stability, 60 points for signal strength, 50 points for effective signal ratio, 50 points for anti-interference capability, and 60 points for semantic accuracy). The audio dimension scores for each preset audio dimension are compared with their corresponding threshold values ​​to determine whether there are any audio anomalies. For example, if the audio dimension score for volume stability is less than 60 points, a "mute" anomaly is determined; if the audio dimension score for signal strength is less than 60 points, a "clipping" anomaly is determined; if the audio dimension score for effective signal ratio is less than 50 points, a "sound quality anomaly" is determined; if the audio dimension score for anti-interference capability is less than 50 points, a "feedback" or "noise" anomaly is determined; and if the audio dimension score for semantic accuracy is less than 60 points, a "speech distortion" anomaly is determined. The results are then summarized to determine whether at least one audio anomaly exists. Next, the preset first weighting coefficients for each audio dimension score (semantic accuracy dimension 30%, anti-interference ability dimension 25%, effective signal ratio dimension 20%, volume stability dimension 15%, signal strength dimension 10%) are called, and the initial audio quality score is calculated by the weighted summation formula (initial audio quality score = Σ (score of each dimension × corresponding first weighting coefficient)).Next, based on the anomaly detection results, the audio quality score and audio anomaly type label are determined: if at least one audio anomaly exists, the first preset score (preset to 50 points) is directly assigned as the audio quality score, and all anomaly types are integrated into an audio anomaly type label (e.g., "noise + speech distortion"); if no audio anomalies exist, the initial audio quality score is used as the final audio quality score, and the audio anomaly type label is set to empty. Finally, a preset audio pass threshold is called (e.g., 70 points, which is greater than the first preset score of 50 points, ensuring that the score in anomaly scenarios will necessarily be lower than the pass threshold). If the audio quality score is greater than or equal to the audio pass threshold, the detection conclusion is "pass"; otherwise, it is "fail". Finally, the audio anomaly type label, audio quality score, scores of each dimension, and detection conclusion are structurally integrated to form a complete audio detection result. Through quantitative scoring and multi-layer judgment logic, standardized evaluation of audio quality and accurate positioning of anomaly types are achieved, providing reliable audio dimension data support for the overall inspection and evaluation of target video terminal equipment.

[0040] In some embodiments, the step of performing mainstream video anomaly detection on the mainstream video stream to obtain the mainstream video detection result includes: Extract keyframes from mainstream video streams to generate corresponding mainstream keyframe sequences; Based on each mainstream keyframe, static image quality detection is performed on mainstream video streams to generate static image detection results. The mainstream video stream is divided into multiple consecutive mainstream video segments according to a preset duration; Based on various mainstream video clips, dynamic smoothness is applied to the mainstream video stream to generate dynamic smoothness detection results; By combining the static image detection results and the dynamic smoothness detection results, the mainstream video detection results corresponding to the mainstream video streams are determined.

[0041] Specifically, static image quality detection is a specialized test targeting the visual quality characteristics of a single frame (mainstream keyframe), primarily identifying static mainstream video anomalies such as black screen, white screen, distorted screen, image occlusion, abnormal brightness, or blurred image. Dynamic smoothness detection targets the continuity of inter-frame transmission within mainstream video segments, primarily identifying dynamic mainstream video anomalies such as frame skipping and stuttering, reflecting the real-time transmission quality of the video stream.

[0042] As an example, keyframe extraction can be performed on mainstream video streams using inter-frame differencing to generate a sequence of mainstream keyframes covering the core content of the entire duration. Based on each keyframe, static image quality can be detected and static image detection results can be generated by statistically analyzing the black / white screen pixel ratio, calculating the sharpness using the Brenner gradient function, and analyzing color deviation using the CIE76 formula. Then, the mainstream video stream is divided into multiple continuous mainstream video segments according to a preset duration. Dynamic smoothness of each mainstream video segment can be detected and dynamic smoothness detection results can be generated by statistically analyzing frame interval fluctuations, calculating inter-frame SSIM similarity, and comparing actual and theoretical frame counts. Finally, by combining the static image detection results and dynamic smoothness detection results, a weighted fusion or rule-based judgment method is used to determine the mainstream video detection result corresponding to the mainstream video stream. Keyframe extraction simplifies the amount of static detection data, and multiple indicators are combined to accurately identify static anomalies such as black screens and insufficient sharpness; simultaneously, fragmented processing adapts to dynamic detection needs, effectively capturing dynamic anomalies such as frame skipping and stuttering. Finally, by integrating the static image detection results and the dynamic smoothness detection results, the mainstream video detection results are obtained. This approach balances detection efficiency and accuracy while achieving full-scene coverage of mainstream video quality. It solves the problems of traditional detection methods, such as having a single dimension and being prone to missing anomalies, and provides a reliable basis for evaluating the video quality of equipment.

[0043] In some embodiments, the step of performing static image quality detection on the mainstream video stream based on each mainstream keyframe and generating static image detection results includes: Each mainstream keyframe in the mainstream keyframe sequence is input into a pre-trained first vision big model. The first vision anomaly big model determines whether the mainstream keyframe has at least one mainstream video anomaly, such as black screen, white screen, distorted screen, screen occlusion, abnormal brightness, or blurred screen. The anomaly detection result corresponding to the mainstream keyframe is generated, and the number of static anomaly frames corresponding to the mainstream keyframe sequence is counted based on the anomaly detection result. The anomaly detection result is used to indicate whether the mainstream keyframe has at least one mainstream video anomaly and the mainstream video anomaly type of the mainstream keyframe. Based on the preset mainstream video anomaly detection rules and the image anomaly detection results corresponding to each mainstream keyframe, the mainstream video stream is mapped according to the rules to determine the static image quality score corresponding to the mainstream video stream. The static image detection results are defined as the image anomaly detection results, static image quality scores, and the number of static anomaly frames corresponding to each mainstream keyframe.

[0044] Specifically, the first-level vision model is a deep vision model pre-trained on large-scale image data, such as a dedicated anomaly detection model based on the Vision Transformer (ViT) or Shifted Window Transformer (Swin Transformer) architecture, capable of discriminating various image quality anomalies. Mainstream video anomalies are typical anomaly types that may appear in static images, including black screens (completely black), white screens (completely white), distorted images (color blocks / noise caused by encoding errors), image occlusion (lens obstructed by stickers, hands, etc.), brightness anomalies (overexposure or underexposure), and image blur (focus failure or motion blur). The pre-defined mainstream video anomaly detection rules are not pre-configured and are standardized rules that map single-frame detection results (image anomaly detection results corresponding to mainstream keyframes) to an overall static image quality score. These rules may include parameters such as anomaly weights and score ranges. The static image quality score is a quantized score (0-100 points) calculated based on single-frame detection results (image anomaly detection results corresponding to mainstream keyframes) and the number of anomaly frames, which can characterize the overall static image quality of mainstream video streams. The static abnormal frame count represents the total number of keyframes in the mainstream keyframe sequence that are identified as containing at least one mainstream video anomaly, reflecting the proportion of abnormal frames. The image anomaly detection result is the analysis conclusion of the first-person vision model on a single keyframe, presented as a structured output, for example: Frame 1: {Exception: false} (Normal frame) Frame 2:{Exception: true, Mainstream video exception type: ["Screen flickering"]} Frame 3:{Exception: true, Mainstream video exception types: ["Blurred image", "Abnormal brightness"]} As an example, each mainstream keyframe from the mainstream keyframe sequence extracted from the mainstream video stream is sequentially input into a pre-trained first-vision large-scale model. This model has been fine-tuned on a dataset containing a large number of labeled samples and can simultaneously identify multiple typical mainstream video anomalies. For each mainstream keyframe, the first-vision large-scale model outputs a structured image anomaly detection result, explicitly indicating whether the mainstream keyframe contains at least one mainstream video anomaly among black screen, white screen, distorted screen, image occlusion, abnormal brightness, or blurred image (lack of high-frequency information at the edges), and labeling the specific mainstream video anomaly type (i.e., black screen, white screen, distorted screen, image occlusion, abnormal brightness, or blurred image). Subsequently, the image anomaly detection results of all mainstream keyframes are traversed, and the number of frames identified as containing at least one mainstream video anomaly is counted to obtain the number of static anomaly frames. Based on this, the overall static quality of mainstream keyframe sequences is quantitatively evaluated according to preset mainstream video anomaly detection rules: for example, if the number of static anomaly frames accounts for more than 30% of the total number of keyframes, the static image quality score is directly set to 0; if the proportion of anomaly frames is between 10% and 30%, the score is calculated according to the linear decay formula; if there are no anomaly frames, a high score (such as 100 points) is assigned. These preset mainstream video anomaly detection rules can also be finely adjusted by combining anomaly type weights (such as "black screen" deducting more points than "slight blur"). Finally, the anomaly detection results (including anomaly type labels) corresponding to each mainstream keyframe, the static image quality score of each mainstream keyframe, and the statistically obtained number of static anomaly frames are jointly encapsulated to form a complete static image detection result. This static image detection result not only provides an objective quantitative score for the static quality of mainstream videos but also retains frame-level anomaly distribution and type details. It can provide interpretable and traceable diagnostic basis for subsequent judgments on maintenance scenarios such as whether the camera is obstructed, whether the lens is dirty, or whether the device is in a black screen state, significantly improving the automation level and accuracy of video terminal static image quality evaluation.

[0045] In some embodiments, the step of performing dynamic smoothness detection on the mainstream video stream based on each mainstream video segment and generating dynamic smoothness detection results includes: For each mainstream video segment, the SSIM algorithm is used to calculate the structural similarity index between any adjacent mainstream video frames in the mainstream video segment. When the structural similarity index is less than the preset similarity threshold, it is determined that there are frame skips between adjacent mainstream video frames, and the number of frame skips corresponding to the mainstream video segment is counted. Extract the timestamps of all mainstream video frames in the mainstream video segment, and calculate the frame interval between any adjacent mainstream video frames in the mainstream video segment based on the timestamps of each mainstream video segment. When the frame interval is greater than or equal to a preset time threshold, it is determined that there is a stutter between adjacent mainstream video frames, and the number of stutters corresponding to the mainstream video segment is counted. If the number of frame skips in the mainstream video segment is greater than or equal to the first preset number, or the number of stutters in the mainstream video segment is greater than or equal to the second preset number, the mainstream video segment is identified as an abnormal video segment, and the abnormality of the mainstream video is determined to be at least one of frame skips and stutters. Calculate the total number of frame skips and stutters for the mainstream keyframe sequence by using the number of frame skips and stutters for each mainstream video segment; Based on the total number of frame skips, the total number of stutters, and the preset mainstream video anomaly detection rules, rule mapping is performed on the mainstream video streams to determine the dynamic smoothness score corresponding to the mainstream video streams. The dynamic smoothness score, total number of frame skips, and total number of stutters are defined as the dynamic smoothness detection results.

[0046] Specifically, the Structural Similarity Index (SSIM) algorithm is a classic algorithm used to measure the structural similarity of two images. It calculates the similarity of three dimensions—brightness, contrast, and structure—and sums them by weight, outputting an index between 0 and 1 (the closer to 1, the more similar the structure). It is suitable for inter-frame consistency detection. A preset similarity threshold (e.g., 0.7) is not a critical value for determining whether adjacent frames are skipped. When the structural similarity index is less than this preset similarity threshold, it is determined that the content of adjacent frames has abruptly changed, indicating a skipped frame. The frame interval is the difference in timestamps between two adjacent mainstream video frames, reflecting the temporal continuity of frame transmission (normally it should be close to the reciprocal of the nominal frame rate, e.g., 40ms for 25fps). A preset time threshold (e.g., 100ms) is not a critical value for determining whether adjacent frames are stuttering. When the frame interval is greater than or equal to this threshold, the image dwell time is considered too long, indicating stuttering between adjacent frames. The first / second preset number of attempts (e.g., 3 for the first preset number of attempts and 2 for the second preset number of attempts) are not critical numbers for determining whether mainstream video segments are abnormal, corresponding to the abnormal judgment criteria for skipped frames and stuttering, respectively. A mainstream video segment with an abnormal number of non-frame skips greater than or equal to the first preset number or a number of stutters greater than or equal to the second preset number indicates that the dynamic smoothness of the segment does not meet the standard. The dynamic smoothness score is a quantitative score (0-100 points) calculated based on the total number of frame skips, the total number of stutters, and preset rules, and represents the dynamic transmission continuity of the mainstream video stream.

[0047] As an example, for each mainstream video segment already divided according to a preset duration, the Structural Similarity Index (SSIM) algorithm is performed on any adjacent mainstream video frames within the segment: Based on the normalized image data of two adjacent frames, the SSIM algorithm extracts the mean brightness, contrast variance, and structural covariance features of the two frames respectively, and calculates the structural similarity index in the 0-1 range using a standard formula. A preset similarity threshold (e.g., 0.7) is then applied. If the calculated structural similarity index between adjacent mainstream video frames is less than the preset similarity threshold, it is determined that there is a frame skip (abrupt change in image content) in this group of adjacent frames, and the number of frame skips corresponding to the current mainstream video segment is accumulated. Subsequently, the timestamps of all mainstream video frames in the current mainstream video segment are extracted. The difference in timestamps between any two adjacent mainstream video frames, i.e., the frame interval, is calculated. A preset time threshold (e.g., 100ms, based on a standard frame interval of 40ms corresponding to a nominal video frame rate of 25fps, which is 2.5 times the standard value) is applied. If the frame interval of a group of adjacent frames is greater than or equal to the preset time threshold, it is determined that there is stuttering (image freeze) in that group of adjacent frames, and the number of stutters corresponding to the current mainstream video segment is accumulated. Next, the first preset number of times (e.g., 3 times) and the second preset number of times (e.g., 2 times) are applied to perform anomaly judgment on each mainstream video segment: if the number of frame skips in the mainstream video segment is greater than or equal to the first preset number of times, or the number of stutters is greater than or equal to the second preset number of times, the mainstream video segment is marked as an abnormal video segment, and the corresponding mainstream video anomaly type (frame skips / stuttering / both) is recorded simultaneously; if the number of frame skips in the mainstream video segment is less than the first preset number of times and the number of stutters is less than the second preset number of times, the mainstream video segment is determined to be a normal video segment. Next, all mainstream video segments are traversed, and the number of frame skips in each mainstream video segment is summed to obtain the total number of frame skips in the mainstream video stream. The number of stutters in each mainstream video segment is also summed to obtain the total number of stutters, reflecting the overall accumulation of dynamic anomalies in the video stream. Then, a pre-defined mainstream video anomaly detection rule is invoked to perform a scoring mapping. For example, if the total number of frame skips > 5 or the total number of stutters > 8, the dynamic smoothness score is 40; if both the total number of frame skips and the total number of stutters are ≤ 2, the dynamic smoothness score is 90. Intermediate cases can be calculated using a piecewise linear function or a weighted attenuation model, and a rule mapping is performed on the mainstream video stream to generate a dynamic smoothness score out of 0–100. Finally, the calculated dynamic smoothness scores (0-100 points), the total number of frame skips, and the total number of stutters are structurally integrated to form a complete dynamic smoothness detection result. By using the SSIM algorithm and timestamp analysis for dual-dimensional detection, it accurately covers two types of dynamic mainstream video anomalies: frame skipping and stuttering. Combined with rule-based scoring, it achieves a quantitative assessment of smoothness, providing reliable dynamic dimension data support for subsequent comprehensive quality judgment of mainstream video streams and accurately quantifying the dynamic smoothness of videos.

[0048] Optionally, the steps to determine the mainstream video detection result corresponding to the mainstream video stream by combining the static image detection results and the dynamic smoothness detection results include: The static image quality score and the dynamic smoothness score are weighted and summed to obtain the initial mainstream video quality score; If the number of static abnormal frames is greater than or equal to the first preset number, or if there is at least one abnormal video segment, the second preset score will be determined as the mainstream video quality score, and the mainstream video abnormality will be used as the mainstream video abnormality type label. If the number of static abnormal frames is less than the first preset number and there are no abnormal video segments, the initial mainstream video quality score will be determined as the mainstream video quality score. If the mainstream video quality score is less than the mainstream video pass threshold, the detection conclusion of the mainstream video detection result is determined to be unqualified; if the mainstream video quality score is greater than or equal to the audio pass threshold, the detection conclusion of the mainstream video detection result is determined to be qualified. If the mainstream video pass threshold is greater than the second preset score, the mainstream video abnormal type label is set to empty. The mainstream video detection results are defined as the following: mainstream video anomaly type labels, mainstream video quality scores, static image quality scores, image anomaly detection results for each mainstream keyframe, number of static anomaly frames, dynamic smoothness scores, total number of frame skips, total number of stutters, and the detection conclusions of the mainstream video detection results.

[0049] Specifically, the mainstream video anomaly type tag does not record the specific type of anomaly in the mainstream video stream. The mainstream video anomaly type tag can include black screen, white screen, distorted screen, screen obstruction, abnormal brightness, blurry screen, frame skipping, and stuttering.

[0050] As an example, the static image quality score from the static image detection results and the dynamic smoothness score from the dynamic smoothness detection results are retrieved. The static image quality score (reflecting image quality issues such as black screens, blurriness, and occlusion) and the dynamic smoothness score (reflecting timing anomalies such as frame skipping and stuttering) are weighted and summed (e.g., static weight 0.6, dynamic weight 0.4) to obtain the initial mainstream video quality score. Then, the number of static abnormal frames in the static image detection results is extracted and compared with a first preset number (e.g., 5 frames); simultaneously, the presence of at least one abnormal video segment in the dynamic smoothness detection results is checked. If either the "number of static abnormal frames is greater than or equal to the first preset number" or "at least one abnormal video segment exists," the mainstream video stream is determined to have serious quality defects. The second preset score (e.g., 40 points) is directly determined as the mainstream video quality score, and mainstream video anomalies (black screen, white screen, distorted screen, image occlusion, abnormal brightness, blurriness, frame skipping, and stuttering) are integrated as mainstream video anomaly type labels. If the number of static abnormal frames is less than the first preset number and there are no abnormal video segments, it is determined that there is no serious abnormality, and the initial mainstream video quality score is directly used as the final mainstream video quality score. Then, the preset mainstream video pass threshold (e.g., 60 points, which is greater than the second preset score of 40 points to ensure that the score will be lower than the pass threshold in serious abnormality scenarios) is called to perform a pass / fail judgment. If the mainstream video quality score is greater than or equal to the audio pass threshold, the detection conclusion is determined to be "pass," and the mainstream video abnormality type label is set to empty; if the mainstream video quality score is less than the mainstream video pass threshold, the detection conclusion is determined to be "unpass." Finally, the mainstream video abnormality type label, mainstream video quality score, and full details of static dimensions (static image quality score, image abnormality detection results of each mainstream keyframe, number of static abnormal frames), full details of dynamic dimensions (dynamic smoothness score, total number of frame skips, total number of stutters), and the final detection conclusion are structurally integrated to form a complete mainstream video detection result. By integrating static and dynamic data, identifying serious anomalies, and verifying compliance, we can achieve a comprehensive assessment of video quality, retain complete detection details, and provide reliable evidence for maintenance personnel to accurately locate the root cause of anomalies.

[0051] In some embodiments, the step of performing auxiliary stream video anomaly detection on the auxiliary stream video stream and obtaining the auxiliary stream video detection result includes: Keyframes are extracted from the auxiliary video stream to generate the corresponding auxiliary keyframe sequence. Each auxiliary stream keyframe and the preset standard reference frame sequence in the auxiliary stream keyframe sequence are input into the pre-trained second vision large model. The auxiliary stream keyframe sequence and the preset standard reference frame sequence are compared by the second vision abnormal large model to generate the auxiliary stream keyframe score in each preset dimension of the auxiliary stream video. The auxiliary stream keyframes are compared with the dimensional thresholds of each preset dimension of the auxiliary stream video to determine whether each auxiliary stream keyframe has at least one auxiliary stream video abnormality, such as black screen, white screen, distorted screen, abnormal color, abnormal text presentation, or missing content. The number of defective frames corresponding to the auxiliary stream keyframe sequence is also counted. For each auxiliary stream keyframe, the second weighting coefficient corresponding to the video dimension score of each auxiliary stream is obtained. The second weighting coefficient and the video score of each auxiliary stream are weighted and summed to obtain the initial video frame quality score of the auxiliary stream keyframe. Based on whether there are auxiliary stream video anomalies in the auxiliary stream keyframes and the initial video frame quality score of the auxiliary stream keyframes, determine the video frame quality score of each auxiliary stream keyframe and the auxiliary stream video quality score of the auxiliary stream video stream, and determine the detection conclusion of the auxiliary stream video detection result based on the auxiliary stream video quality score. The auxiliary stream video detection result is determined by the video frame quality score of each auxiliary stream keyframe, the auxiliary stream video anomaly type label, the auxiliary stream video quality score, and the detection conclusion of the auxiliary stream video detection result.

[0052] Optionally, the steps of determining the video frame quality score of each auxiliary stream keyframe and the auxiliary stream video quality score of the auxiliary stream video stream based on whether there are auxiliary stream video anomalies in the auxiliary stream keyframes and the initial video frame quality score of the auxiliary stream keyframes, and determining the detection conclusion of the auxiliary stream video detection result based on the auxiliary stream video quality score include: If at least one auxiliary stream video anomaly exists in the auxiliary stream keyframe, the video frame quality score of the auxiliary stream keyframe will be set to 0. If there are no auxiliary stream video anomalies in the auxiliary stream keyframes, the initial video frame quality score will be determined as the video frame quality score. The auxiliary stream video quality score is obtained by averaging the video frame quality scores of each auxiliary stream keyframe. If the auxiliary stream video quality score is less than the auxiliary stream video pass threshold, or the number of defective frames is greater than or equal to the second preset number, the detection conclusion of the auxiliary stream video detection result is determined to be unqualified, and the auxiliary stream video anomalies of each auxiliary stream key frame are used as auxiliary stream video anomaly type labels. If the auxiliary stream video quality score is greater than or equal to the auxiliary stream video pass threshold, and the number of defective frames is less than the second preset number, the detection conclusion of the auxiliary stream video detection result is determined to be qualified.

[0053] Specifically, the preset standard reference frame sequence is a pre-configured set of standard frame images (such as standard PPT presentation frames and standard image display frames) that match the content of the auxiliary stream video stream, serving as the benchmark for quality comparison of auxiliary stream keyframes. The second vision big model is a computer vision model pre-trained based on deep learning (such as the Transformer architecture), which can be optimized for quality detection tasks of auxiliary stream images (documents, images, etc.) and has the ability to perform inter-frame comparison and multi-dimensional quality scoring. The preset dimensions of the auxiliary stream video are the core dimensions for evaluating the quality of auxiliary stream images. The preset dimensions of the auxiliary stream video can be image validity, encoding / transmission integrity, color accuracy, text readability, content integrity, etc., adapting to the quality requirements of auxiliary stream document / image display. The auxiliary stream video dimension score is a quantitative score (0-100 points) output by the second vision big model after comparing the auxiliary stream keyframes with the preset standard reference frames, representing the quality level of the corresponding auxiliary stream video preset dimension. Auxiliary stream video anomalies refer to quality defects in auxiliary stream keyframes. These can include black screens, white screens, distorted screens, color anomalies (significant color deviation from the standard frame), text rendering anomalies (blurred text, missing fonts, misplaced characters, font size too small, or OCR recognition failure, resulting in inaccurate information delivery), and content omissions (partial or complete absence of content, such as missing pages in a PPT, blank charts, or cropped key areas, indicating inconsistency or missing content compared to the standard reference frame). Auxiliary stream video anomalies correspond to preset dimensions. For example, black screens and white screens represent auxiliary stream video anomalies in the image validity dimension, distorted screens in the encoding / transmission integrity dimension, color anomalies in the color accuracy dimension, text rendering anomalies in the text readability dimension, and content omissions in the content integrity dimension. The number of defective frames is the total number of keyframes in the auxiliary stream keyframe sequence that are determined to have at least one auxiliary stream video anomaly, reflecting the cumulative degree of anomaly frames in the auxiliary stream video stream. The second weighting coefficient is the weight value assigned to the preset dimensions of the auxiliary stream video score, reflecting the degree of influence of different preset dimensions of the auxiliary stream video on the quality of the auxiliary stream keyframes. The video frame quality score is the final quality score of a single auxiliary stream keyframe. A score of 0 is set when there are auxiliary stream video anomalies in the keyframe; otherwise, the initial video frame quality score is used. The second preset quantity is the critical number of defective frames (e.g., 4 frames) required to determine if the auxiliary stream video stream has serious anomalies. The auxiliary stream video pass threshold is the critical value (e.g., 70 points) for the auxiliary stream video stream to pass quality standards, used to determine whether the overall quality of the auxiliary stream video stream meets the requirements. The auxiliary stream video anomaly type label is an identifier that records all anomaly types in the auxiliary stream keyframes, integrating types such as black screens and text rendering anomalies present in each frame; it can be left blank when there are no anomalies.

[0054] As an example, the inter-frame difference method can be used to extract keyframes from the separated auxiliary stream video stream, resulting in a time-ordered sequence of auxiliary stream keyframes. Subsequently, each auxiliary stream keyframe in the sequence is normalized to a preset resolution (e.g., 1920×1080) and simultaneously input into a pre-trained second-vision big model along with a pre-configured preset standard reference frame sequence (precisely matching the auxiliary stream content, such as standard PPT frames or standard image display frames corresponding to a presentation document). This second-vision big model is trained and optimized using massive amounts of anomalous auxiliary stream samples (including black screens, white screens, distorted screens, color anomalies, text rendering anomalies, or missing content frames). It can automatically extract core features such as color distribution, text outlines, content completeness, and structural consistency between the auxiliary stream keyframes and their corresponding standard reference frames, perform inter-frame comparisons, and output an auxiliary stream video dimension score for each keyframe in each preset dimension of the auxiliary stream video (image validity dimension, encoding / transmission integrity dimension, color accuracy dimension, text readability dimension, and content completeness dimension). Next, the dimensional thresholds corresponding to the preset dimensions of each auxiliary stream video are extracted (for example, the threshold for the picture validity dimension is 60 points, the threshold for the encoding / transmission integrity dimension is 60 points, the threshold for the color accuracy dimension is 70 points, the threshold for the text readability dimension is 65 points, and the threshold for the content integrity dimension is 60 points). The score of each auxiliary stream video dimension of each auxiliary stream keyframe is compared with the corresponding dimensional threshold one by one, and a single-frame anomaly judgment is performed: if the auxiliary stream video dimension score for the picture validity dimension is <60 points, it is determined that there is a black screen in the auxiliary stream keyframe. The auxiliary stream video is considered abnormal if it displays a blank screen. If the auxiliary stream video score for encoding / transmission integrity is less than 60, the auxiliary stream keyframe is considered to have a screen-distorted display. If the auxiliary stream video score for color accuracy is less than 70, the auxiliary stream keyframe is considered to have a color error. If the auxiliary stream video score for text readability is less than 65, the auxiliary stream keyframe is considered to have a text rendering error. If the auxiliary stream video score for content integrity is less than 60, the auxiliary stream keyframe is considered to have a content missing error. The results of determining whether each auxiliary stream keyframe has at least one type of auxiliary stream video abnormality are summarized, and the total number of frames with abnormalities (i.e., the number of defective frames) is counted by iterating through all auxiliary stream keyframes. For each auxiliary stream keyframe, the preset second weighting coefficients corresponding to the scores of each auxiliary stream video dimension are called. The initial video frame quality score of the auxiliary stream keyframe is calculated using the weighted summation formula: Initial Video Frame Quality Score = Σ (Scores of each auxiliary stream video dimension × corresponding second weighting coefficients). If the auxiliary stream keyframe has at least one auxiliary stream video anomaly, the video frame quality score of the auxiliary stream keyframe is directly set to 0 (highlighting the fatal impact of the anomaly on the quality of a single frame). If the auxiliary stream keyframe does not have any auxiliary stream video anomalies, the initial video frame quality score of the auxiliary stream keyframe is directly determined as the final video frame quality score of the auxiliary stream keyframe.Next, the arithmetic mean of the video frame quality scores of all auxiliary stream keyframes is calculated to obtain the overall auxiliary stream video quality score. Then, a preset auxiliary stream video quality threshold (70 points) and a second preset number (e.g., 4 frames, set based on the proportion of the total number of keyframes to ensure accurate identification when abnormal frames accumulate to a certain level) are used to determine the overall detection conclusion: If the auxiliary stream video quality score is less than the auxiliary stream video quality threshold, or the number of defective frames is greater than or equal to the second preset number, the auxiliary stream video quality is deemed substandard, and the detection conclusion is "unqualified." Auxiliary stream video anomaly types from all auxiliary stream keyframes are integrated to generate auxiliary stream video anomaly type tags (e.g., "text rendering anomaly + color anomaly + content missing"). If the auxiliary stream video quality score is greater than or equal to the auxiliary stream video quality threshold, and the number of defective frames is less than the second preset number, the auxiliary stream video quality is deemed qualified, and the detection conclusion is "qualified." Finally, the video frame quality scores of each auxiliary stream keyframe, the generated auxiliary stream video anomaly type tags, the calculated auxiliary stream video quality score, and the final detection conclusion are structurally integrated to form a complete auxiliary stream video detection result. By leveraging the precise feature comparison capabilities of the large visual model, comprehensive identification of anomalies unique to auxiliary flow scenarios can be achieved. Combined with hierarchical scoring and judgment logic, this ensures both the accuracy and standardization of detection, and provides reliable auxiliary flow dimension data support for subsequent overall equipment inspection and evaluation.

[0055] In some embodiments, the step of obtaining the network quality detection result of the target video terminal device includes: Obtain the network quality parameters of the target video terminal device, including network transmission latency, packet loss rate, transmission latency jitter, and available bandwidth; If any network quality parameter exceeds the corresponding threshold range, the network quality test result is determined to be unqualified. If all network quality parameters are determined to be within the corresponding parameter threshold range, the network quality test result is deemed qualified. The conclusions of all network quality parameters and network quality test results are defined as the network quality test results.

[0056] Specifically, the parameter threshold range is a preset "qualified" value range for each network quality parameter. For example: network transmission latency <100ms, packet loss rate <1%, transmission latency jitter <30ms, and available bandwidth >2Mbps.

[0057] As an example, network quality parameters of the target video terminal device can be obtained through active probing or passive monitoring. These parameters include network transmission delay, packet loss rate, transmission delay jitter, and available bandwidth. Network transmission delay can be obtained by sending an Internet Control Message Protocol Echo (ICMP Echo) request to the terminal or by measuring the round-trip time (RTT) using a Session Initiation Protocol OPTIONS Request (SIP OPTIONS) message. Packet loss rate and transmission delay jitter can be calculated by parsing the sequence number and timestamp from the Real-time Transport Protocol (RTP) / Real-time Transport Control Protocol (RTCP) messages returned by the terminal. Available bandwidth can be estimated using dedicated bandwidth probing tools (such as iperf3) or based on receiver rate adaptive feedback (such as Google Congestion Control). Subsequently, the system compares each network quality parameter with a preset threshold range. If any network quality parameter exceeds its threshold range (e.g., RTT = 200 ms > 100 ms, or packet loss rate = 2% > 1%), the network quality detection result is immediately determined to be "unqualified". Only when all network quality parameters are within their respective threshold ranges is the network quality detection result determined to be "qualified". Finally, all acquired network quality parameters (including specific values) and the detection conclusions are packaged together into a structured network quality detection result. This result can not only be used to independently assess network health but also serve as a key input for subsequent multimodal fusion diagnostics.

[0058] In some embodiments, the step of generating an inspection report for the target video terminal device based on network quality detection results, audio detection results, mainstream video detection results, and auxiliary stream video detection results includes: When any one of the network quality detection results, audio detection results, mainstream video detection results, and auxiliary stream video detection results is determined to be unqualified, the network quality detection results, audio detection results, mainstream video detection results, and auxiliary stream video detection results are input into the pre-trained inspection anomaly root cause analysis model. Based on the inspection anomaly root cause analysis model, the root cause analysis of the anomaly source in the current inspection process is performed to obtain the anomaly root cause analysis results. The inspection anomaly root cause analysis model is constructed based on the Bidirectional Encoder Representations from Transformers (BERT) model based on the converter architecture. Based on network quality detection results, audio detection results, mainstream video detection results, auxiliary stream video detection results, and anomaly root cause analysis results, the preset inspection and detection template is filled in to generate an inspection and detection report.

[0059] Specifically, the inspection report is a structured document integrating comprehensive inspection data from equipment networks, audio, mainstream video, and auxiliary video streams, along with root cause analysis results. It provides a clear overview of the equipment's inspection quality status and detailed problem information. The inspection anomaly root cause analysis model is a deep learning model built on the BERT model. Trained and optimized with massive amounts of inspection anomaly samples (including anomaly data from each dimension and corresponding root causes), it possesses the ability to mine anomaly sources from multi-dimensional inspection data. Anomaly sources are the core causes leading to unqualified inspection conclusions, including equipment hardware failures (such as camera damage or microphone failure), network failures (such as insufficient bandwidth or excessive transmission latency), and software configuration anomalies (such as incorrect encoding parameters). The anomaly root cause analysis results are structured outputs from the inspection anomaly root cause analysis model, including the core root cause of the anomaly, root cause confidence (e.g., 95%), associated anomaly dimensions, and preliminary solution suggestions, providing precise guidance for problem troubleshooting. The preset inspection and testing template is a pre-configured standardized report template that can include fixed modules such as basic equipment information, test data columns for various dimensions, anomaly details, root cause analysis, and test conclusion summary, ensuring a consistent report format.

[0060] As an example, the system first retrieves the generated network quality test results, audio test results, mainstream video test results, and auxiliary stream video test results. It then verifies the "test conclusion" field in each of the four test results to determine if any of the test results has a "non-compliant" conclusion.

[0061] If the verification finds that at least one test result is "unqualified", the anomaly root cause analysis process is initiated. Specifically, the complete structured data of the four test results are uniformly converted into a text format that can be recognized by the inspection anomaly root cause analysis model, and input into the pre-trained inspection anomaly root cause analysis model. This inspection anomaly root cause analysis model is built on the architecture of a Bidirectional Encoder Representations from Transformers (BERT) model, which has been trained and optimized with massive historical inspection data (covering detection data and corresponding root cause labels for various scenarios such as equipment hardware failure, network failure, and software configuration anomalies). It has the ability to mine multi-dimensional data semantic associations. By analyzing the correlation of anomaly dimensions (such as "abnormal audio noise + high network packet loss rate" can be associated with the root cause of "insufficient network bandwidth"), quantifying the degree of parameter deviation, etc., the anomaly root cause analysis results are output. The content includes the core root cause of the anomaly (such as "insufficient available network bandwidth causes mainstream video to stutter and audio to distort"), the confidence of the root cause, the list of associated anomaly dimensions, and preliminary solution suggestions (such as "it is recommended to expand the network bandwidth to more than 10Mbps"). If the verification finds that all four test results are "qualified", then there is no need to perform anomaly root cause analysis, and the anomaly root cause analysis result field is marked as "no anomaly".

[0062] Subsequently, the pre-configured inspection template filling process is initiated. Specifically, a pre-configured standardized pre-configured inspection template is retrieved. This template may include fixed modules such as basic information fields (equipment model, inspection time, etc.), quantitative parameter fields for various dimensions of network, audio, mainstream video, and auxiliary video, detailed scoring fields, inspection conclusion fields, anomaly detail summary fields (filled with "None" if no anomalies are found), root cause analysis fields (filled with "None" if no anomalies are found), and a final inspection conclusion summary field. The core data from the four inspection results are filled in one by one according to the template field correspondence. If there are anomaly root cause analysis results, the root cause content, confidence level, and solution suggestions are filled into the "Root Cause Analysis" field. If there are no non-compliance conclusions, "No anomalies" is filled into both the anomaly detail field and the root cause analysis field. Finally, the completeness of the filled content of each field in the template is verified to ensure that there is no data omission or filling error. After the verification is passed, the final inspection report is generated. The report has a unified format and intuitive data, which can be directly used by maintenance personnel to view the equipment quality status and troubleshoot anomalies. By integrating network, audio, and main / auxiliary video streaming data across all dimensions, a comprehensive view of equipment inspection status is achieved, avoiding the problem of missing data from a single dimension. For any non-compliant scenario, a root cause analysis model based on the BERT architecture is used to uncover anomaly correlations, accurately locating the root cause of the fault and solving the problem that traditional inspections can only detect anomalies but cannot trace their root causes. Standardized reports are generated by filling in preset templates, ensuring a consistent output format, improving the efficiency of troubleshooting for maintenance personnel, and providing reliable data support for equipment quality assessment and decision-making.

[0063] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the intelligent inspection method of the video terminal equipment of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0064] This application also provides an intelligent inspection device for video terminal equipment; please refer to [reference needed]. Figure 2 The intelligent inspection device for video terminal equipment includes: The video acquisition module 201 is used to initiate an inspection call command to the target video terminal device. The target video terminal device responds to the inspection call command by acquiring detection data including audio track, main video track and auxiliary video track, and sends the detection data to the intelligent inspection system of the video terminal device. The main video track is the video image captured by the camera of the target video terminal device, and the auxiliary video track is the screen-shared video image of the target video terminal device. The multi-track data separation module 202 is used to extract audio stream, main video stream and auxiliary video stream based on detection data; The audio detection module 203 is used to perform audio anomaly detection on the audio stream and obtain the audio detection result; The video detection module 204 is used to perform mainstream video anomaly detection on the mainstream video stream and obtain the mainstream video detection result, and to perform auxiliary video anomaly detection on the auxiliary video stream and obtain the auxiliary video detection result. The comprehensive evaluation module 205 is used to obtain the network quality detection results of the target video terminal device, and generate an inspection report of the target video terminal device based on the network quality detection results, audio detection results, mainstream video detection results and auxiliary stream video detection results.

[0065] The intelligent inspection device for video terminal equipment provided in this application, employing the intelligent inspection method for video terminal equipment described in the above embodiments, can solve the technical problem that existing inspection methods cannot automatically and comprehensively evaluate the audio and video quality of video terminal equipment during operation. Compared with the prior art, the beneficial effects of the intelligent inspection device for video terminal equipment provided in this application are the same as those of the intelligent inspection method for video terminal equipment provided in the above embodiments, and other technical features in the intelligent inspection device for video terminal equipment are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0066] This application provides an intelligent inspection device for a video terminal device. The intelligent inspection device for a video terminal device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the intelligent inspection method for the video terminal device in the first embodiment described above.

[0067] The following is for reference. Figure 3 The diagram illustrates a structural schematic of an intelligent inspection device suitable for implementing the video terminal device of the embodiments of this application. The intelligent inspection device for the video terminal device in the embodiments of this application may include, but is not limited to, mobile terminals such as laptops, tablets (PADs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 3 The intelligent inspection device for video terminal equipment shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0068] like Figure 3As shown, the intelligent inspection device of the video terminal equipment may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the intelligent inspection device of the video terminal equipment. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the intelligent inspection device of the video terminal equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows an intelligent inspection device of a video terminal equipment with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented or possessed alternatively.

[0069] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0070] The intelligent inspection device for video terminal equipment provided in this application, employing the intelligent inspection method for video terminal equipment described in the above embodiments, can solve the technical problem that existing inspection methods cannot automatically and comprehensively evaluate the audio and video quality of video terminal equipment during operation. Compared with the prior art, the beneficial effects of the intelligent inspection device for video terminal equipment provided in this application are the same as those of the intelligent inspection method for video terminal equipment provided in the above embodiments, and other technical features of this intelligent inspection device for video terminal equipment are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0071] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0072] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0073] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the intelligent inspection method of the video terminal device in the above embodiments.

[0074] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0075] The aforementioned computer-readable storage medium may be included in the intelligent inspection device of the video terminal equipment; or it may exist independently and not be installed in the intelligent inspection device of the video terminal equipment.

[0076] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the intelligent inspection device of the video terminal device, the intelligent inspection device of the video terminal device causes the following: It initiates an inspection call command to the target video terminal device, wherein the target video terminal device, in response to the inspection call command, collects detection data including audio track, main video track, and auxiliary video track, and sends the detection data to the intelligent inspection system of the video terminal device. The main video track is the video image captured by the camera of the target video terminal device, and the auxiliary video track is the screen-shared video image of the target video terminal device; it extracts the audio stream, main video stream, and auxiliary video stream based on the detection data; it performs audio anomaly detection on the audio stream to obtain audio detection results; it performs main video anomaly detection on the main video stream to obtain main video detection results, and it performs auxiliary video anomaly detection on the auxiliary video stream to obtain auxiliary video detection results; it obtains the network quality detection results of the target video terminal device, and generates an inspection report for the target video terminal device based on the network quality detection results, audio detection results, main video detection results, and auxiliary video detection results.

[0077] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0078] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0079] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0080] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the intelligent inspection method for the video terminal device described above. This solves the technical problem that existing inspection methods cannot automatically and comprehensively evaluate the audio and video quality of video terminal devices during operation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the intelligent inspection method for video terminal devices provided in the above embodiments, and will not be elaborated upon here.

[0081] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the intelligent inspection method for video terminal devices as described above.

[0082] The computer program product provided in this application can solve the technical problem that existing inspection methods cannot automatically and comprehensively evaluate the audio and video quality of video terminal equipment during operation. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the intelligent inspection method for video terminal equipment provided in the above embodiments, and will not be repeated here.

[0083] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for intelligent inspection of a video terminal device, characterized in that, The application discloses an intelligent inspection system applied to a video terminal device, and an intelligent inspection method of the video terminal device. A target video terminal device is initiated with an inspection call instruction, wherein the target video terminal device collects detection data containing an audio track, a main stream video track and an auxiliary stream video track in response to the inspection call instruction, and sends the detection data to the intelligent inspection system of the video terminal device, the main stream video track is a video picture collected by a camera of the target video terminal device, and the auxiliary stream video track is a screen sharing video picture of the target video terminal device; An audio stream, a main stream video stream and an auxiliary stream video stream are extracted based on the detection data; Audio anomaly detection is performed on the audio stream to obtain an audio detection result; Main stream video anomaly detection is performed on the main stream video stream to obtain a main stream video detection result, and auxiliary stream video anomaly detection is performed on the auxiliary stream video stream to obtain an auxiliary stream video detection result; A network quality detection result of the target video terminal device is obtained, and an inspection detection report of the target video terminal device is generated according to the network quality detection result, the audio detection result, the main stream video detection result and the auxiliary stream video detection result.

2. The intelligent inspection method of video terminal equipment according to claim 1, wherein, The step of performing audio anomaly detection on the audio stream to obtain an audio detection result comprises: The peak amplitude sequence, the decibel value sequence and the audio energy value sequence of the audio stream in the time domain are calculated; Short-time Fourier transform is performed on the audio stream to obtain an audio energy spectrum corresponding to the audio stream, and the energy proportion of a human voice frequency band is determined based on the audio energy spectrum; The signal-to-noise ratio corresponding to the audio stream is calculated based on the audio energy spectrum and the audio energy value sequence; The audio stream is converted into corresponding audio recognition text, and the character-level text matching degree between the audio recognition text and a preset standard text in an inspection task is calculated; The audio detection result corresponding to the audio stream is determined according to the decibel value sequence, the peak amplitude sequence, the energy proportion of the human voice frequency band, the signal-to-noise ratio and the character-level text matching degree.

3. The intelligent inspection method of video terminal equipment according to claim 2, wherein, The step of determining the audio detection result corresponding to the audio stream according to the decibel value sequence, the peak amplitude sequence, the energy proportion of the human voice frequency band, the signal-to-noise ratio and the character-level text matching degree comprises: The decibel value sequence, the peak amplitude sequence, the energy proportion of the human voice frequency band, the signal-to-noise ratio and the character-level text matching degree are subjected to rule mapping according to a preset audio anomaly detection rule to generate audio dimension scores of the audio stream in each audio preset dimension; The audio dimension scores of the audio stream in each audio preset dimension are compared with dimension threshold values of each audio preset dimension respectively to determine whether the audio stream has at least one audio anomaly in silence, clipping, howling, audio quality anomaly, noise or speech distortion; First weighting coefficients corresponding to the audio dimension scores are obtained, and each first weighting coefficient and each audio dimension score are subjected to weighted summation to obtain an audio quality initial score; If there is at least one of the audio abnormalities, determining a first preset score as the audio quality score, and taking the audio abnormality as an audio abnormality type label; If there is no any of the audio abnormalities, determining the audio quality initial score as the audio quality score, and setting the audio abnormality type label as empty; If the audio quality score is less than an audio qualified threshold, determining a detection conclusion of the audio detection result as unqualified, and if the audio quality score is greater than or equal to the audio qualified threshold, determining the detection conclusion of the audio detection result as qualified, the audio qualified threshold being greater than the first preset score; determining the audio abnormality type label, the audio quality score, the audio dimension score of each audio preset dimension, and the detection conclusion of the audio detection result as the audio detection result.

4. The intelligent inspection method of video terminal equipment according to claim 1, wherein, The step of performing main stream video abnormality detection on the main stream video stream to obtain a main stream video detection result comprises: performing key frame extraction on the main stream video stream to generate a corresponding main stream key frame sequence; performing static picture quality detection on the main stream video stream based on each main stream key frame to generate a static picture detection result; dividing the main stream video stream into a plurality of continuous main stream video segments according to a preset time length; performing dynamic fluency detection on the main stream video stream based on each main stream video segment to generate a dynamic fluency detection result; comprehensively determining the main stream video detection result corresponding to the main stream video stream based on the static picture detection result and the dynamic fluency detection result.

5. The intelligent inspection method of video terminal equipment according to claim 4, wherein, The step of performing static picture quality detection on the main stream video stream based on each main stream key frame to generate a static picture detection result comprises: inputting each main stream key frame in the main stream key frame sequence into a pre-trained first visual large model, judging whether the main stream key frame has at least one of a main stream video abnormality including black screen, white screen, screen, picture occlusion, brightness abnormality or picture blur through the first visual abnormality large model, generating a picture abnormality detection result corresponding to the main stream key frame, and according to the picture abnormality detection result, the static abnormal frame quantity corresponding to the main stream key frame sequence is counted, the picture abnormality detection result is used to indicate whether the main stream key frame has at least one of the main stream video abnormality and the main stream video abnormality type of the main stream key frame; performing rule mapping on the main stream video stream based on a preset main stream video abnormality detection rule and the picture abnormality detection result corresponding to each main stream key frame to determine a static picture quality score corresponding to the main stream video stream; determining the picture abnormality detection result corresponding to each main stream key frame, the static picture quality score and the static abnormal frame quantity as the static picture detection result.

6. The intelligent inspection method of video terminal equipment according to claim 4, wherein, The step of performing dynamic fluency detection on the main stream video stream based on each main stream video segment to generate a dynamic fluency detection result comprises: For each of the main stream video segments, a structural similarity index SSIM algorithm is used to calculate the structural similarity index between any adjacent main stream video frames in the main stream video segment, and when the structural similarity index is less than a preset similarity threshold, it is determined that there is a frame skipping between adjacent main stream video frames, and the number of frame skipping corresponding to the main stream video segment is counted. The timestamps of all the main stream video frames in the main stream video segment are extracted, and the frame interval between any adjacent main stream video frames in the main stream video segment is calculated based on the timestamps of each of the main stream video segments, and when the frame interval is greater than or equal to a preset time threshold, it is determined that there is a stall between adjacent main stream video frames, and the number of stalls corresponding to the main stream video segment is counted. If the number of frame skipping of the main stream video segment is greater than or equal to a first preset number, or the number of stalls of the main stream video segment is greater than or equal to a second preset number, the main stream video segment is determined as an abnormal video segment, and it is determined that the main stream video anomaly is at least one of frame skipping and stall. The total number of frame skipping and the total number of stalls corresponding to the main stream key frame sequence are calculated through the number of frame skipping corresponding to each of the main stream video segments and the number of stalls corresponding to each of the main stream video segments. Based on the total number of frame skipping and the total number of stalls and a preset main stream video anomaly detection rule, the main stream video stream is mapped by the rule to determine the dynamic streaming smoothness score corresponding to the main stream video stream. The dynamic streaming smoothness score, the total number of frame skipping and the total number of stalls are determined as the dynamic streaming smoothness detection result.

7. The intelligent inspection method of video terminal equipment according to claim 1, wherein, The step of performing auxiliary stream video anomaly detection on the auxiliary stream video stream to obtain auxiliary stream video detection result comprises: Key frames are extracted from the auxiliary stream video stream to generate a corresponding auxiliary stream key frame sequence; Each auxiliary stream key frame in the auxiliary stream key frame sequence and a preset standard reference frame sequence are input into a pre-trained second visual large model, and the second visual anomaly large model is used to compare the auxiliary stream key frame sequence and the preset standard reference frame sequence to generate auxiliary stream video dimension scores of the auxiliary stream key frame in each auxiliary stream video preset dimension; The auxiliary stream video dimension scores of the auxiliary stream key frame in each of the auxiliary stream video preset dimensions are compared with the dimension threshold values of each auxiliary stream video preset dimension respectively to determine whether each of the auxiliary stream key frames has an auxiliary stream video anomaly, and the number of defective frames corresponding to the auxiliary stream key frame sequence is counted. For each of the auxiliary stream key frames, a second weighting coefficient corresponding to each of the auxiliary stream video dimension scores is obtained, and each of the second weighting coefficients and each of the auxiliary stream video scores is weighted and summed to obtain a video frame quality initial score of the auxiliary stream key frame. Based on whether the auxiliary stream key frame has the auxiliary stream video anomaly and the video frame quality initial score of the auxiliary stream key frame, a video frame quality score of each of the auxiliary stream key frames and an auxiliary stream video quality score of the auxiliary stream video stream are determined, and a detection conclusion of the auxiliary stream video detection result is determined based on the auxiliary stream video quality score. The video frame quality score, the auxiliary stream video anomaly type label, the auxiliary stream video quality score, and the detection conclusion of the auxiliary stream video detection result of each of the auxiliary stream key frames are determined as auxiliary stream video detection results.

8. The intelligent inspection method of video terminal equipment according to claim 1, wherein, The step of generating the inspection detection report of the target video terminal device according to the network quality detection result, the audio detection result, the main stream video detection result, and the auxiliary stream video detection result comprises: When at least one of the network quality detection result, the audio detection result, the main stream video detection result, and the auxiliary stream video detection result is unqualified, the network quality detection result, the audio detection result, the main stream video detection result, and the auxiliary stream video detection result are input into a pre-trained inspection abnormal root cause analysis model to obtain an abnormal root cause analysis result, the inspection abnormal root cause analysis model being constructed based on a bidirectional encoder representation from transformers (BERT) model; Based on the network quality detection result, the audio detection result, the main stream video detection result, the auxiliary stream video detection result, and the abnormal root cause analysis result, a preset inspection detection template is filled to generate the inspection detection report.

9. An intelligent inspection device for a video terminal device, characterized by The device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the intelligent inspection method of the video terminal device according to any one of claims 1 to 8.

10. A storage medium, characterized by The storage medium is a computer readable storage medium, and the storage medium stores a computer program, the computer program being executed by the processor to implement the steps of the intelligent inspection method of the video terminal device according to any one of claims 1 to 8. The storage medium is a computer readable storage medium, and the storage medium stores a computer program, the computer program being executed by the processor to implement the steps of the intelligent inspection method of the video terminal device according to any one of claims 1 to 8.