A television broadcast quality monitoring system based on a multi-modal model
By designing a television broadcast quality monitoring system based on multimodal model, the problems of low accuracy and high false alarm rate in the existing technology are solved, and more accurate fault detection and broadcast quality assurance are achieved.
Patent Information
- Application Number
- CN202510239915.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing television broadcast quality monitoring system relies on single mode data, resulting in inaccurate fault detection, high false alarm rate, lack of contextual understanding and insufficient ability to handle compound faults.
Design a television broadcast quality monitoring system based on multimodal model. Through data acquisition, preprocessing, multimodal fusion, quality evaluation and feedback control modules, multimodal data such as video, audio, text and other multimodal data are collected and processed in real time, data alignment, feature extraction and feature fusion, and broadcast quality evaluation and fault detection.
It improves the accuracy and robustness of TV broadcast quality monitoring, can more accurately identify and classify various faults in TV broadcast, and realizes real-time monitoring and alarms to ensure broadcast quality.
Smart Images

Figure CN119728959B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal information fusion, and specifically to a television broadcast quality monitoring system based on a multimodal model. Background Art
[0002] With the continuous development of television media, the broadcast quality of television programs has received increasing attention. Existing television broadcast quality monitoring systems mostly rely on single-modal data (such as audio, images) for quality assessment. However, television signals contain multiple-modal data (such as video, audio, text captions, etc.), and there are complex correlation relationships between these data. The technical problems mainly existing in the analysis of single-modal data are as follows:
[0003] (1) Inaccurate fault detection: Existing single-modal systems may overlook problems that are obvious in other modalities. For example, the video may look normal, but the audio may have noise, or the captions may be out of sync.
[0004] (2) High false alarm rate: Systems that rely only on single-modal data may issue alarms without actual faults. For example, a sudden change in audio may be wrongly marked as a fault when there is a normal sound effect in the scene.
[0005] (3) Lack of context understanding: Single-modal data cannot provide the context information of multi-modal data. For example, caption text can help understand whether the speech quality in the audio matches the lip movements in the video.
[0006] (4) Insufficient ability to handle complex scenarios: Abnormalities in the broadcast quality of real-world televisions may involve complex problems in multiple aspects, such as picture freezing, frame skipping, audio loss, audio interruption, audio-visual out-of-sync, and a combination of incorrect captions. Single-modal systems are difficult to handle such compound faults.
[0007] In view of the above problems, how to design a television broadcast quality monitoring system based on a multimodal model is an urgent problem to be solved in this field. Summary of the Invention
[0008] The purpose of the present invention is to provide a television broadcast quality monitoring system based on a multimodal model, which solves the problems of low accuracy and high false alarm rate in monitoring with single-modal data in the prior art.
[0009] To achieve the above object, the present invention is realized through the following technical solutions:
[0010] Provide a television broadcast quality monitoring system based on a multimodal model, including:
[0011] A data acquisition module for real-time acquisition of multi-modal data in the television broadcast signal;
[0012] A data preprocessing module for preprocessing the collected multimodal data;
[0013] A multimodal fusion module for data alignment, feature extraction, and feature fusion of the preprocessed multimodal data;
[0014] A quality assessment module for evaluating the quality of TV broadcasts based on the fused multimodal features and determining whether there are broadcast failures;
[0015] A feedback control module for generating corresponding control instructions according to the quality assessment results and alarming the TV broadcast.
[0016] Preferably, the data acquisition module collects multimodal data during the TV broadcast in real time through a TV signal receiver and a web crawler, including: video data, audio data, text subtitle data, closed caption data, and EPG information.
[0017] Preferably, the data preprocessing module performs the following steps:
[0018] Extract key frames and scene transition frames from the video;
[0019] Perform duplicate removal, denoising, and enhancement preprocessing operations on the extracted frame images to remove noise and interference factors;
[0020] Perform enhancement and noise reduction processing on the audio signal to improve the audio quality;
[0021] Clean and segment the text subtitles to extract useful information.
[0022] Preferably, the multimodal fusion module includes:
[0023] A modality alignment unit, a feature extraction unit, and a feature fusion unit;
[0024] The modality alignment unit is used to align the timestamps of different modality data and specifically performs the following process:
[0025] Alignment of video data and audio data, alignment of text data, unifying the time reference, weight assignment of text data, and generating a unified time series.
[0026] Preferably, set the timestamp of the video frame to and the timestamp of the audio to where represents the index of the video frame, represents the index of the audio frame; the alignment process of the video data and the audio data is specifically as follows:
[0027] Calculate the time offset:
[0028] ;
[0029] Perform temporal adjustment on all video frames:
[0030] ;
[0031] Search for audio frames:
[0032] .
[0033] Preferably, the alignment of the text data includes: text with timestamps and text without timestamps;
[0034] The specific execution process for the text with timestamps is as follows:
[0035] Align the time interval of the text segment with the video frame. If , then the video frame corresponds to the text segment ;
[0036] The specific execution process for the text without timestamps is as follows:
[0037] Use dynamic time warping for text alignment:
[0038] ;
[0039] where, when , otherwise , 0 indicates text match, 1 indicates text mismatch; represents the index of the manually input text, represents the index of the recognized text.
[0040] Preferably, the specific execution process for the unified time reference is as follows:
[0041] Select the time step of the video frame as the unified reference, and perform interpolation or downsampling on the audio frames to match the video frame rate:
[0042] ;
[0043] where, is the sampling rate of the audio, is the frame rate of the video, is the index of the video frame, is the index of the audio frame.
[0044] Preferably, the specific execution process for the weight assignment of the text data is as follows:
[0045] For a text segment covering multiple video frames , assign weights :
[0046] ;
[0047] Or classify according to video frame features and perform weight normalization:
[0048] ;
[0049] Among them, , are respectively the indexes of the start and end of the video frames covered by the text segment.
[0050] Preferably, the feature extraction unit is used to extract the feature representations of video data, audio data, and text data respectively;
[0051] The process of extracting the features of the video data is specifically as follows:
[0052] Set the video frame sequence as , where each is the image of the th frame, and use a convolutional neural network to extract the visual features of each frame:
[0053] ;
[0054] Input the visual features of each frame extracted into a recurrent neural network to capture the temporal information between video frames:
[0055] ;
[0056] The process of extracting the features of the audio data is specifically as follows:
[0057] Set the audio signal as , where each is a sub-frame of the audio, generate a Mel spectrogram , and use a one-dimensional CNN to extract the local features of the spectrogram:
[0058] ;
[0059] Input the local features of the spectrogram extracted into an RNN to extract the temporal features of the audio:
[0060] ;
[0061] The process of extracting the features of the text data is specifically as follows:
[0062] Set the text data as , where each is a text segment, and the semantic features of each text segment are extracted using the pre-trained language model BERT:
[0063] ;
[0064] Select the [CLS] token of BERT as the representation of the entire text segment, and take the average of all word embeddings:
[0065] .
[0066] Preferably, the feature fusion unit is used to fuse multi-modal features, and specifically performs the following process:
[0067] Use the ReLU model to perform dimensional mapping on the features of each modality to make them have the same dimension. For the video features, the mapping is specifically:
[0068] ;
[0069] For the audio features, the mapping is specifically:
[0070] ;
[0071] For the text features, the mapping is specifically:
[0072] ;
[0073] Among them, , , are learnable weights, , , are bias terms;
[0074] Concatenate the features of all modalities and calculate the weights of each modality through a fully connected layer; the calculation method for concatenating the features of all modalities is:
[0075] ;
[0076] The calculation of the weights of each modality is specifically:
[0077] ;
[0078] Among them, is the weight matrix, is the bias term, is the modality weight vector;
[0079] Weighted sum the modality features according to the weights to obtain the fused features:
[0080] ;
[0081] Input the fused features at each time step into an RNN or Transformer for sequence-level processing to capture temporal information.
[0082] Preferably, the quality assessment module includes:
[0083] A fault detection unit for detecting faults in television broadcasts based on multimodal features, specifically: using an anomaly detection algorithm to analyze the fused multimodal features to determine whether there are anomalies;
[0084] A fault classification unit for classifying the detected faults to determine the fault type and severity, specifically: using a classifier to classify the fault types;
[0085] A quality scoring unit for comprehensively evaluating the quality of television broadcasts using a scoring model based on the results of fault detection and classification, and generating a detailed quality assessment report.
[0086] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0087] (1) Multimodal data fusion: By fusing data of multiple modalities such as video, audio, and text, the complementarity between modalities is fully utilized to improve the accuracy and robustness of television broadcast quality monitoring;
[0088] (2) Fault detection and classification: Based on multimodal features for fault detection and classification, various faults in television broadcasts can be more accurately identified and classified;
[0089] (3) Real-time monitoring and feedback: Real-time monitoring of the quality of television broadcasts, and generating corresponding control instructions according to the monitoring results to achieve real-time adjustment and alarm of television broadcasts, ensuring the broadcast quality. Description of the Drawings
[0090] Figure 1 is a structural block diagram of a television broadcast quality monitoring system based on a multimodal model provided by an embodiment of the present invention;
[0091] Figure 2 is a flowchart of a television broadcast quality monitoring system based on a multimodal model provided by an embodiment of the present invention;
[0092] Figure 3 is a video preprocessing flowchart of a television broadcast quality monitoring system based on a multimodal model provided by an embodiment of the present invention;
[0093] Figure 4It is a text preprocessing flowchart of a TV broadcast quality monitoring system based on a multi-modal model provided by an embodiment of the present invention. Detailed implementation manners
[0094] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.
[0095] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relationship terms determined for the convenience of describing the structural relationship of each component or element of the present invention, and do not specifically refer to any component or element of the present invention, and should not be construed as a limitation to the present invention.
[0096] Embodiment:
[0097] As Figure 1 shown, this embodiment provides a TV broadcast quality monitoring system based on a multi-modal model, including:
[0098] A data acquisition module for real-time acquisition of multi-modal data in TV broadcast signals;
[0099] A data preprocessing module for preprocessing the acquired multi-modal data;
[0100] A multi-modal fusion module for data alignment, feature extraction, and feature fusion of the preprocessed multi-modal data;
[0101] A quality assessment module for evaluating the TV broadcast quality based on the fused multi-modal features and determining whether there are broadcast faults;
[0102] A feedback control module for generating corresponding control instructions according to the quality assessment results and alarming the TV broadcast.
[0103] Based on the above-mentioned various modules, the monitoring process executed by the monitoring system of this embodiment is as Figure 2 shown.
[0104] 1. The data acquisition module includes: real-time acquisition of multi-modal data during TV broadcast through means such as TV signal receivers and web crawlers, including video, audio, text subtitles, closed captions, EPG information, etc. In this embodiment, the specific acquisition process is:
[0105] Video data acquisition: Use a TV signal receiver to collect TV program video data in real time from satellite signals, with a resolution of 1920x1080 and a frame rate of 30fps;
[0106] Audio data acquisition: Synchronously collect the audio data of the TV program, with a sampling rate of 44.1kHz and stereo;
[0107] Text data acquisition: Obtain the subtitle data of the TV program through a web crawler, in the format of an SRT file, containing timestamps and corresponding text content;
[0108] Sample data:
[0109] Video data: The length of the collected video clip is 10 seconds, with a total of 300 frames.
[0110] Audio data: The length of the collected audio clip is 10 seconds, with a total of 441,000 sampling points.
[0111] Text data: The collected subtitle data contains 2 subtitle segments, corresponding to the time intervals [0.0, 5.0] and [5.0, 10.0] seconds respectively.
[0112] 2. Data preprocessing module, used to preprocess the collected multi-modal source data, including:
[0113] (1) Extract key frames and scene transition frames from the video;
[0114] (2) Perform preprocessing operations such as duplicate removal, denoising, and enhancement on the extracted frame images to remove noise and interference factors;
[0115] (3) Process the audio signal for enhancement, noise reduction, etc. to improve the audio quality;
[0116] (4) Clean and segment the text subtitles to extract useful information;
[0117] In this implementation, the preprocessing process is as follows:
[0118] As Figure 3 shown, the video preprocessing process is as follows:
[0119] Key frame extraction: Use the OpenCV library or FFmpeg to extract 1 key frame per second;
[0120] Image enhancement: Apply an image enhancement algorithm (such as histogram equalization) to remove noise and improve image clarity;
[0121] The audio preprocessing process is as follows:
[0122] Noise reduction processing: Use Audacity software for audio noise reduction, and set the noise reduction intensity to 20dB;
[0123] As Figure 4 shown, the text preprocessing process is as follows:
[0124] Text cleaning: Remove special characters and spaces in the subtitles and unify the text format;
[0125] Word segmentation: Use the Jieba word segmentation tool to segment the Chinese subtitles.
[0126] 3. The multimodal fusion module includes: a modality alignment unit, a feature extraction unit, and a feature fusion unit. Among them, the modality alignment unit is used to align the timestamps of different modality data to ensure the consistency of multimodal data in the time dimension; specifically as follows:
[0127] 1) Video and audio alignment:
[0128] The timestamp of the video frame is and the timestamp of the audio is where represents the index of the video frame, represents the index of the audio frame, and the alignment is performed through the following steps:
[0129] (1) Calculate the time offset:
[0130] ;
[0131] (2) Adjust the time of all video frames:
[0132] ;
[0133] (3) Find the closest audio frame:
[0134] ;
[0135] Using the data collected by the data acquisition module in this embodiment, after calculation, the video frame timestamp is =[0.0, 0.033, 0.066,..., 10.0] seconds;
[0136] The audio frame timestamp is =[0.0, 0.022, 0.044,..., 10.0] seconds;
[0137] Calculate the time offset and perform interpolation processing on the video frames to match the audio frame rate;
[0138] 2) Text data alignment
[0139] Case 1: The text has timestamps (such as subtitles)
[0140] Align the time interval of the text segment with the video frames. If , then the video frame corresponds to the text segment ;
[0141] Case 2: The text has no timestamps (such as manually entered commentary)
[0142] Use dynamic time warping (DTW) for text alignment:
[0143] ;
[0144] where, when , otherwise , 0 indicates text match and 1 indicates text mismatch; represents the index of the manually entered text, represents the index of the recognized text;
[0145] 3) Unify the time reference
[0146] Select the time step of the video frames as the unified reference. Interpolate or downsample the audio frames to match the video frame rate:
[0147] ;
[0148] where, is the sampling rate of the audio (unit: Hz, such as 44.1 kHz); is the frame rate of the video (unit: fps, such as 30 fps); is the index of the video frame (such as the 0th frame, the 1st frame, etc.); is the index of the audio frame (the index of the aligned audio data);
[0149] 4) Text weight assignment
[0150] For a text segment that covers multiple video frames , assign a weight :
[0151] ;
[0152] Or classify according to video frame features to ensure weight normalization:
[0153] ;
[0154] where, , are the start and end indices of the video frames covered by the text segment;
[0155] 5) Final alignment result
[0156] Generate a unified time series, with each time step containing the corresponding video frame, audio frame, and text segment:
[0157] {
[0158] "timestep":0,
[0159] "video_frame":frame_0,
[0160] "audio_frame":audio_0,
[0161] "text_segment":text_0,
[0162] },
[0163] {
[0164] "timestep":1,
[0165] "video_frame":frame_1,
[0166] "audio_frame":audio_1,
[0167] "text_segment":text_1,
[0168] },
[0169] The feature extraction unit is used to extract the feature representations of video, audio, and text data respectively:
[0170] 1) Video feature extraction:
[0171] Use a convolutional neural network (CNN) to extract the visual features of video frames, and use a recurrent neural network to extract the temporal features between video frames:
[0172] Assume the video frame sequence is , where each is the image of the -th frame:
[0173] ① Use a convolutional neural network (CNN) to extract the visual features of each frame:
[0174] ;
[0175] ② Input these features into a recurrent neural network (RNN) to capture the temporal information between video frames:
[0176] ;
[0177] 2) Audio feature extraction:
[0178] The acoustic features of the audio are extracted by combining the spectrogram with a recurrent neural network (RNN). Assume the audio signal is , where each is a frame of the audio:
[0179] ① Generate the Mel spectrogram , and then use a one-dimensional CNN to extract the local features of the spectrogram:
[0180] ;
[0181] ② Input these features into the RNN to extract the temporal features of the audio:
[0182] ;
[0183] 3) Text feature extraction:
[0184] The semantic features of the text are extracted using a pre-trained language model (such as BERT). Assume the text data is , where each is a text segment:
[0185] ① Use a pre-trained language model such as BERT to extract the semantic features of each text segment:
[0186] ;
[0187] ② Select the [CLS] token of BERT as the representation of the entire text segment and take the average of all word embeddings:
[0188] ;
[0189] The feature fusion unit is used to fuse multi-modal features:
[0190] 1) The modal feature dimension mapping is specifically as follows. Use the ReLU model to map the features of each modality to have the same dimension:
[0191] ① Video feature mapping:
[0192] ;
[0193] ② Audio feature mapping:
[0194] ;
[0195] ③ Text feature mapping:
[0196] ;
[0197] Among them, 、 、 are learnable weights, 、 、 are bias terms;
[0198] 2) Calculate the attention weights. Specifically: Concatenate the features of all modalities and calculate the weights of each modality through a fully connected layer, including:
[0199] ① Concatenate features
[0200] ;
[0201] ② Calculate weights
[0202] ;
[0203] Among them, is the weight matrix, is the bias term, is the modality weight vector;
[0204] 3) Fuse features. Specifically: Perform weighted summation on the modality features according to the weights to obtain the fused features:
[0205] ;
[0206] 4) Time series level processing. Specifically: Input the fused features of each time step into an RNN or Transformer for sequence-level processing to capture temporal information.
[0207] 4. The quality assessment module includes:
[0208] (1) Fault detection unit:
[0209] Used to detect faults in TV broadcasts based on multi-modal features, such as video stuttering, audio distortion, subtitle disorder, etc.; Specifically: Use anomaly detection algorithms (such as isolation forest, autoencoder, etc.) to analyze the fused multi-modal features to determine whether there are anomalies;
[0210] (2) Fault classification unit:
[0211] Used to classify the detected faults to determine the fault type and severity; Specifically: Use a classifier (such as support vector machine, deep neural network, etc.) to classify the fault type;
[0212] (3) Quality scoring unit:
[0213] Based on the results of fault detection and classification, the quality scoring unit comprehensively evaluates the quality of TV broadcasts using a scoring model and generates a detailed quality assessment report, which includes the scoring results and fault details.
[0214] 5. Feedback control module:
[0215] Based on the quality assessment results, generate control instructions, send alarm messages to the monitoring center, and record the fault details.
[0216] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A television broadcast quality monitoring system based on a multimodal model, characterized in that: include: Data acquisition module, used for real-time acquisition of multi-modal data from television broadcast signals; A data preprocessing module, used for preprocessing the collected multimodal data; Multimodal fusion module, used to perform data alignment, feature extraction and feature fusion on the preprocessed multimodal data; The quality assessment module is used to assess the quality of TV broadcasting based on the fused multimodal features and determine whether there are any broadcasting failures; A feedback control module is used to generate corresponding control instructions according to the quality assessment results and issue an alarm for TV broadcasting; The multimodal fusion module comprises: Modal alignment unit, feature extraction unit and feature fusion unit; The modality alignment unit is used to align the timestamps of data of different modalities, and specifically performs the following process: Alignment of video and audio data, alignment of text data, unified time base, weight assignment of text data, and generation of unified time series; The specific alignment of video data and audio data is as follows: Set the timestamp of the video frame to , the timestamp of the audio is ,in, Represents the index of the video frame, Indicates the index of the audio frame; the alignment of the video data and the audio data, the specific process is: Calculate the time offset: ; Temporally adjust all video frames: ; Find audio frames: ; The alignment of the text data includes: text with a timestamp and text without a timestamp; The specific execution process of the text with a timestamp is as follows: The text segment The time interval and Video frame alignment, if , then the video frame Corresponding text segment ; The specific execution process of the text without timestamp is as follows: Use Dynamic Time Warping to align text: ; Among them, when hour, ,otherwise , 0 means the text matches, 1 means the text does not match; represents the index of manually entered text, Indicates the index of the recognized text; The specific implementation process of the unified time base is as follows: Select the time step of video frames As a baseline, audio frames are interpolated or downsampled to match the video frame rate: ; in, is the audio sampling rate, is the frame rate of the video, is the index of the video frame, is the index of the audio frame; The specific implementation process of the weight distribution of the text data is as follows: For text segments covering multiple video frames , assign weights : ; Or classify the video frame features and normalize the weights: ; in, , The indices of the start and end of the video frames covered by the text segment, respectively.
2. A television broadcast quality monitoring system based on a multimodal model according to claim 1, characterized in that: The data acquisition module collects multimodal data in the process of television broadcasting in real time through a television signal receiver and a web crawler, including: video data, audio data, text subtitle data, closed subtitle data and EPG information.
3. The television broadcast quality monitoring system based on a multimodal model according to claim 1, characterized in that: The data preprocessing module performs the following steps: Extract key frames and scene transition frames from the video; De-duplication, denoising and enhancement pre-processing operations are performed on the extracted frame images to remove noise and interference factors; Enhance and reduce noise on audio signals to improve audio quality; Clean and segment text subtitles to extract useful information.
4. A television broadcast quality monitoring system based on a multimodal model according to claim 1, characterized in that: The feature extraction unit is used to extract feature representations of video data, audio data and text data respectively; The feature extraction of the video data specifically performs the following process: Set the video frame sequence to , where each It is Frame images, using convolutional neural networks to extract visual features of each frame: ; The visual features extracted from each frame are input into the recurrent neural network to capture the timing information between video frames: ; Extract the features of audio data and perform the following process: Set the audio signal to , where each It is a sub-frame of audio, generating a Mel spectrum graph , use one-dimensional CNN to extract local features of the spectrogram: ; The local features of the extracted spectrogram are input into the RNN to extract the temporal features of the audio: ; Extract the features of text data and perform the following process: Set the text data to , where each is a text segment, and the pre-trained language model BERT is used to extract the semantic features of each text segment: ; Select BERT's [CLS] token as the representation of the entire text segment and average all word embeddings: 。 5. The television broadcast quality monitoring system based on a multimodal model according to claim 1, characterized in that: The feature fusion unit is used to fuse multimodal features, and specifically performs the following process: Use the ReLU model to map the dimensions of the features of each modality so that they have the same dimension and map the video features. Specifically: ; Map the audio features, specifically: ; Map the text features, specifically: ; in, , , are learnable weights, , , is the bias term; The features of all modalities are concatenated, and the weight of each modality is calculated through a fully connected layer; the calculation method for concatenating the features of all modalities is: ; The weight of each modality is calculated as follows: ; in, is the weight matrix, is the bias term, is the modal weight vector; The modal features are weighted and summed according to the weights to obtain the fusion features: ; The fusion features of each time step Input into RNN or Transformer for sequence-level processing to capture timing information.
6. A television broadcast quality monitoring system based on a multimodal model according to claim 1, characterized in that: The quality assessment module includes: A fault detection unit is used to detect faults in television broadcasting based on multimodal features, specifically: using an anomaly detection algorithm to analyze the fused multimodal features to determine whether there is an anomaly; A fault classification unit is used to classify the detected faults and determine the fault type and severity, specifically: using a classifier to classify the fault type; The quality scoring unit is used to comprehensively evaluate the quality of television broadcasting based on the results of fault detection and classification using a scoring model and to generate a detailed quality evaluation report.
Citation Information
Patent Citations
Method and device for detecting playing anomaly of video polyphonic ringtone
CN118764559A
Multi-modal model and method for fusing characters, images and audios
CN118861988A
Unstructured video data automatic classification method
CN119516266A