Multi-dimensional news broadcast program relay monitoring method, device, equipment and medium
By introducing audio duration and host location recognition models, the problems of dynamic duration and visual tampering in the news broadcast were solved, achieving high-precision, fully automated broadcast monitoring and reducing reliance on manual labor and monitoring costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-07
AI Technical Summary
Existing technologies are ill-suited to the dynamic duration changes of the news broadcast, lack spatial visual verification mechanisms, and rely on high costs for manual monitoring, resulting in low accuracy of broadcast monitoring and frequent omissions and errors.
By employing an audio duration recognition model and a host location recognition model based on speech recognition, combined with customized hot word matching and target detection technologies, multi-dimensional monitoring of audio and video signals is achieved, including dynamic duration extraction and visual morphology verification, forming a dual comprehensive judgment logic.
It improves the accuracy and fault tolerance of broadcast duration monitoring, keenly identifies image tampering behavior, achieves fully automated monitoring, reduces labor costs, and enhances monitoring timeliness and reliability.
Smart Images

Figure CN122349030A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of broadcast television monitoring technology, and more specifically, to a multi-dimensional method, apparatus, equipment, and medium for monitoring the rebroadcast of CCTV's news program. Background Technology
[0002] In the modern broadcasting and television media field, the unified news broadcast program at a specific time is a carrier of important information transmission value. Typically, broadcasting networks require all levels of relay organizations (such as local relay channels) to synchronously and accurately rebroadcast the news broadcast program from their superiors at designated times. The standardization and completeness of this rebroadcast directly affect the accurate delivery rate of core news information. Currently, monitoring the program rebroadcast quality for a massive number of relay channels mainly relies on traditional broadcasting operation and maintenance systems or manual inspections, but this approach has revealed many technical shortcomings in practical applications: I. Existing technologies are ill-suited for the precise monitoring of dynamic broadcast durations. While large-scale news broadcasts, uniformly distributed, have a basic, conventional duration (e.g., a standard 30 minutes), their actual broadcast duration is not absolutely fixed. In the face of sudden major events or special news arrangements, the actual broadcast duration often dynamically extends or shortens. This specific duration is usually announced verbally by the program host at the beginning of the broadcast or during the broadcast. Most existing broadcast monitoring systems employ a mechanical matching method based on preset fixed duration thresholds, failing to intelligently extract dynamic duration notification information from the audio and video streams. This technological limitation makes the system highly prone to misjudgments, such as misreporting normally extended programs as "abnormal rebroadcast timeouts," or failing to identify "rebroadcast duration reductions" caused by local illegal cut-offs, resulting in a significant decrease in the accuracy of duration-based monitoring.
[0003] Second, existing technologies lack spatial visual verification mechanisms specific to particular program formats. Standardized news broadcasts often possess highly fixed visual characteristics, such as the spatial position of news anchors within the broadcast frame and the compositional proportions, which are distinct and unique. This is a crucial visual criterion for determining whether the rebroadcast footage belongs to the original broadcast source and whether there has been malicious tampering (such as local channels illegally inserting local commercials or replacing parts of the footage). However, most existing video monitoring technologies remain at the level of basic anomaly identification, such as video frame freezing and black / color bar detection, and have not introduced high-precision computer vision-specific verification mechanisms for the unique characteristic of "anchor spatial position." Therefore, when rebroadcast footage is partially tampered with or replaced, existing systems cannot effectively identify it, making it difficult to guarantee the authenticity and compliance of the rebroadcast footage.
[0004] Third, traditional broadcast monitoring solutions suffer from misaligned priorities and excessively high manual monitoring costs. Existing broadcast industry monitoring systems primarily focus on monitoring the operational status of underlying equipment and the stability of physical transmission links, emphasizing solutions to communication-level issues such as signal interruptions and signal-to-noise ratio degradation. They lack application-layer monitoring mechanisms specifically for the "News Broadcast" scenario. Facing hundreds or even thousands of downstream broadcast channels, relying solely on manual, 24 / 7, multi-channel audio and video content comparison not only incurs enormous human resource costs but is also susceptible to human fatigue, easily leading to missed or incorrect reports. This fundamentally fails to meet the demands of today's broadcast television networks for massive concurrency and high real-time intelligent monitoring.
[0005] In summary, there is an urgent need in this field for a multi-dimensional monitoring method for key news program broadcasts, in order to effectively overcome the technical shortcomings of existing technologies, such as low accuracy in identifying the dynamic duration of programs, lack of spatial visual anti-tampering verification mechanisms, and over-reliance on manual inspections, thereby achieving fully automated and high-precision monitoring of the standardization and integrity of broadcasts on lower-level channels. Summary of the Invention
[0006] The present invention aims to solve at least one of the aforementioned technical problems existing in the prior art.
[0007] Therefore, the first aspect of the present invention provides a multi-dimensional method for monitoring the rebroadcast of CCTV News programs.
[0008] A second aspect of the present invention provides a multi-dimensional monitoring device for the broadcast of CCTV's news program.
[0009] A third aspect of the present invention provides an electronic device.
[0010] A fourth aspect of the present invention provides a computer-readable storage medium.
[0011] This invention provides a multi-dimensional method for monitoring the rebroadcast of CCTV's "News Broadcast" program, comprising: During the preset program broadcast time period, audio and video signals of the target channel are collected synchronously, and the audio and video signals are preprocessed to separate the audio signals and video frame sequences. The audio signal is input into a pre-trained audio duration recognition model, and the program duration value in the audio signal is extracted through speech recognition and used as the standard duration for the day. The video frame sequence is input into a pre-trained host position recognition model for target detection. The host position coordinates in each frame are identified, and the host position coordinates are compared with a preset host standard position threshold to obtain the percentage of frames with normal positions. Obtain the actual broadcast duration of the target channel; Based on the difference between the actual broadcast duration and the standard duration of the day, and the percentage of normal frames at the location, a broadcast monitoring and judgment result for the target channel is generated.
[0012] The multi-dimensional news broadcast monitoring method according to the above-described technical solution of the present invention may also have the following additional technical features: In the above technical solution, the step of inputting the audio signal into a pre-trained audio duration recognition model, extracting the program duration value from the audio signal through speech recognition, and using it as the standard duration for the day includes: The audio signal is input into an audio duration recognition model that supports single-inference of long audio; The audio signal is matched and extracted using a pre-configured custom hot word library, which includes keywords related to program duration notifications. If a duration notification containing the keywords is extracted, the duration value in the duration notification is parsed and extracted as the standard duration for the day. If the duration notification information is not extracted, the preset default duration will be used as the temporary standard duration, and the temporary standard duration will be corrected based on the extracted program end characteristics to obtain the standard duration for the day.
[0013] In the above technical solution, the preset host standard position threshold includes a horizontal position range, a vertical position range, and a position offset threshold; The step of inputting the video frame sequence into a pre-trained host position recognition model for target detection, identifying the host's position coordinates in each frame, and comparing the host's position coordinates with a preset host standard position threshold includes: The coordinates of the center point of the host's head in the current video frame are extracted using the host position recognition model. Determine whether the coordinates of the head center point fall within the horizontal position range and the vertical position range, and obtain the coordinate offset of the current video frame relative to the previous video frame, and determine whether the coordinate offset is less than the position offset threshold. If all conditions are met, the host's position in the current video frame is determined to be normal.
[0014] In the above technical solution, generating a broadcast monitoring and judgment result for the target channel based on the difference between the actual broadcast duration and the standard duration of the day, and the proportion of normal frames at the location, includes: Obtain the actual start and end times of the program broadcast on the target channel, and calculate the actual broadcast duration. Calculate the absolute value of the difference between the actual broadcast duration and the standard broadcast duration for that day; If the absolute value of the difference is less than or equal to the preset error range, the current broadcast duration is determined to be normal. If the absolute value of the difference is greater than the preset error range, the current broadcast duration is determined to be abnormal, and an abnormality type marker in the duration dimension is generated.
[0015] In the above technical solution, the step of generating a broadcast monitoring and judgment result for the target channel based on the difference between the actual broadcast duration and the standard duration of the day, and the proportion of normal frames at the location, further includes: If the percentage of frames with the specified location being normal is not lower than a preset percentage threshold, and there is no situation where the host's location coordinates are not identified for a consecutive preset number of frames, then the current broadcast is determined to be normal. If the percentage of normal frames at the specified location is lower than the preset percentage threshold, or if the host's location coordinates are not identified for a consecutive preset number of frames, then the current broadcast is determined to be abnormal, and an abnormality type marker for the frame dimension is generated. If the current broadcast duration and the current broadcast video are both normal, a broadcast monitoring pass result is generated for the target channel; otherwise, a broadcast failure result is generated and an early warning signal is triggered.
[0016] In the above technical solution, the preprocessing of the audio and video signals to separate the audio signals and video frame sequences includes: The original audio was denoised using an adaptive noise reduction algorithm, and the audio was segmented after unifying the sampling rate to obtain the audio signal. The original video is denoising the separated video using a filtering algorithm. Video frames are extracted according to a preset frequency and normalized to a preset pixel size to obtain the video frame sequence.
[0017] The above technical solution also includes a model optimization step: Periodically collect new audio samples containing program duration notifications, as well as new video frame samples containing host broadcasts; The audio duration recognition model is iteratively trained using the newly added audio samples, and the customized hot word library is updated. The host position recognition model is iteratively trained using the newly added video frame samples, and the host standard position threshold is dynamically adjusted.
[0018] This invention provides a multi-dimensional news broadcast program rebroadcast monitoring device, comprising: The acquisition and preprocessing module is used to synchronously acquire audio and video signals of the target channel during a preset program broadcast period, and to preprocess the audio and video signals to separate the audio signals and video frame sequences. The duration recognition module is used to input the audio signal into a pre-trained audio duration recognition model, extract the program duration value from the audio signal through speech recognition, and use it as the standard duration for the day. The video frame monitoring module is used to input the video frame sequence into a pre-trained host position recognition model for target detection, identify the host position coordinates in each frame, and compare the host position coordinates with a preset host standard position threshold to obtain the percentage of frames with normal positions. The duration calculation module is used to obtain the actual broadcast duration of the target channel broadcast; The comprehensive judgment module is used to generate a rebroadcast monitoring judgment result for the target channel based on the difference between the actual rebroadcast duration and the standard duration of the day, as well as the proportion of normal frames at the location.
[0019] The present invention provides an electronic device comprising: a processor and a memory communicatively connected to the processor; The memory stores a computer program that can be executed by the processor, and when the processor executes the computer program, it implements the method as described in any of the above technical solutions.
[0020] The present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method as described in any one of the above technical solutions.
[0021] In summary, due to the adoption of the above-mentioned technical features, the beneficial effects of the present invention are: This invention introduces an audio duration recognition model based on speech recognition technology, combined with a customized hot word matching mechanism, to accurately and in real-time extract dynamic program duration notifications spoken by the host from the audio stream. This design breaks away from the traditional mindset of relying heavily on preset fixed durations for broadcast monitoring, effectively solving the problem of system misjudgment caused by temporary adjustments (extensions or shortenings) to program durations due to major events, and significantly improving the monitoring accuracy and fault tolerance of broadcast durations during special periods.
[0022] In response to the highly standardized visual format of important news broadcasts, this invention innovatively applies target detection technology to the business-level review of broadcast compliance. By extracting and comparing the highly distinctive visual feature of "the host's spatial position (horizontal, vertical, and offset)," this solution breaks through the limitations of traditional video monitoring, which only focuses on low-level signal anomalies such as black screens and still frames. It can keenly and accurately identify deeper malicious tampering behaviors such as illegal insertion of local advertisements, partial replacement or obstruction of images, etc., effectively ensuring the authenticity and standardization of broadcast content.
[0023] This invention ingeniously combines "dynamic standard duration calculation" in the time dimension with "video frame position normality percentage statistics" in the spatial dimension, forming a dual comprehensive judgment logic for program broadcasting-specific scenarios. This multi-dimensional cross-verification mechanism successfully extends the monitoring dimension of the broadcast monitoring system from the physical transmission link status to the deep level of program content standardization, effectively filling the technical gap in the compliance review of complex business operations in existing equipment-level operation and maintenance systems.
[0024] This solution achieves fully automated parallel monitoring of a massive number of downstream broadcast channels within preset broadcast periods by simultaneously collecting data and performing pipelined AI analysis. This mechanism completely eliminates the reliance on manual inspections inherent in traditional methods, significantly reducing the manpower and operational costs for broadcast monitoring agencies. Furthermore, it fundamentally avoids the risks of missed or false reports due to human fatigue and blind spots, thereby improving the overall timeliness and reliability of monitoring.
[0025] This invention provides a continuous model optimization process, supporting iterative training with new samples and dynamic updates to the hot word library and position thresholds. This allows the monitoring system to adapt to subtle changes in program hosts, broadcasting styles, or studio formats, exhibiting extremely high system robustness. Simultaneously, the solution can immediately output fine-grained anomaly type markers (such as insufficient duration, image tampering, etc.) after anomaly detection, and retain the detection data and audio / video clips, providing a complete and objective chain of data evidence for subsequent violation verification, rectification, and accountability.
[0026] Additional aspects and advantages of the invention will become apparent in the following description or may be learned by practice of the invention. Attached Figure Description
[0027] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart of a multi-dimensional news broadcast program rebroadcast monitoring method according to an embodiment of the present invention; Figure 2 This is a flowchart of the training and recognition process of the audio duration recognition model in a multi-dimensional news broadcast monitoring method according to an embodiment of the present invention. Figure 3 This is a flowchart of the training and recognition process of the presenter position recognition model in a multi-dimensional news broadcast monitoring method according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the standard position area of the host in a multi-dimensional news broadcast program rebroadcast monitoring method according to an embodiment of the present invention. Detailed Implementation
[0028] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0029] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0030] The following reference Figures 1 to 4 This describes a multi-dimensional news broadcast monitoring method, apparatus, equipment, and medium provided according to some embodiments of the present invention.
[0031] Some embodiments of this application provide a multi-dimensional method for monitoring the rebroadcast of CCTV's news program.
[0032] like Figure 1 As shown, the first embodiment of the present invention proposes a multi-dimensional news broadcast program rebroadcast monitoring method, including the following steps S1 to S7, wherein S6 and S7 are not mandatory and can be flexibly adjusted as needed.
[0033] S1. During the preset program broadcast time period, the audio and video signals of the target channel are collected synchronously, and the audio and video signals are preprocessed to separate the audio signals and video frame sequences.
[0034] In some embodiments, the preset program broadcast time slot can be set based on a uniformly issued broadcast schedule, for example, setting the monitoring trigger time to 19:00:00 daily. The monitoring system simultaneously acquires signals from multiple target broadcast channels by building a multi-channel audio and video acquisition module. To ensure the timeliness of monitoring, in one specific embodiment, the signal acquisition delay can be controlled within a preset time, such as 1 second. After acquisition, the audio and video signals are preprocessed. Specifically, the audio preprocessing process is as follows: the separated original audio is denoised using an adaptive noise reduction algorithm to remove environmental noise and signal interference; then the sampling rate is unified, and in one specific embodiment, it is converted into 44.1kHz sampling rate, 16-bit quantized digital audio, and the audio is segmented according to a preset time length, such as every 30 seconds, to obtain the audio signal for subsequent parallel processing by the model. The video preprocessing process is as follows: using filtering algorithms such as Gaussian filtering to reduce image noise in the separated original video and eliminate image blur; then, according to a preset frequency, in a specific embodiment, video frames are extracted at a frequency of 1 to 3 frames per second, such as 2 frames per second, and the extracted video frames are normalized to a preset pixel size, such as uniformly scaling to a pixel size of 1920×1080, to obtain the video frame sequence.
[0035] S2. Input the audio signal into a pre-trained audio duration recognition model, extract the program duration value from the audio signal through speech recognition, and use it as the standard duration for the day.
[0036] like Figure 2 As shown, for the recognition of dynamic duration, this embodiment adopts an audio duration recognition model based on speech recognition technology. In some embodiments, this model can be optimized using the VibeVoice-ASR related architecture to support single inference for audio files up to 60 minutes long, thereby avoiding contextual breaks and recognition errors caused by audio slicing. Specifically, during the recognition process, the system matches and extracts the audio signal using a pre-configured customized hot word library. The customized hot word library includes keywords related to program duration notifications, such as "News Broadcast," "program duration," "minutes," and "today." If duration notification information containing the keywords is extracted, for example, recognizing the speech "Today's News Broadcast program duration is 32 minutes," the duration value of 32 minutes is parsed and extracted, and used as the standard duration for that day. If the duration notification information is not extracted, a preset default duration is used, which in some embodiments is set to a regular 30 minutes, as a temporary standard duration. Meanwhile, based on the extracted program end features, such as identifying a specific end credits audio fingerprint or end credits freeze frame, the actual end point of the program is determined, and the temporary standard duration is corrected accordingly to finally obtain the standard duration for the day.
[0037] For the audio duration recognition model, during the model training phase, the system collects audio samples of presenters announcing duration notifications in important news programs within a preset time period (e.g., the past 3 years) as positive examples, for example, collecting 1000 positive examples and labeling them with duration values and broadcast time periods; simultaneously, it collects audio samples of non-duration notifications such as news content broadcasts, background music, and noise as negative examples, for example, collecting 5000 negative examples. After injecting customized hot words into the model, a preset fine-tuning script, such as the LoRA fine-tuning script, is used to train the model. In one specific embodiment, the model's recognition accuracy is optimized through 100 iterations, so that the final model's duration notification recognition accuracy reaches over 99%, and the word error rate is controlled within 1%. After training is completed, the audio signal is input into the pre-trained audio duration recognition model, and the program duration value in the audio signal is extracted through speech recognition.
[0038] S3. Input the video frame sequence into a pre-trained host position recognition model for target detection, identify the host position coordinates in each frame, and compare the host position coordinates with a preset host standard position threshold to obtain the percentage of frames with normal positions.
[0039] like Figure 3As shown, for visual verification of broadcast footage, this embodiment constructs a host position recognition model using a two-stage detection method combining YOLO series object detection algorithms and cascaded classifiers. The preset host standard position thresholds include a horizontal position range, a vertical position range, and a position offset threshold. After extracting the coordinates of the host's head center point and body contour in the current video frame using the host position recognition model, it is determined whether the head center point coordinates fall within the horizontal and vertical position ranges. In a specific embodiment, for a 1080P normalized image, the set standard position thresholds are: a horizontal position range of 500 to 1420 pixels and a vertical position range of 300 to 900 pixels. Furthermore, to enhance the model's adaptability to video images of different resolutions, the preset host standard position thresholds can be defined not only using the absolute pixel coordinates mentioned above but also using a ratio relative to the image edge. For example, refer to... Figure 4 As shown, in a preferred embodiment, the vertical position thresholds are set as follows: the upper boundary is 25% from the top edge of the screen, and the lower boundary is 15% from the bottom edge of the screen; the horizontal position thresholds are set as follows: the left boundary is 30% from the left edge of the screen, and the right boundary is 30% from the right edge of the screen. As long as the host's feature coordinates fall within the area defined by these proportions, the position is considered normal. Simultaneously, the system obtains the host's coordinate offset from the previous video frame and determines whether this offset is less than the preset position offset threshold. In one specific embodiment, the position offset threshold is set to ±50 pixels. If all the above conditions are met, the host's position in the current video frame is determined to be normal. Finally, the proportion of frames with normal positions in the total number of frames in the current broadcast video is calculated to obtain the percentage of frames with normal positions. If the coordinates are outside the range or the offset is too large, it indicates an abnormality such as screen scaling, partial occlusion, or replacement / tampering.
[0040] For the presenter position recognition model, during the model training phase, the system collects video frames from important news programs where presenters are broadcasting normally as positive examples, for example, 20,000 positive examples covering different presenters and different broadcasting scenarios, and labels the center point coordinates of the presenter's head, the coordinates of their entire body outline, and their position region. Simultaneously, it collects non-presenter footage, such as news footage, subtitle footage, and advertising footage, as negative examples, for example, 50,000 negative examples. All samples are normalized to the same size and then input into the model for training. In one specific embodiment, through 80 iterations to optimize the model's anti-interference ability, the final presenter position recognition accuracy reaches over 98.5%, effectively resisting interference such as logo occlusion and image scaling. After training, the video frame sequence is input into the pre-trained presenter position recognition model for target detection.
[0041] S4. Obtain the actual broadcast duration of the target channel.
[0042] Specifically, in determining the broadcast duration, the actual start time and actual end time of the target channel are obtained, and the actual broadcast duration is calculated.
[0043] S5. Based on the difference between the actual broadcast duration and the standard duration of the day, and the proportion of normal frames at the location, generate a broadcast monitoring and judgment result for the target channel.
[0044] In step S5, a multi-dimensional comprehensive judgment is made by combining the duration recognition results in the time dimension and the location monitoring results in the spatial dimension. First, the absolute value of the difference between the actual broadcast duration and the standard duration of the day is calculated. If the absolute value of the difference is less than or equal to a preset error range (in some embodiments, this error range can be set to ±5 seconds), the current broadcast duration is determined to be normal. If it exceeds the preset error range, the current broadcast duration is determined to be abnormal, and an abnormality type mark in the duration dimension such as not starting on time, not ending on time, insufficient duration, or excessive duration is generated. In terms of broadcast image judgment, it is determined whether the percentage of frames with normal position is not lower than a preset percentage threshold (in one specific embodiment, this threshold is set to 95%), and there is no situation where the host's position coordinates are not identified for a consecutive preset number of frames (e.g., 10 consecutive frames). If the condition is met, the current broadcast image is determined to be normal; otherwise, the current broadcast image is determined to be abnormal, and an abnormality type mark in the image dimension such as image tampering, missing host, or abnormal position is generated. A broadcast monitoring pass result for the target channel is generated only if the current broadcast duration and the current broadcast video are both normal; if any abnormality exists, a broadcast fail result is generated and an early warning signal is triggered.
[0045] In one specific embodiment, the actual start time (e.g., 19:00:02) and actual end time (e.g., 19:32:01) of a local TV station's broadcast are recorded, and the actual broadcast duration is calculated to be 31 minutes and 59 seconds. Step S2 identifies that the standard duration for the day is 32 minutes, and the difference between the actual broadcast duration and the standard duration is 1 second, which is within the preset error range (±5 seconds), and determines that the broadcast duration is normal.
[0046] The total number of broadcast video frames on that day was counted (3000 frames). Among them, 2985 frames were in the correct position for the host, accounting for 99.5%, which is not lower than the preset threshold (95%), and the broadcast was judged to be normal.
[0047] The overall assessment result is that the city-level television station's broadcast duration and broadcast image are normal, and the broadcast is deemed qualified for that day.
[0048] In another specific embodiment, when a county-level TV station broadcast, the actual start time was 19:00:15, the actual end time was 19:28:00, and the actual broadcast duration was 27 minutes and 45 seconds. Step S2 identified that the standard duration for the day was 30 minutes, and the difference between the actual broadcast duration and the standard duration was 2 minutes and 15 seconds, which exceeded the preset error range. At the same time, the percentage of frames in the video where the host's position was normal was 88%, which was lower than the preset threshold. Therefore, the broadcast was deemed unqualified, and the abnormality type was marked as "not starting on time, insufficient duration, and abnormal picture".
[0049] In some embodiments, the method further includes an anomaly warning mechanism, namely step S6. When a broadcast failure result is generated, the system immediately issues an audible and visual alarm to the monitoring personnel. At the same time, it automatically captures and retains audio and video clips, duration identification records, and abnormal frame data of the host's position during the abnormal period, archives and stores them to form a monitoring log, which serves as an objective basis for subsequent verification.
[0050] For example, in the event of the above-mentioned non-compliance, an audible and visual alarm will be immediately triggered to notify the monitoring personnel. At the same time, audio and video clips of the abnormal period (19:00:15-19:28:00), duration identification records (standard duration 30 minutes, actual duration 27 minutes and 45 seconds), and host position monitoring records (abnormal frame number and specific abnormal frame) will be retained. All monitoring data (including qualified and unqualified data) will be archived and stored for a period of one year to form a monitoring log for subsequent verification, rectification, and statistical analysis.
[0051] In some embodiments, the method further includes a model optimization mechanism, namely step S7. Specifically, to adapt to the evolution of program formats, the system periodically (usually monthly) collects new audio samples containing notifications of new program durations and new video frame samples containing footage of new hosts broadcasting. The new samples are used to iteratively train the audio duration recognition model and the host position recognition model, while simultaneously updating the customized hot word library and dynamically adjusting the host's standard position threshold according to the actual studio environment. This continuously improves the model's recognition accuracy, ensuring that the model adapts to subtle adjustments in the format of the news broadcast (such as minor adjustments to the host's broadcasting position and broadcasting script).
[0052] Other embodiments of the present invention provide a multi-dimensional news broadcast program rebroadcast monitoring device, which includes a data acquisition and preprocessing module, a duration recognition module, a screen monitoring module, a duration calculation module, and a comprehensive judgment module, for performing the corresponding functions of the above steps.
[0053] Specifically, the acquisition and preprocessing module is used to synchronously acquire audio and video signals of the target channel within a preset program broadcast time period, and preprocess the audio and video signals to separate audio signals and video frame sequences; the duration recognition module is used to input the audio signals into a pre-trained audio duration recognition model, extract the program duration value in the audio signals through speech recognition, and use it as the standard duration for the day; the image monitoring module is used to input the video frame sequence into a pre-trained host position recognition model for target detection, identify the host position coordinates in each frame, and compare the host position coordinates with a preset host standard position threshold to obtain the percentage of frames with normal positions; the duration calculation module is used to obtain the actual broadcast duration of the target channel; and the comprehensive judgment module is used to generate a broadcast monitoring judgment result for the target channel based on the difference between the actual broadcast duration and the standard duration for the day, as well as the percentage of frames with normal positions.
[0054] Furthermore, this disclosure also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the steps of the above-described multi-dimensional news broadcast program rebroadcast monitoring method.
[0055] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned related methods.
[0056] In this specification, the illustrative expressions of the terms used do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0057] Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this invention shall be included within the scope of protection of this invention.
Claims
1. A multi-dimensional method for monitoring the rebroadcast of CCTV's "News Broadcast" program, characterized in that, include: During the preset program broadcast time period, audio and video signals of the target channel are collected synchronously, and the audio and video signals are preprocessed to separate the audio signals and video frame sequences. The audio signal is input into a pre-trained audio duration recognition model, and the program duration value in the audio signal is extracted through speech recognition and used as the standard duration for the day. The video frame sequence is input into a pre-trained host position recognition model for target detection. The host position coordinates in each frame are identified, and the host position coordinates are compared with a preset host standard position threshold to obtain the percentage of frames with normal positions. Obtain the actual broadcast duration of the target channel; Based on the difference between the actual broadcast duration and the standard duration of the day, and the percentage of normal frames at the location, a broadcast monitoring and judgment result for the target channel is generated.
2. The multi-dimensional news broadcast monitoring method according to claim 1, characterized in that, The step of inputting the audio signal into a pre-trained audio duration recognition model, extracting the program duration value from the audio signal through speech recognition, and using it as the standard duration for the day includes: The audio signal is input into an audio duration recognition model that supports single-inference of long audio; The audio signal is matched and extracted using a pre-configured custom hot word library, which includes keywords related to program duration notifications. If a duration notification containing the keywords is extracted, the duration value in the duration notification is parsed and extracted as the standard duration for the day. If the duration notification information is not extracted, the preset default duration will be used as the temporary standard duration, and the temporary standard duration will be corrected based on the extracted program end characteristics to obtain the standard duration for the day.
3. The multi-dimensional news broadcast monitoring method according to claim 1, characterized in that, The preset host standard position threshold includes a horizontal position range, a vertical position range, and a position offset threshold; The step of inputting the video frame sequence into a pre-trained host position recognition model for target detection, identifying the host's position coordinates in each frame, and comparing the host's position coordinates with a preset host standard position threshold includes: The coordinates of the center point of the host's head in the current video frame are extracted using the host position recognition model. Determine whether the coordinates of the head center point fall within the horizontal position range and the vertical position range, and obtain the coordinate offset of the current video frame relative to the previous video frame, and determine whether the coordinate offset is less than the position offset threshold. If all conditions are met, the host's position in the current video frame is determined to be normal.
4. The multi-dimensional news broadcast monitoring method according to claim 1, characterized in that, The generation of a broadcast monitoring and judgment result for the target channel based on the difference between the actual broadcast duration and the standard duration of the day, and the percentage of normal frames at the location, includes: Obtain the actual start and end times of the program broadcast on the target channel, and calculate the actual broadcast duration. Calculate the absolute value of the difference between the actual broadcast duration and the standard broadcast duration for that day; If the absolute value of the difference is less than or equal to the preset error range, the current broadcast duration is determined to be normal. If the absolute value of the difference is greater than the preset error range, the current broadcast duration is determined to be abnormal, and an abnormality type marker in the duration dimension is generated.
5. The multi-dimensional news broadcast monitoring method according to claim 4, characterized in that, The step of generating a broadcast monitoring and judgment result for the target channel based on the difference between the actual broadcast duration and the standard duration of the day, and the percentage of normal frames at the location, further includes: If the percentage of frames with the specified location being normal is not lower than a preset percentage threshold, and there is no situation where the host's location coordinates are not identified for a consecutive preset number of frames, then the current broadcast is determined to be normal. If the percentage of normal frames at the specified location is lower than the preset percentage threshold, or if the host's location coordinates are not identified for a consecutive preset number of frames, then the current broadcast is determined to be abnormal, and an abnormality type marker for the frame dimension is generated. If the current broadcast duration and the current broadcast video are both normal, a broadcast monitoring pass result is generated for the target channel; otherwise, a broadcast failure result is generated and an early warning signal is triggered.
6. The multi-dimensional news broadcast monitoring method according to claim 1, characterized in that, The preprocessing of the audio and video signals to separate the audio signals and video frame sequences includes: The original audio was denoised using an adaptive noise reduction algorithm, and the audio was segmented after unifying the sampling rate to obtain the audio signal. The original video is denoising the separated video using a filtering algorithm. Video frames are extracted according to a preset frequency and normalized to a preset pixel size to obtain the video frame sequence.
7. The multi-dimensional news broadcast monitoring method according to claim 1, characterized in that, It also includes model optimization steps: Periodically collect new audio samples containing program duration notifications, as well as new video frame samples containing host broadcasts; The audio duration recognition model is iteratively trained using the newly added audio samples, and the customized hot word library is updated. The host position recognition model is iteratively trained using the newly added video frame samples, and the host standard position threshold is dynamically adjusted.
8. A multi-dimensional news broadcast program rebroadcast monitoring device, characterized in that, include: The acquisition and preprocessing module is used to synchronously acquire audio and video signals of the target channel during a preset program broadcast period, and to preprocess the audio and video signals to separate the audio signals and video frame sequences. The duration recognition module is used to input the audio signal into a pre-trained audio duration recognition model, extract the program duration value from the audio signal through speech recognition, and use it as the standard duration for the day. The video frame monitoring module is used to input the video frame sequence into a pre-trained host position recognition model for target detection, identify the host position coordinates in each frame, and compare the host position coordinates with a preset host standard position threshold to obtain the percentage of frames with normal positions. The duration calculation module is used to obtain the actual broadcast duration of the target channel broadcast; The comprehensive judgment module is used to generate a rebroadcast monitoring judgment result for the target channel based on the difference between the actual rebroadcast duration and the standard duration of the day, as well as the proportion of normal frames at the location.
9. An electronic device, characterized in that, include: A processor and a memory communicatively connected to the processor; The memory stores a computer program that can be executed by the processor, which, when executing the computer program, implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.