Broadcast audio quality discrimination method and device

The broadcast audio quality discrimination method, which uses multi-dimensional feature extraction and dynamic weighted fusion calculation, solves the problem that traditional monitoring methods are difficult to identify the causes of audio quality degradation, and achieves accurate assessment and fault tracing, thereby improving operation and maintenance efficiency and user experience.

CN121725822AActive Publication Date: 2026-03-24BEIJING HIZHI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Traditional methods for monitoring broadcast audio quality are insufficient to meet the needs of refined management in complex environments, and cannot accurately identify the causes of audio quality degradation, resulting in low operation and maintenance efficiency.

Method used

By employing multi-dimensional feature extraction and dynamic weighted fusion calculation methods, a comprehensive audibility score is generated. This score is then combined with user feedback for closed-loop optimization, enabling quality assessment and root cause localization, and outputting real-time alarm information and visualized data.

Benefits of technology

It enables accurate assessment of broadcast audio quality and fault tracing, improving operation and maintenance efficiency and user experience. It has high accuracy, robustness and scalability, and can achieve 24/7 unattended quality monitoring in complex networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725822A_ABST
    Figure CN121725822A_ABST
Patent Text Reader

Abstract

The invention relates to a broadcast audio quality discrimination method and device, and belongs to the technical field of broadcast audio monitoring, and the discrimination method comprises the steps: receiving and preprocessing a broadcast audio signal stream, and outputting a framed signal sequence; extracting multi-dimensional features of the framed signal sequence in parallel, and generating a multi-dimensional feature set comprising a signal strength feature, an interference feature, a noise feature and a propagation disturbance feature; performing dynamic weighted fusion calculation based on the multi-dimensional feature set to generate a comprehensive audibility score; performing quality discrimination and root cause positioning according to the comprehensive audibility score and the multi-dimensional feature set, and outputting a quality grade and a fault type identifier; and when the quality grade is lower than a preset grade threshold value or the multi-dimensional feature set satisfies an alarm triggering condition, generating real-time alarm information and visual data. According to the invention, accurate evaluation and accurate traceability of the broadcast audio quality are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of broadcast audio monitoring technology, and in particular to a method and apparatus for judging broadcast audio quality. Background Technology

[0002] As modern broadcasting systems rapidly evolve from traditional analog transmission towards digitalization, networking, and intelligence, the transmission channels for broadcast audio are becoming increasingly diversified, including FM, AM, digital audio broadcasting (DAB), and IP-based network streaming media, widely used in key scenarios such as public broadcasting, emergency communications, traffic information broadcasting, and media dissemination. In these applications, the quality of the audio signal directly affects the accuracy, audibility, and user experience of information transmission. Especially during emergency broadcasts or the release of important information, any audio distortion, interruption, or decreased intelligibility can lead to misinterpretation of information or even public safety risks. Therefore, establishing an efficient, accurate, and fault-based broadcast audio quality assessment mechanism has become a core technical requirement in the operation and maintenance management of modern broadcasting systems.

[0003] However, in real-world operating environments, broadcast audio signals are highly susceptible to interference and damage from various internal and external factors during transmission. Because broadcast links typically cover complex topologies involving long distances and multiple nodes, varying degrees of quality degradation can occur during transmission, relaying, wireless propagation, and reception. For example, in densely populated urban areas, the complex electromagnetic environment allows radio frequency interference from various electronic devices to intrude into the broadcast frequency band, causing co-channel or adjacent-channel interference, resulting in noise, howling, or signal aliasing in the audio. In remote or mountainous areas, terrain obstruction and changes in atmospheric conditions can easily trigger signal attenuation or multipath propagation effects, leading to echoes, ghosting, or discontinuities at the receiving end. Furthermore, in older equipment or low-quality links, problems such as amplifier nonlinear distortion, power supply frequency noise, and analog-to-digital conversion errors are common, further exacerbating audio quality instability.

[0004] Against this backdrop, traditional methods for monitoring broadcast audio quality are no longer sufficient to meet the demands of refined management in today's complex environment. Even if quality degradation is detected, the root cause cannot be identified, leading to maintenance personnel spending a significant amount of time on manual troubleshooting and resulting in low response efficiency. Summary of the Invention

[0005] In order to achieve accurate assessment and precise traceability of broadcast audio quality, this application provides a method and apparatus for determining broadcast audio quality.

[0006] Firstly, this application provides a method for judging broadcast audio quality, which adopts the following technical solution: A method for judging broadcast audio quality, the method comprising: It receives broadcast audio signal streams, performs preprocessing, and outputs a framed signal sequence. Multi-dimensional features are extracted in parallel from the framed signal sequence to generate a multi-dimensional feature set containing signal strength features, interference features, noise features, and propagation disturbance features; Based on the multi-dimensional feature set, a dynamic weighted fusion calculation is performed to generate a comprehensive audibility score; Based on the comprehensive audibility score and the multi-dimensional feature set, quality judgment and root cause localization are performed, and the quality level and fault type identifier are output. When the quality level is lower than a preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions, real-time alarm information and visualization data are generated.

[0007] By adopting the above technical solutions, anomaly alarms and visual outputs are achieved, forming a fully automated processing capability from perception to decision-making. Different physical mechanisms of broadcast quality degradation are decoupled through multi-dimensional features, and a dynamic weighting mechanism is introduced to improve the scenario adaptability of the scoring model. Combined with contribution analysis, accurate root cause tracing is achieved, overcoming the limitations of traditional single-indicator evaluation. The entire system not only possesses high accuracy and robustness but also good scalability, enabling 24 / 7 unattended quality monitoring in complex broadcast networks, improving the operational efficiency and user experience of broadcast systems.

[0008] Optionally, the step of extracting multi-dimensional features in parallel from the framed signal sequence to generate a multi-dimensional feature set including signal strength features, interference features, noise features, and propagation disturbance features includes: Receive the signal sequence after frame processing; Feature extraction is performed in parallel on each frame of the signal sequence; Among them, the signal strength feature is calculated as the first feature value, the interference feature is identified as the second feature value, the noise feature is measured as the third feature value, and the propagation disturbance feature is evaluated as the fourth feature value. Aggregate the first, second, third, and fourth feature values ​​of all frames to generate a multi-dimensional feature set that includes signal strength feature set, interference feature set, noise feature set, and propagation disturbance feature set.

[0009] By adopting the above technical solutions, independent and coordinated quantitative characterization of the four core degradation mechanisms of signal strength, interference, noise and propagation disturbance is achieved, breaking through the limitations of traditional single-dimensional monitoring and enhancing the intelligence level of broadcast audio quality assessment and operation and maintenance guidance capabilities.

[0010] Optionally, the step of generating a comprehensive audibility score by performing dynamic weighted fusion calculation based on the multi-dimensional feature set includes: Normalize the feature values ​​of the multi-dimensional feature set to output a unified scale feature set; Select the weight allocation method according to the preset scenario strategy, and generate a set of feature weight coefficients; Based on the unified scale feature set and the feature weight coefficient set, a dynamic weighted fusion calculation is performed to output a frame-level audibility score sequence. The frame-level audibility score sequence is aggregated to generate a comprehensive audibility score.

[0011] By adopting the above technical solutions, a refined and psychoacoustic-based objective evaluation of broadcast audio quality is achieved. The traditional static weighting model is upgraded to a "basic + dynamic" dual-mode weighting mechanism, and a non-linear normalization strategy is combined to improve the consistency of feature perception, overcoming the technical bottleneck that a single weight setting cannot handle complex scenarios. By automatically optimizing the scoring logic under different broadcast environments and operational needs, the intelligence level and practical value of audio quality assessment are significantly improved.

[0012] Optionally, the following steps may be included after generating the overall audibility score: Receive subjective audio quality ratings from user terminals; The deviation value is calculated between the subjective audio quality score data and the comprehensive audibility score for the corresponding time period. When the deviation value continues to exceed the optimization trigger threshold, the multi-dimensional feature set of the time period is extracted; Based on the mapping relationship between the multi-dimensional feature set and the subjective audio quality rating data, the weight allocation strategy of the dynamic weighted fusion calculation is updated.

[0013] By adopting the above technical solution, a closed-loop optimization mechanism driven by user subjective feedback is introduced into the traditional broadcast audio quality assessment process, constructing a complete adaptive system of "perceptual correction—feature tracing—weight evolution." By receiving real user scores and performing spatiotemporal alignment and deviation analysis with the system scores, the system can accurately identify systematic deviations between the algorithm and human auditory perception; by setting continuous triggering conditions and extracting associated feature sets, it effectively distinguishes between accidental fluctuations and structural defects; finally, based on data-driven mapping relationships, it dynamically adjusts the weight allocation strategy, enabling the assessment model to have environmental adaptability. This application's solution not only improves the psychoacoustic accuracy of the comprehensive audibility score but also significantly reduces the frequency of manual parameter tuning, enhancing the system's robustness in dealing with complex and ever-changing broadcast environments.

[0014] Optionally, the steps of performing quality assessment and root cause localization based on the comprehensive audibility score and the multi-dimensional feature set, and outputting the quality level and fault type identifier, include: The comprehensive audibility score is compared with a preset grading threshold, and a quality level identifier is output. The contribution measure of each feature is calculated based on the multi-dimensional feature set and the feature weight coefficient set, and a contribution sequence is generated. Identify the dominant contribution features in the contribution sequence, map them to a preset fault type library, and output fault type identifiers.

[0015] By adopting the above technical solution, quality grading is achieved based on a comprehensive audibility score. A contribution measurement model reveals the influence weight of each degradation factor, and the results are mapped to a pre-defined fault type library based on the principle of maximum contribution, ultimately forming a structured diagnostic result containing both quality level and fault type. This technical solution not only accurately identifies quality degradation states but also scientifically locates the root causes, improving the intelligent operation and maintenance level of the broadcasting system and providing strong support for ensuring broadcast service quality.

[0016] Optionally, when the quality level is lower than a preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions, the step of generating real-time alarm information and visualization data includes: Receive quality level identifiers, fault type identifiers, and multi-dimensional feature sets; When the quality level identifier is lower than the preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions, an alarm message containing a fault type identifier, timestamp, and feature summary is assembled. A visual data stream is generated based on the multi-dimensional feature set and fault type identifier, including a scoring time-series curve and a fault location map. The alarm messages and visualized data streams are distributed to the target management terminal, and the associated data is stored in the diagnostic log library.

[0017] By adopting the above technical solutions, the system can not only accurately identify quality problems, but also clearly reveal their causes, scope of impact, and spatiotemporal distribution characteristics. It also ensures efficient information transmission through structured messages and hierarchical push mechanisms, and provides a complete data chain for subsequent analysis by combining time-series log storage, thereby improving the intelligent operation and maintenance capabilities and service quality assurance level of the broadcasting system.

[0018] Optionally, the discrimination method further includes: Receive the quality level and fault type identifiers from multiple broadcast nodes; Associate the location coordinates of the multiple broadcast nodes in a preset geographic information system; Based on the quality level and fault type identifier, calculate the spatial impact range of the fault node; The real-time alarm information is prioritized based on the spatial influence range, and a hierarchical alarm distribution sequence is output.

[0019] By adopting the above technical solutions, this approach deeply integrates broadcast quality diagnosis and geospatial analysis. It achieves wide-area monitoring data aggregation by receiving quality and fault identifiers from multiple nodes; assigns precise spatial attributes to fault information through GIS coordinate association; quantifies the impact range of broadcast faults by dynamically calculating coverage radius and user density values; and generates hierarchical alarm sequences through a multi-factor comprehensive decision-making model, driving a shift in operation and maintenance strategies from passive response to proactive optimization. This solution not only improves the accuracy of fault location and the timeliness of response but also upgrades traditional audio quality monitoring into a smart operation and maintenance platform with spatial cognition, impact prediction, and intelligent sorting capabilities, providing strong technical support for ensuring the high availability of modern broadcast systems.

[0020] Secondly, this application provides a broadcast audio quality discrimination device, which adopts the following technical solution: A broadcast audio quality discrimination device, the discrimination device comprising: The preprocessing module is used to receive the broadcast audio signal stream, perform preprocessing, and output the signal sequence after frame processing. The multidimensional feature extraction module is used to extract multidimensional features in parallel from the signal sequence after the frame processing, and generate a multidimensional feature set including signal strength features, interference features, noise features and propagation disturbance features; The weighted fusion module is used to perform dynamic weighted fusion calculation based on the multi-dimensional feature set to generate a comprehensive audibility score. The quality discrimination module is used to perform quality discrimination and root cause localization based on the comprehensive audibility score and the multi-dimensional feature set, and output the quality level and fault type identifier. The alarm module is used to generate real-time alarm information and visualization data when the quality level is lower than a preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions.

[0021] Thirdly, this application provides a computer device, which adopts the following technical solution: A computer device includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to perform the steps of the method as described in the first aspect.

[0022] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect. Attached Figure Description

[0023] Figure 1 This is a first flowchart illustrating a broadcast audio quality discrimination method according to one embodiment of this application.

[0024] Figure 2 This is a second flowchart illustrating a broadcast audio quality discrimination method according to one embodiment of this application.

[0025] Figure 3 This is a schematic diagram of the third process of a broadcast audio quality discrimination method according to one embodiment of this application.

[0026] Figure 4 This is a schematic diagram of the fourth process of a broadcast audio quality discrimination method according to one embodiment of this application.

[0027] Figure 5 This is a schematic diagram of the fifth process of a broadcast audio quality discrimination method according to one embodiment of this application.

[0028] Figure 6 This is a schematic diagram of the sixth process of a broadcast audio quality discrimination method according to one embodiment of this application.

[0029] Figure 7 This is a schematic diagram of the seventh process of a broadcast audio quality discrimination method according to one embodiment of this application. Detailed Implementation

[0030] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figure 1-7 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0031] This application discloses a method for judging the quality of broadcast audio.

[0032] Reference Figure 1 A method for judging broadcast audio quality, the method comprising: Step S101: Receive the broadcast audio signal stream and perform preprocessing, then output the signal sequence after frame processing; Broadcast audio signal streams typically originate from transmission channels such as FM, AM, DAB, or network streaming media, and are presented as continuous analog or digital audio signals. During transmission, these signals may be affected by various environmental factors and system defects, such as electromagnetic interference, multipath effects, equipment aging, and link interruptions, leading to a decline in audible quality. Therefore, real-time acquisition of broadcast audio signal streams is not only fundamental to quality monitoring but also a prerequisite for subsequent intelligent diagnostics. Signal stream acquisition must ensure temporal continuity and high fidelity to avoid additional errors introduced by insufficient sampling rate or quantization distortion. Generally, standard sampling frequencies (such as 44.1kHz or 48kHz) are used for digital sampling, and synchronization mechanisms ensure time alignment of signal frames, providing a stable and consistent data source for subsequent processing.

[0033] Furthermore, the core objective of preprocessing is to eliminate interference from non-speech content and standardize the signal morphology, thereby creating favorable conditions for subsequent framing and feature extraction. First, noise reduction filtering aims to remove background noise from the signal that is unrelated to the broadcast speech, such as power system frequency interference (50 / 60Hz), impulse noise, or broadband white noise.

[0034] In some embodiments, commonly used filtering techniques include wavelet transform denoising, adaptive filtering (such as the LMS algorithm), or speech enhancement methods based on spectral subtraction. These methods can effectively suppress steady-state and transient noise while preserving the main spectral structure of the speech. Secondly, amplitude normalization compresses the dynamic range of the filtered signal to a uniform interval (e.g., [-1,1] or [0,1]) to prevent amplitude differences caused by different signal sources or gain settings from affecting the consistency comparison of subsequent features. Normalization not only improves the stability of feature extraction but also avoids the problem of strong signals masking weak features. Finally, framing processing divides the continuous signal into discrete frames of short time intervals, typically 20-50 milliseconds long (e.g., 25ms), with frames overlapping (e.g., 10ms overlap) to balance time resolution and spectral stability. Each frame is considered an approximately stationary signal, facilitating frequency domain analysis. Simultaneously, to support time series analysis and alarm tracing, each frame is appended with a precise timestamp, forming a time-stamped sequence of framed signals.

[0035] Step S102: Extract multi-dimensional features in parallel from the signal sequence after frame processing to generate a multi-dimensional feature set containing signal strength features, interference features, noise features and propagation disturbance features; Parallel extraction means that multiple feature channels operate independently and simultaneously without interference, which improves computational efficiency while ensuring the orthogonality and independence of each feature, avoiding misjudgments caused by information coupling. The four types of extracted core features characterize the quality degradation mechanism of broadcast audio from different physical dimensions.

[0036] Specifically, signal strength features reflect the strength of audio energy, typically characterized by calculating the energy (i.e., the sum of squares) of each frame of signal and converting it to a logarithmic scale (e.g., dB). Low intensity often corresponds to signal attenuation at the receiver or insufficient transmission power. Interference features focus on external electromagnetic or co-channel signal intrusion. By performing spectral analysis on each frame of signal and matching it with preset interference templates (such as known adjacent-channel interference, harmonic components, or illegal broadcast frequencies), the presence and intensity of abnormal spectral components are identified. Noise features focus on the degree of background noise intrusion. Typically, noise components in silent segments are separated using voice activity detection (VAD) to estimate the signal-to-noise ratio (SNR). Low SNR indicates impaired speech clarity. Propagation disturbance features target the multipath effect unique to wireless propagation. Autocorrelation functions or short-time Fourier transforms are used to detect the presence of delayed echoes, and the time difference and energy ratio between the main signal and the reflected signal are calculated to determine whether there is severe distortion or "ghosting" phenomenon. These four features together constitute a quality feature set covering the entire "signal-interference-noise-propagation" link, which can comprehensively reflect the degradation path of broadcast audio during transmission.

[0037] Step S103: Perform dynamic weighted fusion calculation based on multi-dimensional feature set to generate comprehensive audibility score; Since the dimensions, distribution range, and degree of influence on subjective listening experience of various features differ, direct addition can lead to certain features dominating the scoring result. Therefore, normalization is necessary to map all feature values ​​to the same numerical range (such as [0,1] or [-1,1]) to make them comparable. After normalization, the system selects an appropriate weight allocation strategy based on the current application scenario to achieve differentiated feature fusion.

[0038] In the embodiments of this application, the basic weight model uses fixed coefficient allocation. For example, signal strength and noise features are given higher weights (0.2–0.4) because they have a significant impact on human auditory perception; while interference and propagation disturbance features are given slightly lower weights (0.1–0.3), which are suitable for stability evaluation in conventional broadcast environments. However, in complex and ever-changing real-world scenarios, fixed weights are difficult to adapt to different channel conditions or user needs. Therefore, a dynamic weight model based on machine learning can be introduced as an advanced mode.

[0039] Specifically, this dynamic weighting model uses a pre-trained nonlinear classifier such as a random forest as its core. It takes the complete feature set of the current frame as input and outputs the optimal weight combination for the given scene. It can automatically learn the relative importance of different features in specific environments. For example, in densely populated urban areas where multipath propagation is more severe, the system can automatically increase the weight of propagation interference features. The fusion process typically uses a weighted linear combination method to generate frame-level audibility scores, and then averages the scores of all frames to obtain the final comprehensive audibility score. This score can be considered an objective quantification of the subjective listening quality of the entire audio segment.

[0040] Step S104: Based on the comprehensive audibility score and multi-dimensional feature set, perform quality judgment and root cause localization, and output quality level and fault type identifier; The quality assessment first compares the overall score with preset grading thresholds to classify it into three levels: "Excellent," "Medium," and "Poor," providing maintenance personnel with an intuitive indication of the quality status. However, knowing only the quality level is insufficient to guide repairs; therefore, the system also needs to perform root cause analysis, that is, identify the fundamental reasons causing the quality degradation.

[0041] Therefore, the system calculates the contribution of each feature to the overall score. The contribution is determined by the product of the feature weight coefficient and (1 - normalized feature value): the higher the weight and the worse the feature's performance (the smaller the normalized value), the greater its negative impact on the overall score. For example, if the normalized value of a noise feature in an audio segment is only 0.2, and its weight is 0.35, then its contribution is 0.35 × (1 - 0.2) = 0.28, significantly higher than other features, indicating that noise is the main cause of the current quality problem. Based on this, the system marks "noise exceeding the standard" as a fault type identifier.

[0042] Furthermore, by associating the audio signal with the source device information (such as transmitter number, relay station ID, or IP stream address), the system can also output a specific responsible device identifier, enabling problem tracing back to the physical node. This three-level output mechanism of "scoring + attribution + location" not only improves the transparency of diagnosis but also provides a basis for decision-making in automated operation and maintenance.

[0043] Step S105: When the quality level is lower than the preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions, generate real-time alarm information and visualization data.

[0044] The alarm mechanism is designed to handle both steady-state degradation and sudden anomalies: on the one hand, long-standing quality problems (such as interference characteristics consistently exceeding the first threshold for a certain period) will be detected to prevent chronic faults from being overlooked; on the other hand, short-term drastic fluctuations (such as sudden changes in the overall audibility score between adjacent frames exceeding the threshold) can also be responded to promptly, applicable to transient faults such as popping sounds, intermittent noises, and frequency hopping. The propagation interference characteristics exceeding the second threshold specifically targets severe multipath problems in wireless links. Once an alarm is triggered, the system not only pushes text or voice alarm information to the monitoring platform but also generates multi-dimensional visual data to assist in manual analysis.

[0045] Specifically, the multidimensional visualization data includes: a trend curve of the overall audibility score over time, helping to observe the quality evolution process; a heatmap matrix composed of various feature values, showing the distribution patterns of different features in the time-frequency dimension, facilitating the detection of periodic interference or noise accumulation; and highlighted markers of faulty devices in the broadcast link topology diagram, intuitively presenting the location of the problem. These visualization methods greatly enhance the interpretability and user-friendliness of the system, enabling technicians to quickly grasp the overall picture of the problem without delving into the underlying data.

[0046] The above implementation achieves anomaly alarms and visual output, forming a fully automated processing capability from perception to decision-making. By decoupling different physical mechanisms of broadcast quality degradation through multi-dimensional features and introducing a dynamic weighting mechanism to improve the scenario adaptability of the scoring model, combined with contribution analysis, accurate root cause tracing is achieved, overcoming the limitations of traditional single-indicator evaluation. The entire system not only possesses high accuracy and robustness but also good scalability, enabling 24 / 7 unattended quality monitoring in complex broadcast networks, improving the operational efficiency and user experience of broadcast systems.

[0047] Reference Figure 2 As one implementation of step S102, the step of extracting multi-dimensional features in parallel from the framed signal sequence to generate a multi-dimensional feature set containing signal strength features, interference features, noise features, and propagation disturbance features includes: Step S201: Receive the signal sequence after frame processing; The signal sequence originates from the broadcast audio stream after preprocessing stages including noise reduction, normalization, and framing. Each frame typically represents a short speech segment of 20 to 50 milliseconds, with a timestamp added to maintain temporal continuity. Because broadcast audio exhibits non-stationary characteristics during transmission—meaning its statistical properties change over time—direct global analysis of the entire signal is insufficient to accurately capture local anomalies. Therefore, it is necessary to divide the signal into several short time frames, assuming approximate stationarity within each frame to satisfy the fundamental premises of frequency domain analysis and feature modeling.

[0048] Understandably, this framing structure not only improves temporal resolution but also enables the system to accurately identify transient interference or sudden distortions. Simultaneously, the framed signal eliminates significant interference from amplitude fluctuations and background noise, ensuring stable and reliable input quality for feature extraction.

[0049] Step S202: Perform feature extraction operations in parallel on each frame of the signal sequence; Among them, the signal strength feature is calculated as the first feature value, the interference feature is identified as the second feature value, the noise feature is measured as the third feature value, and the propagation disturbance feature is evaluated as the fourth feature value. Specifically, the signal strength feature is calculated as the first feature value to quantify the energy level of the current frame's audio, reflecting the strength of the broadcast signal at the receiver. The system extracts the time-domain energy value of the signal within the frame, which is the sum of the squares of the amplitudes of all sampling points. This measure is directly related to the perceived loudness of the signal. This energy value is then converted to a logarithmic scale (e.g., in decibels, dB) to match the non-linear perception of sound intensity by the human ear; small changes are more sensitive at low volumes and relatively less sensitive at high volumes. Through logarithmic transformation, the feature value's numerical distribution is made closer to subjective listening perception, enhancing its correlation with human audibility assessment. For example, when a frame's signal suffers severe attenuation due to long-distance transmission, its energy decreases significantly, and the first feature value after logarithmic transformation will also decrease significantly, effectively identifying the risk of signal weakening.

[0050] Simultaneously, the system performs spectral component analysis on the same frame signal and matches it with interference templates to generate a second feature value, used to identify interference problems caused by external electromagnetic or co-frequency signal intrusion. Broadcast bands often face electromagnetic interference from industrial equipment, illegal radio stations, or other communication systems. This interference often manifests as abnormal energy accumulation at specific frequencies, such as power frequency harmonics (50Hz and its multiples), adjacent channel crosstalk, or bursty pulse interference. To accurately identify this type of interference, the system first performs frequency domain transformation (e.g., Fast Fourier Transform, FFT) on the frame signal to generate a spectrum, showing the energy distribution of the signal at different frequencies. Then, it performs kurtosis matching on this spectrum with a pre-set interference template library, which stores the spectral characteristics of known typical interference patterns, such as fixed frequency peak values, bandwidth characteristics, and energy distribution patterns. Through correlation analysis or pattern recognition algorithms (e.g., Dynamic Time Warping (DTW) or Euclidean distance comparison), it determines whether the current spectrum is highly similar to a certain template, and determines the type of interference and its energy proportion based on the matching results. For example, if a continuous high-energy narrowband signal is detected near 107.9MHz and it successfully matches an illegal broadcast template, the system determines it to be co-channel interference and generates a second characteristic value based on the proportion of its energy to the total spectrum energy. The lower the value, the more severe the interference.

[0051] In another parallel path, the system generates a third eigenvalue by separating background noise and quantifying the signal-to-noise ratio (SNR) to assess the degree of noise intrusion in broadcast audio. Background noise may originate from the receiving environment (such as traffic noise or wind noise) or the electronic equipment itself (such as amplifier thermal noise), and its presence can mask speech details and reduce speech clarity. To accurately estimate the noise components, the system employs spectral subtraction, a classic speech enhancement technique: first, the noise spectrum is estimated in silent frames without speech activity, and then it is subtracted from the spectrum of the current frame to restore the "clean" speech components. Based on this, the variance ratio of the total energy of the original signal to the energy of the noise components is calculated and converted into a decibel value as an SNR indicator. The higher this third eigenvalue, the clearer the speech is relative to the noise; conversely, a lower value indicates severe noise pollution. The advantage of spectral subtraction lies in its good estimation ability for steady-state noise and the fact that it does not require additional sensors or reference signals, making it suitable for practical broadcast monitoring scenarios. Furthermore, by dynamically updating the noise model, the system can adapt to changes in environmental noise, improving long-term operational stability. For example, in a nighttime urban environment, the overall ambient noise is reduced, and the system can automatically adjust the noise floor to avoid misinterpreting low-volume programs as high noise.

[0052] In addition, the system detects multipath propagation components to generate a fourth feature value, specifically designed to assess propagation interference caused by reflected waves in radio broadcasting. In complex electromagnetic environments such as densely populated urban areas or mountainous regions, broadcast signals often experience multiple reflections due to obstacles such as buildings and mountains, creating multiple signal copies with slightly different arrival times at the receiver—a phenomenon known as multipath effect. These delayed echoes, when superimposed on the main signal, can cause frequency-selective fading, phase distortion, and even speech ghosting, severely impacting the auditory experience. To identify such problems, the system calculates the autocorrelation function sequence of the frame signal, which reflects the similarity between the signal and itself at different time delays. Under ideal single-path propagation, the autocorrelation function shows a significant peak only at zero delay; however, when multipath is present, additional peaks appear at non-zero delay locations. The system sets a threshold delay (e.g., 1 millisecond, corresponding to approximately 300 meters of reflection path difference), identifies peak components exceeding this threshold, and analyzes their energy proportion. If the energy of a certain peak exceeds the proportion of the main peak energy exceeding a preset threshold, it is determined to be significant multipath interference, and a fourth feature value is generated accordingly. The smaller this value, the more severe the multipath effect. For example, at tunnel exits or in areas with tall buildings, multiple strong reflection paths are often observed, resulting in multiple long-distance peaks in the autocorrelation function. The system can then output a low fourth eigenvalue, indicating a serious defect in the propagation link.

[0053] Step S203: Aggregate the first feature value, second feature value, third feature value and fourth feature value of all frames to generate a multi-dimensional feature set including signal strength feature set, interference feature set, noise feature set and propagation disturbance feature set.

[0054] Each feature set retains the temporal order of the original frames, forming a time series structure that facilitates subsequent trend analysis, abrupt change detection, or sliding window statistics. For example, the signal strength feature set can be used to plot signal fading curves, the interference feature set can reveal the occurrence patterns of periodic interference, the noise feature set can reflect diurnal noise variations, and the propagation disturbance feature set can help locate periods of high multipath occurrence.

[0055] The above implementation method achieves independent and coordinated quantitative characterization of the four core degradation mechanisms: signal strength, interference, noise, and propagation disturbance. It breaks through the limitations of traditional single-dimensional monitoring and enhances the intelligence level and operation and maintenance guidance capabilities of broadcast audio quality assessment.

[0056] Reference Figure 3 As one implementation of step S103, the step of generating a comprehensive audibility score by performing dynamic weighted fusion calculation based on a multi-dimensional feature set includes: Step S301: Normalize the feature values ​​of the multi-dimensional feature set and output a unified scale feature set; Among them, since the original value range and variation law of various features are significantly different, for example, the signal strength may be distributed in the range of [-30,0] in decibels, while the interference feature may be expressed in the form of energy percentage between [0,100%]. If they are directly involved in the weighted calculation, it will inevitably cause some features to dominate the scoring results, which violates the principle of fair integration.

[0057] Therefore, this solution adopts a differentiated normalization strategy to customize the processing based on the physical characteristics and perceptual sensitivity of different types of features: For the signal strength feature set and noise feature set, both are approximately linearly related to speech intelligibility, that is, the higher the intensity or the lower the noise, the more significant the improvement in hearing. Therefore, a linear normalization method is used to map them to the first numerical interval (such as [0,1]) to ensure the consistency of the change trend; For the interference feature set and propagation disturbance feature set, their impact on subjective hearing often presents a nonlinear threshold effect, that is, when the interference is small, it is almost imperceptible, but once it exceeds a certain threshold, the hearing experience deteriorates rapidly. Therefore, a nonlinear function (such as Sigmoid or logarithmic compression) is used to transform the original value to the second numerical interval (such as [0,1]), so that weak interference is moderately suppressed, while severe interference is significantly amplified, which is more in line with the perceptual characteristics of the human ear.

[0058] Step S302: Select the weight allocation method according to the preset scenario strategy and generate a set of feature weight coefficients; The core objective of weight allocation is to determine the relative importance of the four types of features in the final score. This importance is not fixed but dynamically adjusted according to the application scenario.

[0059] In this application embodiment, the system provides two modes for selection: one is to call the pre-configured basic weight model, which is suitable for stability evaluation in a normal broadcast environment. This model sets a fixed weight ratio based on long-term operation and maintenance experience. For example, signal strength and noise features are given higher weights (0.2–0.4) because they have a significant impact on speech clarity, while interference and propagation disturbance features are given slightly lower weights (0.1–0.3) because they occur less frequently.

[0060] Secondly, it utilizes a machine learning model to achieve dynamic weight adjustment, making it particularly suitable for complex and ever-changing real-world operating environments. This machine learning model, based on a random forest algorithm and trained on a large amount of historical broadcast data, takes a unified scale feature set of the current audio segment as input and outputs the optimal weight combination adapted to the current scenario. For example, in densely populated urban areas where multipath effects occur frequently, the model can automatically increase the weight of propagation interference features; while in industrial areas where electromagnetic interference is severe, the weight of interference features will be increased. This data-driven dynamic weight mechanism overcomes the limitations of traditional fixed weights, enabling the scoring system to have self-learning and environmental adaptability, thus improving the accuracy and robustness of the scoring results.

[0061] Step S303: Perform dynamic weighted fusion calculation based on the unified scale feature set and feature weight coefficient set to output a frame-level audibility score sequence; Specifically, the process iterates through the entire audio segment frame by frame, performing a weighted summation operation on each frame: the four normalized feature values ​​corresponding to the frame are multiplied by their respective weight coefficients and then summed to obtain the audibility score for that frame. This score can be regarded as an objective quantitative expression of the subjective listening quality of the broadcast audio within that short period of time; a higher score indicates a better listening experience.

[0062] Understandably, because each frame participates in the calculation independently, the system can capture local fluctuations in audio quality, such as sudden pops, brief interruptions, or transient anomalies like multipath jumps, avoiding the smoothing masking problem caused by averaging across the entire segment. Simultaneously, the frame-level scoring sequence retains temporal information, forming a time-series curve reflecting the trend of audibility changes, providing data support for subsequent abrupt change detection, trend analysis, and alarm triggering.

[0063] For example, if a frame experiences a sharp drop in interference feature values ​​due to strong interference, and these values ​​have a high weight, its frame-level score will be significantly lower than the preceding and following frames, forming a clear "valley" and indicating the presence of a transient degradation event. This fine-grained scoring mechanism not only improves the system's response sensitivity but also provides a temporal anchor for root cause localization.

[0064] Step S304: Perform aggregation operation on the frame-level audibility score sequence to generate a comprehensive audibility score.

[0065] The purpose of aggregation is to condense the audibility of the entire audio segment into a representative value, facilitating overall quality rating and horizontal comparison. Common aggregation methods include arithmetic mean, weighted average, or median fusion. The optimal solution combines the advantages of the mean and median, reflecting the overall average level while suppressing interference from extreme abnormal frames.

[0066] For example, in a 10-second broadcast audio clip, if the vast majority of frames have a score consistently above 0.8, with only a few frames dropping below 0.3 due to transient interference, using the median or truncated mean can prevent these outliers from excessively lowering the overall score, thus more accurately reflecting the listener's actual listening experience. The comprehensive audibility score, as the final output of this method, can not only be used to determine the current quality level of the broadcast program (e.g., excellent / average / poor), but also as historical data for trend analysis, equipment performance evaluation, or service quality assessment.

[0067] The above implementation achieves a refined and psychoacoustic-based objective evaluation of broadcast audio quality. It upgrades the traditional static weighting model to a "basic + dynamic" dual-mode weighting mechanism, and combines it with a non-linear normalization strategy to improve feature perception consistency, overcoming the technical bottleneck that a single weight setting cannot handle complex scenarios. By automatically optimizing the scoring logic under different broadcast environments and operational needs, it significantly improves the intelligence level and practical value of audio quality assessment.

[0068] Reference Figure 4 As a further implementation of the broadcast audio quality assessment method, after the step of generating a comprehensive audibility score, the method further includes: Step S401: Receive subjective audio quality rating data from the user terminal; In this process, user terminals collect subjective ratings from users in actual listening environments through dedicated apps, smart radios, or embedded SDKs. These ratings typically use the MOS (Mean Opinion Score) standard, a semantic scoring system from 1 to 5, where 1 represents "very poor" and 5 represents "excellent." These ratings are not isolated data points but are uploaded along with rich metadata, including precise timestamps, device models, network status, geographical location, and environmental noise levels, ensuring that each rating can be traced back to a real user experience in a specific time and space context.

[0069] Furthermore, to ensure data quality, the system performs a rigorous cleaning and screening process upon receiving data. For example, it uses clustering algorithms such as DBSCAN to identify and remove abnormal behavior patterns, such as invalid data where a single user submits the exact same extremely low scores consecutively within a short period, or where the score distribution deviates significantly from the group trend. More importantly, the system uses NTP (Network Time Protocol) to synchronize the time between user devices and the central server, precisely aligning subjective scores to the timeline of the comprehensive audibility score output by the preceding modules, avoiding mismatches caused by device clock drift.

[0070] Step S402: Calculate the deviation value between the subjective audio quality score data and the comprehensive audibility score for the corresponding time period; Among them, the comprehensive audibility score is usually output continuously at a fixed period (such as generating a batch every 5 minutes), which has a high time density; while the user subjective score shows sparsity and non-uniform distribution characteristics, and may have only a few scores or even no scores at certain times.

[0071] To this end, this application employs Dynamic Time Warping (DTW) technology to align two non-equal length sequences. By elastically stretching or compressing the time axis, the optimal matching path is found, thereby accurately associating the audio content segments corresponding to the user rating and the system rating.

[0072] Subsequently, the system calculates the deviation between the mean subjective score and the mean overall audibility score for each time period. Since MOS is based on a 5-point scale while the system score is based on a percentage or other normalized scale, a scaling factor k is introduced for unified mapping. The final deviation δ is defined as the absolute difference between the two. For example, if the average user MOS is 3.6 (equivalent to 72 points) within a 5-minute time period, and the system score is 85 points, if the scaling factor k = 0.8, then the deviation is |72−85×0.8| = 8.

[0073] It should be noted that the system can also introduce a perception compensation factor to compensate for user ratings in high-noise environments (such as subway stations and shopping malls). This is because background noise reduces speech intelligibility in such environments, and even if the broadcast signal itself is of acceptable quality, users tend to give lower subjective ratings. If such ratings are not adjusted for environmental adaptability, the system may mistakenly attribute low user ratings to its own high ratings, thereby misjudging that the model has a systematic bias and triggering unnecessary weight updates.

[0074] Step S403: When the deviation value continues to exceed the optimization trigger threshold, extract the multi-dimensional feature set of the time period; The system does not respond to a single instance of exceeding the limit; instead, it employs a dual-judgment logic to avoid false triggering caused by momentary anomalies. The system sets an optimized trigger threshold θ (e.g., 10 minutes) and requires that the deviation δ > θ be satisfied for N consecutive time periods (N≥3) before it can be determined as a systematic deviation.

[0075] To enhance robustness, the system often employs a sliding window mechanism for verification. For example, a window length of 5 time periods is set; if at least 4 of these time periods exceed the limit, a persistent scoring deviation is confirmed. This design effectively filters out temporary deviations caused by individual user errors, brief signal interruptions, or sudden changes in the local environment, ensuring that model adjustments are only initiated when the overall algorithm performance deviates from user perception over a prolonged period. Once the triggering condition is met, the system immediately extracts the complete multi-dimensional feature set for the time period exceeding the limit, including time-series data of four types of features: signal strength, interference, noise, and propagation disturbance.

[0076] Furthermore, the system extends the extraction of feature data from preceding and following buffer periods (e.g., moving forward one period and backward one period) to capture the gradual trend and recovery process before the fault occurs, forming a more complete "problem window" feature matrix. For example, if deviations exceed the standard for three consecutive periods from 12:00 to 12:15, feature data from six periods—11:45 to 12:30—are actually extracted for subsequent attribution analysis. This extended data acquisition strategy helps reveal potential patterns hidden in the time series, such as slowly increasing noise or periodic interference, providing a more comprehensive contextual basis for weight updates.

[0077] Step S404: Based on the mapping relationship between the multi-dimensional feature set and the subjective audio quality rating data, update the weight allocation strategy for dynamic weighted fusion calculation.

[0078] This process is essentially a supervised learning-driven weighting mechanism. Its goal is to analyze which features contribute the most to subjective scoring bias and then adjust their influence weight in weighted fusion so that future scores are closer to the real listening experience.

[0079] Specifically, the system first constructs a feature-rating regression model, using a multi-dimensional feature set as input variables and the average user subjective rating as the target label. It then trains machine learning models such as random forests and utilizes its built-in feature importance assessment function to quantify the contribution of the four types of features to the rating prediction error. For example, if the analysis reveals that the importance of the noise feature is as high as 0.52, far exceeding that of other features, it indicates that the current system is not sensitive enough to noise, leading to inflated ratings in high-noise scenarios. Accordingly, the system dynamically adjusts the weight coefficients of each feature according to a preset weight update rule, combined with the learning rate and decay coefficient, increasing the weight of the noise feature from 0.3 to 0.396, thus enhancing its influence in subsequent ratings.

[0080] Subsequently, the newly generated set of weight coefficients will be packaged into a strategy package and pushed to the front-end dynamic weighted fusion module. The strategy will then be verified through an A / B testing mechanism: the new strategy will be deployed on some nodes and compared with the old strategy in terms of deviation reduction rate, user satisfaction improvement, etc. The strategy will be fully deployed after the key indicators meet the requirements, thus ensuring the stability of the system and realizing the continuous evolution of the model.

[0081] In the above implementation, a closed-loop optimization mechanism driven by user subjective feedback is introduced based on the traditional broadcast audio quality assessment process, constructing a complete adaptive system of "perceptual correction—feature tracing—weight evolution." By receiving real user scores and performing spatiotemporal alignment and deviation analysis with the system scores, the system can accurately identify systematic deviations between the algorithm and human auditory perception; by setting continuous triggering conditions and extracting associated feature sets, it effectively distinguishes between accidental fluctuations and structural defects; finally, based on data-driven mapping relationships, it dynamically adjusts the weight allocation strategy, enabling the assessment model to have environmental adaptability. This application's solution not only improves the psychoacoustic accuracy of comprehensive audibility scoring but also significantly reduces the frequency of manual parameter tuning, enhancing the system's robustness in dealing with complex and ever-changing broadcast environments.

[0082] Reference Figure 5 As one implementation of step S104, the steps of performing quality discrimination and root cause localization based on the comprehensive audibility score and multi-dimensional feature set, and outputting the quality level and fault type identifier, include: Step S501: Compare the comprehensive audibility score with the preset grading threshold and output the quality level label; This process is essentially a rule-driven quality status discrimination mechanism, designed to map continuous score values ​​to discrete semantic levels, facilitating operators' quick understanding of the current audio quality level. The system typically sets two key thresholds: a first threshold and a second threshold, with the first threshold being higher than the second, forming a three-tiered rating system: "Excellent," "Medium," and "Poor."

[0083] Specifically, when the overall audibility score is not lower than the first threshold, it indicates that the overall audio performance is good, the speech is clear, and there is no obvious distortion, and the system outputs an "Excellent" rating. When the score is between the second and first thresholds, it indicates that the audio quality has deteriorated to some extent, such as slight noise or brief interference, but it has not seriously affected the listening experience, and the system outputs a "Medium" rating. When the score is lower than the second threshold, it means that the audio quality has seriously deteriorated, and there may be intermittent, popping, severe distortion, or complete inaudibility, and the system judges it as a "Poor" rating. This three-tiered classification method is more flexible and practical than the traditional binary judgment (pass / fail), and can more finely reflect the changes in quality gradient, adapting to the actual needs of the broadcasting industry for service quality grading management. For example, in an emergency broadcasting system, a "Medium" rating can trigger an early warning mechanism to remind technicians to pay attention to potential problems, while a "Poor" rating will immediately initiate a fault response procedure.

[0084] Step S502: Calculate the contribution measure of each feature based on the multi-dimensional feature set and the feature weight coefficient set, and generate a contribution sequence. Specifically, the system first extracts the feature weight coefficients determined in the preceding dynamic weighted fusion stage. These coefficients reflect the relative importance of each feature to the audibility score in the current scenario and embody the system's sensitivity configuration to different degradation mechanisms. At the same time, the system obtains the normalized value of each feature in the multi-dimensional feature set. This value represents the performance level of the current feature (the closer to 1, the better the state; the closer to 0, the worse the state).

[0085] Based on this, the system performs a contribution metric calculation for each feature. The core idea is that the negative impact of a feature on the overall quality score depends not only on its own degree of degradation (i.e., 1 - normalized value) but also on its weight in the scoring model. Therefore, the contribution metric = weight coefficient × (1 - normalized feature value). The larger the value, the more significant the role that feature plays in the current quality decline.

[0086] For example, if the normalized value of a noise feature in an audio segment is only 0.3, and its weight is 0.35, then its contribution metric is 0.35×(1−0.3)=0.245, which is significantly higher than other features, indicating that noise is the main factor causing quality degradation. In this way, the system generates a contribution sequence containing four indicators: signal strength, interference, noise, and propagation disturbance, providing a quantitative basis for subsequent root cause identification.

[0087] Step S503: Identify the dominant contribution features in the contribution sequence, map them to a preset fault type library, and output the fault type identifier.

[0088] The dominant contribution feature is determined by the maximum value in the contribution sequence; that is, the feature with the highest contribution metric is considered the primary cause of the current quality problem. This decision-making mechanism based on quantitative comparison is objective and repeatable, avoiding subjective biases caused by human experience-based judgment.

[0089] Subsequently, the system associates and matches the dominant feature type with a preset fault type library: the signal strength feature corresponds to the "weak signal" flag, which usually indicates insufficient transmission power, excessive transmission distance, or degraded receiving antenna performance; the interference feature corresponds to the "interference intrusion" flag, reflecting the presence of co-channel, adjacent-channel, or other electromagnetic interference sources; the noise feature corresponds to the "excessive noise" flag, indicating increased background noise or deterioration of the equipment's signal-to-noise ratio; and the propagation disturbance feature corresponds to the "propagation disturbance" flag, pointing to problems such as multipath effects, reflected wave interference, or channel instability.

[0090] In this embodiment, the fault type library is designed to be scalable, allowing for continuous enrichment of fault modes and mapping rules based on actual operation and maintenance experience. For example, when the system detects that the contribution of propagation disturbance characteristics is the highest, it can automatically output a "propagation disturbance" flag, prompting technicians to check the geographical environment around the launch station or adjust modulation parameters, thereby achieving precise problem-oriented maintenance.

[0091] In the above implementation, quality grading is achieved based on a comprehensive audibility score. A contribution measurement model reveals the influence weights of each degradation factor, and the results are mapped to a pre-defined fault type library based on the principle of maximum contribution, ultimately forming a structured diagnostic result containing both quality level and fault type. This technical solution not only accurately identifies quality degradation states but also scientifically locates the root causes, improving the intelligent operation and maintenance level of the broadcasting system and providing strong support for ensuring broadcast service quality.

[0092] Reference Figure 6 As one implementation of step S105, the step of generating real-time alarm information and visualization data when the quality level is lower than a preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions includes: Step S601: Receive the quality level identifier, fault type identifier, and multi-dimensional feature set; Among them, the quality level label comes from the graded judgment of the comprehensive audibility score in the quality discrimination stage, and is usually expressed as semantic labels such as "excellent", "medium" and "poor", representing the overall auditory performance level of the current broadcast audio; the fault type label comes from the contribution analysis and fault mapping in the root cause localization process, such as "weak signal", "interference intrusion", "excessive noise" or "propagation disturbance", revealing the specific technical causes of quality degradation; and the multi-dimensional feature set, as the underlying data support, contains four time-aligned numerical sequences of signal strength feature set, interference feature set, noise feature set and propagation disturbance feature set, which record the dynamic evolution of audio quality at the frame scale.

[0093] Step S602: When the quality level identifier is lower than the preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions, assemble an alarm message containing the fault type identifier, timestamp and feature summary. This application introduces "dynamic alarm conditions" to achieve more sensitive and adaptive detection capabilities. On the one hand, the system monitors in real time whether the quality level indicator is lower than the preset quality threshold (such as the "poor" level). Once confirmed, it is regarded as a serious quality problem and the alarm process is immediately initiated. On the other hand, the system also actively analyzes the abnormal patterns in the multi-dimensional feature set, that is, it judges whether there is a specific deterioration behavior based on the alarm triggering conditions.

[0094] In this application embodiment, the alarm triggering conditions include, but are not limited to, the following: when the interference feature value exceeds the interference threshold for N consecutive frames, it indicates that there is persistent electromagnetic interference or illegal signal intrusion, rather than accidental noise, and the system determines it as a valid threat; when the propagation disturbance feature set shows a sudden change in amplitude between adjacent frames that exceeds the preset gradient, it indicates that the wireless channel has changed drastically, such as a sudden increase in multipath or a change in the reflection environment; when the noise feature set deviates from the historical baseline by more than the tolerance range, it indicates that the background noise level has shifted significantly, which may be due to equipment aging or environmental changes.

[0095] Understandably, setting these dynamic conditions not only improves the accuracy of alarms but also effectively avoids false alarms caused by short-term fluctuations, enhancing the system's robustness and practicality. By combining static threshold judgment with dynamic pattern recognition, the system can achieve comprehensive coverage across different time scales and degradation types, reducing missed alarms and false alarms.

[0096] Furthermore, the system associates fault type identifiers with a predefined alarm level mapping table to determine the urgency of the current event. For example, "signal loss" corresponds to the "urgent" level, "interference lasting more than 5 frames" corresponds to the "serious" level, and "noise exceeding the limit" is marked as the "warning" level. Different levels will determine the subsequent push strategy and response priority. Subsequently, the system extracts key statistics from a multi-dimensional feature set as feature summaries, such as the mean, peak value, abrupt change location, or degradation duration of each feature, to quantify the severity and impact range of the fault. At the same time, the system embeds signal source device identifiers (such as transmitter number, IP address) and geographical location information (such as latitude and longitude or area code) to achieve accurate mapping from logical faults to physical nodes.

[0097] The generated alarm message contains multiple fields such as fault type, alarm level, timestamp, feature summary, device ID, and location information, forming a complete and self-consistent fault record. It can be used for instant push or stored as a log for a long time, greatly improving the readability and traceability of alarm information.

[0098] Step S603: Generate a visual data stream based on the multi-dimensional feature set and fault type identifier, including a scoring time series curve and a fault location map; Specifically, the visualized data stream is a multimodal, spatiotemporally correlated graphical combination, encompassing various forms such as scoring time-series curves, fault location maps, and feature heatmaps. The time-series curve of the comprehensive audibility score clearly displays the evolution trend of audio quality with time on the horizontal axis and score on the vertical axis, and overlays alarm trigger points on the curve to help identify the onset of degradation and the recovery process. The broadcast network topology map, based on geospatial data, marks faulty device nodes and renders their affected coverage area using a signal propagation model, enabling maintenance personnel to intuitively determine the range of affected listeners. The heatmap matrix of multi-dimensional feature sets, aligned frame by frame, uses color intensity to represent the intensity of changes in various feature values, forming a two-dimensional "time-feature" map, facilitating the discovery of latent patterns such as periodic interference, noise aggregation, or propagation anomalies.

[0099] Step S604: Distribute alarm messages and visual data streams to the target management terminal and store the associated data in the diagnostic log library.

[0100] The distribution mechanism intelligently selects the push channel based on the alarm level: for emergency alarms (such as signal interruption), the system calls the SMS API or instant messaging interface to ensure that on-duty personnel are notified as soon as possible; for general alarms (such as excessive noise), they are written to the message queue and processed asynchronously by the background service to avoid resource contention. This hierarchical push strategy takes into account both response speed and system stability.

[0101] Meanwhile, all raw data (including multi-dimensional feature sets, alarm messages, visualized data snapshots, and timestamp information) is stored in the diagnostic log library in a time-partitioned manner. Time-partitioned storage not only improves query efficiency in a big data environment but also facilitates subsequent historical backtracking, trend analysis, and model training. For example, by retrieving all alarm logs within a certain time period, the frequency of interference in a specific area can be analyzed, thereby optimizing spectrum planning; by comparing the characteristic baseline changes of different devices, equipment aging trends can be predicted, enabling preventative maintenance.

[0102] In the above implementation, the system can not only accurately identify quality problems, but also clearly reveal their causes, scope of impact and spatiotemporal distribution characteristics. It also ensures efficient information transmission through structured messages and hierarchical push mechanisms, and provides a complete data chain for subsequent analysis by combining time-series log storage, thereby improving the intelligent operation and maintenance capabilities and service quality assurance level of the broadcasting system.

[0103] Reference Figure 7 As a further implementation of the broadcast audio quality discrimination method, the discrimination method also includes: Step S701: Receive quality level and fault type identifiers from multiple broadcast nodes; Each broadcast node, as an independent sensing unit in the distributed monitoring network, integrates complete audio quality discrimination and root cause localization capabilities. It can autonomously complete the entire process of analysis, including signal acquisition, feature extraction, dynamic weighted fusion, quality grading, and fault mapping, and output two types of highly structured diagnostic results: quality level identifier and fault type identifier.

[0104] Specifically, the quality level label is an overall evaluation of the audibility of the broadcast audio received by the node, usually presented as semantic labels such as "excellent", "medium", and "poor". It is generated based on the comparison between the comprehensive audibility score and the preset grading threshold, reflecting the instantaneous audio quality status of the node at the current moment. The fault type label is the root cause label formed by standardizing and encoding the main degradation factors identified through contribution analysis, such as "interference intrusion-12", "propagation disturbance-05" or "weak signal-08". These codes not only indicate the fault category, but may also include sub-category numbers or severity levels, which facilitates subsequent classification statistics and pattern recognition.

[0105] In this embodiment, by aggregating such identification information from multiple nodes at different geographical locations and topological levels, the system constructs a spatiotemporally aligned fault sample matrix, realizing a shift from "point-based perception" to "area-based cognition." This multi-node data aggregation mechanism enhances the monitoring coverage and provides a rich and representative raw data foundation for subsequent spatial correlation analysis, enabling the system to identify regional common problems rather than isolated events.

[0106] Step S702: Associate the location coordinates of multiple broadcast nodes in a preset geographic information system; Each broadcast node registers its latitude and longitude coordinates, altitude, antenna orientation, and transmission power in the GIS database during the initial deployment phase, and binds them to its unique device identifier (such as device ID or IP address) to form a "device-location" index table. When the system receives a fault alarm from a node, it can quickly query the GIS database through its device ID to transform the abstract logical event "Node A: Interference Intrusion" into a specific geographic entity, such as: "Interference event occurred at 39.9042°N, 116.4072°E".

[0107] This spatial mapping goes beyond simple coordinate labeling; it incorporates the topology information of the broadcast network, including the link relationships between signal sources and relay stations, the hierarchical division of coverage areas, and the dependency paths of upstream and downstream devices. For example, if a relay station experiences a "weak signal" fault, the system can trace its upstream transmission tower based on the topology and predict whether the fault may affect multiple downstream receiving nodes, thereby identifying potential cascading failure risks.

[0108] Step S703: Calculate the spatial impact range of the fault node based on the quality level and fault type identifier; The spatial impact range includes the coverage radius and the user density value affected; Specifically, the calculation of the coverage radius fully considers the actual impact of fault type on signal propagation capability, rather than relying solely on the nominal parameters of the equipment. For example, when the fault type is identified as "propagation disturbance," it indicates the presence of severe multipath effects or reflection interference, leading to a decrease in signal diffraction capability in complex terrain. The system will correct the theoretical coverage radius according to the weighting factor of this fault type, which may reduce the effective coverage area by more than 30%. When the fault is "excessive noise," especially in mountainous or urban canyon environments, the rise in background noise baseline will significantly reduce the signal-to-noise ratio. The system combines the elevation data and environmental noise feature set provided by GIS, and uses a modified free-space path loss model to re-estimate the signal attenuation curve, thereby deriving the dynamically adjusted actual coverage boundary.

[0109] Simultaneously, the system also calculates the impact user density value, using the number of potential listeners within the coverage radius as a key indicator of social impact. This value is derived by integrating real-time terminal access data provided by operators, GIS heat maps generated based on census data, and a gradient model of signal strength attenuation with distance. For example, in a high-density residential area, even if the signal is only slightly degraded, the impact density value will still be significantly amplified due to the large number of affected users; while in remote areas, even if the signal is completely interrupted, the impact value will be relatively low if there are few users.

[0110] Step S704: Prioritize real-time alarm information based on spatial influence range and output a hierarchical alarm distribution sequence.

[0111] Traditional alarm systems often simply rank faults based on their severity level (e.g., "urgent" or "critical"), which is insufficient to address the resource optimization needs in complex scenarios. This application introduces a multi-dimensional comprehensive evaluation mechanism, using coverage radius reduction ratio, user density impact, and fault propagation risk as core ranking factors to construct a dynamic prioritization model that considers technical impact, social impact, and system stability. Specifically, the coverage radius reduction ratio reflects the degree of technical damage the fault inflicts on broadcast service capabilities; the user density value reflects the breadth of its impact on the public's listening experience; and the fault propagation risk, based on broadcast link topology analysis, assesses whether a fault at a node could trigger a cascading failure of downstream equipment, thus determining the potential depth of its harm.

[0112] In some embodiments, to achieve scientific prioritization, the system can employ a composite decision-making algorithm combining entropy weighting and TOPSIS: entropy weighting automatically determines the objective weights of each factor based on historical data (avoiding subjective weighting bias), while TOPSIS outputs a normalized priority index by calculating the relative proximity of each fault event to the ideal optimal solution. For example, a central node in a city experiences a 40% reduction in coverage radius due to strong interference, affecting 80,000 users, and is located upstream of the backbone link; its priority index can reach 0.85. Meanwhile, a suburban node, although suffering a 60% loss in coverage radius, only affects 5,000 users and has no downstream dependencies, resulting in a priority of only 0.62. Based on this, the system generates an alarm distribution sequence arranged in descending order of priority, ensuring that maintenance resources are prioritized for the most impactful and highest-risk fault points, greatly improving the efficiency and rationality of emergency response.

[0113] In the above implementation, broadcast quality diagnosis and geospatial analysis are deeply integrated. By receiving quality and fault identifiers from multiple nodes, wide-area monitoring data is aggregated; precise spatial attributes are assigned to fault information through GIS coordinate association; the impact range of broadcast faults is quantitatively expressed by dynamically calculating coverage radius and user density values; and a hierarchical alarm sequence is generated through a multi-factor comprehensive decision-making model, driving the transformation of operation and maintenance strategies from passive response to proactive optimization. This solution not only improves the accuracy of fault location and the timeliness of response, but also upgrades traditional audio quality monitoring into a smart operation and maintenance platform with spatial cognition, impact prediction, and intelligent sorting capabilities, providing strong technical support for ensuring the high availability of modern broadcast systems.

[0114] This application also discloses a broadcast audio quality discrimination device.

[0115] A broadcast audio quality discrimination device, the discrimination device comprising: The preprocessing module is used to receive the broadcast audio signal stream, perform preprocessing, and output the signal sequence after frame processing. The multidimensional feature extraction module is used to extract multidimensional features in parallel from the signal sequence after frame processing, and generate a multidimensional feature set containing signal strength features, interference features, noise features and propagation disturbance features; The weighted fusion module is used to perform dynamic weighted fusion calculations based on multi-dimensional feature sets to generate a comprehensive audibility score. The quality discrimination module is used to perform quality discrimination and root cause localization based on the comprehensive audibility score and multi-dimensional feature set, and output the quality level and fault type identifier. The alarm module is used to generate real-time alarm information and visualization data when the quality level is lower than the preset level threshold or when the multi-dimensional feature set meets the alarm triggering conditions.

[0116] The broadcast audio quality discrimination device according to the present application embodiment can implement any of the above-described broadcast audio quality discrimination methods, and the specific working process of each module in the broadcast audio quality discrimination device can refer to the corresponding process in the above-described method embodiment.

[0117] In the several embodiments provided in this application, it should be understood that the provided methods and apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0118] This application also discloses a computer device.

[0119] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a broadcast audio quality discrimination method as described above.

[0120] This application also discloses a computer-readable storage medium.

[0121] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the broadcast audio quality discrimination methods.

[0122] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0123] In this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0124] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.

[0125] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A method for judging the quality of broadcast audio, characterized in that, The discrimination method includes: It receives broadcast audio signal streams, performs preprocessing, and outputs a framed signal sequence. Multi-dimensional features are extracted in parallel from the framed signal sequence to generate a multi-dimensional feature set containing signal strength features, interference features, noise features, and propagation disturbance features; Based on the multi-dimensional feature set, a dynamic weighted fusion calculation is performed to generate a comprehensive audibility score; Based on the comprehensive audibility score and the multi-dimensional feature set, quality judgment and root cause localization are performed, and the quality level and fault type identifier are output. When the quality level is lower than a preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions, real-time alarm information and visualization data are generated.

2. The broadcast audio quality discrimination method according to claim 1, characterized in that, The steps of extracting multi-dimensional features in parallel from the framed signal sequence to generate a multi-dimensional feature set containing signal strength features, interference features, noise features, and propagation disturbance features include: Receive the signal sequence after frame processing; Feature extraction is performed in parallel on each frame of the signal sequence; Among them, the signal strength feature is calculated as the first feature value, the interference feature is identified as the second feature value, the noise feature is measured as the third feature value, and the propagation disturbance feature is evaluated as the fourth feature value. Aggregate the first, second, third, and fourth feature values ​​of all frames to generate a multi-dimensional feature set that includes signal strength feature set, interference feature set, noise feature set, and propagation disturbance feature set.

3. The broadcast audio quality discrimination method according to claim 2, characterized in that, The steps for generating a comprehensive audibility score by performing dynamic weighted fusion calculation based on the multi-dimensional feature set include: Normalize the feature values ​​of the multi-dimensional feature set to output a unified scale feature set; Select the weight allocation method according to the preset scenario strategy, and generate a set of feature weight coefficients; Based on the unified scale feature set and the feature weight coefficient set, a dynamic weighted fusion calculation is performed to output a frame-level audibility score sequence. The frame-level audibility score sequence is aggregated to generate a comprehensive audibility score.

4. The broadcast audio quality discrimination method according to claim 3, characterized in that, Following the step of generating a comprehensive audibility score are: Receive subjective audio quality ratings from user terminals; The deviation value is calculated between the subjective audio quality score data and the comprehensive audibility score for the corresponding time period. When the deviation value continues to exceed the optimization trigger threshold, the multi-dimensional feature set of the time period is extracted; Based on the mapping relationship between the multi-dimensional feature set and the subjective audio quality rating data, the weight allocation strategy of the dynamic weighted fusion calculation is updated.

5. The broadcast audio quality discrimination method according to claim 4, characterized in that, The steps for performing quality assessment and root cause localization based on the comprehensive audibility score and the multi-dimensional feature set, and outputting quality level and fault type identifier, include: The comprehensive audibility score is compared with a preset grading threshold, and a quality level identifier is output. The contribution measure of each feature is calculated based on the multi-dimensional feature set and the feature weight coefficient set, and a contribution sequence is generated. Identify the dominant contribution features in the contribution sequence, map them to a preset fault type library, and output fault type identifiers.

6. The broadcast audio quality discrimination method according to claim 1, characterized in that, The steps for generating real-time alarm information and visualization data when the quality level is lower than a preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions include: Receive quality level identifiers, fault type identifiers, and multi-dimensional feature sets; When the quality level identifier is lower than the preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions, an alarm message containing a fault type identifier, timestamp, and feature summary is assembled. A visual data stream is generated based on the multi-dimensional feature set and fault type identifier, including a scoring time-series curve and a fault location map. The alarm messages and visualized data streams are distributed to the target management terminal, and the associated data is stored in the diagnostic log library.

7. A broadcast audio quality discrimination method according to any one of claims 1 to 6, characterized in that, The discrimination method also includes: Receive the quality level and fault type identifiers from multiple broadcast nodes; Associate the location coordinates of the multiple broadcast nodes in a preset geographic information system; Based on the quality level and fault type identifier, calculate the spatial impact range of the fault node; The real-time alarm information is prioritized based on the spatial influence range, and a hierarchical alarm distribution sequence is output.

8. A broadcast audio quality discrimination device, characterized in that, The discrimination device includes: The preprocessing module is used to receive the broadcast audio signal stream, perform preprocessing, and output the signal sequence after frame processing. The multidimensional feature extraction module is used to extract multidimensional features in parallel from the signal sequence after the frame processing, and generate a multidimensional feature set including signal strength features, interference features, noise features and propagation disturbance features; The weighted fusion module is used to perform dynamic weighted fusion calculation based on the multi-dimensional feature set to generate a comprehensive audibility score. The quality discrimination module is used to perform quality discrimination and root cause localization based on the comprehensive audibility score and the multi-dimensional feature set, and output the quality level and fault type identifier. The alarm module is used to generate real-time alarm information and visualization data when the quality level is lower than a preset level threshold or the multi-dimensional feature set meets the alarm triggering conditions.

9. A computer device, characterized in that: The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Real-time audio stream comparison method and system

    CN114495984A

  • Audio quality analysis method and device, electronic equipment and storage medium

    CN116013367A

  • Multi-mode voiceprint fault diagnosis method for converter transformer

    CN120412647A

  • Broadcast audio processing method and device of broadcast intelligent audio mixing technology

    CN120510858A

  • Broadcast signal interference correction method, system, medium and equipment

    CN120748422A

Cited By

  • News live voice anomaly real-time monitoring correction method and system thereof

    CN122135742A