Nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion
Through a multi-algorithm fusion monitoring method, combined with intelligent sensors and deep learning models, the problems of low efficiency and high misjudgment rate of traditional inspections of nuclear power plant speakers have been solved, and real-time health diagnosis and high-accuracy fault detection of speakers have been achieved, reducing maintenance costs and safety risks.
Patent Information
- Application Number
- CN202511103565.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-07
AI Technical Summary
In existing technologies, traditional manual inspections of nuclear power plant speakers are inefficient, costly, and difficult to detect faults in real time. Single detection indicators cannot adapt to complex acoustic environments and lack intelligent analysis, resulting in a high misjudgment rate.
A monitoring method based on multi-algorithm fusion is adopted to obtain audio data through intelligent sensor modules. Combined with acoustic feature extraction and deep learning models, sound event detection, voiceprint similarity calculation and speech recognition are performed to achieve real-time health diagnosis of speakers, reduce maintenance costs and improve fault detection accuracy.
It realizes real-time health diagnosis of speaker status, reduces maintenance costs, improves the accuracy of fault detection, shortens fault response time, and reduces safety risks.
Smart Images

Figure CN120602881B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of loudspeaker state monitoring, more particularly, to a nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion. BACKGROUND
[0002] The loudspeaker of the broadcast public address system in a nuclear power plant is an important tool for daily operation and emergency communication. The loudspeaker is widely distributed and the environment is complex. The traditional manual inspection method is low in efficiency and high in cost, and it is difficult to find faults in time. In the prior art, audio monitoring relies on a single indicator (such as loudness) or simple comparison, which cannot adapt to the diversified acoustic environment in the nuclear power plant and cannot meet the complex needs of the nuclear power plant. The prior art has the following problems:
[0003] Manual inspection depends on: manual inspection, which is time-consuming and labor-intensive, and cannot find sudden faults in real time. Manual detection is easily affected by subjective factors, especially in complex acoustic environments (such as high-noise turbine workshops), where subtle faults are difficult to detect.
[0004] Single detection index limitation: the existing technology is mostly based on a single indicator (such as loudness or signal-to-noise ratio), lacks multi-dimensional data fusion analysis, and has a high misjudgment rate.
[0005] Lack of intelligent analysis: without introducing intelligent algorithms such as deep learning, the detection logic cannot be automatically optimized according to the acoustic scene.
[0006] Therefore, the prior art has defects and needs to be improved. SUMMARY
[0007] In view of the above problems, the purpose of the present application is to provide a nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion, which realizes real-time health diagnosis of the loudspeaker in the nuclear power plant by deploying an intelligent sensor module, combining acoustic feature extraction and a deep learning model, reduces maintenance costs, and improves the reliability of the broadcast system. Through SED, audio similarity, and ASR triple verification, the fault detection accuracy is significantly improved. Wireless deployment is flexible, reducing installation and maintenance costs. Real-time generation of alarm information shortens the fault response time and reduces safety risks.
[0008] The first aspect of the present application provides a nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion, comprising:
[0009] acquiring collected audio data through a monitoring module;
[0010] acquiring original audio data played by a broadcast system;
[0011] determining a SED comprehensive score based on the original audio data and the collected audio data;
[0012] inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score;
[0013] inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score;
[0014] performing weighted calculation on the SED comprehensive score, the voiceprint similarity score and the speech recognition ASR score to determine a loudspeaker comprehensive score;
[0015] comparing the loudspeaker comprehensive score with a preset evaluation threshold to determine loudspeaker state information.
[0016] In the scheme, the SED comprehensive score is determined based on the original audio data and the sound event detection of the collected audio data, comprising:
[0017] performing root mean square calculation on the collected audio data to determine an RMS value of the collected audio data;
[0018] calculating a ratio of the RMS value of the collected audio data to a preset RMS value threshold to determine a loudness score;
[0019] performing center frequency calculation on the collected audio data and the original audio data to determine a collected audio data center frequency and an original audio data center frequency;
[0020] calculating a center frequency score through the collected audio data center frequency and the original audio data center frequency;
[0021] ;
[0022] wherein P is the center frequency score, B is a preset center frequency threshold, f1 is the collected audio data center frequency, and f0 is the original audio data center frequency;
[0023] performing flatness calculation on the collected audio data and the original audio data to determine a flatness score;
[0024] performing signal-to-noise ratio calculation on the collected audio data and the original audio data to determine a signal-to-noise ratio score;
[0025] multiplying the loudness score, the center frequency score, the flatness score and the signal-to-noise ratio score by corresponding first score weights respectively, and accumulating the calculation results to determine the SED comprehensive score.
[0026] In the scheme, it further comprises:
[0027] Based on the detected environment, the first scoring weights of the loudness score, the center frequency score, the flatness score, and the signal-to-noise ratio score are dynamically adjusted by acoustic scene classification ASC.
[0028] In this scheme, the collected audio data and the original audio data are input into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score, including:
[0029] The collected audio data and the original audio data are input into a preset voiceprint similarity calculation model, and voiceprint features of the collected audio data and the original audio data are extracted by the preset voiceprint similarity calculation model.
[0030] The voiceprint features of the collected audio data and the original audio data are calculated for similarity to determine a voiceprint similarity score.
[0031] In this scheme, the collected audio data and the original audio data are input into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score, including:
[0032] The collected audio data and the original audio data are input into a preset voiceprint similarity calculation model, and voiceprint features of the collected audio data and the original audio data are extracted by the preset voiceprint similarity calculation model.
[0033] The collected audio data and the original audio data are input into a preset voiceprint similarity calculation model, and voiceprint features of the collected audio data and the original audio data are extracted by the preset voiceprint similarity calculation model.
[0034] The ratio of the number of correctly recognized words to the number of original audio data recognized words is calculated to determine a speech recognition ASR score.
[0035] In this scheme, the loudspeaker comprehensive score is compared with a preset evaluation threshold to determine loudspeaker state information, including:
[0036] The preset evaluation threshold includes a first preset evaluation threshold and a second preset evaluation threshold.
[0037] When the loudspeaker comprehensive score is less than or equal to the first preset evaluation threshold, a loudspeaker fault is determined, and a red warning is given.
[0038] When the loudspeaker comprehensive score is between the first preset evaluation threshold and the second preset evaluation threshold, a loudspeaker suspected fault is determined, and an orange warning is given.
[0039] When the loudspeaker comprehensive score is greater than or equal to the second preset evaluation threshold, the loudspeaker is determined to be in good condition.
[0040] In this scheme, it also includes:
[0041] Obtaining auxiliary audio data through an adjacent monitoring module;
[0042] Segmenting the collected audio data according to a preset time length to obtain a plurality of sub-collected audio data;
[0043] Segmenting the auxiliary audio data according to a preset time length to obtain a plurality of sub-auxiliary audio data;
[0044] Comparing the collected audio data and the original audio data based on time sequence to determine a first error value of each sub-collected audio data in the collected audio data, and drawing a first error value change curve;
[0045] Determining a first error interval through the distribution of the first error values of all the sub-collected audio data in the collected audio data;
[0046] Segmenting the first error value change curve through the first error interval, and determining a time interval outside the first error interval as an audio influence interval;
[0047] Comparing the auxiliary audio data and the original audio data based on time sequence to determine a second error value of each sub-auxiliary audio data in the auxiliary audio data, and drawing a second error value change curve;
[0048] Determining a second error interval through the distribution of the second error values of all the sub-auxiliary audio data in the auxiliary audio data;
[0049] Determining whether the auxiliary audio data in the audio influence interval meets the second error interval;
[0050] If yes, replacing the collected audio data in the audio influence interval with the auxiliary audio data;
[0051] Otherwise, no processing is performed.
[0052] In the scheme, the replacing the collected audio data in the audio influence interval with the auxiliary audio data comprises:
[0053] Determining a middle value of the first error interval as a first correction coefficient;
[0054] Determining a middle value of the second error interval as a second correction coefficient;
[0055] Multiplying the sub-auxiliary audio data in the audio influence interval by a ratio of the first correction coefficient and the second correction coefficient to determine corrected audio data;
[0056] Replacing the audio signal of the collected audio data in the audio influence interval with the corrected audio data.
[0057] The method comprises the following steps of: comparing the collected audio data with the original audio data based on time sequence, determining a first error value of each sub-collected audio data in the collected audio data, and the determination comprises the following steps of:
[0058] sequentially analyzing each sub-collected audio data, determining an audio time corresponding to each wave peak and each wave trough in the sub-collected audio data as a verification time;
[0059] calculating a change rate of energy intensity corresponding to each preset frequency at the verification time, and drawing a change rate curve of energy intensity at the verification time;
[0060] calculating an average curve of the change rate curves of energy intensity at all verification times;
[0061] extracting the change rate of energy intensity corresponding to each preset frequency from the average curve, performing weighted calculation, and determining the first error value of the current sub-collected audio data.
[0062] The method further comprises the following steps of:
[0063] abnormal marking is performed on coordinate points with a change rate of energy intensity greater than a preset change rate threshold in a spectrum diagram, and abnormal coordinates are determined;
[0064] traversal is performed based on the abnormal coordinates, and other abnormal coordinates with a closest pixel distance to the abnormal coordinates are selected to form a candidate box;
[0065] the number of abnormal coordinates in the candidate box is counted, and an abnormal coordinate proportion of the candidate box is determined;
[0066] a corresponding preset pixel proportion threshold is determined according to a pixel area of the candidate box;
[0067] when the abnormal coordinate proportion of the candidate box is greater than or equal to the corresponding preset pixel proportion threshold, the traversal is continued, other abnormal coordinates with a closest pixel distance to the candidate box are selected, and the candidate box is updated;
[0068] when the abnormal coordinate proportion of the candidate box is less than the corresponding preset pixel proportion threshold, the abnormal coordinates are filtered;
[0069] when the abnormal coordinate proportion of the updated candidate box is greater than or equal to the corresponding preset pixel proportion threshold, the traversal is continued, and the candidate box is updated; when the abnormal coordinate proportion of the updated candidate box is less than the corresponding preset pixel proportion threshold, the traversal is ended, and an area in the candidate box is determined as an abnormal area;
[0070] the spectrum diagram in the abnormal area is locally replaced;
[0071] In the first error value process of the sub-acquisition audio data, the energy intensity change rate corresponding to the preset frequency in the abnormal area is filtered.
[0072] The audio signal of the collected audio data in the audio influence interval is replaced by the corrected audio data. The application discloses a nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion, and the method comprises the following steps: acquiring collected audio data through a monitoring module; acquiring original audio data played by a broadcast system; performing sound event detection on the collected audio data based on the original audio data to determine a SED comprehensive score; inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score; inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition ASR score; performing weighted calculation on the SED comprehensive score, the voiceprint similarity score and the speech recognition ASR score to determine a loudspeaker comprehensive score; and comparing the loudspeaker comprehensive score with a preset evaluation threshold to determine loudspeaker state information. The application can improve the loudspeaker fault detection accuracy through the triple verification of the sound event detection, the audio similarity and the speech recognition algorithm. BRIEF DESCRIPTION OF DRAWINGS
[0073] Figure 1 A flowchart of a nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion is shown.
[0074] Figure 2 A flowchart of a voiceprint similarity score calculation method is shown.
[0075] Figure 3 A flowchart of a speech recognition ASR score calculation method is shown. DETAILED DESCRIPTION
[0076] In order to more clearly understand the above-mentioned purposes, features and advantages of the application, the application will be further described in detail below with reference to the drawings and specific embodiments. It should be noted that the embodiments of the application and the features in the embodiments can be combined with each other without conflict.
[0077] In the following description, many specific details are set forth in order to provide a thorough understanding of the application, but the application can also be practiced without the other ways different from those described herein, therefore, the scope of protection of the application is not limited by the specific embodiments disclosed below.
[0078] Figure 1 A flowchart of a nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion is shown.
[0079] As Figure 1The application discloses a loudspeaker state monitoring method based on multi-algorithm fusion for a nuclear power plant, which comprises the following steps:
[0080] In S102, the monitoring module is used to acquire the collected audio data.
[0081] In S104, the original audio data played by the broadcasting system is acquired.
[0082] In S106, the collected audio data is subjected to sound event detection based on the original audio data to determine a SED comprehensive score.
[0083] In S108, the collected audio data and the original audio data are input into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score.
[0084] In S110, the collected audio data and the original audio data are input into a preset speech recognition model for analysis to determine a speech recognition ASR score.
[0085] In S112, the SED comprehensive score, the voiceprint similarity score and the speech recognition ASR score are subjected to weighted calculation to determine a loudspeaker comprehensive score.
[0086] In S114, the loudspeaker comprehensive score is compared with a preset evaluation threshold to determine loudspeaker state information.
[0087] According to the application, the monitoring module configured on site is used to collect the audio signal played by the loudspeaker to determine the collected audio data. The collected audio data and the original audio data (i.e. the audio signal played by the broadcasting system) are compared by the system, and the health state of the loudspeaker is scored according to the comparison result, and the equipment operation condition of the loudspeaker is determined according to the score. The application adopts three audio signal processing and analysis methods, namely sound event detection SED (Sound Event Detection, SED), voiceprint similarity and speech recognition algorithm, and adopts a functional value scoring method (i.e. 0-1 scoring method). The collected audio data and the original audio data are compared to obtain the SED comprehensive score, the voiceprint similarity score and the speech recognition ASR score. The SED comprehensive score, the voiceprint similarity score and the speech recognition ASR score are multiplied by the corresponding second score weight, and the calculation results are accumulated to determine the loudspeaker comprehensive score. The second score weight of the SED comprehensive score, the voiceprint similarity score and the speech recognition ASR score is determined by the acoustic scene classification ASC. The loudspeaker comprehensive score is compared with the preset evaluation threshold of the system to determine the loudspeaker state information.
[0088] The quality of the loudspeaker can be directly reflected in the SED, and the quality of the loudspeaker can be basically judged through the related feature combination, therefore, the second score weight of the SED comprehensive score is set to be relatively high. In the public broadcasting system, due to the characteristics of the loudspeaker sound, the voiceprint similarity of the original audio data and the collected audio data cannot achieve a high score, therefore, the second score weight of the voiceprint similarity score is set to be relatively low. The speech recognition technology can assist in evaluating the performance of the loudspeaker. For monitoring of a single loudspeaker, if the recognition accuracy is low, it may be caused by distortion of the loudspeaker, non-smooth frequency response or low signal-to-noise ratio, which can indicate that the performance of the monitored loudspeaker is poor. In the public broadcasting system, the density of the loudspeaker deployment is often large, when a single loudspeaker is damaged, the sound of other loudspeakers within the test point range can still be received, so that the content of the collected audio data is recognized, therefore, the second score weight of the speech recognition ASR score is set to be relatively low in the environment of multi-loudspeaker deployment.
[0089] The preset voiceprint similarity calculation model and the preset speech recognition model are both trained by historical audio data obtained in a historical monitoring process.
[0090] According to the embodiment of the present application, the SED comprehensive score is determined based on the sound event detection of the collected audio data on the original audio data, which includes:
[0091] The RMS value of the collected audio data is determined by performing root mean square calculation on the collected audio data;
[0092] The loudness score is determined by calculating the ratio of the RMS value of the collected audio data to the preset RMS value threshold;
[0093] The center frequency of the collected audio data and the center frequency of the original audio data are determined by performing center frequency calculation on the collected audio data and the original audio data;
[0094] The center frequency score is calculated by the center frequency of the collected audio data and the center frequency of the original audio data;
[0095] ;
[0096] Wherein, P is the center frequency score, B is the preset center frequency threshold, f1 is the center frequency of the collected audio data, f0 is the center frequency of the original audio data;
[0097] The flatness score is determined by performing flatness calculation on the collected audio data and the original audio data;
[0098] The signal-to-noise ratio score is determined by performing signal-to-noise ratio calculation on the collected audio data and the original audio data;
[0099] The loudness score, the center frequency score, the flatness score and the signal-to-noise ratio score are multiplied by the corresponding first score weight, respectively, and the calculation results are accumulated to determine the SED comprehensive score.
[0100] It should be noted that the sound event detection SED is a method for identifying and locating specific sound events in an audio signal. In sound event detection, the original audio data signal is converted into meaningful information through feature extraction, so that the machine learning model can understand and classify the sound event. In the audio basic signal operation, through the use of extraction tools, the extracted features usually include loudness, center frequency, flatness, signal-to-noise ratio, sampling rate, etc. These features are very important for audio quality evaluation, noise suppression, audio device calibration, and health status monitoring of audio signals.
[0101] In the audio processing software, the indicator of loudness usually uses RMS (Root Mean Square) to measure the average level or average power of the audio signal, which is used to quantify the loudness or volume of the audio data, so as to standardize the playback volume of different audio and ensure consistency on different platforms and devices. The loudness of the collected audio data is calculated by the system preset root mean square calculation method, and the calculation result is standardized by 0-1 scoring method and the like to obtain the RMS value of the collected audio data. In order to measure whether the loudness indicator can reach 1 point, the same type of loudspeaker needs to be run at full power in the broadcast system, and the maximum RMS value of the audio data collected by the loudspeaker monitoring module is taken as the preset RMS value threshold of this type of loudspeaker.
[0102] The audio center frequency is usually measured in Hz. In audio processing, it is considered as the "center" of the frequency spectrum energy of the audio data signal, and the pre-prepared audio data and the collected audio data should be theoretically consistent. When the center frequency deviates, it means that the loudspeaker or power amplifier system audio reproduction is not accurate, the sound is unbalanced, some frequency bands are missing, and the sound quality experience is affected. The collected audio data center frequency of the collected audio data and the original audio data center frequency of the original audio data are calculated by the system preset center frequency calculation method, the center frequency score is calculated by the collected audio data center frequency and the original audio data center frequency, and the center frequency score is standardized by 0-1 scoring method and the like. The final center frequency score is determined. When the center frequency score is 1, the original audio data and the collected audio data center frequency of the loudspeaker played by the monitoring module are consistent. The preset center frequency threshold B is determined based on the center frequency deviation frequency range that can be accepted by the human ear.
[0103] The flatness is an index for quantifying the uniformity of the frequency spectrum distribution of the audio data. The closer the value is to 1, the more the frequency spectrum of the sample is like white noise, that is, the flatter. On the contrary, if it is much smaller than 1, it means that the sample spectrum has obvious peaks, and the spectrum is more diverse and complex. The flatness of the collected audio data and the flatness of the original audio data are calculated by using a preset flatness calculation method of the system, the flatness score is determined by calculating the ratio of the flatness of the original audio data to the flatness of the collected audio data, and the flatness score is standardized by using a 0-1 scoring method or the like. When the collected audio data is less than or equal to the original audio data, the flatness score is 1.
[0104] The unit of the signal-to-noise ratio is decibels (dB), which is very suitable for representing the relative size between two orders of magnitude, especially when it involves a contrast between very large or very small values, such as the ratio of audio noise intensity. Influenced by factors such as power amplifier systems and the environment, the signal-to-noise ratio of the audio signal received at the speaker end is usually smaller than that of the played audio signal. The signal-to-noise ratio of the collected audio data and the signal-to-noise ratio of the original audio data are calculated by using a preset signal-to-noise ratio calculation method of the system, the signal-to-noise ratio score is determined by calculating the ratio of the signal-to-noise ratio of the original audio data to the signal-to-noise ratio of the collected audio data, and the signal-to-noise ratio score is standardized by using a 0-1 scoring method or the like. When the signal-to-noise ratio of the collected audio data is equal to the signal-to-noise ratio of the original audio data, the signal-to-noise ratio score is 1; when the signal-to-noise ratio of the collected audio data is 0, the signal-to-noise ratio score is 0; when the signal-to-noise ratio of the collected audio data is greater than the signal-to-noise ratio of the original audio data, it can be considered that the system is disturbed or the speaker is not working properly.
[0105] The first scoring weight of the loudness score, the center frequency score, the flatness score and the signal-to-noise ratio score is set according to the perception of a person to these four characteristics when the speaker amplifies. The initial value of the first scoring weight of the loudness score, the center frequency score, the flatness score and the signal-to-noise ratio score is 60%, 10%, 10% and 20% respectively. In a public place sound amplification system, in order to ensure the effectiveness of the broadcast system, it is important to listen clearly and intuitively, so the first scoring weight of the loudness score accounts for a larger proportion; the first scoring weight of the center frequency score accounts for a relatively low proportion, but it has a basic role in system design and unit matching; the flatness has a lower impact on the detection of the state of the speaker, so the first scoring weight of the flatness score also accounts for a relatively low proportion, but the flatness score can be used to judge whether the speaker can reproduce all frequencies of sound evenly, avoid some frequencies being too strong or too weak, and thus ensure the naturalness and consistency of the sound quality. Clean sound is essential for high-quality audio experience, so the first scoring weight of the signal-to-noise ratio accounts for a relatively higher proportion than the first scoring weight of the center frequency score and the flatness score.
[0106] According to the embodiments of the present application, the method further comprises:
[0107] The first scoring weight of the loudness score, the center frequency score, the flatness score and the signal-to-noise ratio score is dynamically adjusted by the acoustic scene classification ASC based on the detected environment.
[0108] It should be noted that the acoustic scene classification ASC is constructed based on a deep learning model, and the acoustic scene classification ASC is used to analyze the collected audio data, identify different factory buildings (for example, production factory building, auxiliary production factory building, power plant building, storage building, etc.), and give the corresponding first scoring weight of the loudness score, the center frequency score, the flatness score and the signal-to-noise ratio score based on the identified scene.
[0109] Figure 2 A flowchart of the voiceprint similarity score calculation method provided by the present application is shown.
[0110] As shown in Figure 2 According to the embodiment of the present application, the collected audio data and the original audio data are input into the preset voiceprint similarity calculation model for analysis to determine the voiceprint similarity score, which includes:
[0111] S202, input the collected audio data and the original audio data into the preset voiceprint similarity calculation model, and extract the voiceprint features of the collected audio data and the original audio data respectively through the preset voiceprint similarity calculation model;
[0112] S204, calculate the similarity of the voiceprint features of the collected audio data and the original audio data to determine the voiceprint similarity score.
[0113] It should be noted that first, the original audio data and the collected audio data collected by the monitoring module need to be extracted by the preset voiceprint similarity calculation model to extract the voiceprint features. Then, the similarity of the voiceprint features of the collected audio data and the original audio data is calculated by the preset voiceprint similarity calculation model. The voiceprint feature similarity calculation method includes Euclidean distance and cosine similarity. The Euclidean distance calculates the Euclidean distance between two features, the smaller the distance, the higher the similarity; the cosine similarity calculates the cosine similarity between two features, the value closer to 1, the higher the similarity. Finally, based on the data standardization method such as 0-1 scoring method, the voiceprint similarity score of the original audio data and the collected audio data is output by the preset voiceprint similarity calculation model. The value of the voiceprint similarity score is 0-1. When the voiceprint similarity score is 1, it can be determined that they are the same audio file, and when the voiceprint similarity score is 0, the original audio data and the collected audio data are irrelevant.
[0114] Figure 3 A flowchart of the voiceprint similarity score calculation method provided by the present application is shown.
[0115] like Figure 3 As shown, according to an embodiment of the present invention, the collected audio data and the original audio data are input into a preset speech recognition model for analysis to determine the speech recognition ASR score, including:
[0116] S302, performing speech recognition on the collected audio data and the original audio data respectively using a speech recognition algorithm ASR to determine the recognized text of the collected audio data and the recognized text of the original audio data;
[0117] S304, comparing the recognized text from the collected audio data with the recognized text from the original audio data to determine the number of correctly recognized texts;
[0118] S306, calculating the ratio of the number of correctly recognized characters to the number of recognized characters in the original audio data, and determining the speech recognition ASR score.
[0119] It should be noted that the collected audio data and the original audio data are input into a preset speech recognition model, and the collected audio data and the original audio data are processed by ASR technology to obtain the collected audio data recognized text and the original audio data recognized text recognized and converted respectively, and the correct number of the collected audio data recognized text is found according to the original audio data recognized text to determine the number of correctly recognized text.
[0120] Among them, when the text recognized by the collected audio data is completely consistent with the text recognized by the original audio data, the ASR score is 1; when the text recognized by the collected audio data is completely inconsistent with the text recognized by the original audio data, the ASR score is 0.
[0121] According to an embodiment of the present invention, comparing the speaker comprehensive score with a preset evaluation threshold to determine the speaker status information includes:
[0122] The preset evaluation threshold includes a first preset evaluation threshold and a second preset evaluation threshold;
[0123] When the speaker comprehensive score is less than or equal to the first preset evaluation threshold, the speaker is determined to be faulty and a red alarm is issued;
[0124] When the speaker comprehensive score is between the first preset evaluation threshold and the second preset evaluation threshold, the speaker is determined to be suspected of failure and an orange alarm is issued;
[0125] When the comprehensive score of the speaker is greater than or equal to the second preset evaluation threshold, it is determined that the state of the speaker is normal.
[0126] It should be noted that the speaker status information includes the speaker being in good condition, suspected fault, and fault.
[0127] According to the comparison result of the loudspeaker comprehensive score and the first preset evaluation threshold and the second preset evaluation threshold, a fault warning light arranged in the scene is monitored to give a warning reminder, when the loudspeaker fails, the fault warning light is controlled to give a red warning, and the loudspeaker needs to be repaired immediately; when the loudspeaker is suspected to fail, the fault warning light is controlled to give an orange warning, that is, it is suggested that the loudspeaker be repaired; and when the loudspeaker is in good condition, the fault warning light keeps displaying green, and no repair is needed.
[0128] The first preset evaluation threshold and the second preset evaluation threshold are both set by a person skilled in the art according to actual needs, and the first preset evaluation threshold is less than the second preset evaluation threshold.
[0129] According to the embodiment of the application, the method further comprises:
[0130] The auxiliary audio data is acquired by the adjacent monitoring module.
[0131] The collected audio data is segmented according to a preset time length to obtain a plurality of sub-collected audio data.
[0132] The auxiliary audio data is segmented according to a preset time length to obtain a plurality of sub-auxiliary audio data.
[0133] The collected audio data and the original audio data are compared based on time sequence to determine a first error value of each sub-collected audio data in the collected audio data, and a first error value change curve is drawn.
[0134] The first error interval is determined by the distribution of the first error values of all the sub-collected audio data in the collected audio data.
[0135] The first error value change curve is segmented by the first error interval, and a time interval outside the first error interval is determined as an audio influence interval (the audio influence interval is marked with data, and the data marking includes an audio start time and an audio end time).
[0136] The auxiliary audio data and the original audio data are compared based on time sequence to determine a second error value of each sub-auxiliary audio data in the auxiliary audio data, and a second error value change curve is drawn (calculated based on the calculation method of the first error value of the sub-collected audio data).
[0137] The second error interval is determined by the distribution of the second error values of all the sub-auxiliary audio data in the auxiliary audio data.
[0138] It is judged whether the auxiliary audio data in the audio influence interval meets the second error interval.
[0139] If yes, the collected audio data in the audio influence interval is replaced by the auxiliary audio data.
[0140] Otherwise, no processing is performed.
[0141] It should be noted that the adjacent monitoring module can be other monitoring modules in the current monitoring scene, or it can be a monitoring module in an adjacent monitoring scene around the current monitoring scene. The number of adjacent monitoring modules selected is one or more, and the auxiliary audio data obtained by each adjacent monitoring module is not necessarily the same.
[0142] The collected audio data is often affected by ambient noise, which affects the accuracy of the speaker's comprehensive score calculation. For regular ambient noise, such as wind and rain, the monitoring module can collect ambient noise audio data before playing the original audio data, and use the noise audio data to perform noise reduction on the collected audio data. However, for irregular ambient noise, such as speech and irregular sounds generated by the operation of mechanical equipment in the venue, the noise reduction effect of collecting ambient noise is not obvious. Therefore, the noise reduction step for the collected audio data includes performing preliminary noise reduction on the collected audio data by collecting ambient noise audio data in advance, and then using auxiliary audio data collected by adjacent monitoring modules to perform auxiliary noise reduction on the collected audio data.
[0143] In the auxiliary noise reduction process, first, the collected audio data and the original audio data are time-synchronized and aligned. By comparing the frequency spectra of the same sub-collected audio data in the collected audio data and the original audio data, the energy intensity changes at different frequencies are calculated, and the first error value of each sub-collected audio data is determined. A first coordinate system is constructed with the x-axis as time and the y-axis as the first error value. The first error values of all sub-collected audio data are input into the first coordinate system. The coordinate points corresponding to all the first error values are fitted to determine the first error value change curve. The y-axis of the first coordinate system is intercepted based on the error interval range preset by the system, and the y-axis interval with the most coordinate points corresponding to the first error value is selected as the first error interval. The x-axis interval corresponding to the coordinate points outside the first error interval is determined as the audio impact interval. The audio impact interval is data-marked, and the data markers include the audio start time and the audio end time.
[0144] According to the first error value and the first error interval of the sub-acquired audio data, the auxiliary audio data and the original audio data are continuously processed, the second error value of each sub-auxiliary audio data is determined, a second coordinate system with the x-axis as time and the y-axis as the second error value is constructed, a second error value change curve is drawn based on the second coordinate system, the y-axis of the second coordinate system is intercepted based on the system preset error interval, and the y-axis interval with the most coordinate points corresponding to the second error value is selected as the second error interval. The auxiliary audio data is intercepted based on the audio influence interval, and it is judged whether the auxiliary audio data of the intercepted part is in the second error interval. If it is, it means that the auxiliary audio data of the intercepted part is less affected or not affected by environmental factors, and the collected audio data of the audio influence interval is replaced by the auxiliary audio data, so as to eliminate the influence of environmental factors on the collected audio data.
[0145] The preset time length is set by a person skilled in the art according to actual needs.
[0146] In addition, after the replacement of the collected audio data of the audio influence interval by the auxiliary audio data is completed, if there is still an audio influence interval to be replaced, the remaining audio influence data that has not been replaced is spliced, the spliced audio data is played again through the broadcast system, and the corresponding audio data is collected for supplementary verification. After each supplementary verification is completed, the replacement audio proportion (the ratio of the replacement audio time length to the collected audio data time length) of the current supplementary verification is counted. If the replacement audio proportion of the current supplementary verification is greater than the system preset audio proportion threshold, the next supplementary verification is performed; otherwise, the supplementary verification is ended.
[0147] According to the embodiment of the application, the collected audio data of the audio influence interval is replaced by the auxiliary audio data, including:
[0148] The middle value of the first error interval is determined as the first correction coefficient;
[0149] The middle value of the second error interval is determined as the second correction coefficient;
[0150] The audio signal of the auxiliary audio data in the audio influence interval is multiplied by the ratio of the first correction coefficient and the second correction coefficient to determine the corrected audio data;
[0151] The audio signal of the collected audio data in the audio influence interval is replaced by the corrected audio data.
[0152] It should be noted that the audio data obtained by different monitoring devices has certain differences due to factors such as the distance between the monitored device and the loudspeaker. In the process of replacing the collected audio data in the audio influence interval by the auxiliary audio data, the audio signal of the auxiliary audio data is modified by the first correction coefficient and the second correction coefficient to eliminate the influence of environmental factors on the collected audio data, so as to ensure that the audio signal of the replaced audio influence interval is consistent with the loudness and other parameters of the other part of the collected audio data to the greatest extent, and does not affect the final evaluation score of the loudspeaker state.
[0153] According to the embodiment of the present application, the collected audio data and the original audio data are compared based on time sequence to determine the first error value of each sub-collected audio data in the collected audio data, including:
[0154] Each sub-collected audio data is analyzed in sequence, and the audio time corresponding to each peak and each trough in the sub-collected audio data is determined as the verification time;
[0155] The energy intensity change rate corresponding to each preset frequency at the verification time is calculated, and the energy intensity change rate curve of the current verification time is drawn;
[0156] The average curve of the energy intensity change rate curves of all verification times is calculated;
[0157] The energy intensity change rate corresponding to each preset frequency is extracted from the average curve, and weighted calculation is performed to determine the first error value of the current sub-collected audio data.
[0158] It should be noted that in the process of determining the verification time, the peaks and troughs with a loudness less than a preset loudness threshold can be filtered. The preset loudness threshold is determined by the system based on the noise audio data of the environmental noise. When the loudness of the audio signal is less than the preset loudness threshold, it can be judged that the audio signal is generated by environmental noise, and it can be filtered, thereby reducing the operation data and improving the data processing speed.
[0159] Due to the influence of the distance between the monitored module and the loudspeaker, the loudness of the collected audio data and the original audio data may have certain deviations, resulting in differences in the waveform graph. When environmental noise data closer to the monitoring module and with higher loudness is collected, it will cover or affect the audio data played by the loudspeaker. However, in the frequency spectrum graph, the amplitude value of the frequency component of the audio signal will be enlarged or reduced in proportion. The energy intensity of each frequency in the frequency spectrum graph will be shifted as a whole, but the relative position and distribution mode of the energy intensity peak remain unchanged.
[0160] The preset frequencies are set by a person skilled in the art according to actual needs, for example, 1 kHz, 2 kHz, 4 kHz, 6 kHz, 8 kHz and 10 kHz, etc., the energy intensity change rates corresponding to the preset frequencies are determined by calculating the ratio of the energy intensity difference of the collected audio data and the sample audio data at the preset frequencies and the energy intensity of the sample audio data at the preset frequencies. The energy intensity change rates corresponding to the preset frequencies at the current verification time are input into the coordinate axis with the frequency as the horizontal axis and the energy intensity change rate as the vertical axis to determine the energy intensity change rate curve of the current verification time. When the energy intensity change rate curves of all verification times are determined, the average curve of the energy intensity change rate curves of all verification times is calculated by matlab, the energy intensity change rates corresponding to the preset frequencies are extracted from the obtained average curve, the energy intensity change rates corresponding to the preset frequencies are multiplied by the corresponding influence weights respectively, and the first error value of the current sub-collected audio data is determined by accumulating the calculation results. The influence weights corresponding to the energy intensity change rates corresponding to the preset frequencies are set by a person skilled in the art based on the frequency range that can be heard by the human ear, and the higher the sound sensitivity, the higher the corresponding influence weight.
[0161] According to the embodiment of the application, the method further comprises:
[0162] In the spectrogram, the coordinate points with the energy intensity change rate greater than the preset change rate threshold are marked as abnormal, and the abnormal coordinates are determined;
[0163] Based on the abnormal coordinates, other abnormal coordinates closest to the pixels of the abnormal coordinates are selected to form a candidate box through iteration;
[0164] The number of abnormal coordinates in the candidate box is counted to determine the abnormal coordinate proportion of the candidate box;
[0165] The corresponding preset pixel proportion threshold is determined according to the pixel area of the candidate box;
[0166] When the abnormal coordinate proportion of the candidate box is greater than or equal to the corresponding preset pixel proportion threshold, the iteration is continued, other abnormal coordinates closest to the pixels of the candidate box are selected, and the candidate box is updated;
[0167] When the abnormal coordinate proportion of the candidate box is less than the corresponding preset pixel proportion threshold, the abnormal coordinates are filtered;
[0168] When the abnormal coordinate proportion of the updated candidate box is greater than or equal to the corresponding preset pixel proportion threshold, the iteration is continued, and the candidate box is updated; when the abnormal coordinate proportion of the updated candidate box is less than the corresponding preset pixel proportion threshold, the iteration is ended, and the area in the candidate box is determined as an abnormal area;
[0169] The spectrogram in the abnormal area is replaced locally;
[0170] In the first error value process of the sub-acquired audio data, the energy intensity change rate corresponding to the preset frequency in the abnormal area is filtered.
[0171] It should be noted that the collected audio data includes a waveform graph and a frequency spectrum graph. In the frequency spectrum graph, when the energy intensity change rate of a certain preset frequency at a certain verification time is greater than a preset change rate threshold, it can be determined that the frequency signal is generated by environmental noise, not audio data played by the broadcast system. The coordinate point corresponding to the verification time and the preset frequency band is marked as an abnormal coordinate. By traversing the abnormal coordinates existing in the collected audio data, one or more abnormal areas are determined based on the system set abnormal area framing rule. According to the replacement method of the audio influence interval of the collected audio data by the auxiliary audio data, a part of the frequency spectrum graph corresponding to the abnormal area is selected from the auxiliary audio data, and the frequency spectrum graph in the abnormal area is locally replaced. The influence of environmental noise can be eliminated, and the modification of the original collected audio data is reduced.
[0172] The preset change rate threshold and the preset pixel ratio threshold are set by those skilled in the art according to actual needs. The preset pixel ratio threshold is determined according to the pixel area of the candidate box. The larger the pixel area of the candidate box, the higher the corresponding preset pixel ratio threshold. For example, when the pixel area of the candidate box is 30, the preset pixel ratio threshold is 0.5; when the pixel area of the candidate box is 50, the preset pixel ratio threshold is 0.58.
[0173] In addition, since the abnormal area in the frequency spectrum graph of the collected audio data is modified by local replacement, when calculating the first error value of the sub-collected audio data containing the abnormal area, the energy intensity change rate corresponding to the preset frequency in the abnormal area is filtered. Similarly, when calculating the second error value of the sub-auxiliary audio data containing the abnormal area, the energy intensity change rate corresponding to the preset frequency in the abnormal area is also filtered.
[0174] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant national and regional laws, regulations and standards. For example, the "collected audio data", "original audio data" and the like involved in the present disclosure are obtained under full authorization.
[0175] The application discloses a nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion, and the method comprises the following steps: acquiring collected audio data through a monitoring module; acquiring original audio data played by a broadcast system; performing sound event detection on the collected audio data based on the original audio data to determine a SED comprehensive score; inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score; inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition ASR score; performing weighted calculation on the SED comprehensive score, the voiceprint similarity score and the speech recognition ASR score to determine a loudspeaker comprehensive score; and comparing the loudspeaker comprehensive score with a preset evaluation threshold to determine loudspeaker state information. The sound event detection, audio similarity and speech recognition algorithm triple verification can improve the loudspeaker fault detection accuracy.
[0176] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. The above described device embodiments are only schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling, direct coupling or communication connection between the components can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0177] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units; they can be located in one place, or distributed on multiple network units; and part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0178] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.
[0179] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by relevant hardware of program instructions, and the foregoing program can be stored in a computer readable storage medium, and the program performs the steps of the above-mentioned method embodiments when executed; and the foregoing storage medium includes: a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0180] Alternatively, the integrated unit of the present application can also be stored in a computer readable storage medium if it is realized in the form of a software function module and sold or used as an independent product. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: a mobile storage device, a ROM, a RAM, a magnetic disk or an optical disk, and various media that can store program codes.
Claims
1. A method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion, characterized in that: include: Acquire collected audio data through the monitoring module; Get the original audio data played by the broadcasting system; Performing sound event detection on the collected audio data based on the original audio data to determine a SED comprehensive score; Inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score; Inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition ASR score; Perform weighted calculation on the SED comprehensive score, voiceprint similarity score, and speech recognition ASR score to determine the speaker comprehensive score; The speaker's comprehensive score is compared with a preset evaluation threshold to determine the speaker status information.
2. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: The performing sound event detection on the collected audio data based on the original audio data to determine the SED comprehensive score includes: Performing a root mean square calculation on the collected audio data to determine an RMS value of the collected audio data; Calculating a ratio of an RMS value of the collected audio data to a preset RMS value threshold to determine a loudness score; Performing center frequency calculation on the collected audio data and the original audio data to determine a center frequency of the collected audio data and a center frequency of the original audio data; Calculate the center frequency score by using the center frequency of the collected audio data and the center frequency of the original audio data; ; Where P is the center frequency score, B is the preset center frequency threshold, f1 is the center frequency of the collected audio data, and f0 is the center frequency of the original audio data. Performing flatness calculation on the collected audio data and the original audio data to determine a flatness score; Calculating a signal-to-noise ratio using the collected audio data and the original audio data to determine a signal-to-noise ratio score; The loudness score, center frequency score, flatness score, and signal-to-noise ratio score are multiplied by the corresponding first scoring weights, and the calculation results are accumulated to determine the SED comprehensive score.
3. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 2, characterized in that: Also includes: Based on the detection environment, the first scoring weights of the loudness score, center frequency score, flatness score and signal-to-noise ratio score are dynamically adjusted through acoustic scene classification ASC.
4. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: The inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score includes: Inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model, and extracting the voiceprint features of the collected audio data and the voiceprint features of the original audio data respectively through the preset voiceprint similarity calculation model; A similarity calculation is performed on the voiceprint features of the collected audio data and the voiceprint features of the original audio data to determine a voiceprint similarity score.
5. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: The step of inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition ASR score includes: Performing speech recognition on the collected audio data and the original audio data respectively by using a speech recognition algorithm ASR to determine the recognized text of the collected audio data and the recognized text of the original audio data; Comparing the recognized text in the collected audio data with the recognized text in the original audio data to determine the number of correctly recognized texts; The ratio of the number of correctly recognized characters to the number of recognized characters in the original audio data is calculated to determine the speech recognition ASR score.
6. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: The step of comparing the speaker comprehensive score with a preset evaluation threshold to determine the speaker status information includes: The preset evaluation threshold includes a first preset evaluation threshold and a second preset evaluation threshold; When the comprehensive score of the speaker is less than or equal to a first preset evaluation threshold, the speaker is determined to be faulty and a red alarm is issued; When the speaker comprehensive score is between a first preset evaluation threshold and a second preset evaluation threshold, the speaker is determined to be suspected of failure and an orange alarm is issued; When the comprehensive score of the speaker is greater than or equal to a second preset evaluation threshold, it is determined that the speaker is in good condition.
7. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: Also includes: Acquire auxiliary audio data through adjacent monitoring modules; Segment the collected audio data according to a preset time length to obtain a plurality of sub-collected audio data; Splitting the auxiliary audio data according to a preset time length to obtain a plurality of sub-auxiliary audio data; Comparing the collected audio data with the original audio data based on a time sequence, determining a first error value of each sub-collected audio data in the collected audio data, and drawing a first error value change curve; determining a first error interval based on a distribution of first error values of all sub-collected audio data in the collected audio data; Segmenting the first error value variation curve according to the first error interval, and determining a time interval outside the first error interval as an audio impact interval; Comparing the auxiliary audio data with the original audio data based on a time sequence, determining a second error value of each sub-auxiliary audio data in the auxiliary audio data, and drawing a second error value change curve; determining a second error interval according to a distribution of second error values of all sub-auxiliary audio data in the auxiliary audio data; determining whether the auxiliary audio data within the audio impact interval satisfies a second error interval; If so, the collected audio data of the audio impact interval is replaced by the auxiliary audio data; Otherwise, no processing is performed.
8. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 7, characterized in that: The replacing of the collected audio data in the audio impact interval by the auxiliary audio data includes: determining a middle value of the first error interval as a first correction coefficient; determining a middle value of the second error interval as a second correction coefficient; multiplying the auxiliary audio sub-data of the audio impact interval by the ratio of the first correction coefficient to the second correction coefficient to determine the corrected audio data; The audio signal of the collected audio data in the audio impact interval is replaced by correcting the audio data.
9. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 7, characterized in that: The comparing the collected audio data with the original audio data based on the time sequence to determine the first error value of each sub-collected audio data in the collected audio data includes: Analyze each sub-collected audio data in turn, and determine the audio time corresponding to each peak and each trough in the sub-collected audio data as the verification time; Calculate the energy intensity change rate corresponding to each preset frequency under the verification time, and draw the energy intensity change rate curve of the current verification time; Calculate the average curve of the energy intensity change rate curves for all verification times; The energy intensity change rate corresponding to each preset frequency is extracted from the average curve, and weighted calculation is performed to determine a first error value of the current sub-collected audio data.
10. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 9, characterized in that: Also includes: In the spectrum graph, the coordinate points whose energy intensity change rate is greater than the preset change rate threshold are marked as abnormal, and the abnormal coordinates are determined; Traverse based on the abnormal coordinates and select other abnormal coordinates with the closest pixel distance to the abnormal coordinates to form a candidate frame; Counting the number of abnormal coordinates in the candidate frame and determining the abnormal coordinate ratio of the candidate frame; Determine a corresponding preset pixel ratio threshold according to the pixel area of the candidate frame; When the abnormal coordinate ratio of the candidate frame is greater than or equal to the corresponding preset pixel ratio threshold, continue traversing, select other abnormal coordinates closest to the candidate frame pixels, and update the candidate frame; When the abnormal coordinate ratio of the candidate frame is less than the corresponding preset pixel ratio threshold, the abnormal coordinates are filtered; When the abnormal coordinate ratio of the updated candidate frame is greater than or equal to the corresponding preset pixel ratio threshold, continue traversing and update the candidate frame; When the abnormal coordinate ratio of the updated candidate frame is less than the corresponding preset pixel ratio threshold, the traversal ends and the area within the candidate frame is determined as an abnormal area; Partially replacing the spectrum graph within the abnormal area; In the first error value process of sub-collecting audio data, the energy intensity change rate corresponding to the preset frequency in the abnormal area is filtered.
Citation Information
Patent Citations
Loudspeaker defect identification method based on voice broadcast tone quality
CN118413800A
Disaster prevention broadcasting evaluation device and disaster prevention broadcasting evaluation method
JP2023142816A