Nuclear power plant loudspeaker state monitoring method based on multi-algorithm fusion
Through the multi-algorithm fusion method of nuclear power plant speaker status monitoring, combined with intelligent sensors and deep learning models, the problems of low efficiency of traditional manual inspections and high misjudgment rate of single detection indicators are solved, and real-time detection and accurate assessment of speaker faults are achieved.
Patent Information
- Application Number
- CN202511103565.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-07
AI Technical Summary
In existing technologies, traditional manual inspections of nuclear power plant speakers are inefficient, costly, and difficult to detect faults in real time. Single detection indicators cannot adapt to complex acoustic environments and lack intelligent analysis, resulting in a high misjudgment rate.
A nuclear power plant speaker status monitoring method based on multi-algorithm fusion is adopted. By deploying intelligent sensor modules and combining acoustic feature extraction with deep learning models, triple verification of sound event detection, audio similarity and speech recognition is performed to obtain a comprehensive speaker score, and dynamic evaluation and alarm are carried out.
It achieves real-time detection of speaker faults, reduces maintenance costs, improves fault detection accuracy, shortens fault response time, and reduces safety risks.
Smart Images

Figure CN120602881A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speaker status monitoring, and more specifically, to a method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion. Background Art
[0002] Nuclear power plant broadcast and sound reinforcement systems are essential tools for daily operations and emergency communications. Their speakers are widely distributed and located in complex environments. Traditional manual inspections are inefficient, costly, and difficult to detect faults in a timely manner. Existing audio monitoring technologies often rely on single metrics (such as loudness) or simple comparisons, failing to adapt to the diverse acoustic environments found in nuclear power plants and meeting the complex needs of nuclear power plants. Existing technologies suffer from the following issues: Reliance on manual inspections: Regular inspections are time-consuming and labor-intensive, and cannot detect sudden faults in real time. Manual inspections are easily affected by subjective factors, especially in complex acoustic environments (such as noisy turbine plants), where subtle faults are difficult to detect.
[0003] Limitations of a single detection indicator: Existing technologies are mostly based on a single indicator (such as loudness or signal-to-noise ratio) and lack multi-dimensional data fusion analysis, resulting in a high misjudgment rate.
[0004] Lack of intelligent analysis: Intelligent algorithms such as deep learning have not been introduced, and the detection logic cannot be automatically optimized according to the acoustic scenario.
[0005] Therefore, the prior art has defects and is in urgent need of improvement. Summary of the Invention
[0006] In response to the above challenges, the present invention aims to provide a nuclear power plant loudspeaker status monitoring method based on multi-algorithm fusion. By deploying intelligent sensor modules, combined with acoustic feature extraction and deep learning models, this method enables real-time health diagnosis of nuclear power plant loudspeakers, reducing maintenance costs and improving the reliability of the broadcasting system. Through triple verification using SED, audio similarity, and ASR, fault detection accuracy is significantly improved. Flexible wireless deployment reduces installation and maintenance costs. Real-time alarm generation shortens fault response time and mitigates safety risks.
[0007] A first aspect of the present invention provides a method for monitoring the status of a speaker in a nuclear power plant based on multi-algorithm fusion, comprising: Acquire collected audio data through the monitoring module; Get the original audio data played by the broadcasting system; Performing sound event detection on the collected audio data based on the original audio data to determine a SED comprehensive score; Inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score; Inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition ASR score; Perform weighted calculation on the SED comprehensive score, voiceprint similarity score, and speech recognition ASR score to determine the speaker comprehensive score; The speaker comprehensive score is compared with a preset evaluation threshold to determine the speaker status information.
[0008] In this solution, performing sound event detection on the collected audio data based on the original audio data to determine the SED comprehensive score includes: Performing a root mean square calculation on the collected audio data to determine an RMS value of the collected audio data; Calculating a ratio of an RMS value of the collected audio data to a preset RMS value threshold to determine a loudness score; Performing center frequency calculation on the collected audio data and the original audio data to determine a center frequency of the collected audio data and a center frequency of the original audio data; Calculate the center frequency score by using the center frequency of the collected audio data and the center frequency of the original audio data; ; Where P is the center frequency score, B is the preset center frequency threshold, f1 is the center frequency of the collected audio data, and f0 is the center frequency of the original audio data. Performing flatness calculation on the collected audio data and the original audio data to determine a flatness score; Calculating a signal-to-noise ratio using the collected audio data and the original audio data to determine a signal-to-noise ratio score; The loudness score, center frequency score, flatness score, and signal-to-noise ratio score are multiplied by the corresponding first scoring weights, and the calculation results are accumulated to determine the SED comprehensive score.
[0009] This plan also includes: Based on the detection environment, the first scoring weights of the loudness score, center frequency score, flatness score and signal-to-noise ratio score are dynamically adjusted through acoustic scene classification ASC.
[0010] In this solution, the collected audio data and the original audio data are input into a preset voiceprint similarity calculation model for analysis to determine the voiceprint similarity score, including: Inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model, and extracting the voiceprint features of the collected audio data and the voiceprint features of the original audio data respectively through the preset voiceprint similarity calculation model; A similarity calculation is performed on the voiceprint features of the collected audio data and the voiceprint features of the original audio data to determine a voiceprint similarity score.
[0011] In this solution, the collected audio data and the original audio data are input into a preset speech recognition model for analysis to determine the speech recognition ASR score, including: Performing speech recognition on the collected audio data and the original audio data respectively by using a speech recognition algorithm ASR to determine the recognized text of the collected audio data and the recognized text of the original audio data; Comparing the recognized text in the collected audio data with the recognized text in the original audio data to determine the number of correctly recognized texts; The ratio of the number of correctly recognized characters to the number of recognized characters in the original audio data is calculated to determine the speech recognition ASR score.
[0012] In this solution, comparing the speaker comprehensive score with a preset evaluation threshold to determine the speaker status information includes: The preset evaluation threshold includes a first preset evaluation threshold and a second preset evaluation threshold; When the comprehensive score of the speaker is less than or equal to a first preset evaluation threshold, the speaker is determined to be faulty and a red alarm is issued; When the speaker comprehensive score is between a first preset evaluation threshold and a second preset evaluation threshold, the speaker is determined to be suspected of failure and an orange alarm is issued; When the comprehensive score of the speaker is greater than or equal to a second preset evaluation threshold, it is determined that the speaker is in good condition.
[0013] This plan also includes: Acquire auxiliary audio data through adjacent monitoring modules; Segment the collected audio data according to a preset time length to obtain a plurality of sub-collected audio data; Splitting the auxiliary audio data according to a preset time length to obtain a plurality of sub-auxiliary audio data; Comparing the collected audio data with the original audio data based on a time sequence, determining a first error value of each sub-collected audio data in the collected audio data, and drawing a first error value change curve; determining a first error interval based on a distribution of first error values of all sub-collected audio data in the collected audio data; Segmenting the first error value variation curve according to the first error interval, and determining a time interval outside the first error interval as an audio impact interval; Comparing the auxiliary audio data with the original audio data based on a time sequence, determining a second error value of each sub-auxiliary audio data in the auxiliary audio data, and drawing a second error value change curve; determining a second error interval according to a distribution of second error values of all sub-auxiliary audio data in the auxiliary audio data; determining whether the auxiliary audio data within the audio impact interval satisfies a second error interval; If so, the collected audio data of the audio impact interval is replaced by the auxiliary audio data; Otherwise, no processing is performed.
[0014] In this solution, the replacement of the collected audio data in the audio impact interval with the auxiliary audio data includes: determining a middle value of the first error interval as a first correction coefficient; determining a middle value of the second error interval as a second correction coefficient; multiplying the auxiliary audio sub-data of the audio impact interval by the ratio of the first correction coefficient to the second correction coefficient to determine the corrected audio data; The audio signal of the collected audio data in the audio impact interval is replaced by correcting the audio data.
[0015] In this solution, comparing the collected audio data with the original audio data based on the time sequence to determine the first error value of each sub-collected audio data in the collected audio data includes: Analyze each sub-collected audio data in turn, and determine the audio time corresponding to each peak and each trough in the sub-collected audio data as the verification time; Calculate the energy intensity change rate corresponding to each preset frequency under the verification time, and draw the energy intensity change rate curve of the current verification time; Calculate the average curve of the energy intensity change rate curves for all verification times; The energy intensity change rate corresponding to each preset frequency is extracted from the average curve, and weighted calculation is performed to determine a first error value of the current sub-collected audio data.
[0016] This plan also includes: In the spectrum graph, the coordinate points whose energy intensity change rate is greater than the preset change rate threshold are marked as abnormal, and the abnormal coordinates are determined; Traverse based on the abnormal coordinates and select other abnormal coordinates with the closest pixel distance to the abnormal coordinates to form a candidate frame; Counting the number of abnormal coordinates in the candidate frame and determining the abnormal coordinate ratio of the candidate frame; Determine a corresponding preset pixel ratio threshold according to the pixel area of the candidate frame; When the abnormal coordinate ratio of the candidate frame is greater than or equal to the corresponding preset pixel ratio threshold, continue traversing, select other abnormal coordinates closest to the candidate frame pixels, and update the candidate frame; When the abnormal coordinate ratio of the candidate frame is less than the corresponding preset pixel ratio threshold, the abnormal coordinates are filtered; When the abnormal coordinate ratio of the updated candidate frame is greater than or equal to the corresponding preset pixel ratio threshold, continue traversing and update the candidate frame; when the abnormal coordinate ratio of the updated candidate frame is less than the corresponding preset pixel ratio threshold, end the traversal and determine the area within the candidate frame as an abnormal area; Partially replacing the spectrum graph within the abnormal area; In the first error value process of sub-collecting audio data, the energy intensity change rate corresponding to the preset frequency in the abnormal area is filtered.
[0017] The audio signal of the collected audio data in the audio influence interval is replaced by the corrected audio data. The present invention discloses a method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion, the method comprising: acquiring collected audio data through a monitoring module; acquiring original audio data played by a broadcasting system; performing sound event detection on the collected audio data based on the original audio data to determine the SED comprehensive score; inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine the voiceprint similarity score; inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine the speech recognition ASR score; performing weighted calculation on the SED comprehensive score, the voiceprint similarity score and the speech recognition ASR score to determine the speaker comprehensive score; comparing the speaker comprehensive score with a preset evaluation threshold to determine the speaker status information. The present invention can improve the accuracy of speaker fault detection through triple verification of sound event detection, audio similarity and speech recognition algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flowchart of a method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion provided by the present invention is shown; Figure 2 A flow chart showing the voiceprint similarity score calculation method provided by the present invention is shown; Figure 3 The flowchart of the speech recognition ASR score calculation method provided by the present invention is shown. DETAILED DESCRIPTION
[0019] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.
[0020] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0021] Figure 1 The flowchart of the method for monitoring the status of a loudspeaker in a nuclear power plant based on multi-algorithm fusion provided by the present invention is shown.
[0022] like Figure 1 As shown, the present invention discloses a method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion, comprising: S102, acquiring collected audio data through a monitoring module; S104, obtaining original audio data played by the broadcasting system; S106, performing sound event detection on the collected audio data based on the original audio data to determine a SED comprehensive score; S108, inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score; S110, inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition ASR score; S112, performing weighted calculation on the SED comprehensive score, the voiceprint similarity score, and the speech recognition ASR score to determine a comprehensive speaker score; S114 , comparing the speaker comprehensive score with a preset evaluation threshold to determine speaker status information.
[0023] According to an embodiment of the present invention, a monitoring module deployed on-site collects audio signals played by a speaker to determine collected audio data. The system compares the collected audio data with the original audio data (i.e., the audio signal played through the broadcasting system). Based on the comparison results, the speaker's health status is scored, and the score is used to determine whether the speaker's equipment is operating normally. This invention utilizes three audio signal processing and analysis methods: sound event detection (SED), voiceprint similarity, and speech recognition algorithms. It also employs a functional value scoring method (i.e., a 0-1 scoring method). The collected audio data is compared with the original audio data to obtain a comprehensive SED score, a voiceprint similarity score, and an ASR score. The SED comprehensive score, voiceprint similarity score, and ASR score are each multiplied by a corresponding second scoring weight, and the calculated results are accumulated to determine a comprehensive speaker score. The second scoring weights for the SED comprehensive score, voiceprint similarity score, and ASR score are all determined by scene recognition using acoustic scene classification (ASC). The comprehensive speaker score is compared with a preset evaluation threshold set by the system to determine speaker status information.
[0024] The quality of a speaker can often be intuitively reflected in its SED. A basic judgment of the speaker's quality can be made through a combination of relevant features. Therefore, the second scoring weight of the SED comprehensive score is set relatively high. In public address systems, due to the characteristics of speaker sound, the voiceprint similarity between the original audio data and the collected audio data cannot achieve a high score. Therefore, the second scoring weight of the voiceprint similarity score is set relatively low. Speech recognition technology can assist in evaluating speaker performance. For monitoring a single speaker, low recognition accuracy may be due to speaker distortion, uneven frequency response, or a low signal-to-noise ratio, which can indicate poor performance of the monitored speaker. In public address systems, speakers are often deployed densely. Even if a single speaker is damaged, sounds from other speakers within the test point may still be detected, allowing the content of the collected audio data to be recognized. Therefore, in environments with multiple speakers, the second scoring weight of the speech recognition (ASR) score is set relatively low.
[0025] Among them, the preset voiceprint similarity calculation model and the preset speech recognition model are both trained through historical audio data obtained during the historical monitoring process.
[0026] According to an embodiment of the present invention, performing sound event detection on collected audio data based on original audio data to determine a SED comprehensive score includes: Performing root mean square calculation on the collected audio data to determine the RMS value of the collected audio data; Calculate the ratio of the RMS value of the collected audio data to a preset RMS value threshold to determine the loudness score; Calculate the center frequency of the collected audio data and the original audio data to determine the center frequency of the collected audio data and the center frequency of the original audio data; Calculate the center frequency score by collecting the center frequency of the audio data and the center frequency of the original audio data; ; Where P is the center frequency score, B is the preset center frequency threshold, f1 is the center frequency of the collected audio data, and f0 is the center frequency of the original audio data. Performing flatness calculation on the collected audio data and the original audio data to determine a flatness score; The signal-to-noise ratio score is determined by calculating the signal-to-noise ratio of the collected audio data and the original audio data; The loudness score, center frequency score, flatness score, and signal-to-noise ratio score are multiplied by the corresponding first scoring weights, and the calculation results are accumulated to determine the SED comprehensive score.
[0027] It should be noted that sound event detection (SED) is a method for identifying and localizing specific sound events in audio signals. In sound event detection, feature extraction is used to convert raw audio data signals into meaningful information, enabling machine learning models to understand and classify sound events. In basic audio signal operations, extraction tools are used to extract features such as loudness, center frequency, flatness, signal-to-noise ratio, and sampling rate. These features are crucial for audio quality assessment, noise suppression, audio equipment calibration, and audio signal health monitoring.
[0028] In audio processing software, the loudness metric typically uses RMS (Root Mean Square) to measure the average level or average power of an audio signal. This is used to quantify the loudness or volume of audio data, allowing for standardization of playback volume across different audio formats and ensuring consistency across different platforms and devices. The loudness of the collected audio data is calculated using the system's preset RMS calculation method, and the results are standardized using methods such as a 0-1 scoring system to obtain the RMS value of the collected audio data. To determine whether the loudness metric can reach 1 point, the same type of speakers must be operated at full power in the broadcast system. The maximum RMS value of the audio data collected by the speaker monitoring module serves as the preset RMS value threshold for this type of speaker.
[0029] The audio center frequency is usually measured in Hz. In audio processing, it is considered the "center" of the spectral energy of the audio data signal. Generally, pre-produced audio data and collected audio data should theoretically be consistent. When the center frequency is offset, it indicates that the speaker or amplifier system's audio reproduction is inaccurate, the sound is unbalanced, and responses in certain frequency bands are missing, affecting the sound quality experience. The system's preset center frequency calculation method is used to calculate the center frequency of the collected audio data and the original audio data, respectively. The center frequency score is calculated based on the collected audio data center frequency and the original audio data center frequency. The calculated center frequency score is normalized using methods such as the 0-1 scoring method to determine the final center frequency score. When the center frequency score is 1, the original audio data and the collected audio data played by the speaker collected by the monitoring module are consistent in center frequency. The preset center frequency threshold B is determined based on the center frequency deviation frequency range acceptable to the human ear.
[0030] Flatness is a metric that quantifies the uniformity of the spectral distribution of audio data. The closer this value is to 1, the more similar the sample's spectrum is to white noise, meaning it's flatter. Conversely, if it's far less than 1, the sample's spectrum has significant peaks, making it more diverse and complex. The system's preset flatness calculation method calculates the flatness of the collected audio data and the original audio data, respectively. The ratio of the original audio data flatness to the collected audio data flatness is then used to determine the flatness score. The flatness score is then normalized using methods such as a 0-1 scoring system. When the collected audio data is less than or equal to the original audio data, the flatness score is 1.
[0031] The signal-to-noise ratio (SNR) is measured in decibels (dB), making it ideal for expressing the relative magnitude between two quantities, particularly when comparing very large or very small values, such as the ratio of audio noise intensities. Due to factors such as the amplifier system and the environment, the SNR of the audio signal received at the speaker is often lower than the SNR of the played audio signal. The system's pre-defined SNR calculation method calculates the SNR of the collected audio data and the SNR of the original audio data. The SNR score is then determined by calculating the ratio of the SNR of the original audio data to the SNR of the collected audio data. The SNR score is then normalized using methods such as a 0-1 scale. When the SNR of the collected audio data equals the SNR of the original audio data, the SNR score is 1; when the SNR of the collected audio data is 0, the SNR score is 0; and when the SNR of the collected audio data is greater than the SNR of the original audio data, it is considered that the system is experiencing interference or the speaker is not functioning properly.
[0032] The first-level weightings for loudness, center frequency, flatness, and signal-to-noise ratio are set based on how people perceive these four characteristics when a loudspeaker is amplified. The initial values for the first-level weightings for loudness, center frequency, flatness, and signal-to-noise ratio are 60%, 10%, 10%, and 20%, respectively. In public sound reinforcement systems, clarity and intuitive listening are crucial to ensuring system effectiveness, so the loudness score carries a higher weighting. The center frequency score has a relatively lower weighting, but it plays a fundamental role in system design and driver matching. Flatness has a relatively low impact on loudspeaker status detection, so its weighting is also relatively low. However, the flatness score can be used to determine whether a loudspeaker can reproduce all frequencies evenly, avoiding overly strong or weak frequencies, thereby ensuring natural and consistent sound quality. Clean sound is crucial to a high-quality audio experience, so the signal-to-noise ratio has a higher weighting than the center frequency and flatness scores.
[0033] According to an embodiment of the present invention, the further embodiment includes: Based on the detection environment, the first scoring weights of the loudness score, center frequency score, flatness score and signal-to-noise ratio score are dynamically adjusted through acoustic scene classification ASC.
[0034] It should be noted that the acoustic scene classification ASC is built based on a deep learning model. Through the acoustic scene classification ASC, the collected audio data is analyzed, different factories are identified (such as production factories, auxiliary production factories, power factories, storage warehouses, etc.), and the corresponding loudness score, center frequency score, flatness score and signal-to-noise ratio score are given based on the identified scene. The first scoring weight.
[0035] Figure 2 The flowchart of the voiceprint similarity score calculation method provided by the present invention is shown.
[0036] like Figure 2 As shown, according to an embodiment of the present invention, the collected audio data and the original audio data are input into a preset voiceprint similarity calculation model for analysis to determine the voiceprint similarity score, including: S202, inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model, and extracting the voiceprint features of the collected audio data and the voiceprint features of the original audio data respectively through the preset voiceprint similarity calculation model; S204, performing similarity calculation on the voiceprint features of the collected audio data and the voiceprint features of the original audio data to determine a voiceprint similarity score.
[0037] It should be noted that, first, the voiceprint features of the original audio data and the collected audio data collected by the monitoring module are extracted using a preset voiceprint similarity calculation model. Then, the preset voiceprint similarity calculation model is used to calculate the similarity between the voiceprint features of the collected audio data and the original audio data. Voiceprint feature similarity calculation methods include Euclidean distance and cosine similarity. Euclidean distance calculates the Euclidean distance between two features; smaller distances indicate higher similarity; cosine similarity calculates the cosine similarity between two features; closer to 1, higher similarity. Finally, based on data normalization methods such as the 0-1 scoring method, the preset voiceprint similarity calculation model outputs a voiceprint similarity score for the original audio data and the collected audio data. The voiceprint similarity score ranges from 0 to 1. A voiceprint similarity score of 1 indicates that the audio files are identical. A voiceprint similarity score of 0 indicates that the original and collected audio data are unrelated.
[0038] Figure 3 The flowchart of the speech recognition ASR score calculation method provided by the present invention is shown.
[0039] like Figure 3 As shown, according to an embodiment of the present invention, the collected audio data and the original audio data are input into a preset speech recognition model for analysis to determine the speech recognition ASR score, including: S302, performing speech recognition on the collected audio data and the original audio data respectively using a speech recognition algorithm ASR to determine the recognized text of the collected audio data and the recognized text of the original audio data; S304, comparing the recognized text from the collected audio data with the recognized text from the original audio data to determine the number of correctly recognized texts; S306, calculating the ratio of the number of correctly recognized characters to the number of recognized characters in the original audio data, and determining the speech recognition ASR score.
[0040] It should be noted that the collected audio data and the original audio data are input into a preset speech recognition model, and the collected audio data and the original audio data are processed by ASR technology to obtain the collected audio data recognized text and the original audio data recognized text recognized and converted respectively, and the correct number of the collected audio data recognized text is found according to the original audio data recognized text to determine the number of correctly recognized text.
[0041] Among them, when the text recognized by the collected audio data is completely consistent with the text recognized by the original audio data, the ASR score is 1; when the text recognized by the collected audio data is completely inconsistent with the text recognized by the original audio data, the ASR score is 0.
[0042] According to an embodiment of the present invention, comparing the speaker comprehensive score with a preset evaluation threshold to determine the speaker status information includes: The preset evaluation threshold includes a first preset evaluation threshold and a second preset evaluation threshold; When the speaker comprehensive score is less than or equal to the first preset evaluation threshold, the speaker is determined to be faulty and a red alarm is issued; When the speaker comprehensive score is between the first preset evaluation threshold and the second preset evaluation threshold, the speaker is determined to be suspected of failure and an orange alarm is issued; When the comprehensive score of the speaker is greater than or equal to the second preset evaluation threshold, it is determined that the state of the speaker is normal.
[0043] It should be noted that the speaker status information includes the speaker being in good condition, suspected fault, and fault.
[0044] Based on the comparison results of the comprehensive score of the speaker with the first preset evaluation threshold and the second preset evaluation threshold, a warning reminder is issued through the fault warning light set in the monitoring scene. When the speaker fails, the fault warning light is controlled to give a red alarm, and the speaker needs to be repaired immediately; when the speaker is suspected of failing, the fault warning light is controlled to give an orange alarm, which means that it is recommended to repair the speaker; when the speaker is in good condition, the fault warning light remains green and no repair is required.
[0045] Among them, the first preset evaluation threshold and the second preset evaluation threshold are both set by those skilled in the art according to actual needs, and the first preset evaluation threshold is smaller than the second preset evaluation threshold.
[0046] According to an embodiment of the present invention, the further embodiment includes: Acquire auxiliary audio data through adjacent monitoring modules; Segment the collected audio data according to a preset time length to obtain a plurality of sub-collected audio data; Splitting the auxiliary audio data according to a preset time length to obtain a plurality of sub-auxiliary audio data; Comparing the collected audio data with the original audio data based on a time sequence, determining a first error value of each sub-collected audio data in the collected audio data, and drawing a first error value change curve; Determining a first error interval based on a distribution of first error values of all sub-collected audio data in the collected audio data; Segmenting the first error value variation curve by the first error interval, and determining the time interval outside the first error interval as the audio impact interval (data marking is performed on the audio impact interval, and the data marking includes the audio start time and the audio end time); Comparing the auxiliary audio data with the original audio data based on a time sequence, determining a second error value of each sub-auxiliary audio data in the auxiliary audio data, and drawing a second error value change curve (calculated based on the calculation method of the first error value of the sub-collected audio data); determining a second error interval based on a distribution of second error values of all sub-auxiliary audio data in the auxiliary audio data; determining whether the auxiliary audio data within the audio impact interval satisfies a second error interval; If so, the collected audio data of the audio impact interval is replaced by the auxiliary audio data; Otherwise, no processing is performed.
[0047] It should be noted that the adjacent monitoring module can be other monitoring modules in the current monitoring scene, or it can be a monitoring module in an adjacent monitoring scene around the current monitoring scene. The number of adjacent monitoring modules selected is one or more, and the auxiliary audio data obtained by each adjacent monitoring module is not necessarily the same.
[0048] The collected audio data is often affected by ambient noise, which affects the accuracy of the speaker's comprehensive score calculation. For regular ambient noise, such as wind and rain, the monitoring module can collect ambient noise audio data before playing the original audio data, and use the noise audio data to perform noise reduction on the collected audio data. However, for irregular ambient noise, such as speech and irregular sounds generated by the operation of mechanical equipment in the venue, the noise reduction effect of collecting ambient noise is not obvious. Therefore, the noise reduction step for the collected audio data includes performing preliminary noise reduction on the collected audio data by collecting ambient noise audio data in advance, and then using auxiliary audio data collected by adjacent monitoring modules to perform auxiliary noise reduction on the collected audio data.
[0049] In the auxiliary noise reduction process, first, the collected audio data and the original audio data are time-synchronized and aligned. By comparing the frequency spectra of the same sub-collected audio data in the collected audio data and the original audio data, the energy intensity changes at different frequencies are calculated, and the first error value of each sub-collected audio data is determined. A first coordinate system is constructed with the x-axis as time and the y-axis as the first error value. The first error values of all sub-collected audio data are input into the first coordinate system. The coordinate points corresponding to all the first error values are fitted to determine the first error value change curve. The y-axis of the first coordinate system is intercepted based on the error interval range preset by the system, and the y-axis interval with the most coordinate points corresponding to the first error value is selected as the first error interval. The x-axis interval corresponding to the coordinate points outside the first error interval is determined as the audio impact interval. The audio impact interval is data-marked, and the data markers include the audio start time and the audio end time.
[0050] Continue to process the auxiliary audio data and the original audio data according to the calculation steps of the first error value and the first error interval of the sub-collected audio data, determine the second error value of each sub-auxiliary audio data, construct a second coordinate system with time on the x-axis and the second error value on the y-axis, draw a second error value change curve based on the second coordinate system, intercept the y-axis of the second coordinate system based on the error interval preset by the system, and select the y-axis interval with the most coordinate points corresponding to the second error value as the second error interval. Intercept the auxiliary audio data based on the audio impact interval, and determine whether the intercepted auxiliary audio data is in the second error interval. If so, it means that the intercepted auxiliary audio data is less affected or not affected by environmental factors. The collected audio data in the audio impact interval is replaced by the auxiliary audio data, thereby eliminating the influence of environmental factors on the collected audio data.
[0051] The preset time length is set by those skilled in the art according to actual needs.
[0052] In addition, after the replacement of the collected audio data of the audio impact interval by the auxiliary audio data is completed, if there is still an audio impact interval to be replaced, the remaining audio impact data that has not been replaced is spliced, and the spliced audio data is played again through the broadcasting system, and the corresponding audio data is collected for supplementary verification. After each supplementary verification is completed, the proportion of replaced audio in the supplementary verification is counted (the ratio of the length of the replaced audio time to the length of the collected audio data time). If the proportion of replaced audio in the supplementary verification is greater than the audio proportion threshold preset by the system, the next supplementary verification is performed; otherwise, the supplementary verification is terminated.
[0053] According to an embodiment of the present invention, replacing the collected audio data of the audio impact interval with the auxiliary audio data includes: determining a middle value of the first error interval as a first correction coefficient; determining a middle value of the second error interval as a second correction coefficient; multiplying the audio signal of the auxiliary audio data in the audio influence interval by the ratio of the first correction coefficient to the second correction coefficient to determine the corrected audio data; The audio signal of the collected audio data in the audio impact interval is replaced by correcting the audio data.
[0054] It should be noted that due to factors such as the distance between the monitoring device and the speaker, there are certain differences in the audio data obtained by different monitoring devices. In the process of replacing the collected audio data of the audio impact interval with auxiliary audio data, the audio signal of the auxiliary audio data is corrected by the first correction coefficient and the second correction coefficient to eliminate the influence of environmental factors on the collected audio data, and to ensure to the greatest extent that the audio signal of the replaced audio impact interval is consistent with the loudness and other parameters of other parts of the collected audio data, without affecting the final evaluation score of the speaker status.
[0055] According to an embodiment of the present invention, comparing the collected audio data with the original audio data based on a time sequence to determine a first error value of each sub-collected audio data in the collected audio data includes: Analyze each sub-collected audio data in turn, and determine the audio time corresponding to each peak and each trough in the sub-collected audio data as the verification time; Calculate the energy intensity change rate corresponding to each preset frequency under the verification time, and draw the energy intensity change rate curve of the current verification time; Calculate the average curve of the energy intensity change rate curves for all verification times; The energy intensity change rate corresponding to each preset frequency is extracted from the average curve, and a weighted calculation is performed to determine a first error value of the current sub-collected audio data.
[0056] It should be noted that, in the process of determining the verification time, peaks and troughs whose loudness is less than a preset loudness threshold can be filtered. The preset loudness threshold is determined by the system based on the noise audio data of the ambient noise. When the loudness of the audio signal is less than the preset loudness threshold, it can be determined that the audio signal is generated by the ambient noise and can be filtered, thereby reducing the calculation data and improving the data processing speed.
[0057] Due to the distance between the monitoring module and the speaker, the loudness of the collected audio data may differ slightly from the original audio data, resulting in differences in the waveforms. When louder ambient noise data is collected closer to the monitoring module, it can overwrite or affect the audio data played by the speaker. However, in the spectrogram, the amplitude values of the audio signal's frequency components will be proportionally amplified or reduced. The energy intensity of each frequency in the spectrogram will shift overall, but the relative position and distribution pattern of the energy intensity peaks remain unchanged.
[0058] The preset frequencies are set by those skilled in the art based on practical needs, for example, 1kHz, 2kHz, 4kHz, 6kHz, 8kHz, and 10kHz. The energy intensity change rate corresponding to each preset frequency is determined by calculating the difference in energy intensity between the collected audio data and the sample audio data at the preset frequencies, and the ratio of the energy intensity of the sample audio data at the preset frequencies. The energy intensity change rate corresponding to each preset frequency at the current verification time is input into a coordinate system with frequency on the horizontal axis and energy intensity change rate on the vertical axis to determine the energy intensity change rate curve for the current verification time. Once the energy intensity change rate curve for all verification times is determined, the average curve of the energy intensity change rate curves for all verification times is calculated using MATLAB. The energy intensity change rate corresponding to each preset frequency is extracted from the obtained average curve. The energy intensity change rate corresponding to each preset frequency is multiplied by the corresponding influence weight, and the calculated results are accumulated to determine the first error value of the current sub-collected audio data. The influence weight corresponding to the energy intensity change rate corresponding to each preset frequency is set by those skilled in the art based on the frequency range audible to the human ear. Higher sound sensitivity corresponds to higher influence weights.
[0059] According to an embodiment of the present invention, the further embodiment includes: In the spectrum graph, the coordinate points whose energy intensity change rate is greater than the preset change rate threshold are marked as abnormal, and the abnormal coordinates are determined; Traverse based on the abnormal coordinates and select other abnormal coordinates with the closest pixel distance to the abnormal coordinates to form a candidate box; Count the number of abnormal coordinates in the candidate frame and determine the proportion of abnormal coordinates in the candidate frame; Determine the corresponding preset pixel ratio threshold according to the pixel area of the candidate frame; When the abnormal coordinate ratio of the candidate frame is greater than or equal to the corresponding preset pixel ratio threshold, continue traversing, select other abnormal coordinates closest to the candidate frame pixel, and update the candidate frame; When the abnormal coordinate ratio of the candidate frame is less than the corresponding preset pixel ratio threshold, the abnormal coordinates are filtered; When the abnormal coordinate ratio of the updated candidate frame is greater than or equal to the corresponding preset pixel ratio threshold, continue traversing and update the candidate frame; when the abnormal coordinate ratio of the updated candidate frame is less than the corresponding preset pixel ratio threshold, end the traversal and determine the area within the candidate frame as an abnormal area; Perform local replacement on the spectrum graph within the abnormal area; In the first error value process of sub-collecting audio data, the energy intensity change rate corresponding to the preset frequency in the abnormal area is filtered.
[0060] It should be noted that the collected audio data includes a waveform diagram and a spectrum diagram. In the spectrum diagram, when the energy intensity change rate of a preset frequency at a certain verification time is greater than the preset change rate threshold, it can be determined that the frequency signal is generated by environmental noise, not the audio data played by the broadcasting system, and the coordinate points corresponding to the verification time and the preset frequency band are marked as abnormal coordinates. By traversing the abnormal coordinates existing in the collected audio data, one or more abnormal areas are determined based on the abnormal area selection rules set by the system, and according to the method of replacing the collected audio data of the audio impact interval through auxiliary audio data, the partial spectrum diagram corresponding to the abnormal area is selected from the auxiliary audio data, and the spectrum diagram in the abnormal area is partially replaced, which can eliminate the impact of environmental noise and reduce the modification of the original collected audio data.
[0061] The preset rate of change threshold and the preset pixel ratio threshold are set by those skilled in the art based on actual needs. The preset pixel ratio threshold is determined based on the pixel area of the candidate box. The larger the pixel area of the candidate box, the higher the corresponding preset pixel ratio threshold. For example, when the pixel area of the candidate box is 30, the preset pixel ratio threshold is 0.5; when the pixel area of the candidate box is 50, the preset pixel ratio threshold is 0.58.
[0062] Furthermore, because the abnormal region in the spectrogram of the acquired audio data is modified through local replacement, the energy intensity change rate corresponding to the preset frequency within the abnormal region is filtered when calculating the first error value of the sub-acquired audio data containing the abnormal region. Similarly, when calculating the second error value of the auxiliary audio data containing the abnormal region, the energy intensity change rate corresponding to the preset frequency within the abnormal region is also filtered.
[0063] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the "collected audio data" and "raw audio data" involved in this disclosure are all obtained with full authorization.
[0064] The present invention discloses a method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion. The method comprises: obtaining collected audio data through a monitoring module; obtaining original audio data played by a broadcasting system; performing sound event detection on the collected audio data based on the original audio data to determine a SED comprehensive score; inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score; inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition (ASR) score; performing a weighted calculation on the SED comprehensive score, the voiceprint similarity score, and the speech recognition (ASR) score to determine a comprehensive speaker score; and comparing the comprehensive speaker score with a preset evaluation threshold to determine speaker status information. The present invention improves the accuracy of speaker fault detection through triple verification using sound event detection, audio similarity, and speech recognition algorithms.
[0065] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0066] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0067] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0068] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0069] Alternatively, if the integrated units described above are implemented as software modules and sold or used as standalone products, they can also be stored on a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product, stored on a storage medium, includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute all or part of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as removable storage devices, ROM, RAM, magnetic disks, or optical disks.
Claims
1. A method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion, characterized in that: include: Acquire collected audio data through the monitoring module; Get the original audio data played by the broadcasting system; Performing sound event detection on the collected audio data based on the original audio data to determine a SED comprehensive score; Inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score; Inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition ASR score; Perform weighted calculation on the SED comprehensive score, voiceprint similarity score, and speech recognition ASR score to determine the speaker comprehensive score; The speaker comprehensive score is compared with a preset evaluation threshold to determine the speaker status information.
2. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: The performing sound event detection on the collected audio data based on the original audio data to determine the SED comprehensive score includes: Performing a root mean square calculation on the collected audio data to determine an RMS value of the collected audio data; Calculating a ratio of an RMS value of the collected audio data to a preset RMS value threshold to determine a loudness score; Performing center frequency calculation on the collected audio data and the original audio data to determine a center frequency of the collected audio data and a center frequency of the original audio data; Calculate the center frequency score by using the center frequency of the collected audio data and the center frequency of the original audio data; ; Where P is the center frequency score, B is the preset center frequency threshold, f1 is the center frequency of the collected audio data, and f0 is the center frequency of the original audio data. Performing flatness calculation on the collected audio data and the original audio data to determine a flatness score; Calculating a signal-to-noise ratio using the collected audio data and the original audio data to determine a signal-to-noise ratio score; The loudness score, center frequency score, flatness score, and signal-to-noise ratio score are multiplied by the corresponding first scoring weights, and the calculation results are accumulated to determine the SED comprehensive score.
3. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 2, characterized in that: Also includes: Based on the detection environment, the first scoring weights of the loudness score, center frequency score, flatness score and signal-to-noise ratio score are dynamically adjusted through acoustic scene classification ASC.
4. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: The inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model for analysis to determine a voiceprint similarity score includes: Inputting the collected audio data and the original audio data into a preset voiceprint similarity calculation model, and extracting the voiceprint features of the collected audio data and the voiceprint features of the original audio data respectively through the preset voiceprint similarity calculation model; A similarity calculation is performed on the voiceprint features of the collected audio data and the voiceprint features of the original audio data to determine a voiceprint similarity score.
5. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: The step of inputting the collected audio data and the original audio data into a preset speech recognition model for analysis to determine a speech recognition ASR score includes: Performing speech recognition on the collected audio data and the original audio data respectively by using a speech recognition algorithm ASR to determine the recognized text of the collected audio data and the recognized text of the original audio data; Comparing the recognized text in the collected audio data with the recognized text in the original audio data to determine the number of correctly recognized texts; The ratio of the number of correctly recognized characters to the number of recognized characters in the original audio data is calculated to determine the speech recognition ASR score.
6. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: The step of comparing the speaker comprehensive score with a preset evaluation threshold to determine the speaker status information includes: The preset evaluation threshold includes a first preset evaluation threshold and a second preset evaluation threshold; When the comprehensive score of the speaker is less than or equal to a first preset evaluation threshold, the speaker is determined to be faulty and a red alarm is issued; When the speaker comprehensive score is between a first preset evaluation threshold and a second preset evaluation threshold, the speaker is determined to be suspected of failure and an orange alarm is issued; When the comprehensive score of the speaker is greater than or equal to a second preset evaluation threshold, it is determined that the speaker is in good condition.
7. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 1, characterized in that: Also includes: Acquire auxiliary audio data through adjacent monitoring modules; Segment the collected audio data according to a preset time length to obtain a plurality of sub-collected audio data; Splitting the auxiliary audio data according to a preset time length to obtain a plurality of sub-auxiliary audio data; Comparing the collected audio data with the original audio data based on a time sequence, determining a first error value of each sub-collected audio data in the collected audio data, and drawing a first error value change curve; determining a first error interval based on a distribution of first error values of all sub-collected audio data in the collected audio data; Segmenting the first error value variation curve according to the first error interval, and determining a time interval outside the first error interval as an audio impact interval; Comparing the auxiliary audio data with the original audio data based on a time sequence, determining a second error value of each sub-auxiliary audio data in the auxiliary audio data, and drawing a second error value change curve; determining a second error interval according to a distribution of second error values of all sub-auxiliary audio data in the auxiliary audio data; determining whether the auxiliary audio data within the audio impact interval satisfies a second error interval; If so, the collected audio data of the audio impact interval is replaced by the auxiliary audio data; Otherwise, no processing is done.
8. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 7, characterized in that: The replacing of the collected audio data in the audio impact interval by the auxiliary audio data includes: determining a middle value of the first error interval as a first correction coefficient; determining a middle value of the second error interval as a second correction coefficient; multiplying the auxiliary audio sub-data of the audio impact interval by the ratio of the first correction coefficient to the second correction coefficient to determine the corrected audio data; The audio signal of the collected audio data in the audio impact interval is replaced by correcting the audio data.
9. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 7, characterized in that: The step of comparing the collected audio data with the original audio data based on a time sequence to determine a first error value of each sub-collected audio data in the collected audio data includes: Analyze each sub-collected audio data in turn, and determine the audio time corresponding to each peak and each trough in the sub-collected audio data as the verification time; Calculate the energy intensity change rate corresponding to each preset frequency under the verification time, and draw the energy intensity change rate curve of the current verification time; Calculate the average curve of the energy intensity change rate curves for all verification times; The energy intensity change rate corresponding to each preset frequency is extracted from the average curve, and weighted calculation is performed to determine a first error value of the current sub-collected audio data.
10. The method for monitoring the status of a nuclear power plant speaker based on multi-algorithm fusion according to claim 9, characterized in that: Also includes: In the spectrum graph, the coordinate points whose energy intensity change rate is greater than the preset change rate threshold are marked as abnormal, and the abnormal coordinates are determined; Traverse based on the abnormal coordinates and select other abnormal coordinates with the closest pixel distance to the abnormal coordinates to form a candidate frame; Counting the number of abnormal coordinates in the candidate frame and determining the abnormal coordinate ratio of the candidate frame; Determine a corresponding preset pixel ratio threshold according to the pixel area of the candidate frame; When the abnormal coordinate ratio of the candidate frame is greater than or equal to the corresponding preset pixel ratio threshold, continue traversing, select other abnormal coordinates closest to the candidate frame pixels, and update the candidate frame; When the abnormal coordinate ratio of the candidate frame is less than the corresponding preset pixel ratio threshold, the abnormal coordinates are filtered; When the abnormal coordinate ratio of the updated candidate frame is greater than or equal to the corresponding preset pixel ratio threshold, continue traversing and update the candidate frame; When the abnormal coordinate ratio of the updated candidate frame is less than the corresponding preset pixel ratio threshold, the traversal ends and the area within the candidate frame is determined as an abnormal area; Partially replacing the spectrum graph within the abnormal area; In the first error value process of sub-collecting audio data, the energy intensity change rate corresponding to the preset frequency in the abnormal area is filtered.
Citation Information
Patent Citations
Detection and classification of siren signals and localization of siren signal sources
CN113176537A
Method and system for detecting state of loudspeaker by using machine auditory sense
CN117119355A
Loudspeaker defect identification method based on voice broadcast tone quality
CN118413800A
Voice interaction method and device of Bluetooth headset, equipment and storage medium
CN118506782A
Automatic audio tuning and compensation program
CN118679756A