A Method for Detecting English Listening Anomalies Based on Reference Signal Alignment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-14
AI Technical Summary
首先,在大规模考试中,不同考点的播放设备可能存在秒级的同步误差,有的考点晚于其它考点播放结束,压缩了考生剩余答题时间,一定程度上影响了考试的公平性
[0023](1)本发明通过“参考信号对齐、声学特征比对、语义关键词检测、信号质量分析”的多指标联合判定方式,不再依赖人工连续监听即可自动发现播放错题、播放滞后、播放超前、局部静音和信噪比过低等异常情况。
Smart Images

Figure CN122575407A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of educational examination monitoring technology, specifically relating to a method for detecting English listening comprehension anomalies based on reference signal alignment. More specifically, this invention relates to a method for detecting English listening comprehension anomalies that combines reference signal alignment, acoustic feature comparison, speech recognition semantic analysis, and anomaly evidence retention. It can be used for real-time monitoring, anomaly alerting, and post-event traceability of audio playback status in multiple examination rooms during standardized testing. Background Technology
[0002] With the widespread adoption of standardized tests, English listening tests, as a crucial component of selective examinations, directly impact the fairness and seriousness of the test. Currently, monitoring of listening tests relies primarily on the subjective judgment of invigilators, which has significant limitations. First, in large-scale examinations, synchronization errors at different test centers can occur within seconds, with some centers finishing playback later than others, compressing the remaining answering time for candidates and affecting fairness to some extent. Second, for serious incidents such as "playing incorrect files" or "incomplete content," human invigilators often detect them late, unable to respond promptly in the initial seconds. Furthermore, some test centers experience excessive background noise from playback equipment or interference from environmental noise, affecting candidates' comprehension of the listening content; a quantitative evaluation standard is needed to address this.
[0003] While existing technologies include silent alarm devices based on sound intensity detection, they cannot identify errors at the content level or address the synchronization issues between different test centers. Therefore, establishing a method using audio signal processing technology to align and compare live audio with standard audio in real-time, automatically, and accurately, effectively ensuring the fairness of the examination, has become an urgent technical challenge. Based on this, this invention designs an English listening comprehension anomaly detection method based on reference signal alignment.
[0004] This invention employs a multi-indicator joint judgment system, including time alignment, acoustic content comparison, signal loss detection, signal-to-noise ratio (SNR) evaluation, and semantic keyword detection, to form a cascaded anomaly detection process encompassing "time synchronization verification, content consistency verification, audio quality verification, and examination process verification." Specifically, time alignment eliminates the impact of overall timing errors on content comparison; acoustic similarity and signal loss masking identify playback errors or partial omissions; SNR evaluates on-site playback quality; and semantic keyword detection constructs a timeline of key examination process nodes. These steps work together to address the limitations of existing technologies in identifying content-level errors, synchronization discrepancies between different examination rooms, and anomalies at examination process nodes, while also reducing the false alarm rate caused by relying on single indicators. Summary of the Invention
[0005] A method for detecting English listening comprehension anomalies based on reference signal alignment, characterized by the following steps:
[0006] S1: Acquire at least one real-time audio stream from the examination room as the signal to be tested, and acquire a reference signal; the reference signal includes at least one of the following: a standard listening audio stream that is pre-imported and played at a set time, and a reference audio stream dynamically selected from multiple real-time audio streams from the examination room; S1-1: When using the pre-imported standard listening audio stream as the reference signal, record the planned playback time and the actual playback start time of the reference audio, and extract a reference segment corresponding to the current monitoring window from the reference audio stream according to the time difference between the current time and the playback start time in each round of analysis;
[0007] S1-2: When standard listening audio is not imported or the reference segment does not meet the valid speech conditions, a benchmark audio stream is dynamically selected from multiple real-time audio streams in the examination room; the dynamic selection method includes at least one of the following: selecting an audio stream with higher root mean square energy, selecting an audio stream with the highest average similarity to other audio streams, or selecting an audio stream with a high signal-to-noise ratio and a speech activity ratio that meets the threshold.
[0008] S2: Preprocess the signal to be tested and the reference signal. The preprocessing includes at least mono conversion, uniform sampling rate, DC removal, bandpass filtering, amplitude normalization, and speech activity detection to remove invalid silence segments and retain valid speech segments.
[0009] S2-1: Perform bandpass filtering on the input audio, retaining the frequency band components from 300 Hz to 4000 Hz; when the maximum amplitude of the processed audio is lower than the preset near-silence threshold, the audio segment is regarded as an invalid reference segment and subsequent synchronization determination is stopped.
[0010] S3: Time-align the preprocessed test signal and reference signal, obtain the time offset between them using cross-correlation, and extract the effective overlapping segment based on the time offset; S3-1: Perform fast Fourier transform cross-correlation calculation on the test signal and reference signal after downsampling. When implemented in the frequency domain, it can be represented as ;in, Indicates the reference signal. The term represents the signal to be measured, and the time delay term represents the deviation between the two audio segments. Indicates Fourier transform, S3-2: The lag sample size corresponding to the cross-correlation peak is denoted as the peak lag sample number. Then the time offset can be expressed as... ;in, Indicates the sampling rate; when A value greater than zero indicates that the signal being measured is ahead of the reference signal; when... When the value is less than zero, it indicates that the signal under test lags behind the reference signal;
[0011] S3-3: To reduce the interference of silence and occasional noise on synchronization determination, a pre-synchronization check is performed on the current reference segment and the segment to be tested before outputting a synchronization anomaly. The check indicators include at least the effective speech overlap duration, the normalized cross-correlation peak value, and the peak prominence. The normalized cross-correlation peak value can be expressed as... Peak prominence can be expressed as .in, This represents the maximum peak value in the cross-correlation curve. This indicates the second largest peak value obtained after removing the area near the main peak. This represents a very small constant used to prevent the denominator from being zero, and its range is [value range missing]. to Only when the effective speech overlap duration, and Only when the preset thresholds are met simultaneously is an output of a synchronization-related exception allowed.
[0012] S4: Extract acoustic features from the time-aligned test signal and reference signal, and calculate the acoustic similarity between them to determine whether the content played in the current examination room is consistent with the standard content.
[0013] S4-1: The acoustic feature is the Mel-frequency cepstral coefficients (MFCC); the MFCC feature vectors of the reference signal and the signal under test are respectively denoted as... and Then the acoustic similarity is calculated as follows: ;when If the similarity is less than the preset similarity threshold, it is determined that the playback content is incorrect or there is a missing word.
[0014] S4-2: Construct a signal loss mask using short-time root-mean-square energy. For the first... The frame, signal loss indication function can be expressed as: ;in, Indicates the reference signal number Frame normalized energy, Indicates the signal to be measured. Frame normalized energy, This indicates that there is a threshold in the reference speech. Indicates the mute threshold of the signal to be measured. This is an indicator function. The duration of continuous signal loss can be expressed as: ;in, and They represent the first The start and end frames of a segment of continuous signal loss interval. Indicates frame shift. When When the duration of signal loss exceeds the preset threshold, the output will show either missing words or silence.
[0015] S5: Perform speech recognition on the test signal and the reference signal respectively, detect whether they contain preset exam process keywords, and calculate the keyword trigger time based on the relative position of the recognized text in the speech segment, so as to construct or update the exam key node timeline;
[0016] S5-1: Construct a keyword database for the examination process, which includes key phrases representing the examination process such as "start of listening", "start of section one", "start of question one", "end of section one", "start of section two", "end of section two", and "end of listening"; switch different keyword mapping rules for question segments according to the current examination paper type;
[0017] S5-2: If the start time of a certain recognized text segment is The end time is The text to be recognized is numbered sequentially by character, starting from 0. The position of the target keyword within the recognized text is then determined. The total length of the identified text is Then the linear interpolation at the keyword trigger time is ,in This allows for a more precise understanding of key time points, which in turn enables the construction of an exam process timeline.
[0018] S6: Using the time offset output in step S3, the acoustic similarity and maximum continuous signal loss duration output in step S4, the signal-to-noise ratio of the speech segment of the signal under test, and the keyword triggering time output in step S5 as inputs, anomaly judgment is performed on the examination room playback state corresponding to the signal under test according to the cascaded judgment logic of synchronization verification, content consistency verification, audio quality verification, and key node verification. Specifically, when the absolute value of the time offset exceeds the preset synchronization threshold and the synchronization judgment pre-verification passes, a synchronization-related anomaly is output; when the maximum continuous signal loss duration exceeds the preset loss duration threshold, or the acoustic similarity is lower than the preset similarity threshold, a content consistency-related anomaly is output; when the signal-to-noise ratio of the speech segment of the signal under test is lower than the preset signal-to-noise ratio threshold, an audio quality-related anomaly is output; when the deviation of the keyword triggering time from the reference key node time axis exceeds the preset node deviation threshold, a key node-related anomaly is output.
[0019] S6-1: The signal-to-noise ratio can be calculated based on the ratio of speech segment power to noise power. ;in, Indicates the total power of the speech segment. This represents the noise power estimated from low-energy samples. When the signal-to-noise ratio (SNR) is below a preset threshold, the output SNR is abnormal.
[0020] S6-2: The priority for anomaly detection is set as follows: continuous signal loss anomalies take precedence over similarity anomalies, and similarity anomalies take precedence over synchronization anomalies. This is because continuous signal loss usually directly corresponds to local silence, dropped audio, or missed playback, representing a more direct content loss affecting test takers. Similarity anomalies typically reflect errors or inconsistencies in the played content. Synchronization anomalies should only be output when both the reference segment and the segment under test have sufficiently effective speech and the cross-correlation peak is reliable. This priority system avoids false alarms of synchronization anomalies due to large cross-correlation offsets when the actual dropped audio or missed playback occurs in the signal under test, thus improving the accuracy and verifiability of anomaly type detection. When the preconditions for synchronization detection are not met, even if the time offset exceeds the threshold, synchronization anomalies are not output temporarily, thereby reducing the false alarm rate caused by silence or weak audio segments.
[0021] S7: When an anomaly is detected, save the audio segment to be tested and the corresponding reference audio segment within the time window of the anomaly occurrence, generate an anomaly evidence directory, and write the anomaly type, anomaly description, examination room identifier, key node time, and anomaly determination status to the database or log file; S7-1: The anomaly evidence directory shall at least contain the abnormal audio file to be tested, the corresponding reference audio file, the anomaly type, the anomaly timestamp, and the anomaly description; the anomaly record table shall at least include the stream address, examination room name, anomaly type, anomaly description, creation time, and manual review status.
[0022] The beneficial effects of this invention are as follows:
[0023] (1) This invention uses a multi-index joint judgment method of “reference signal alignment, acoustic feature comparison, semantic keyword detection and signal quality analysis” to automatically detect abnormal situations such as incorrect playback, playback delay, playback ahead, local silence and low signal-to-noise ratio without relying on continuous manual monitoring.
[0024] (2) The present invention sets a synchronization judgment precondition. Synchronization anomaly is only output when there is valid speech in both the reference segment and the segment to be tested and the main peak of cross-correlation is significant. This can effectively reduce false alarms caused by silent segments, noisy segments and weak speech segments.
[0025] (3) This invention supports two reference construction modes: preset standard audio benchmark and multi-examination room dynamic benchmark. It is applicable to formal examination scenarios with standard audio files, as well as on-site environments that lack preset benchmarks or need to temporarily switch reference sources, making it more adaptable.
[0026] (4) This invention identifies key moments in the examination process by identifying keywords and using time interpolation to form a structured examination timeline, which is then written into the database along with abnormal events, making it easier for managers to conduct in-process analysis and post-event auditing.
[0027] (5) The present invention saves abnormal audio segments and corresponding reference segments when an anomaly occurs, forming a complete electronic evidence chain, which can significantly improve the traceability and review efficiency of anomaly handling.
[0028] The core of this invention is not simply to use single technical means such as cross-correlation, acoustic feature comparison, speech recognition and signal-to-noise ratio calculation in parallel, but to combine the above means according to the causal relationship and judgment priority in the scenario of abnormal hearing test, forming a cascaded multi-indicator joint judgment mechanism of "time alignment, semantic comparison, quality assessment and abnormal evidence collection".
[0029] Specifically, when using cross-correlation for time alignment alone, unstable correlation peaks can easily occur in silent segments, weak speech segments, or segments with occasional noise, leading to false alarms of synchronization anomalies. Therefore, this invention introduces pre-checking indicators for synchronization determination, such as effective speech overlap duration, normalized cross-correlation peak value, and peak prominence, before outputting synchronization anomalies, and combines these with a signal loss mask to determine whether the current segment meets reliable synchronization determination conditions.
[0030] Meanwhile, relying solely on acoustic similarity cannot accurately distinguish between "incorrect playback content" and "overall time misalignment." To address this, the present invention first calculates the time offset through cross-correlation and extracts effective overlapping segments. Then, it performs acoustic similarity calculation, continuous signal loss detection, and speech recognition keyword detection on the aligned segments. This ensures that the content comparison results truly reflect the differences in the playback content, rather than superficial differences caused by overall timing discrepancies.
[0031] Furthermore, this invention sets up a priority-based judgment process for continuous signal loss anomalies, similarity anomalies, synchronization anomalies, and signal-to-noise ratio anomalies: continuous signal loss is prioritized for identifying partial audio drops, missed playbacks, or silence; acoustic similarity is used to identify errors or inconsistencies in the playback content; synchronization anomalies are only output when the preconditions for synchronization judgment are met; and signal-to-noise ratio anomalies are used to evaluate the on-site playback quality. This cascaded judgment logic reduces cross-judgments caused by silent segments, weak speech segments, noisy segments, and time offsets, improving the accuracy and verifiability of anomaly type identification.
[0032] Therefore, through the synergistic cooperation between the various steps, this invention achieves a lower false alarm rate, higher anomaly localization accuracy, and more complete evidence retention capability than single-index detection in the specific application scenario of exam audio anomaly detection. It can effectively solve the problem that existing technologies cannot simultaneously identify synchronization deviation, content errors, partial omissions, and audio quality anomalies.
[0033] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0035] Figure 1 This is a flowchart of the method in a specific embodiment of the present invention;
[0036] Figure 2 This is a schematic diagram of reference signal alignment in a specific embodiment of the present invention; Detailed Implementation
[0037] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. The technical solution of the present invention will be further described below in conjunction with specific embodiments and accompanying drawings.
[0038] This invention proposes a method for detecting English listening comprehension anomalies based on reference signal alignment. Please refer to [link to relevant documentation]. Figure 1 As shown, the specific algorithm is explained below:
[0039] Step 1: The real-time audio stream in the examination room is an RTSP audio stream, or it can be other network audio streams or locally acquired streams capable of outputting digital audio data in real time. The reference signal includes at least one of the following: standard listening audio that is pre-imported and played at a set time, or a reference audio stream dynamically selected from multiple real-time audio streams in the examination room.
[0040] It should be noted that the "test signal" represents the exam audio that needs to be determined to have playback abnormalities, while the "reference signal" represents the audio used as a benchmark for time alignment and content comparison. The reference signal can come from either a preset standard audio or a dynamic representative stream from multiple exam streams; therefore, this invention is not limited to implementation scenarios where a standard listening audio file must be provided in advance.
[0041] Step 1-1: When using pre-imported standard hearing audio as a reference signal, record the planned playback time and the actual playback start time of the reference audio, and extract the reference segment corresponding to the current monitoring window from the reference audio according to the time difference between the current time and the playback start time in each round of analysis.
[0042] It should be noted that the standard listening audio file needs to be imported before the exam begins, and the corresponding scheduled playback time needs to be configured. After startup, the standard listening audio is output through the local player, and an independent acquisition task is created for each real-time audio stream in the exam room. Here, FFmpeg is used to connect to the RTSP source via TCP, and each audio stream is decoded into a mono PCM data stream in real time.
[0043] Furthermore, the actual start time of the standard listening audio is recorded to avoid the deviation between the planned playback time and the actual playback time affecting the localization of the reference segment. Let the actual start time of the standard listening audio be... The current monitoring time is The analysis window length is Sampling rate Then the first The reference segment index range corresponding to the round monitoring can be represented as: ,in, Indicates the first The reference audio sampling intervals corresponding to each round of analysis are defined. Extracting reference segments from the standard listening audio based on these intervals ensures that the progress of the reference signal matches the current playback progress in the examination room, thereby reducing the computational overhead of locating the reference content from scratch in each round of analysis.
[0044] Steps 1-2: When standard listening audio is not imported or the reference segment does not meet the valid speech conditions, a benchmark audio stream is dynamically selected from multiple real-time audio streams from the examination rooms.
[0045] It should be noted that in actual examination scenarios, the following situations may occur: First, the monitoring system may not have pre-installed standard listening audio; second, although standard listening audio may be imported, the reference segment extracted within the current analysis window may be in a silent or near-silent segment, making it unsuitable for direct use in synchronous judgment; third, the standard audio playback link may be temporarily interrupted, requiring rapid restoration of the reference baseline from the on-site audio. To address these situations, this invention allows entry into a dynamic reference election mode.
[0046] Furthermore, when the number of available test streams is small, the audio stream with higher root mean square energy and a larger proportion of speech activity is prioritized as the dynamic reference; when the number of available test streams is large, the stream with the highest representativeness is selected as the benchmark audio stream by calculating the average similarity or centrality between each audio stream and other audio streams. Let the first... The centrality score of the audio in the road test was The current number of available exam room streams is Then the centrality score can be expressed as ,in Indicates the first Road test audio, Indicates the first Road audio and the first Similarity between audio signals. Final selection. The largest audio stream is used as the dynamic reference stream.
[0047] Step 2: The preprocessing includes at least mono conversion, uniform sampling rate, DC removal, bandpass filtering, amplitude normalization, and voice activity detection. Its purpose is to eliminate amplitude and spectrum differences caused by different audio acquisition devices, different transmission links, and different environmental noise, and improve the stability of subsequent time alignment and content comparison.
[0048] Step 2-1: Perform bandpass filtering on the input audio, retaining the frequency band components from 300 Hz to 4000 Hz; when the maximum amplitude of the processed audio is lower than the preset near-silence threshold, the audio segment is regarded as an invalid reference segment and subsequent synchronization determination is stopped.
[0049] It should be noted that, firstly, both the signal to be tested and the reference signal are uniformly converted into audio sequences with a sampling rate of 16000 Hz, mono, and 16-bit depth. Then, DC removal, bandpass filtering, and amplitude normalization are performed sequentially. The normalized audio can be represented as... ,in, This represents the filtered audio sample. This represents the mean of the filtered audio. This represents a very small constant to prevent the denominator from being zero. This represents the normalized audio.
[0050] Furthermore, speech activity detection is performed on the normalized audio based on a fixed frame length, retaining only valid speech segments to eliminate invalid silence fragments and reduce the impact of background noise on subsequent analysis. To determine whether the current reference segment meets the valid speech criteria, its speech activity percentage can be calculated. ,in Indicates the percentage of voice activity. This indicates the number of frames detected as speech. This indicates the total number of frames in the current segment.
[0051] It should be noted that when the maximum amplitude, root mean square energy, and proportion of speech activity of a segment are all higher than the preset thresholds, the reference segment can be considered to meet the valid speech conditions; otherwise, it is regarded as a silent segment, a near-silent segment, or an invalid reference segment, and the determination of synchronization anomalies based on the segment is stopped.
[0052] Step 3: Time alignment ensures that the signal under test and the reference signal are on the same time base, preventing distortion in content similarity calculations caused by overall playback errors, delays, or link latency. It's important to note that without time alignment, even if the content of the signal under test is perfectly correct, it may still be misjudged as abnormal due to overall time offset.
[0053] Step 3-1: To reduce computational complexity, the test signal and reference signal can be appropriately downsampled before performing cross-correlation calculation based on Fast Fourier Transform. The cross-correlation function can be expressed as... When implemented in the frequency domain, it can be represented as ,in Indicates the reference signal. Indicates the signal to be measured. Indicates time delay. Indicates Fourier transform, This represents the conjugate form of the spectrum of the signal under test. The hysteresis sample size corresponding to the peak is obtained by searching for the correlation peak with the largest absolute value in the cross-correlation curve.
[0054] Step 3-2: Record the lag sample size corresponding to the peak cross-correlation value as the peak lag sample size, and calculate the time offset accordingly. Let the lag sample size corresponding to the peak cross-correlation value be... Then the time offset can be expressed as ,in Indicates the sampling rate. When When, it indicates that the signal to be measured is ahead of the reference signal; when The time offset indicates that the signal under test lags behind the reference signal. Based on the time offset, the reference segment and the segment under test are synchronously trimmed to obtain the effective overlap interval between them.
[0055] Step 3-3: To reduce the interference of silence and occasional noise on synchronization determination, a pre-synchronization check is performed on the current reference segment and the segment to be tested before outputting a synchronization anomaly. The pre-check indicators include at least the effective speech overlap duration, the normalized cross-correlation peak value, and the peak prominence. The normalized cross-correlation peak value can be expressed as... Peak prominence can be expressed as ,in This represents the maximum peak value in the cross-correlation curve. This indicates the second largest peak value obtained after removing the area near the main peak. This represents a minimal constant used to prevent the denominator from being zero, and its range is [value range missing]. to .
[0056] Furthermore, the synchronization reliability determination condition can be expressed as: ,in Indicates the effective speech overlap duration. Indicates the threshold for overlap duration. This represents the normalized correlation peak threshold. This represents the peak prominence threshold. When... When, it indicates that the current segment meets the preconditions for determining synchronization anomalies; when Even if the time offset exceeds the threshold, synchronization anomalies will not be output temporarily. This effectively reduces false alarms caused by silent segments, weak speech segments, and occasional impulse noise.
[0057] Step 4: Extract acoustic features from the time-aligned test signal and reference signal, and calculate the acoustic similarity between them to determine whether the content played in the current examination room is consistent with the standard content. It should be noted that Step 3 primarily addresses the consistency of the time reference, while Step 4 primarily addresses the consistency of the played content. By aligning first and then comparing, the acoustic feature comparison can more accurately reflect differences at the content level, rather than superficial differences caused by time offset.
[0058] Step 4-1: The acoustic feature is the Mel-frequency cepstral coefficients (MFCC); the MFCC feature vectors of the reference signal and the signal under test are respectively denoted as... and Then the acoustic similarity is calculated as follows: ,when If the similarity is less than the preset similarity threshold, it is determined that there is a significant difference between the content being played in the current examination room and the reference content, and the error in the played content, missing content, or missing words can be output.
[0059] Step 4-2: Construct a signal loss mask using short-time root-mean-square energy and calculate the duration of continuous signal loss. For the first... The frame, signal loss indication function can be expressed as: ,in, Indicates the reference signal number Frame normalized energy, Indicates the signal to be measured. Frame normalized energy, This indicates that there is a threshold in the reference speech. This indicates the mute threshold of the signal to be tested.
[0060] Furthermore, the duration of continuous signal loss can be expressed as: ,in and They represent the first The start and end frames of a segment of continuous signal loss interval. This indicates frame shift. It should be noted that when... When the duration exceeds the preset signal loss threshold, it indicates that the signal under test has experienced continuous silence, partial dropouts, or obvious omissions in the presence of reference speech. In this case, the omission of words can be output first.
[0061] Step 5: Perform speech recognition on both the test signal and the reference signal, detect whether they contain preset exam process keywords, and calculate the keyword trigger time based on the relative position of the recognized text in the speech segment to construct or update the exam key node timeline. It should be noted that simple acoustic similarity can only determine "whether the sound is similar", while the keyword timeline can further answer "which stage of the exam process is currently in", thus providing stronger semantic support for anomaly detection.
[0062] Step 5-1: Input the speech segments segmented by speech activity detection into the automatic speech recognition model to obtain the corresponding transcribed text. The automatic speech recognition model can be an end-to-end speech recognition model based on CTC, Attention, Transformer, or Conformer structures. Here, a lightweight speech recognition model based on deep learning is used to transcribe the speech segments, and text cleaning, keyword normalization, and exam paper-type keyword mapping are performed on the transcription results to detect exam process keywords such as "listening begins," "first section begins," "first question begins," "first section ends," "second section begins," "second section ends," and "listening ends."
[0063] Step 5-2: If the start time of a certain recognized text segment is The end time is The text to be recognized is numbered sequentially by character, starting from 0. The position of the target keyword within the recognized text is then determined. The total length of the identified text is Then the linear interpolation at the keyword trigger time is ,in This allows for a more precise determination of key node times than directly using the start time of a speech segment, thus enabling the construction of an examination process timeline. It should be noted that the linear interpolation method described above is a simple and highly accurate time estimation method; in other implementations, keyword trigger times can be directly determined based on character-level timestamps, forced alignment results, or word-level timestamps.
[0064] Step 6: Using the time offset output in Step S3, the acoustic similarity and maximum continuous signal loss duration output in Step S4, the signal-to-noise ratio of the speech segment of the signal under test, and the keyword triggering time output in Step S5 as inputs, perform anomaly judgment on the examination room playback status corresponding to the signal under test according to the cascaded judgment logic of synchronization verification, content consistency verification, audio quality verification, and key node verification. Specifically, when the absolute value of the time offset exceeds the preset synchronization threshold and the synchronization judgment pre-verification passes, a synchronization-related anomaly is output; when the maximum continuous signal loss duration exceeds the preset loss duration threshold, or the acoustic similarity is lower than the preset similarity threshold, a content consistency-related anomaly is output; when the signal-to-noise ratio of the speech segment of the signal under test is lower than the preset signal-to-noise ratio threshold, an audio quality-related anomaly is output; when the deviation of the keyword triggering time from the reference key node time axis exceeds the preset node deviation threshold, a key node-related anomaly is output.
[0065] It should be noted that this invention does not rely on a single indicator to draw anomaly conclusions, but rather makes a joint judgment based on multiple dimensions such as time, content, quality, and semantics. This can avoid false positives and false negatives caused by judging based solely on a single similarity or a single silence duration.
[0066] Step 6-1: The signal-to-noise ratio can be calculated based on the ratio of speech segment power to noise power. ,in, Indicates the total power of the speech segment. This represents the noise power estimated from low-energy samples. When the signal-to-noise ratio (SNR) is lower than the preset threshold, it indicates that the audio quality in the current examination room is poor, with significant environmental noise or equipment background noise, and an abnormal SNR should be output.
[0067] Step 6-2: The anomaly detection priority is set as follows: continuous signal loss anomalies take precedence over similarity anomalies, and similarity anomalies take precedence over synchronization anomalies. The decision rule for anomaly types can be expressed as follows: ,in Indicates the threshold for signal loss duration. Indicates the similarity threshold. Indicates the synchronization offset threshold. This represents the signal-to-noise ratio (SNR) threshold. It should be noted that if the preconditions for synchronization determination are not met, even if the time offset exceeds the threshold, synchronization anomalies will not be output temporarily, thereby reducing the false alarm rate caused by silent or weak speech segments. Combining the above priority rules, synchronization anomalies, missing word anomalies, and SNR anomalies can ultimately be output in a unified and stable manner.
[0068] It should be noted that the signal loss duration threshold can be set from 0.8s to 3s; the similarity threshold from 0.60 to 0.85; the synchronization offset threshold from 2s to 5s; the minimum effective overlap duration from 2s to 5s; and the signal-to-noise ratio threshold from 10dB to 20dB. For the signal loss mask, the reference signal mute threshold can be set from 0.03 to 0.10, and the test signal mute threshold can be set from 0.10 to 0.25. These thresholds can be adjusted based on the noise level in the examination room, the performance of the playback equipment, and the examination management requirements.
[0069] Step 7: When an anomaly is determined, save the audio segment to be tested and the corresponding reference audio segment within the time window of the anomaly occurrence, generate an anomaly evidence directory, and write the anomaly type, anomaly description, examination room identifier, key node time, and anomaly determination status into the database or log file.
[0070] Step 7-1: The abnormal evidence directory should at least include the abnormal audio file to be tested, the corresponding reference audio file, the abnormal type, the abnormal timestamp, and the abnormal description; the abnormal record table should at least include the stream address, the examination room name, the abnormal type, the abnormal description, the creation time, and the manual review status.
[0071] Furthermore, when an anomaly is detected in a certain analysis round, the audio segment to be tested and the corresponding reference segment within the time window of the anomaly occurrence are immediately extracted, and an evidence catalog is established using the examination room number, the time of the anomaly occurrence, and the anomaly type as indexes. At the same time, the anomaly type, anomaly description, examination room identification, key examination time, and anomaly judgment status are written to the database or log file, thus forming a complete closed loop from anomaly discovery and anomaly evidence collection to manual review.
[0072] In other embodiments, acoustic features are not limited to Mel-frequency cepstral coefficients (MFCCs), but can also include filter bank energy features (FBank), linear predictive cepstral coefficients (LPCCs), spectral centroids, spectral roll-off points, deep speech model embedding vectors, or combinations thereof. Acoustic similarity is not limited to cosine similarity, but can also include Euclidean distance transformation similarity, dynamic time warping distance transformation similarity, or similarity scores based on neural network matching models. The speech recognition model can be a CTC model, Attention model, Transformer model, Conformer model, or an end-to-end automatic speech recognition model. As long as it can output the recognized text or keyword detection results corresponding to the speech segment, it can be used for key node detection in the examination process of this invention. In addition to using preset thresholds and priority rules, the anomaly detection logic can also use support vector machines, random forests, gradient boosting trees, or neural network classifiers, taking time offset, acoustic similarity, maximum continuous signal loss duration, signal-to-noise ratio, keyword time deviation, and other indicators as input, and outputting the corresponding anomaly type and anomaly confidence.
[0073] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for detecting English listening comprehension anomalies based on reference signal alignment, characterized in that, Includes the following steps: S1: Acquire at least one real-time audio stream from the examination room as the signal to be tested, and determine the reference signal according to the preset reference acquisition mode; in the standard reference mode, the reference signal is the standard listening audio that is pre-imported and played at a set time; in the backup dynamic reference mode, when the standard listening audio is not imported, the set playback time has not been reached, or the current standard reference segment does not meet the valid speech conditions, select the audio stream that meets the valid speech conditions and has the highest centrality score from multiple real-time audio streams from the examination room as the reference audio stream; The reference signal is used as a unified comparison object for the signal to be tested in subsequent steps; S2: Preprocess the signal to be tested and the reference signal. The preprocessing includes at least mono conversion, uniform sampling rate, DC removal, bandpass filtering, amplitude normalization, and speech activity detection to remove invalid silence segments and retain valid speech segments. S3: Time alignment is performed on the preprocessed test signal and reference signal, cross-correlation is used to obtain the time offset between them, and effective overlapping segments are extracted based on the time offset. S4: Extract acoustic features from the time-aligned test signal and reference signal, calculate the acoustic similarity between them, and calculate the duration of continuous signal loss to determine whether the content played in the current examination room is consistent with the reference content. S5: Perform speech recognition on the test signal and the reference signal respectively, detect whether they contain preset exam process keywords, and determine the keyword trigger time based on the relative position of the recognized text in the corresponding speech segment to construct or update the exam key node timeline; when the target keyword is detected, number the recognized text character by character, starting from 0, and set the character number of the first character of the target keyword in the recognized text to be 0. The total number of characters in the identified text is The start time of the corresponding speech segment is The end time is The keyword triggering time is determined using linear interpolation. , which is represented as ,in To generate a timeline of key exam milestones; S6: Using the time offset output in step S3, the acoustic similarity and maximum continuous signal loss duration output in step S4, the signal-to-noise ratio of the speech segment of the signal under test, and the keyword triggering time output in step S5 as inputs, anomaly judgment is performed on the examination room playback state corresponding to the signal under test according to the cascaded judgment logic of synchronization verification, content consistency verification, audio quality verification, and key node verification. Specifically, when the absolute value of the time offset exceeds the preset synchronization threshold and the synchronization judgment pre-verification passes, a synchronization-related anomaly is output; when the maximum continuous signal loss duration exceeds the preset loss duration threshold, or the acoustic similarity is lower than the preset similarity threshold, a content consistency-related anomaly is output; when the signal-to-noise ratio of the speech segment of the signal under test is lower than the preset signal-to-noise ratio threshold, an audio quality-related anomaly is output; when the deviation of the keyword triggering time from the reference key node time axis exceeds the preset node deviation threshold, a key node-related anomaly is output. S7: When an anomaly is determined, save the audio segment to be tested and the corresponding reference audio segment within the time window of the anomaly occurrence, generate an anomaly evidence directory, and write the anomaly type, anomaly description, examination room identifier, key node time, and anomaly determination status to the database or log file.
2. The method for detecting English listening comprehension anomalies based on reference signal alignment according to claim 1, characterized in that, In step S1, when entering the backup dynamic reference mode, the real-time audio streams of the examination room that do not meet the valid speech conditions or audio quality conditions are first eliminated based on the speech activity detection results, short-time energy, and signal-to-noise ratio; for the remaining... For each audio stream, acoustic features are extracted, and the similarity between any two audio streams is calculated to obtain the first... Centrality score of audio stream ,in , Indicates the first Audio stream, It represents the similarity between two audio streams; the audio stream with the highest centrality score is selected as the dynamic reference signal.
3. The method for detecting English listening comprehension anomalies based on reference signal alignment according to claim 1, characterized in that, In step S3, after downsampling the test signal and the reference signal, a fast Fourier transform is used to calculate the cross-correlation function. The position of the correlation peak with the largest absolute value is searched in the cross-correlation function to obtain the lag sample size corresponding to the peak. The time offset is calculated based on the lag sample size. The test signal and the reference signal are then truncated in the time domain based on the time offset to obtain an effective overlapping segment.
4. The method for detecting English listening comprehension anomalies based on reference signal alignment according to claim 1, characterized in that, In step S3, before outputting a synchronization anomaly, a synchronization judgment pre-verification is performed between the current reference segment and the segment to be tested. The pre-verification includes at least the following: constructing speech activity masks for a reference segment and a segment to be tested based on the speech activity detection results; calculating the effective speech overlap duration between the two segments; calculating the normalized correlation peak and peak prominence between the reference segment and the segment to be tested; and allowing the output of synchronization-type anomalies only when the effective speech overlap duration, normalized correlation peak, and peak prominence all meet the preset thresholds.
5. The method for detecting English listening comprehension anomalies based on reference signal alignment according to claim 1, characterized in that, In step S4, the acoustic features are feature vectors that can characterize the time-frequency characteristics or content characteristics of speech. These acoustic features include at least one of Mel-frequency cepstral coefficients (MFCC), filter bank energy features (FBank), linear prediction cepstral coefficients (LPCC), spectral features, or speech model embedding features. The acoustic similarity is calculated using at least one of cosine similarity, Euclidean distance transformation similarity, dynamic time warping distance transformation similarity, or matching scores based on neural networks. Then, based on the time-aligned reference signal and the signal to be tested, the frame-by-frame short-time root mean square energy is calculated and normalized according to their respective energy distributions to obtain the reference normalized energy sequence and the signal to be tested normalized energy sequence. For the first... A frame is marked as a signal loss frame when the reference normalized energy is greater than the reference signal silence threshold and the measured normalized energy is less than the measured signal silence threshold; otherwise, it is marked as a non-signal loss frame. This is how a signal loss mask is constructed. The longest duration corresponding to consecutive signal loss frames in the signal loss mask is counted as the maximum consecutive signal loss duration.
6. The method for detecting English listening comprehension anomalies based on reference signal alignment according to claim 1, characterized in that, In step S5, a keyword library for the exam process is constructed, containing "start of listening", "start of section 1", "start of question 1", "end of section 1", "start of section 2", "end of section 2", and "end of listening". The corresponding question segment keyword mapping rules are switched according to the exam paper type. When a target keyword is detected, the keyword triggering time is determined by linear interpolation based on the relative position of the keyword in the recognized text and the start and end times of the corresponding audio segment, so as to generate a timeline of key exam nodes.
7. The method for detecting English listening comprehension anomalies based on reference signal alignment according to claim 1, characterized in that, In step S6, the anomaly determination result includes anomaly type, anomaly description, and anomaly determination status; the anomaly type includes at least one of synchronization anomaly, content consistency anomaly, audio quality anomaly, and key node anomaly, wherein content consistency anomaly includes missing word anomaly, incorrect word anomaly, or inconsistent playback content anomaly, audio quality anomaly includes signal-to-noise ratio anomaly, volume anomaly, or insufficient effective speech anomaly, and key node anomaly includes missing keywords in the examination process, key nodes being advanced, or key nodes being delayed.
8. The method for detecting English listening comprehension anomalies based on reference signal alignment according to claim 1, characterized in that, In step S7, the abnormal audio segment to be tested and the corresponding reference audio segment are saved to the abnormal evidence directory, and the abnormal type, abnormal description, examination room name, stream address, creation time and manual review status are written into the abnormal record table.