A robot voice interaction self-adaptive noise reduction method and system for a strong noise environment of flight training

CN122618973BActive Publication Date: 2026-09-22ZHUHAI XIANG YI AVIATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611103864.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-22
Estimated Expiration
2046-07-24

AI Technical Summary

Technical Problem

[0007]本发明的目的在于提供一种面向飞行训练强噪声环境的机器人语音交互自适应降噪方法及系统,用以解决现有技术中飞行训练强噪声环境下无法根据噪声工况进行自适应处理,难以兼顾告警抑制与中文人声保护,且降噪结果与语音识别需求不匹配,导致机器人语音交互稳定性不足的技术问题

Benefits of technology

本发明公开了一种面向飞行训练强噪声环境的机器人语音交互自适应降噪方法及系统,基于滑动窗口的能量和频谱拼接实时噪声指纹,并通过余弦匹配、连续窗口加权和模板命中次数校验确定当前噪声工况,能够综合利用噪声的能量特征、频谱特征和时间连续性,降低单个窗口中的语音片段、瞬态声或工况切换扰动对工况判断的影响。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618973B_ABST
    Figure CN122618973B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio processing, and discloses a robot voice interaction adaptive noise reduction method and system for a strong noise environment of flight training, which comprises the following steps: acquiring audio and extracting a real-time noise fingerprint; determining a current noise working condition based on cosine matching and continuous window weighting verification; generating a frequency domain attenuation mask, a time domain gating mask and a human voice protection mask, multiplying the masks after taking a larger gain at an overlapping position, and applying the masks to the audio to obtain preprocessed audio; splicing a working condition vector and a log amplitude spectrum to input a convolutional coding and decoding network to generate a multi-gear enhanced audio candidate; and selecting an optimal candidate output based on a comprehensive score of recognition confidence, keyword hit rate and non-speech probability. The system comprises a noise working condition recognition module, a mask preprocessing module, a deep enhancement module and an identification feedback selection module. The application can adaptively reduce noise according to the noise working condition of flight training, protect Chinese human voice features, and improve the stability and accuracy of voice interaction in a complex acoustic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, specifically to an adaptive noise reduction method and system for robot voice interaction in high-noise environments during flight training. Background Technology

[0002] Flight training sites are subject to various types of noise, including engine noise, airflow noise, equipment vibration noise, and periodic alarm sounds. The energy distribution, main frequency bands, harmonic composition, and periodic characteristics of this noise can change depending on the training phase and equipment operating conditions. Voice data collected by the robot can easily be masked by continuous noise or sudden alarm sounds, thus affecting voice interaction and recognition.

[0003] Existing speech denoising methods typically use fixed filtering parameters or a single enhancement model to process audio. Fixed parameters are difficult to adapt to the spectral differences and temporal variations between different flight training noise conditions; when using a uniform enhancement intensity, strong suppression may weaken speech components at the same time, while weak suppression may leave more environmental noise.

[0004] Warning sounds during flight training typically have a relatively concentrated main frequency band, harmonic range, and repetition period. When the warning frequency band overlaps with the human voice frequency band, directly attenuating the warning frequency band can easily damage the human voice component. Judging the noise state based solely on a single audio window is also susceptible to interference from short-duration speech, transient impact sounds, and operational switching processes, causing frequent changes in noise reduction parameters.

[0005] Furthermore, there is a lack of feedback between the single noise reduction output and the subsequent speech recognition results. The degree of noise reduction at the audio signal level does not always align with the speech intelligibility required for speech recognition.

[0006] Therefore, this application provides an adaptive noise reduction method and system for robot voice interaction in high-noise environments for flight training to solve the above-mentioned technical problems. Summary of the Invention

[0007] The purpose of this invention is to provide an adaptive noise reduction method and system for robot voice interaction in high-noise environments for flight training, in order to solve the technical problems in the prior art where the robot voice interaction is not stable due to the inability to adaptively process according to the noise conditions in high-noise environments for flight training, the difficulty in balancing alarm suppression and Chinese voice protection, and the mismatch between the noise reduction results and the requirements of voice recognition.

[0008] To address the aforementioned technical problems, this invention provides an adaptive noise reduction method for robot voice interaction in high-noise environments during flight training, comprising:

[0009] Acquire audio data and construct a real-time noise fingerprint based on energy and spectrum splicing using a sliding window; Cosine matching is performed on the noise template formed by the real-time noise fingerprint and the average fingerprint of the same working condition. The weighted window matching score is calculated and verified according to the number of template hits to determine the current noise working condition. The system reads the parameter set of the current noise condition, as well as the alarm main frequency band, harmonic interval, and repetition period. It generates a frequency domain attenuation mask based on the alarm main frequency band and harmonic interval, a time domain gating mask based on the repetition period, and a human voice protection mask based on the real-time signal-to-noise ratio and the human voice protection coefficient within the parameter set. At the point where the alarm frequency band and the human voice frequency band overlap, it takes the larger gain of the frequency domain attenuation mask and the human voice protection mask. The processed frequency domain mask is multiplied by the time domain gating mask and applied to the audio to obtain the preprocessed audio. The current noise condition is mapped to a vector, the vector is concatenated with the logarithmic magnitude spectrum of the preprocessed audio, and the vector is input into a convolutional encoder-decoder network to generate a time-frequency enhancement mask. The time-frequency enhancement mask is scaled according to the enhancement level within the parameter group and applied to the preprocessed audio to generate multiple enhanced audio candidates. Speech recognition is performed on the enhanced audio candidates, and a comprehensive score is calculated based on the recognition confidence, keyword hit rate, and probability of no speech, according to weights. The enhanced audio candidate with the highest score is then output to the speech interaction recognition system.

[0010] In some specific embodiments, constructing the noise template further includes: Amplitude normalization, fixed gain equalization, framing, and time-frequency transformation are performed on the on-site recordings of various flight training noise conditions. Short-time energy, energy dispersion, peak frequency, centroid, flatness, alarm repetition period, and voice activity percentage are extracted from each recording. Each acoustic quantity is then concatenated with the log-Melbourne spectrum mean and the Mel frequency cepstral coefficient mean to form a sample noise fingerprint. The baseline fingerprint is obtained by averaging the noise fingerprints of similar samples according to the operating condition label. The operating condition label, reference fingerprint, alarm main frequency band, harmonic range, energy envelope, and repetition period are stored together as a noise template.

[0011] In some specific embodiments, the weighted window matching score and verification based on the number of template hits further includes: Arrange a preset number of consecutive windows according to their end time, assign higher weights to later windows, and remove single-window scores that are below the similarity threshold. Based on the accumulated single-window scores retained according to the noise template, candidate operating conditions that reach the cumulative score threshold are filtered. Count the number of times the template of the candidate working condition is hit, and retain the candidate working conditions that reach the hit threshold; For candidate operating conditions that have not reached the hit count threshold, the earliest window is removed and the next window is added. Cosine matching, weighted sum, and hit count verification are then re-executed until the current noisy operating condition is determined.

[0012] In some specific embodiments, the parameter set for the current noise condition includes alarm suppression weight, low-frequency suppression weight, human voice protection factor, and enhancement level; During the offline configuration phase, the alarm suppression weight is configured as the attenuation gain of the alarm main frequency band and harmonic range, the low frequency suppression weight is configured as the attenuation gain of frequencies below the human voice frequency band, the human voice protection coefficient is configured as the lower limit of the gain of the human voice frequency band, and the enhancement level is converted into the scaling factor of the time-frequency enhancement mask. The parameter values ​​are combined and stored according to the operating condition label. During the operation phase, the parameter group with the same operating condition label as the current noise operating condition is read, and the stored parameter values ​​are kept unchanged.

[0013] In some specific embodiments, a frequency-domain attenuation mask is generated based on the alarm main frequency band and harmonic interval, and a time-domain gating mask is generated based on the repetition period, further including: Read the alarm main frequency band, harmonic range, energy envelope and repetition period from the noise template of the current noise condition; Set the attenuation gain of each frequency point in the alarm main frequency band and harmonic interval according to the alarm suppression weight to form a frequency domain attenuation mask; The alarm pulse period is located in each repetition cycle according to the energy envelope. The gating attenuation of the alarm pulse period is configured to be higher than that of the non-alarm pulse period to form a time-domain gating mask.

[0014] In some specific embodiments, generating a voice protection mask based on the real-time signal-to-noise ratio and the voice protection coefficient within the parameter set further includes: Calculate the real-time signal-to-noise ratio of each time-frequency unit within the human voice frequency band, map the segmented intervals of the real-time signal-to-noise ratio to the protection gain, and limit the lower limit of the protection gain based on the human voice protection coefficient to generate a human voice protection mask. Where the alarm main frequency band or harmonic interval overlaps with the human voice frequency band, the gain of the frequency domain attenuation mask and the gain of the human voice protection mask are taken to a larger value in each time-frequency unit. Preserve the frequency domain attenuation mask gain for frequencies that do not fall within the human voice frequency band.

[0015] In some specific embodiments, the current noise condition is mapped as a vector, the vector is concatenated with the logarithmic magnitude spectrum of the preprocessed audio, and the result is input into a convolutional encoder-decoder network to generate a time-frequency enhancement mask, further including: Map the current noise condition to a learnable vector, copy the learnable vector along the time dimension and concatenate it with the logarithmic magnitude spectrum; The concatenated result is subjected to convolutional encoding, normalization, nonlinear activation, and transposed convolutional decoding to output a single-channel time-frequency quantity. A restricted activation function is used to convert the single-channel time-frequency quantity into a time-frequency enhancement mask.

[0016] In some specific embodiments, scaling the time-frequency enhancement mask according to the enhancement level within the parameter group and applying it to the preprocessed audio generates multiple enhanced audio candidates, further including: Each enhancement level is converted into a mask scaling factor, and the time-frequency enhancement mask is scaled and applied to the pre-processed audio to reconstruct the enhanced audio at each level. The live audio without time-frequency enhancement mask is combined with the enhanced audio at each level to form enhanced audio candidates; Extract the recognition confidence, keyword hit rate and no speech probability of each enhanced audio candidate, accumulate the recognition confidence and keyword hit rate with positive weights, and subtract the no speech probability with negative weights to obtain the comprehensive score; Sort by overall score and output the enhanced audio candidates with the highest overall score.

[0017] In some specific embodiments, training the convolutional encoder-decoder network further includes: The existing near-clean speech is segmented into speech segments, and noise segments are extracted from the on-site recordings of various flight training noise conditions. The speech segments and noise segments are mixed according to the preset signal-to-noise ratio range, and condition labels are added to form semi-synthetic training samples. The logarithmic magnitude spectrum of the semi-synthetic training samples and the learnable vector mapped by the working condition labels are input into the convolutional encoder-decoder network, and the speech segments are used as the supervision target. The spectrum reconstruction loss, the human voice frequency band amplitude spectrum weighted mean square loss, and the energy suppression loss above the human voice frequency band are added together with preset weights, and the convolutional encoder-decoder network parameters are updated based on the sum of the losses.

[0018] Based on the same concept, the present invention also provides an adaptive noise reduction system for robot voice interaction in high-noise environments for flight training, comprising: The audio acquisition and fingerprint extraction module is configured to acquire audio and splice real-time noise fingerprints based on the energy and spectrum of a sliding window. The noise condition identification module is configured to perform cosine matching between the real-time noise fingerprint and the noise template formed by the average of the fingerprints under the same condition, weight the window matching score and verify according to the number of times the template is hit, and determine the current noise condition. The mask preprocessing module is configured to read the parameter set of the current noise condition, as well as the alarm main frequency band, harmonic interval, and repetition period; generate a frequency domain attenuation mask based on the alarm main frequency band and harmonic interval; generate a time domain gating mask based on the repetition period; generate a human voice protection mask based on the real-time signal-to-noise ratio and the human voice protection coefficient within the parameter set; take the larger gain of the frequency domain attenuation mask and the human voice protection mask at the overlap of the alarm frequency band and the human voice frequency band; multiply the processed frequency domain mask and the time domain gating mask and apply them to the audio to obtain the preprocessed audio. The depth enhancement and candidate generation module is configured to map the current noise condition into a vector, concatenate the vector with the logarithmic magnitude spectrum of the preprocessed audio, input the vector into a convolutional encoder-decoder network to generate a time-frequency enhancement mask, scale the time-frequency enhancement mask according to the enhancement level within the parameter group and apply it to the preprocessed audio to generate multiple enhanced audio candidates. The recognition feedback selection module is configured to perform speech recognition on enhanced audio candidates, calculate a comprehensive score based on recognition confidence, keyword hit rate, and probability of no speech, and output the enhanced audio candidate with the highest score to the speech interaction recognition.

[0019] Compared with existing technologies, its advantages are as follows: This invention discloses an adaptive noise reduction method and system for robot voice interaction in high-noise environments for flight training. It is based on real-time noise fingerprinting by splicing energy and spectrum of a sliding window, and determines the current noise condition through cosine matching, continuous window weighting and template hit count verification. It can comprehensively utilize the energy characteristics, spectrum characteristics and temporal continuity of noise to reduce the impact of voice segments, transient sounds or condition switching disturbances in a single window on condition judgment.

[0020] By using the alarm main frequency band and harmonic range to form a frequency domain attenuation mask, and using the alarm repetition period and energy envelope to form a time domain gating mask, periodic alarm noise can be processed from the frequency position and the time period of alarm pulse occurrence, respectively, reducing additional attenuation in non-alarm time-frequency regions.

[0021] A voice protection mask is generated based on the real-time signal-to-noise ratio and the voice protection coefficient. At the point where the alarm frequency band and the voice frequency band overlap, the larger gain of the frequency domain attenuation mask and the voice protection mask is taken, which can form a lower limit of gain in the voice frequency band, so that the corresponding voice components are still preserved when the alarm frequency and the voice frequency overlap.

[0022] The current noise condition is mapped to a vector and concatenated with the logarithmic amplitude spectrum of the preprocessed audio. This allows the convolutional encoder-decoder network to simultaneously receive acoustic spectrum and noise condition information when generating the time-frequency enhancement mask, enabling it to generate corresponding time-frequency enhancement processing for different flight training noise conditions. By scaling the time-frequency enhancement mask according to the enhancement level within the parameter group, it is also possible to generate candidate results with different enhancement levels within the same audio window.

[0023] Speech recognition is performed on multiple enhanced audio candidates separately, and a comprehensive score is calculated by combining recognition confidence, keyword hit rate and no speech probability. The enhanced audio can be selected by feedback using the speech interaction recognition results, so that the final output simultaneously considers the recognizability of the speech content, the retention of interactive keywords, and the judgment result of no speech.

[0024] The convolutional codec network uses semi-synthetic training samples with operational condition labels and updates the network parameters together through spectrum reconstruction loss, human voice frequency band amplitude spectrum weighted mean square loss and energy suppression loss above the human voice frequency band. It can simultaneously constrain full spectrum reconstruction, human voice frequency band preservation and high-frequency noise suppression during training, so that the network training objective is consistent with the voice interaction processing requirements in the high-noise environment of flight training. Attached Figure Description

[0025] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating some specific embodiments of the robot voice interaction adaptive noise reduction method for high-noise environments during flight training, as proposed by the present invention. Figure 2 This is a schematic diagram comparing the effect of rule-based noise reduction and deep learning noise reduction in another embodiment of the adaptive noise reduction method for robot voice interaction in high-noise environments for flight training according to the present invention. Figure 3 This is a schematic diagram comparing the ASR recognition effect before and after noise reduction in another embodiment of the adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training according to the present invention. Figure 4 This is a schematic diagram of the cumulative recognition results of five noise states in a continuous window in another embodiment of the adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training according to the present invention. Figure 5 This is a schematic diagram of the structure of a robot voice interaction adaptive noise reduction system for high-noise environments during flight training, according to some specific embodiments of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms "a," "the," and "the" as used in the embodiments of this application are also intended to include the plural forms, unless the context clearly indicates otherwise, and "multiple" generally includes at least two.

[0028] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0029] It should be understood that although the terms first, second, third, etc., may be used in the embodiments of this application, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, first may also be referred to as second without departing from the scope of the embodiments of this application, and similarly, second may also be referred to as first.

[0030] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”

[0031] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.

[0032] It should be noted that any symbols and / or numbers present in the specification that are not marked in the accompanying drawings are not reference numerals.

[0033] Reference Figure 1 An adaptive noise reduction method for robot voice interaction in noisy flight training environments includes: S101, acquire audio, and splice real-time noise fingerprint based on energy and spectrum splicing of a sliding window; S102, perform cosine matching on the noise template formed by the real-time noise fingerprint and the average fingerprint of the same working condition, weight the window matching score and verify according to the number of template hits to determine the current noise working condition; S103: Read the parameter set of the current noise condition, as well as the alarm main frequency band, harmonic interval, and repetition period. Generate a frequency domain attenuation mask based on the alarm main frequency band and harmonic interval, generate a time domain gating mask based on the repetition period, and generate a human voice protection mask based on the real-time signal-to-noise ratio and the human voice protection coefficient in the parameter set. Take the larger gain of the frequency domain attenuation mask and the human voice protection mask at the overlap of the alarm frequency band and the human voice frequency band. Multiply the processed frequency domain mask and the time domain gating mask and apply them to the audio to obtain the preprocessed audio. S104, map the current noise condition into a vector, concatenate the vector with the logarithmic amplitude spectrum of the preprocessed audio, input the vector into the convolutional encoder-decoder network to generate a time-frequency enhancement mask, scale the time-frequency enhancement mask according to the enhancement level within the parameter group and apply it to the preprocessed audio to generate multiple enhanced audio candidates; S105 performs speech recognition on the enhanced audio candidates, calculates a comprehensive score based on recognition confidence, keyword hit rate, and probability of no speech, and outputs the enhanced audio candidate with the highest score to the speech interaction recognition system.

[0034] To provide a clearer explanation, the steps in the embodiments of the present invention are described in detail below: S101, acquire audio, and splice real-time noise fingerprint based on energy and spectrum splicing of a sliding window; Furthermore, a noise template is constructed, including: Amplitude normalization, fixed gain equalization, framing, and time-frequency transformation are performed on the on-site recordings of various flight training noise conditions. Short-time energy, energy dispersion, peak frequency, centroid, flatness, alarm repetition period, and voice activity percentage are extracted from each recording. Each acoustic quantity is then concatenated with the log-Melbourne spectrum mean and the Mel frequency cepstral coefficient mean to form a sample noise fingerprint. The baseline fingerprint is obtained by averaging the noise fingerprints of similar samples according to the operating condition label. The operating condition label, reference fingerprint, alarm main frequency band, harmonic range, energy envelope, and repetition period are stored together as a noise template.

[0035] In this embodiment, audio is received from a microphone or microphone array installed on the flight training robot. The audio can be mono pulse code modulation data with a 16kHz sampling rate and 16-bit quantization precision, or it can be a single-channel audio stream formed after channel selection. The acquisition end writes a frame sequence number and acquisition timestamp for each frame according to the arrival time of the audio frame, so that the audio in the current window belongs to the data already obtained at the current processing time when entering the subsequent working condition identification. When a short-term frame loss occurs, the previous valid frame can be used to fill the gap and write a frame loss flag; if more than 3 frames are lost consecutively, the real-time noise fingerprint update is stopped, the previous valid noise working condition is retained, and a new valid window is waited for.

[0036] This embodiment uses a 25ms frame length and a 10ms frame shift, corresponding to 400 and 160 sampling points respectively at a 16kHz sampling rate. Each frame of audio is multiplied by a Hanning window and then subjected to a 512-point short-time Fourier transform, yielding a complex spectrum of 257 non-negative frequency positions. The sliding window is preferably 6s, with a 2s interval between the start times of adjacent windows and a 4s overlap between adjacent windows. The window parameters can also be set to 4–8s based on the robot processor's buffering capacity, and the window interval can be set to 1–3s, as long as it allows for continuous output of energy and spectrum statistics while maintaining the order of subsequent windows.

[0037] A real-time noise fingerprint can be composed of short-time energy, energy dispersion, peak frequency, centroid, flatness, alarm repetition period, voice activity percentage, log-Mel spectrum mean, and Mel frequency cepstral coefficient mean. The log-Mel spectrum can be configured with 40 frequency bands, and the Mel frequency cepstral coefficients can retain 12 dimensions, thus forming a 59-dimensional real-time noise fingerprint. Each dimension is standardized using the mean and standard deviation obtained from offline statistics before stitching; if the standard deviation of any dimension is less than... When, set the standard deviation of this dimension to To avoid numerical anomalies, the real-time noise fingerprint, along with the window end time, window start time, and effective frame ratio, is written to the window cache.

[0038] Flight training noise conditions can include ground alarms, ground proximity warning system tests, wing not-open alarms, pre-takeoff stationary conditions, and takeoff caution conditions. In the exemplary configuration, five 60-second on-site recordings are collected for each condition. Amplitude normalization adjusts the peak amplitude of each recording to a preset range, such as 0.95 of full scale. Fixed-gain equalization uses the same set of equalization parameters to correct the frequency response of the acquisition channel, and then performs framing and time-frequency transformation according to a 25ms frame length, a 10ms frame shift, and a 512-point transform length. When clipping exists in the on-site recordings, clipped frames are marked and excluded from the template statistics; recordings with an effective frame ratio lower than 0.8 are not included in the baseline fingerprint calculation.

[0039] Short-time energy is the sum of squares of the sampled values ​​of each frame; energy dispersion is the ratio of the standard deviation to the mean of the short-time energy within the window; the peak frequency of the spectrum is the frequency at which the maximum value of the average power spectrum is located; the centroid of the spectrum is the frequency-weighted sum of the spectral energy and then divided by the total spectral energy; the spectral flatness is the ratio of the geometric mean to the arithmetic mean of the spectral amplitude. The alarm repetition period can be obtained from the autocorrelation peak interval of the energy envelope, and the proportion of speech activity is obtained by dividing the number of frames marked as speech by the speech activity detector within the window by the total number of valid frames. When no stable autocorrelation peak is detected, the repetition period field is written to 0 and a periodic flag is set, and the subsequent time-domain gating mask maintains a gain of 1 for the entire time period under this flag.

[0040] Let the first Type of working conditions have The valid sample noise fingerprint, the first Each sample noise fingerprint is Then the reference fingerprint for this working condition for: ; in, Indicates the operating condition label number. Indicates the first The number of valid samples for the same type of working condition The average fingerprint value under the same working conditions.

[0041] Samples whose outlier distance exceeds three times the median distance of similar samples can be excluded first, and then the baseline fingerprint can be recalculated. The obtained baseline fingerprint uses the same feature order and normalization parameters as the real-time noisy fingerprint during operation.

[0042] Each noise template can be stored using a structured record, which includes at least the template version, operating condition label, 59-dimensional baseline fingerprint, upper and lower limits of the alarm main frequency band, upper and lower limits of each harmonic interval, normalized energy envelope, repetition period, standardized mean, and standardized standard deviation. The template can be stored in the robot's read-only configuration area and loaded into the runtime memory during startup. If any template field is missing, the template will not participate in matching; if only the alarm field is missing but the baseline fingerprint is complete, the template can still participate in operating condition identification, and a no-alarm mask will be used for the missing alarm field.

[0043] In this embodiment, a real-time noise fingerprint can also be constructed using several quantiles of the power spectrum, the spectral roll-off frequency, or modulation spectrum statistics, and a noise template can be formed using the same averaging method based on the operating condition labels. This implementation maintains the data link for the real-time noise fingerprint based on energy and spectrum splicing using a sliding window, and the feature dimension can be adjusted according to the selected energy and spectrum quantities. The real-time noise fingerprint and window timing information are used to combine with the noise template to complete the current noise operating condition judgment.

[0044] S102, perform cosine matching on the noise template formed by the real-time noise fingerprint and the average fingerprint of the same working condition, weight the window matching score and verify according to the number of template hits to determine the current noise working condition; Furthermore, the weighted window matching score is validated based on the number of template hits, including: Arrange a preset number of consecutive windows according to their end time, assign higher weights to later windows, and remove single-window scores that are below the similarity threshold. Based on the accumulated single-window scores retained according to the noise template, candidate operating conditions that reach the cumulative score threshold are filtered. Count the number of times the template of the candidate working condition is hit, and retain the candidate working conditions that reach the hit threshold; For candidate operating conditions that have not reached the hit count threshold, the earliest window is removed and the next window is added. Cosine matching, weighted sum, and hit count verification are then re-executed until the current noisy operating condition is determined.

[0045] Further, the validation based on the number of template hits includes: Extract the highest and second highest template matching scores of each consecutive window and calculate the difference. Weight the difference with the number of consecutive hits of the same label in adjacent windows to obtain the working condition separation quantity. When the separation amount of working conditions reaches the preset separation threshold, the candidate working condition identified by the highest template matching score is retained; If the separation amount of working conditions does not reach the preset separation threshold, the window weights are reallocated according to the continuity of the same label and the separation amount of working conditions is recalculated. If the recalculated separation amount still does not reach the preset separation threshold, the earliest window is removed and the next window is added until the current noise condition is output.

[0046] In this embodiment, at the end of each sliding window, the real-time noise fingerprint and all valid noise templates of that window are read. Let the first... The real-time noise fingerprint of a continuous window is , No. The baseline fingerprint of each noise template is The corresponding cosine matching score is: ; in, Indicates the first The window and the first The single-window score for each template. To prevent extremely small positive numbers with a denominator of 0, this embodiment preferably uses... .

[0047] The L2 norm of the real-time noisy fingerprint or the reference fingerprint is lower than When a match is invalid, the corresponding match is marked as invalid. The cosine matching score, along with the window number, template label, and validity flag, constitutes a single-window matching record.

[0048] In this embodiment, five consecutive windows are maintained according to the end time of the window, with window weights of 0.10, 0.15, 0.20, 0.25, and 0.30 respectively, and a similarity threshold of 0.60 is preferred. Later windows reflect the acoustic state closer to the current moment and therefore receive higher weights. In the startup phase where there are fewer than five windows, the relative weights of the existing windows can be renormalized; if there are fewer than three existing windows, the current noise condition is not switched temporarily. Single window scores below 0.60 are recorded as 0 in the cumulative calculation of the corresponding template, while the original scores are retained for debugging and condition separation calculation.

[0049] No. Cumulative score of each noise template It can be represented as: ; in, This indicates the number of consecutive windows participating in the cumulative calculation; in this example, it is 5. Indicates the first Normalized weights for each window; This represents the similarity threshold, which is 0.60 in this example; This is an indicator function that takes a value of 1 when the condition is true and 0 when the condition is false. The preferred cumulative score threshold is 0.80.

[0050] Template labels that reach the threshold are written into the candidate working condition set; when the candidate working condition set is empty, the previous valid working condition is maintained and the system waits for the next window.

[0051] The template with the highest score in a single window is considered the template that has been matched by that window. Number of hits per template for: ; in, Iterate through all valid noise templates. Indicates the first The number of times a template is hit in a continuous window.

[0052] The hit count threshold is preferably 3, meaning that at least 3 out of 5 windows hit the same condition. When multiple candidate conditions reach the threshold simultaneously, the cumulative scores are compared first, and then the hit counts of the 3 later windows are compared; if they are still the same, the previous valid condition is maintained to reduce frequent switching caused by short-term conflicts.

[0053] Let the weighted average score of the highest template within a continuous window be . The weighted average score of the second-highest template is The highest number of consecutive hits with the same tag in the template is Then the separation amount under operating conditions It can be represented as: ; in, The weighting for the score difference is preferably 0.60 in this embodiment; This represents the number of consecutive windows.

[0054] The working condition separation quantity reflects both the score difference of the highest template relative to other templates and the degree of continuity of the labels over time.

[0055] The preset separation threshold is preferably 0.45. When the separation value of the operating condition reaches 0.45, the candidate operating condition with the highest template identifier is written into the current noise operating condition register area, and the end time of the effective window is recorded at the same time.

[0056] During reassignment, the weight of windows that consecutively hit the same label can be increased by 20%, and then all weights are normalized. The weight adjustment only applies to this recalculation and does not overwrite the stored base weights. If two consecutive label segments of the same length appear, the label segment with the later ending time is selected to receive the increased weight.

[0057] While the window is sliding forward, the previous valid operating condition continues to be used for subsequent processing, and a pending confirmation status flag is written. If no valid candidate is obtained after four consecutive slides, the current operating condition can be set to a general strong noise state, using a conservative parameter group, and then switched to a specific operating condition after the matching is restored.

[0058] The window cache is updated using a first-in, first-out (FIFO) method, ensuring that each judgment only uses the audio at the current time and before.

[0059] See Figure 4 In the continuous window cumulative identification process for five noise states, single-window labels may experience short-term fluctuations. The cumulative score, number of hits, and working condition separation amount together form the working condition switching conditions. A set of exemplary data used to represent the evaluation record format can record the single-window recognition rate as 73.2% and the continuous window cumulative recognition rate as 100.0%; these values ​​are used to illustrate the record fields and comparison methods, and will be recalculated based on the actual collected evaluation set during specific deployment.

[0060] In this embodiment, the number of consecutive window hits can be replaced with a duration check, meaning that the condition is only confirmed when the cumulative hit duration of the same template reaches 6 seconds; or an exponentially increasing time weight can be used instead of a fixed weight sequence. All of the above methods are based on cosine matching between the real-time noise fingerprint and the mean template of the same condition, and output the current noise condition, matching score, confirmation time, and pending confirmation status.

[0061] S103: Read the parameter set of the current noise condition, as well as the alarm main frequency band, harmonic interval, and repetition period. Generate a frequency domain attenuation mask based on the alarm main frequency band and harmonic interval, generate a time domain gating mask based on the repetition period, and generate a human voice protection mask based on the real-time signal-to-noise ratio and the human voice protection coefficient in the parameter set. Take the larger gain of the frequency domain attenuation mask and the human voice protection mask at the overlap of the alarm frequency band and the human voice frequency band. Multiply the processed frequency domain mask and the time domain gating mask and apply them to the audio to obtain the preprocessed audio. Furthermore, the parameter set for the current noise condition includes alarm suppression weight, low-frequency suppression weight, human voice protection factor, and enhancement level; During the offline configuration phase, the alarm suppression weight is configured as the attenuation gain of the alarm main frequency band and harmonic range, the low frequency suppression weight is configured as the attenuation gain of frequencies below the human voice frequency band, the human voice protection coefficient is configured as the lower limit of the gain of the human voice frequency band, and the enhancement level is converted into the scaling factor of the time-frequency enhancement mask. The parameter values ​​are combined and stored according to the operating condition label. During the operation phase, the parameter group with the same operating condition label as the current noise operating condition is read, and the stored parameter values ​​are kept unchanged.

[0062] Furthermore, a frequency-domain attenuation mask is generated based on the alarm main frequency band and harmonic interval, and a time-domain gating mask is generated based on the repetition period, including: Read the alarm main frequency band, harmonic range, energy envelope and repetition period from the noise template of the current noise condition; Set the attenuation gain of each frequency point in the alarm main frequency band and harmonic interval according to the alarm suppression weight to form a frequency domain attenuation mask; The alarm pulse period is located in each repetition cycle according to the energy envelope. The gating attenuation of the alarm pulse period is configured to be higher than that of the non-alarm pulse period to form a time-domain gating mask.

[0063] Furthermore, a voice protection mask is generated based on the real-time signal-to-noise ratio and the voice protection coefficient within the parameter set, including: Calculate the real-time signal-to-noise ratio of each time-frequency unit within the human voice frequency band, map the segmented intervals of the real-time signal-to-noise ratio to the protection gain, and limit the lower limit of the protection gain based on the human voice protection coefficient to generate a human voice protection mask. Where the alarm main frequency band or harmonic interval overlaps with the human voice frequency band, the gain of the frequency domain attenuation mask and the gain of the human voice protection mask are taken to a larger value in each time-frequency unit. Preserve the frequency domain attenuation mask gain for frequencies that do not fall within the human voice frequency band.

[0064] Furthermore, where the alarm main frequency band or harmonic range overlaps with the human voice frequency band, the gain of the frequency domain attenuation mask and the gain of the human voice protection mask are taken as larger values ​​on a time-frequency basis, including: Calculate the energy gradient of adjacent frequency points and the phase change of adjacent frames in the time-frequency cells at the overlapping area; Time-frequency units whose energy gradient satisfies the continuous condition of speech formants and whose phase change satisfies the continuous condition of speech period are included in the human voice protection set. The larger value between the frequency domain attenuation mask gain and the human voice protection mask gain is taken for the human voice protection set. For time-frequency units not included in the human voice protection set, retain the frequency domain attenuation mask gain; The gains after time-frequency position merging are used to form an overlapping frequency band mask, which is then written into the processed frequency domain mask.

[0065] In this embodiment, the original audio spectrum of the current noise condition and the buffer is used for subsequent mask generation processing. The alarm description field in the parameter group and noise template is read according to the condition label. An exemplary human voice frequency band is set to 300–3400 Hz. Each mask is represented by a time-frequency unit gain, with the gain range limited to 0–1; a smaller gain indicates a higher degree of attenuation. If the condition label is in a pending confirmation state, the parameter group of the previous valid condition is read; if there is no previous valid condition, a general conservative parameter group is read.

[0066] The alarm suppression weight controls the attenuation level of the alarm's main frequency band and harmonic range; the low-frequency suppression weight controls the attenuation level of low-frequency noise below 300Hz; the voice protection coefficient sets the lower limit of gain within the range of 300-3400Hz; and the enhancement level provides candidates for enhancing audio of different intensities for subsequent steps. All parameters are bound to operating condition labels to avoid mixing parameters from different operating conditions at the same time.

[0067] An example parameter configuration is as follows: Current noise condition alarm suppression weight, low frequency suppression weight, human voice protection coefficient, and enhancement level; Ground alarm levels: 0.65–0.75, 0.25–0.35, 0.78–0.82 (highest level); Ground proximity warning system test settings: 0.55–0.65, 0.25–0.35, 0.80–0.84 (highest setting); Wings not deployed warning: 0.45–0.55, 0.30–0.40, 0.88–0.92 (standard range); Before takeoff, the standard gear is 0.25–0.35, 0.55–0.65, or 0.88–0.92. For takeoff, please use conservative gears: 0.20–0.30, 0.15–0.25, 0.93–0.97. This embodiment can use the median value of each interval as the initial configuration. For example, the parameters for ground alarms can be 0.70, 0.30, 0.80, and the strong setting, respectively. The parameter values ​​are used to form an executable implementation method and do not limit the values ​​within the intervals determined according to the acquisition equipment and training model.

[0068] Parameter records can use operating condition tags as primary keys and be configured with version numbers and checksums. During runtime, the entire parameter set is read at once and a read-only snapshot is generated, remaining unchanged within a 6-second processing window. When the operating condition is switched within a window, the new parameter set takes effect from the next window. If a parameter record is missing or verification fails, a general parameter set consisting of an alarm suppression weight of 0.30, a low-frequency suppression weight of 0.20, a human voice protection coefficient of 0.90, and a conservative setting is used, and a parameter anomaly flag is output.

[0069] The alarm main frequency band and harmonic interval are represented by upper and lower frequency limits, and the energy envelope is represented by a normalized sampling sequence within one repetition period, with the repetition period expressed in milliseconds. After reading, the frequency boundaries are mapped to the frequency index of the short-time Fourier transform, and the energy envelope is resampled to the time index corresponding to the current frame shift.

[0070] Let the alarm suppression weight be . The low-frequency suppression weight is Then the frequency domain attenuation mask It can be formed by the following formula: ; in, Indicates the frame index. Indicates frequency, Indicates the main alarm frequency band of the current operating condition. This represents the set of harmonic intervals.

[0071] For the 2 to 3 frequency points at the edge of the frequency band, a linear transition can be used to reduce spectral artifacts caused by abrupt gain changes.

[0072] The position within the period is obtained by taking the remainder of the current frame time divided by the repetition period, and then aligned with the normalized energy envelope. Positions with an energy envelope higher than 0.60 are marked as alarm pulse periods. Let the pulse period gating coefficient be... The non-pulse period gating coefficient is Then the time-domain gated mask can be expressed as: ; in, This represents the set of alarm pulse frames located based on the energy envelope and repetition period.

[0073] When the repetition period is 0 or the energy envelope is missing, Set it to 1 so that the frequency domain mask can perform preprocessing independently.

[0074] The real-time signal-to-noise ratio (SNR) can be calculated from the current amplitude spectrum energy and the noise energy estimate updated in non-speech frames. The noise energy estimate is initialized with the first 10 frames at startup and then updated only with a historical coefficient of 0.95 when the speech activity probability is below 0.20. Let the real-time SNR be... The protection gain obtained by segmented mapping is The voice protection factor is ,but: ; Among them, when hour, ;when hour, ;when hour, .

[0075] It is established only within the human voice frequency band, and the output is a gain matrix with the same time dimension and corresponding frequency dimension as the original audio spectrum.

[0076] Fusion gain in overlapping frequency bands for: ; in, This indicates the human voice frequency band of 300–3400 Hz.

[0077] By taking a larger gain for each time-frequency unit, the lower limit of the gain formed by the human voice protection coefficient is directly applied to the overlapping region.

[0078] The energy gradient between adjacent frequencies is taken as the average absolute value of the logarithmic energy difference between the current frequency and its left and right adjacent frequencies; the phase change between adjacent frames is taken as the phase difference between the current frame and the previous frame, and the result is folded back to... The phase change in the first frame is initialized to 0, and the frequency band boundaries only use existing adjacent frequencies.

[0079] The exemplary condition for speech formant continuity is that the energy gradients of three adjacent frequency points are all below 6 dB, and the condition for speech period continuity is that the difference in phase change between adjacent frames is below 0.35π. When both conditions are met simultaneously, the corresponding time-frequency unit is assigned to the voice protection set. The set determination only refines the gain selection within overlapping frequency bands and does not change the processing outside the voice frequency band.

[0080] This type of time-frequency unit continues to use the gain formed by alarm suppression weights. When the gradient or phase cannot be calculated, it is also treated as not belonging to the human voice protection set, and a feature missing flag is written.

[0081] During merging, the time and frequency indices remain unchanged. The overlapping frequency band mask covers the corresponding positions in the original frequency domain attenuation mask, while other positions retain the original gain. The resulting processed frequency domain mask is then used in subsequent steps.

[0082] In this embodiment, the probability of voice activity can also be used as an additional protection criterion for overlapping frequency bands. For example, when the probability of voice activity is higher than 0.70, the lower limit of the gain of the overlapping frequency band is directly set to the human voice protection coefficient; when the probability of voice activity is lower than 0.20, the gain of the frequency domain attenuation mask is maintained. This method also completes the gain selection between frequency domain attenuation and human voice protection at the point where the alarm frequency band and the human voice frequency band overlap.

[0083] This yields the complete processed frequency domain mask. The processed frequency domain mask is multiplied by the time-domain gated mask and applied to the audio spectrum, resulting in the preprocessed audio through inverse time-frequency transformation.

[0084] Let the complex spectrum of the original audio be... The processed spectrum is: ; The processed spectrum retains the original spectrum phase. An inverse short-time Fourier transform and overlapping summation are performed using a Hanning window, a 10ms frame shift, and a 512-point transform length to obtain the preprocessed audio. Abnormal values ​​in the mask are first truncated to 0-1. If the proportion of abnormal time-frequency units exceeds 5%, the mask result for that window is discarded, and the unmasked audio is output with a preprocessing abnormality flag set.

[0085] S104, map the current noise condition into a vector, concatenate the vector with the logarithmic amplitude spectrum of the preprocessed audio, input the vector into the convolutional encoder-decoder network to generate a time-frequency enhancement mask, scale the time-frequency enhancement mask according to the enhancement level within the parameter group and apply it to the preprocessed audio to generate multiple enhanced audio candidates; Furthermore, the current noise condition is mapped as a vector, and the vector is concatenated with the logarithmic magnitude spectrum of the preprocessed audio. This concatenation is then input into a convolutional encoder-decoder network to generate a time-frequency enhancement mask, including: Map the current noise condition to a learnable vector, copy the learnable vector along the time dimension and concatenate it with the logarithmic magnitude spectrum; The concatenated result is subjected to convolutional encoding, normalization, nonlinear activation, and transposed convolutional decoding to output a single-channel time-frequency quantity. A restricted activation function is used to convert the single-channel time-frequency quantity into a time-frequency enhancement mask.

[0086] Furthermore, the time-frequency enhancement mask is scaled according to the enhancement level within the parameter group and applied to the preprocessed audio to generate multiple enhanced audio candidates, including: Each enhancement level is converted into a mask scaling factor, and the time-frequency enhancement mask is scaled and applied to the pre-processed audio to reconstruct the enhanced audio at each level. The live audio without time-frequency enhancement mask is combined with the enhanced audio at each level to form enhanced audio candidates; Extract the recognition confidence, keyword hit rate and no speech probability of each enhanced audio candidate, accumulate the recognition confidence and keyword hit rate with positive weights, and subtract the no speech probability with negative weights to obtain the comprehensive score; Sort by overall score and output the enhanced audio candidates with the highest overall score.

[0087] Furthermore, training the convolutional encoder-decoder network includes: The existing near-clean speech is segmented into speech segments, and noise segments are extracted from the on-site recordings of various flight training noise conditions. The speech segments and noise segments are mixed according to the preset signal-to-noise ratio range, and condition labels are added to form semi-synthetic training samples. The logarithmic magnitude spectrum of the semi-synthetic training samples and the learnable vector mapped by the working condition labels are input into the convolutional encoder-decoder network, and the speech segments are used as the supervision target. The spectrum reconstruction loss, the human voice frequency band amplitude spectrum weighted mean square loss, and the energy suppression loss above the human voice frequency band are added together with preset weights, and the convolutional encoder-decoder network parameters are updated based on the sum of the losses.

[0088] In this embodiment, the preprocessed audio, the current noise condition, and the parameter group snapshot are used together for subsequent deep learning enhancement processing. The preprocessed audio is subjected to a 25ms frame length, a 10ms frame shift, a Hanning window, and a 512-point short-time Fourier transform to obtain a 257-dimensional amplitude spectrum. The amplitude spectrum is then enhanced... Then take the natural logarithm to form the number of time frames. The frequency-dimensional logarithmic amplitude spectrum is 257. The current noise condition can be mapped to a condition index of 0 to 4, and then mapped to a 16-dimensional vector by the learnable embedding table. Unknown conditions use a separate general vector to avoid forcibly mapping unknown states to a known training condition.

[0089] 16-dimensional learnable vector copying This is then concatenated with the 257-dimensional logarithmic amplitude spectrum according to the characteristic dimension, forming... The input matrix contains both the acoustic spectrum of that frame and the current operating conditions at each time point, allowing the same spectrum to produce different mask responses under different operating conditions.

[0090] The encoder may include two 2D convolutional layers, with the number of channels decreasing from 1 to 16 and then from 16 to 32. Batch normalization and linear rectified activation are applied after each convolutional layer. The decoder includes two transposed convolutional layers, with the number of channels decreasing from 32 to 16 and then from 16 to 1. The output temporal dimension is made to match the number of input frames through cropping or padding, and the frequency dimension is 257. The convolutional kernel can be 3×3, and the stride can be set to 1 or 2 depending on the encoding and decoding dimensions. The exemplary network has approximately 128,000 parameters, and the network structure and parameter version are saved along with the model file during deployment.

[0091] Let the convolutional encoder-decoder network be... The working condition vector is The preprocessed audio spectrum is Then the time-frequency enhancement mask is: ; in, Represents network parameters, This represents a sigmoid activation function that restricts the output to between 0 and 1. for .

[0092] When the network file is not loaded, the condition vector is missing, or the inference output size is incorrect, the time-frequency enhancement mask is initialized to an all-1 matrix, and it can still generate candidates without further enhancement. When non-numerical elements appear in the inference output, the corresponding elements are set to 1 and the abnormal position is recorded; when the abnormality ratio exceeds 5%, the all-1 mask is used.

[0093] Mask after gear scaling It can be represented as: ; in, This indicates an increased gear. This indicates the corresponding scaling factor. The example scaling factors for Conservative, Standard, and Powerful settings are 0.40, 0.70, and 1.00, respectively.

[0094] Within a parameter group, the enhancement level is used as the priority level, and candidates can be added at positions adjacent to the priority level, so that the speech recognition feedback can compare different enhancement intensities.

[0095] The candidate record includes a candidate number, enhancement level, audio sample array, timestamp range, current noise condition, parameter group version, and network version. An exemplary candidate set includes live audio, conservative enhanced audio, standard enhanced audio, and strong enhanced audio. A candidate is removed if reconstruction of a given enhancement level fails; at least one valid candidate from the live audio and preprocessed audio is retained.

[0096] All candidates are identified using the same speech recognition model, the same keyword list, and the same decoding parameters to ensure comparable scores. Recognition confidence is the average confidence of the valid words in the recognition results; keyword hit rate is the ratio of the number of hit keywords to the number of keywords required for the current interaction command; the probability of no speech is the posterior probability of no speech output by the speech recognition model. When there is a lack of recognized text, the recognition confidence and keyword hit rate are set to 0; when there is a lack of no speech probability, the no speech probability is set to 1.

[0097] When the overall scores are the same, candidates with lower enhancement levels are prioritized for output; when the enhancement levels are still the same, candidates with higher speech activity percentages are prioritized for output. The output record includes the selected candidate number, the overall score, and the corresponding recognition result.

[0098] The exemplary training configuration uses approximately 1 hour of existing near-clean speech, segmented into 2-second segments. Equal-length noise segments are extracted from on-site recordings across five work conditions and randomly mixed within a signal-to-noise ratio range of -5 to 10 dB. The work condition labels of the noise segments serve as the work condition labels for the semi-synthesized training samples. The training samples can be divided into a training set and a validation set in a 7:3 ratio. Segments from the same original recording are placed in the same dataset to reduce data correlation caused by segment overlap.

[0099] The short-time Fourier transform parameters of the input and output remain consistent with those during the runtime phase. During state initialization, the convolutional layer parameters are randomly initialized, and the working condition embedding vector is initialized with zero-mean small-amplitude random numbers. During the forward computation of each batch of samples, a time-frequency enhancement mask is output. The estimated speech amplitude spectrum is obtained through the masking effect, and then the training loss is calculated by comparing it with the corresponding approximate clean speech amplitude spectrum.

[0100] The total loss can be expressed as: ; in, Indicates the full-spectrum reconstruction loss. This represents the mean square loss of the amplitude spectrum in the 300–3400 Hz human voice frequency band. This indicates the residual energy suppression loss in the frequency band above 3400Hz.

[0101] The exemplary training uses an adaptive moment estimation optimizer with an initial learning rate of 0.001 and a batch size of 16. Training stops when the total loss of the validation set fails to decrease for 10 consecutive rounds, and the network parameters, working condition embedding table, spectral parameters, and model version of the round with the lowest total loss of the validation set are saved.

[0102] The inference phase is executed sequentially as follows: audio frame segmentation—generating logarithmic amplitude spectrum—reading current condition vector—concatenating condition vector—convolutional encoding—transposed convolutional decoding—generating time-frequency enhancement mask—level scaling—audio reconstruction. After each window processing is completed, the intermediate features of that window are cleared, retaining only the condition status and necessary logs. An exemplary runtime record can record the single-frame network computation time as 1.2ms, the complete inference time for a 6-second window as 80ms, and 200ms as the acceptable processing time limit for voice interaction; these values ​​are used to illustrate the timing field and time limit judgment method, and the deployment values ​​are remeasured by the target processor.

[0103] Evaluation data can be derived from an evaluation set consisting of unused ambient noise and speech segments. Evaluation metrics include signal-to-noise ratio change, energy retention rate in the 300–3400 Hz human voice band, energy change above 3400 Hz, keyword hit rate, recognition confidence, and character error rate. Evaluation results are used to select the model version and check parameter set suitability; they do not directly modify parameters already loaded during runtime.

[0104] See Figure 2Regular masking and convolutional enhancement can be compared using the average signal-to-noise ratio (SNR) and high-frequency energy fields. A set of example values ​​illustrating the evaluation table record format are: average SNR of input audio 0.00dB, 2.12dB after regular masking, and 3.85dB after convolutional enhancement; average high-frequency energy of input audio -5.99dB, -8.20dB after regular masking, and -14.89dB after convolutional enhancement. Another set of example records can be written with an SNR change of 1.83dB, a high-frequency energy change of -2.56dB, and a vocal frequency band energy retention rate of 0.85.

[0105] In this embodiment, a one-dimensional temporal convolutional codec network or a two-dimensional convolutional codec network containing skip connections can also be used to generate a time-frequency enhancement mask, as long as the network receives the current noise condition vector and the logarithmic amplitude spectrum of the preprocessed audio, and outputs the enhancement mask corresponding to the time-frequency position of the preprocessed audio.

[0106] S105 performs speech recognition on the enhanced audio candidates, calculates a comprehensive score based on recognition confidence, keyword hit rate, and probability of no speech, and outputs the enhanced audio candidate with the highest score to the speech interaction recognition system.

[0107] In this embodiment, an enhanced audio candidate set is obtained, and each candidate is fed into the same speech recognition interface. The recognition interface returns the recognized text, word-level confidence score, probability of no speech, recognition start and end time, and candidate number. The keyword list can be configured from the interactive commands of the current flight training subject, including command words such as "start," "stop," "check," and "confirm." The keyword list is loaded at the start of an interactive session and remains unchanged until the session ends; when the keyword list is empty, the keyword hit rate is set to 0, and the keywords are sorted only based on recognition confidence score and probability of no speech.

[0108] Let the first The confidence level for identifying the enhanced audio candidates is: Keyword hit rate The probability of no voice is The overall score is ,but: ; Among them, the recognition confidence, keyword hit rate and no speech probability are all normalized to 0 to 1; 0.5, 0.4 and -0.3 are the corresponding weights.

[0109] In this embodiment, the recognition confidence weight can be set to 0.4–0.6, the keyword hit rate weight to 0.3–0.5, and the absolute value of the no-voice probability deduction weight to 0.2–0.4. All candidates use the same set of weights, and the calculation results and candidate numbers form a sorting record.

[0110] The recognition confidence score can be formed by averaging the confidence scores of the effective words. If the recognition result contains... The first valid word, the first The confidence level of each word is ,but: ; when hour, Set to 0. Keyword hit rate is calculated by dividing the number of unique hit keywords by the number of target keywords; when the number of target keywords is 0, [the following is omitted as it's not relevant to the initial calculation]. Set to 0. After the comprehensive score is calculated, the audio sample arrays of the candidates with the highest scores and the recognition results are sorted in descending order and output to the voice interaction recognition module. If the score difference between multiple candidates is less than 0.02, the candidate with the lower enhancement level is selected to reduce additional spectral distortion.

[0111] See Figure 3 The speech recognition evaluation before and after noise reduction can record keyword hit rate, average recognition confidence, and character error rate. A set of example values ​​for illustrating the comparison fields and data format can be set as follows: keyword hit rate before processing 0.150, after processing 0.450; average recognition confidence before processing 0.144, after processing 0.223; character error rate before processing 0.816, after processing 0.725. In actual deployment, the above fields are re-obtained using independent evaluation data, and the example values ​​are not used as the running judgment threshold.

[0112] If the highest overall score is below 0.20 in five consecutive interactions, the system can enter a fixed conservative processing mode and output a recognition quality prompt to the interactive control terminal. If the highest overall score reaches 0.40 for two consecutive subsequent interactions, adaptive processing based on the current noise level will resume. If the candidate set is empty, the current window's ambient audio will be output directly, and a candidate empty flag will be set. If the speech recognition interface times out, the ambient audio candidates will be retained and re-called in the next window to avoid interrupting the audio link due to a single recognition failure.

[0113] To enable the internal algorithm processing to be monitored and verified, this embodiment also allows for the setting of a runtime log. The log is written when a working condition changes, a mask anomaly occurs, candidate sorting is completed, or continuous low scores trigger or resume adaptive processing. Fields include window start and end times, real-time noise fingerprint summary, cosine matching scores for each template, cumulative score, number of hits, working condition separation amount, current working condition label, parameter group version, frequency domain mask mean, time domain gating ratio, proportion of human voice protection set, candidate number, enhancement level, recognition confidence, keyword hit rate, probability of no speech, comprehensive score, selected candidate number, and anomaly flag. The log can be saved to the robot debugging area in comma-separated text or structured object format and queried by timestamp through a read-only debugging interface. These fields can be used to verify the intermediate results of each step and the final selection process.

[0114] In this embodiment, keyword confidence, instruction syntax compliance, or recognized text completeness can be used instead of keyword hit rate in the comprehensive score; alternatively, the weights can be normalized while maintaining recognition confidence as a positive weight and the probability of no speech as a negative weight. All these methods rely on the speech recognition feedback from multiple enhanced audio candidates and output the highest-ranked candidate to subsequent speech interaction recognition.

[0115] The following describes this embodiment in conjunction with application scenarios: In the step of acquiring audio and extracting real-time noise fingerprints, 16kHz mono live audio is used as input, processed with a 25ms frame length, a 10ms frame shift, a 512-point short-time Fourier transform, and a 6s sliding window. The current window forms a 59-dimensional real-time noise fingerprint, including 7 acoustic statistics, a 40-dimensional log-Melbourne spectrum mean, and a 12-dimensional Mel frequency cepstral coefficient mean.

[0116] In the step of matching noise conditions, the five most recent consecutive windows are read, with weights of 0.10, 0.15, 0.20, 0.25, and 0.30 respectively. Assuming the five matching scores between the current real-time noise fingerprint and the ground alarm template are 0.72, 0.78, 0.83, 0.88, and 0.94 respectively, all reaching the similarity threshold of 0.60, then the cumulative score of the ground alarm template is: ; The cumulative score reaches 0.80, and the highest score in 4 out of 5 windows corresponds to the ground alarm template, reaching the 3-hit threshold. If the weighted average score of the second-highest template is 0.58, and the score difference weight is 0.60, then the working condition separation quantity is: ; The operating condition separation value reaches 0.45, therefore the ground alarm is output as the current noise operating condition.

[0117] In the mask generation and preprocessing steps, the ground alarm parameter group is read, with an example set of alarm suppression weight 0.70, low-frequency suppression weight 0.30, human voice protection coefficient 0.80, and strong setting. The fundamental frequency domain gain within the alarm main frequency band and harmonic range is... The gain at frequencies below 300Hz is If the real-time signal-to-noise ratio of a certain overlapping time-frequency unit is 2dB, then the segmented protection gain is 0.75. After being limited by the voice protection coefficient, the voice protection mask gain is 0.80. At this position, the larger value between the frequency domain gain of 0.30 and the voice protection gain of 0.80 is taken, resulting in 0.80. If this frame is located during the alarm pulse period, the time-domain gating gain is... The final joint gain is The joint gain is applied to the corresponding complex spectrum and inversely transformed to obtain the preprocessed audio.

[0118] In the deep enhancement candidate generation step, ground alarms are mapped to a 16-dimensional condition vector, which is concatenated with a 257-dimensional logarithmic amplitude spectrum to form a 273-dimensional conditional spectrum input. The convolutional encoder-decoder network outputs a time-frequency enhancement mask. If the network mask for a certain time-frequency unit is 0.40, then the scaling masks for the conservative, standard, and strong modes are 0.76, 0.58, and 0.40, respectively. These masks are applied to the preprocessed audio and reconstructed, forming four enhanced audio candidates together with the on-site audio.

[0119] In the step of identifying feedback and selecting the best output, assuming the "recognition confidence, keyword hit rate, and probability of no speech" for the on-site audio, conservative file, standard file, and strong file candidates are "0.46, 0.50, 0.18", "0.62, 0.75, 0.10", "0.71, 1.00, 0.08", and "0.64, 0.75, 0.15" respectively, the corresponding comprehensive scores are 0.376, 0.580, 0.731, and 0.575 respectively. The standard file candidate has the highest comprehensive score of 0.731, therefore, the enhanced standard file audio and its recognition result are output to the speech interaction recognition, completing the entire process from on-site audio input to enhanced audio candidate selection and output.

[0120] For the purpose of simplicity, the method steps disclosed in the above embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0121] like Figure 5 As shown, the present invention also provides an adaptive noise reduction system for robot voice interaction in high-noise environments for flight training, comprising: The audio acquisition and fingerprint extraction module 201 is configured to acquire audio and splice real-time noise fingerprints based on the energy and spectrum of a sliding window. The noise condition identification module 202 is configured to perform cosine matching on a noise template formed by the real-time noise fingerprint and the average of the fingerprints under the same operating condition, weight the window matching score and verify it according to the number of times the template is hit, and determine the current noise condition. The mask preprocessing module 203 is configured to read the parameter set of the current noise condition, as well as the alarm main frequency band, harmonic interval, and repetition period; generate a frequency domain attenuation mask based on the alarm main frequency band and harmonic interval; generate a time domain gating mask based on the repetition period; generate a human voice protection mask based on the real-time signal-to-noise ratio and the human voice protection coefficient within the parameter set; take the larger gain of the frequency domain attenuation mask and the human voice protection mask at the overlap of the alarm frequency band and the human voice frequency band; multiply the processed frequency domain mask and the time domain gating mask and apply them to the audio to obtain the preprocessed audio. The depth enhancement and candidate generation module 204 is configured to map the current noise condition into a vector, concatenate the vector with the logarithmic magnitude spectrum of the preprocessed audio, input the vector into a convolutional encoder-decoder network to generate a time-frequency enhancement mask, scale the time-frequency enhancement mask according to the enhancement level within the parameter group and apply it to the preprocessed audio to generate multiple enhanced audio candidates. The recognition feedback selection module 205 is configured to perform speech recognition on the enhanced audio candidates, calculate a comprehensive score based on the recognition confidence, keyword hit rate and no speech probability according to weights, and output the enhanced audio candidate with the highest score to the speech interaction recognition.

[0122] It is worth noting that although only some basic functional modules are disclosed in the embodiments of this invention, it does not mean that the composition of this system is limited to the above-mentioned basic functional modules. On the contrary, based on the above-mentioned basic functional modules, those skilled in the art can arbitrarily add one or more functional modules in combination with existing technology to form an infinite number of embodiments or technical solutions. That is to say, this system is open rather than closed. The fact that this embodiment only discloses a few basic functional modules should not be considered as the scope of protection of this invention being limited to the disclosed basic functional modules. At the same time, for the convenience of description, the above devices are described separately according to their functions as various units and modules. Of course, in implementing this invention, the functions of each unit and module can be implemented in one or more software and / or hardware.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An adaptive noise reduction method for robot voice interaction in high-noise environments for flight training, characterized in that, include: Acquire audio data and construct a real-time noise fingerprint based on energy and spectrum splicing using a sliding window; Cosine matching is performed on the noise template formed by the real-time noise fingerprint and the average fingerprint of the same working condition. The weighted window matching score is calculated and verified according to the number of template hits to determine the current noise working condition. The system reads the parameter set of the current noise condition, as well as the alarm main frequency band, harmonic interval, and repetition period. It generates a frequency domain attenuation mask based on the alarm main frequency band and harmonic interval, a time domain gating mask based on the repetition period, and a human voice protection mask based on the real-time signal-to-noise ratio and the human voice protection coefficient within the parameter set. At the point where the alarm frequency band and the human voice frequency band overlap, it takes the larger gain of the frequency domain attenuation mask and the human voice protection mask. The processed frequency domain mask is multiplied by the time domain gating mask and applied to the audio to obtain the preprocessed audio. The current noise condition is mapped to a vector, the vector is concatenated with the logarithmic magnitude spectrum of the preprocessed audio, and the vector is input into a convolutional encoder-decoder network to generate a time-frequency enhancement mask. The time-frequency enhancement mask is scaled according to the enhancement level within the parameter group and applied to the preprocessed audio to generate multiple enhanced audio candidates. Speech recognition is performed on the enhanced audio candidates, and a comprehensive score is calculated based on the recognition confidence, keyword hit rate and no speech probability according to the weight. The enhanced audio candidate with the highest score is output to the speech interaction recognition. Among them, a preset number of consecutive windows are arranged according to the end time, and higher weights are assigned to later windows, while single window scores below the similarity threshold are removed. Based on the accumulated single-window scores retained according to the noise template, candidate operating conditions that reach the cumulative score threshold are filtered. Count the number of times the template of the candidate working condition is hit, and retain the candidate working conditions that reach the hit threshold; For candidate operating conditions that have not reached the hit count threshold, the earliest window is removed and the next window is added. Cosine matching, weighted sum, and hit count verification are then re-executed until the current noisy operating condition is determined.

2. The adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training as described in claim 1, characterized in that, Constructing a noise template, further including: Amplitude normalization, fixed gain equalization, framing, and time-frequency transformation are performed on the on-site recordings of various flight training noise conditions. Short-time energy, energy dispersion, peak frequency, centroid, flatness, alarm repetition period, and proportion of voice activity are extracted from each recording. Each acoustic quantity is then concatenated with the log-Melbourne spectrum mean and the Mel frequency cepstral coefficient mean to form a sample noise fingerprint. The baseline fingerprint is obtained by averaging the noise fingerprints of similar samples according to the operating condition label. The operating condition label, reference fingerprint, alarm main frequency band, harmonic range, energy envelope, and repetition period are stored together as a noise template.

3. The adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training as described in claim 1, characterized in that, The parameter set for the current noise condition includes alarm suppression weight, low-frequency suppression weight, human voice protection factor, and enhancement level; During the offline configuration phase, the alarm suppression weight is configured as the attenuation gain of the alarm main frequency band and harmonic range, the low frequency suppression weight is configured as the attenuation gain of frequencies below the human voice frequency band, the human voice protection coefficient is configured as the lower limit of the gain of the human voice frequency band, and the enhancement level is converted into the scaling factor of the time-frequency enhancement mask. The parameter values ​​are combined and stored according to the operating condition label. During the operation phase, the parameter group with the same operating condition label as the current noise operating condition is read, and the stored parameter values ​​are kept unchanged.

4. The adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training as described in claim 1, characterized in that, A frequency-domain attenuation mask is generated based on the alarm main frequency band and harmonic interval, and a time-domain gated mask is generated based on the repetition period. Further, it includes: Read the alarm main frequency band, harmonic range, energy envelope and repetition period from the noise template of the current noise condition; Set the attenuation gain of each frequency point in the alarm main frequency band and harmonic interval according to the alarm suppression weight to form a frequency domain attenuation mask; The alarm pulse period is located in each repetition cycle according to the energy envelope. The gating attenuation of the alarm pulse period is configured to be higher than that of the non-alarm pulse period to form a time-domain gating mask.

5. The adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training according to claim 1, characterized in that, A voice protection mask is generated based on the real-time signal-to-noise ratio and the voice protection coefficient within the parameter group, and further includes: Calculate the real-time signal-to-noise ratio of each time-frequency unit within the human voice frequency band, map the segmented intervals of the real-time signal-to-noise ratio to the protection gain, and limit the lower limit of the protection gain based on the human voice protection coefficient to generate a human voice protection mask. Where the alarm main frequency band or harmonic interval overlaps with the human voice frequency band, the gain of the frequency domain attenuation mask and the gain of the human voice protection mask are taken to a larger value in each time-frequency unit. Preserve the frequency domain attenuation mask gain for frequencies that do not fall within the human voice frequency band.

6. The adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training according to claim 1, characterized in that, Mapping the current noise condition to a vector, concatenating the vector with the logarithmic magnitude spectrum of the preprocessed audio, and inputting the convolutional encoder-decoder network to generate a time-frequency enhancement mask, further includes: Map the current noise condition to a learnable vector, copy the learnable vector along the time dimension and concatenate it with the logarithmic magnitude spectrum; The concatenated result is subjected to convolutional encoding, normalization, nonlinear activation, and transposed convolutional decoding to output a single-channel time-frequency quantity. A restricted activation function is used to convert the single-channel time-frequency quantity into a time-frequency enhancement mask.

7. The adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training according to claim 1, characterized in that, The time-frequency enhancement mask is scaled according to the enhancement level within the parameter group and applied to the preprocessed audio to generate multiple enhanced audio candidates, further including: Each enhancement level is converted into a mask scaling factor, and the time-frequency enhancement mask is scaled and applied to the pre-processed audio to reconstruct the enhanced audio at each level. The live audio without time-frequency enhancement mask is combined with the enhanced audio at each level to form enhanced audio candidates; Extract the recognition confidence, keyword hit rate and no speech probability of each enhanced audio candidate, accumulate the recognition confidence and keyword hit rate with positive weights, and subtract the no speech probability with negative weights to obtain the comprehensive score; Sort by overall score and output the enhanced audio candidates with the highest overall score.

8. The adaptive noise reduction method for robot voice interaction in a high-noise environment for flight training according to claim 6, characterized in that, Training the convolutional encoder-decoder network further includes: The existing near-clean speech is segmented into speech segments, and noise segments are extracted from the on-site recordings of various flight training noise conditions. The speech segments and noise segments are mixed according to the preset signal-to-noise ratio range, and condition labels are added to form semi-synthetic training samples. The logarithmic magnitude spectrum of the semi-synthetic training samples and the learnable vector mapped by the working condition labels are input into the convolutional encoder-decoder network, and the speech segments are used as the supervision target. The spectrum reconstruction loss, the human voice frequency band amplitude spectrum weighted mean square loss, and the energy suppression loss above the human voice frequency band are added together with preset weights, and the convolutional encoder-decoder network parameters are updated based on the sum of the losses.

9. An adaptive noise reduction system for robot voice interaction in high-noise environments for flight training, characterized in that, include: The audio acquisition and fingerprint extraction module is configured to acquire audio and splice real-time noise fingerprints based on the energy and spectrum of a sliding window. The noise condition identification module is configured to perform cosine matching between the real-time noise fingerprint and the noise template formed by the average of the fingerprints under the same condition, weight the window matching score and verify according to the number of times the template is hit, and determine the current noise condition. The mask preprocessing module is configured to read the parameter set of the current noise condition, as well as the alarm main frequency band, harmonic interval, and repetition period; generate a frequency domain attenuation mask based on the alarm main frequency band and harmonic interval; generate a time domain gating mask based on the repetition period; generate a human voice protection mask based on the real-time signal-to-noise ratio and the human voice protection coefficient within the parameter set; take the larger gain of the frequency domain attenuation mask and the human voice protection mask at the overlap of the alarm frequency band and the human voice frequency band; multiply the processed frequency domain mask and the time domain gating mask and apply them to the audio to obtain the preprocessed audio. The depth enhancement and candidate generation module is configured to map the current noise condition into a vector, concatenate the vector with the logarithmic magnitude spectrum of the preprocessed audio, input the vector into a convolutional encoder-decoder network to generate a time-frequency enhancement mask, scale the time-frequency enhancement mask according to the enhancement level within the parameter group and apply it to the preprocessed audio to generate multiple enhanced audio candidates. The recognition feedback selection module is configured to perform speech recognition on enhanced audio candidates, calculate a comprehensive score based on recognition confidence, keyword hit rate and no speech probability according to weights, and output the enhanced audio candidate with the highest score to the speech interaction recognition. Among them, a preset number of consecutive windows are arranged according to the end time, and higher weights are assigned to later windows, while single window scores below the similarity threshold are removed. Based on the accumulated single-window scores retained according to the noise template, candidate operating conditions that reach the cumulative score threshold are filtered. Count the number of times the template of the candidate working condition is hit, and retain the candidate working conditions that reach the hit threshold; For candidate operating conditions that have not reached the hit count threshold, the earliest window is removed and the next window is added. Cosine matching, weighted sum, and hit count verification are then re-executed until the current noisy operating condition is determined.

Citation Information

Patent Citations

  • Exhaust valve self-adaptive noise reduction method based on environmental noise recognition

    CN120260595A

  • Single-channel speech enhancement method based on time-frequency interaction in vehicle-mounted environment

    CN121331152A