Audio optimization method and system for speech recognition

By evaluating and reselecting the noise spectrum in real time, using the energy hysteresis characteristics and spectrum characteristics of the speech signal, the problem of poor adaptability of fixed noise spectrum in outdoor environments is solved, and the accuracy of speech recognition is improved.

CN120299470BActive Publication Date: 2025-08-12FUZHOU UNIV ZHICHENG COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510787147.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-08-12
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

In outdoor environments, the spectrum subtraction of fixed noise spectrum is poor in adaptability to complex noise signals, resulting in a decrease in speech recognition accuracy.

Method used

By collecting audio data in real time, evaluating the energy changes of neighboring frames and contrast frames, reselecting the noise spectrum, using the energy lag characteristics and spectrum characteristics of the speech signal, extracting noise modes and speech modes, and obtaining new noise spectrums for audio enhancement.

Benefits of technology

It improves the accuracy of noise frame selection, enhances the audio enhancement effect of spectral subtraction in complex noise environments outdoors, and improves the accuracy of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299470B_ABST
    Figure CN120299470B_ABST
Patent Text Reader

Abstract

The present application relates to the field of speech enhancement technology, and specifically to an audio optimization method and system for speech recognition, the method comprising: collecting audio data in real time and evenly dividing it into audio frames; for each audio frame, presetting the neighboring frames of the audio frame, and evaluating whether to reselect the noise spectrum when using spectral subtraction to enhance the audio data; if reselected, obtaining the modes of the audio frame; selecting the noise mode from the modes of the audio frame and obtaining the number of lagged frames of the remaining modes; obtaining the spectrum change characteristic values and speech characteristic values of each mode, and selecting the main speech mode; obtaining the audio mode characteristic values of each mode other than the main speech mode; and then obtaining the noise frame and the new noise spectrum. The present application aims to enhance the audio enhancement effect of spectral subtraction on the speech features of the audio frame by improving the accuracy of noise frame selection, thereby improving the accuracy of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech enhancement technology, and in particular to an audio optimization method and system for speech recognition. Background Art

[0002] Speech recognition is a technology that converts speech signals into text. In speech recognition, noise can interfere with the speech signal, reducing recognition accuracy. Spectral subtraction, a commonly used audio optimization method, utilizes techniques such as Fourier transforms to remove background noise through frequency domain processing, preserving speech phase information and thereby enhancing the speech signal. Spectral subtraction is simple in principle, easy to implement, widely applicable, and minimally degrading to speech. It effectively improves the signal-to-noise ratio of audio data, providing a clearer input signal for the speech recognition system and reducing noise interference, thereby enhancing the accuracy and reliability of speech recognition.

[0003] In practical applications, the choice of noise spectrum directly impacts the effectiveness of spectral subtraction on audio data. Outdoor audio data is characterized by numerous noise sources, resulting in strong time-varying characteristics. When using spectral subtraction with a fixed noise spectrum to enhance audio data in outdoor environments, the method exhibits poor adaptability to complex outdoor noise signals. When noise signal characteristics change, the fixed noise spectrum directly results in the incorrect filtering of useful speech signals, resulting in reduced speech recognition accuracy. Summary of the Invention

[0004] In view of the above, it is necessary to provide an audio optimization method and system for speech recognition. Compared with traditional audio optimization methods and systems for speech recognition, this method improves the accuracy of noise frame selection and enhances the audio enhancement effect of spectral subtraction on the speech features of audio frames, thereby improving the accuracy of speech recognition:

[0005] In a first aspect, an embodiment of the present application provides an audio optimization method for speech recognition, the method comprising the following steps:

[0006] Collect audio data in real time and evenly divide it into audio frames of preset length; enhance the audio data using spectral subtraction;

[0007] For each audio frame, neighboring frames and comparison frames of the audio frame are preset before the audio frame. The energy change degree of all neighboring frames and all comparison frames is used to evaluate whether to reselect the noise spectrum when enhancing the audio data;

[0008] If reselected, each mode of the audio frame is obtained; the modal energy change value of each mode is obtained by the energy change degree of each same mode of all neighboring frames, and the noise mode is selected from the mode of the audio frame; the energy of each same mode of all neighboring frames is arranged in time sequence to form each energy sequence, and the number of lagged frames of the remaining modes is obtained by analyzing the cross-correlation of the energy sequences between the noise mode of the audio frame and the remaining modes;

[0009] The spectral change characteristic value of each mode is obtained by the difference degree of the marginal spectrum of each mode between the audio frame and its neighboring frames; the speech characteristic value of each mode is obtained by combining the proportion of the number of lagged frames of each mode in the neighboring frames with the spectral change characteristic value, and the main speech mode is selected from the modes of the audio frame; the audio modal characteristic values of other modes are obtained by the similarity of the Mel-frequency cepstral coefficients between the main speech mode and other modes;

[0010] The discreteness of the audio modal feature values of all other modes of each audio frame is obtained, and the noise frame and the new noise spectrum are obtained through the discreteness of each audio frame and a preset number of adjacent audio frames thereafter.

[0011] In one embodiment, the evaluating whether to reselect the noise spectrum when enhancing the audio data includes:

[0012] Calculate the mean energy of all neighboring frames of each audio frame;

[0013] Recording the absolute value of the slope of the fitted straight line of the mean value of all neighboring frames of each audio frame as the first absolute value of each audio frame;

[0014] Recording the absolute value of the slope of the fitted straight line of the mean values of all comparison frames of each audio frame as the second absolute value of each audio frame;

[0015] Extracting the upper quartile of the first absolute value of all comparison frames of each audio frame;

[0016] When the mean of the first absolute value and the second absolute value is greater than the upper quartile, the noise spectrum is reselected; otherwise, the noise spectrum is not reselected.

[0017] In one embodiment, obtaining the modal energy change value of each modality and selecting the noise modality from the modality of the audio frame includes:

[0018] The absolute value of the slope of the fitted straight line of the energy of each mode of all neighboring frames is used as the modal energy change value of each mode, and the mode with the largest modal energy change value of the audio frame is used as the noise mode.

[0019] In one embodiment, the process of obtaining the number of delayed frames is as follows:

[0020] Obtaining a cross-correlation sequence of energy sequences between the noise mode of the audio frame and the remaining modes;

[0021] The absolute value of the difference between the sequence number corresponding to the maximum value in the cross-correlation sequence and the center sequence number of the cross-correlation sequence is used as the lag frame number of the remaining modes.

[0022] In one embodiment, the calculation process of the spectrum change characteristic value is:

[0023] Each marginal spectrum sequence is normalized, and the DTW distance between the marginal spectrum sequences of the same mode of any two adjacent audio frames in all neighboring frames is calculated. The spectrum change characteristic value is the average value of the DTW distance between all any two adjacent audio frames in all neighboring frames.

[0024] In one embodiment, the method for obtaining the speech feature value is:

[0025] For each mode, the expression of the speech feature value of the mode is:

[0026] ; Where F represents the speech feature value of the modality; m represents the number of delayed frames of the modality; M represents the number of neighboring frames; ε represents a preset positive integer; Indicates the spectral variation eigenvalue of the mode.

[0027] In one embodiment, the main speech mode is the mode with the largest speech feature value in the audio frame.

[0028] In one embodiment, the method for obtaining the audio modal feature value is:

[0029] The first 12 coefficients of the Mel-frequency cepstral coefficients are used to form a speech feature vector. The audio modality feature value is a normalized value of the cosine similarity of the speech feature vectors between the main speech modality and other modalities.

[0030] In one embodiment, the method for obtaining the noise frame and the new noise spectrum is:

[0031] Each audio frame and the audio frame with the smallest discreteness among a preset number of adjacent audio frames thereafter are used as noise frames, and the frequency spectrum of the noise frame is used as a new noise spectrum.

[0032] In a second aspect, an embodiment of the present application also provides an audio optimization system for speech recognition, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, the system implements the steps of any one of the above-mentioned audio optimization methods for speech recognition.

[0033] This application has at least the following beneficial effects:

[0034] This application utilizes the hysteresis characteristics of speech signals in energy changes and the difference between them and noise signals in spectral energy distribution, adopts time-frequency analysis methods to extract speech features, reduce the interference of outdoor noise time variation on speech feature extraction, improve the accuracy of speech signal feature extraction, and thus improve the accuracy of subsequent noise frame selection, and enhance the adaptability of spectral subtraction for audio enhancement in complex outdoor noise environments;

[0035] Furthermore, based on the independent characteristics of noise signals and speech signals in outdoor environment audio data, different modes of alternative noise frames are estimated, and the distribution differences of different modal features in speech frames and non-speech frames are utilized to avoid the influence of time-varying characteristics of outdoor noise on the selected noise frames, improve the accuracy of noise frame selection, and enhance the audio enhancement effect of spectral subtraction on the speech features of audio frames, so as to improve the accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0037] Figure 1 A flowchart of the steps of an audio optimization method for speech recognition provided in one embodiment of the present application;

[0038] Figure 2 A flowchart for evaluating whether to reselect the noise spectrum;

[0039] Figure 3 Schematic diagram of the noise frame acquisition process. DETAILED DESCRIPTION

[0040] In the description of the embodiments of this application, words such as "exemplary," "or," and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "or," and "for example" is intended to present the relevant concepts in a concrete manner.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application relates. The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. It should be understood that, unless otherwise indicated, " / " represents or.

[0042] It should also be noted that the terms "first" and "second" in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0043] The following describes in detail the specific solutions of the audio optimization method and system for speech recognition provided by this application with reference to the accompanying drawings.

[0044] See also Figure 1 , which shows a flowchart of the steps of an audio optimization method for speech recognition provided by an embodiment of the present application, the method comprising the following steps:

[0045] Step 1: collect audio data in real time and evenly divide it into audio frames of preset length.

[0046] The sampling rate of the audio data in this embodiment is 22.05kHz. To facilitate data processing, the audio data is framed, and the frame length of each audio frame is set to 25 milliseconds. To minimize distortion, this embodiment sets an overlap between two adjacent audio frames, where the frame length of the overlapping portion is 1 / 4 of the audio frame length. The sampling rate, audio frame length, and overlapping portion are preset and can be set by the implementer. This application does not impose any special restrictions.

[0047] At the same time, in order to avoid spectrum leakage and information loss caused by audio data frame processing, this application performs windowing processing on each audio frame separately.

[0048] In this embodiment, a Hamming window function is used to perform windowing processing on each audio frame.

[0049] Step 2: For each audio frame, preset neighboring frames and comparison frames of the audio frame before the audio frame, and evaluate whether to reselect the noise spectrum when enhancing the audio data based on the degree of energy change of all neighboring frames and all comparison frames.

[0050] The energy of speech signals is relatively fixed, while the energy of outdoor noise signals varies more randomly. When the noise signal energy changes, the fixed noise spectrum directly causes the useful speech signal to be incorrectly filtered out, making it impossible to accurately recognize speech.

[0051] For each audio frame, take N-1 audio frames before it, record the N audio frames including itself as the comparison frames for each audio frame, and calculate the energy of each comparison frame. If the number of audio frames before each audio frame is less than N-1, no further calculation is performed.

[0052] The change of outdoor noise signal directly causes the change of its energy. Considering the complex situation of instantaneous change and steady change of outdoor noise signal energy, it is judged whether each audio frame is in the enhanced state of noise signal time-varying.

[0053] Given the strong time-varying characteristics of speech signals over short periods of time, the energy of a single audio frame changes rapidly. Therefore, starting with each audio frame as the last frame, we select M-1 audio frames forward and record the M audio frames, including each audio frame, as its neighboring frames. We then calculate the average energy of all neighboring frames for each audio frame.

[0054] A fitted straight line is obtained for the mean values of all neighboring frames of each audio frame. The absolute value of the slope of the fitted line is recorded as the first absolute value, which reflects the instantaneous energy change near each audio frame. The calculation of the slope is well known in the art and will not be further described in this application.

[0055] A fitting straight line of the mean values of all comparison frames of each audio frame is obtained, and the absolute value of the slope of the fitting straight line is recorded as a second absolute value, which is used to reflect the degree of stable energy change near each audio frame.

[0056] In this embodiment, in order to simultaneously analyze the instantaneous energy change degree and the steady energy change degree near each audio frame, the value of N is set to 120 and the value of M is set to 10. When the value of N is at least 10 times the value of M, the values of N and M can be preset manually and can be set by the implementer according to actual conditions.

[0057] In this embodiment, the least squares method is used to obtain the fitting straight line. The least squares method is a well-known technology and will not be described in detail in this application. As other implementation methods, on the basis of being able to obtain the fitting straight line of the mean of all neighboring frames of each audio frame and the fitting straight line of the mean of all comparison frames of each audio frame, the implementer may adopt other existing technologies, such as the weighted least squares method, etc., and this application does not impose any special restrictions.

[0058] The upper quartile of the first absolute value of all comparison frames of each audio frame is extracted and recorded as the energy change threshold. When the average of the first absolute value and the second absolute value of each audio frame is greater than the energy change threshold, it indicates that the time-varying nature of the noise signal is enhanced. When the spectral subtraction method is used to enhance the audio data, the noise spectrum is reselected. Otherwise, the noise spectrum is not reselected. The flowchart of evaluating whether to reselect the noise spectrum is as follows: Figure 2 shown.

[0059] Step 3: If reselected, obtain each mode of the audio frame; obtain the modal energy change value of each mode through the energy change degree of each identical mode of all neighboring frames, and select the noise mode from the mode of the audio frame; arrange the energy of each identical mode of all neighboring frames in time sequence to form each energy sequence, and obtain the number of delayed frames of the remaining modes by analyzing the cross-correlation of the energy sequences between the noise mode of the audio frame and the remaining modes.

[0060] For each audio frame, N-1 audio frames are taken backward, and the N audio frames including itself are recorded as the alternative noise frames of each audio frame. Taking into account that the number of noise sources and their vibration modes change relatively fixedly in a short period of time, each mode of each alternative noise frame is obtained, and the energy of each mode is calculated. At the same time, the marginal spectrum sequence of each mode is obtained to reflect the spectral energy distribution of each mode. Among them, the calculation process of the marginal spectrum sequence is a well-known technology and will not be repeated in this application. When the number of audio frames after each audio frame is less than N-1, no subsequent calculation is performed.

[0061] In this embodiment, the empirical mode decomposition method is used to obtain each mode of each alternative noise frame. The ensemble empirical mode decomposition algorithm is a well-known technology and will not be described in detail in this application. As other implementation methods, on the basis of being able to obtain each mode of each alternative noise frame, the implementer may adopt other existing technologies, such as the ensemble empirical mode decomposition algorithm, the complementary ensemble empirical mode decomposition method, etc., and this application does not impose any special restrictions.

[0062] In outdoor environments, the energy changes of speech signals lag behind those of noise signals. For example, when reporting outdoors, as the noise energy increases, the speaker will raise their voice volume, and the energy of the speech signal will also increase. However, this energy change lags behind the ambient noise.

[0063] At the same time, as the tone and volume increase, the speaker's voice signal characteristics also change slightly, and it takes a certain amount of time for the speaker to adapt to the new noise signal environment. During this period, the energy change of the voice signal will continue to lag behind the energy change of the noise signal.

[0064] In order to extract the lag characteristics of the speech signal in the audio data, for each candidate noise frame, each candidate noise frame is taken as the last frame, M-1 audio frames are taken forward, and the M audio frames including each candidate noise frame are recorded as the reference frames of each candidate noise frame.

[0065] Obtain a fitted straight line for the energy of any identical mode in all comparison frames of each candidate noise frame, and use the absolute value of the slope of the fitted straight line as the modal energy change value of the identical mode. When using the modal decomposition algorithm to obtain the modes of each candidate noise frame, the modes are numbered in the order in which they were obtained, and modes with the same number in any two candidate noise frames are considered the same mode.

[0066] In this embodiment, the least squares method is used to obtain the fitting straight line. The least squares method is a well-known technology and will not be described in detail in this application. As other implementation methods, on the basis of being able to obtain the fitting straight line of the energy of any same mode of all control frames of each alternative noise frame, the implementer may adopt other existing technologies, such as the weighted least squares method, etc., and this application does not impose any special restrictions.

[0067] For each candidate noise frame, the mode with the largest modal energy change is selected as the noise mode, representing the vibration mode of the noise source where the primary energy change occurs. Furthermore, due to the influence of the noise source system resonance, the energy changes of the remaining vibration modes of the same noise source also exhibit hysteresis, but this hysteresis is less pronounced than that of the speech signal.

[0068] The energy of each identical mode in all control frames of each candidate noise source is arranged in time sequence to form an energy sequence for each mode. A cross-correlation sequence is obtained between the energy sequences of the noise mode of each candidate noise source and the remaining modes. The absolute value of the difference between the sequence number corresponding to the maximum value in the cross-correlation sequence and the center sequence number of the cross-correlation sequence is used as the lag frame number of the remaining modes of each candidate noise source, reflecting the degree of lag in energy changes of the remaining modes compared to the noise mode. The greater the lag frame number of a mode, the greater the lag in energy changes of the mode, and the greater the probability that the mode contains a speech signal.

[0069] Step 4: Obtain the spectrum change characteristic value of each mode by the difference degree of the marginal spectrum of each same mode between the audio frame and its neighboring frames; obtain the speech characteristic value of each mode by the proportion of the number of lagged frames of each mode in the neighboring frames, combined with the said spectrum change characteristic value, and select the main speech mode from the mode of the audio frame; obtain the audio modal characteristic value of other modes by the similarity of the Mel-frequency cepstral coefficients between the main speech mode and other modes.

[0070] The frequency distribution of both the phonemes of speech signals and outdoor noise signals exhibits a certain degree of time-varying properties. However, the time-varying characteristics of speech signals are concentrated within a specific frequency range, and the energy variation between different frequencies is large, resulting in stronger time-varying properties at smaller scales. In contrast, the frequency variation range of outdoor noise signals is large, but the energy variation between different frequencies is small, allowing them to be processed as stationary signals over shorter time periods.

[0071] The marginal spectrum sequence is normalized to control the calculation range. For all reference frames of each candidate noise frame, the DTW (Dynamic Time Warping) distance between the marginal spectrum sequences of the same mode of any two adjacent audio frames in all reference frames is calculated. The average of the DTW distances between all any two adjacent audio frames in all reference frames is used as the spectrum change characteristic value of each mode of each candidate noise frame, which is used to reflect the degree of change in the spectrum energy of each mode in all reference frames of the candidate noise frame. The larger the spectrum change characteristic value, the stronger the time-varying nature of the short-term spectrum energy and the greater the probability that the mode contains a speech signal.

[0072] In this embodiment, the Min-Max normalization method is used to normalize the marginal spectrum sequence.

[0073] Furthermore, the speech feature value of each mode is obtained by combining the proportion of the number of lag frames of each mode of each candidate noise frame in the control frame with the spectrum change feature value of each mode, which is used to reflect the probability of containing speech signals in each mode of the candidate noise frame. For each mode, the expression of the speech feature value of the mode is:

[0074] ; Wherein, F represents the speech feature value of the mode; m represents the number of lag frames of the mode; M represents the number of neighboring frames; ε represents a preset positive integer used to avoid affecting the calculation result of the speech feature value when the numerator is 0. In this embodiment, the value of ε is 1. The value of ε can be preset manually and is not particularly limited in this application; Indicates the spectral variation eigenvalue of the mode.

[0075] Furthermore, the mode with the largest speech feature value of each candidate noise frame is taken as the main speech mode, indicating the mode with the highest probability of containing a speech signal.

[0076] Step 5: Obtain the discreteness of the audio modal feature values of all other modes of each audio frame, and obtain the noise frame and the new noise spectrum through the discreteness of each audio frame and a preset number of neighboring audio frames thereafter.

[0077] Taking into account the inherent voiceprint characteristics of speech signals, this application calculates the Mel-frequency cepstral coefficients for each mode of each candidate noise frame. The Mel-frequency cepstral coefficient is a vector containing 13 coefficients, and the first 12 coefficients form the speech feature vector for each mode. Mel-frequency cepstral coefficients are well known in the art and will not be described in detail in this application.

[0078] The audio modal feature values of the other modalities are obtained by calculating the similarity of the speech feature vectors between the main speech modality and the other modalities of each candidate noise frame, which is used to reflect the degree of difference in speech features between the other modalities and the main speech modality. For each candidate noise frame, the expression of the audio modal feature values of the other modalities of the candidate noise frame is:

[0079] Where, represents the audio modal eigenvalue of the kth mode of the candidate noise frame; norm[ ] represents the arc tangent normalization function; sim( ) represents the cosine similarity function; 、 They represent the main speech mode and the speech feature vector of the kth mode of the alternative noise frame respectively.

[0080] Furthermore, the discreteness of the audio modal feature values for all other modes of each candidate noise frame is calculated to reflect the distribution of speech features within different modalities. The fewer speech components a candidate noise frame contains, the closer the distribution of speech features across the modalities, the smaller the calculated discreteness, and the more accurately it reflects the noise conditions within the time period of each candidate noise frame.

[0081] In this embodiment, the discreteness of the audio modal eigenvalue is the standard deviation. As other implementation methods, on the basis of being able to measure the degree of uneven distribution of the audio modal eigenvalue, the implementer may adopt other existing technologies, such as variance, coefficient of variation, etc., and this application does not impose any special restrictions.

[0082] Furthermore, the candidate noise frame with the smallest discreteness among all candidate noise frames of each audio frame is used as the noise frame, and the spectrum of the noise frame is used as the new noise spectrum. Figure 3 shown.

[0083] The audio signal is enhanced by using the new noise spectrum using spectral subtraction to complete the audio optimization. It should be noted that when there is no need to reselect the noise spectrum, the audio signal is directly enhanced using spectral subtraction. Spectral subtraction is a well-known technique and will not be described in detail in this application.

[0084] Based on the same inventive concept as the above method, an embodiment of the present application also provides an audio optimization system for speech recognition, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned audio optimization methods for speech recognition are implemented.

[0085] In summary, this application utilizes the hysteresis characteristics of speech signals in energy changes and the difference between them and noise signals in spectral energy distribution, adopts time-frequency analysis methods to extract speech features, reduce the interference of outdoor noise time variation on speech feature extraction, improve the accuracy of speech signal feature extraction, and thus improve the accuracy of subsequent noise frame selection, and enhance the adaptability of spectral subtraction for audio enhancement in complex outdoor noise environments;

[0086] Furthermore, based on the independent characteristics of noise signals and speech signals in outdoor environment audio data, different modes of alternative noise frames are estimated, and the distribution differences of different modal features in speech frames and non-speech frames are utilized to avoid the influence of time-varying characteristics of outdoor noise on the selected noise frames, improve the accuracy of noise frame selection, and enhance the audio enhancement effect of spectral subtraction on the speech features of audio frames, so as to improve the accuracy of speech recognition.

[0087] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operations of the systems, methods and computer program products according to the embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.

[0088] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the basic characteristics of the present application. Therefore, from all perspectives, the above embodiments of the present application should be regarded as exemplary and non-restrictive.

Claims

1. An audio optimization method for speech recognition, characterized in that: The method comprises the following steps: Collect audio data in real time and evenly divide it into audio frames of preset length; enhance the audio data using spectral subtraction; For each audio frame, neighboring frames and comparison frames of the audio frame are preset before the audio frame. The energy change degree of all neighboring frames and all comparison frames is used to evaluate whether to reselect the noise spectrum when enhancing the audio data; If reselected, each mode of the audio frame is obtained; the modal energy change value of each mode is obtained by the energy change degree of each same mode of all neighboring frames, and the mode with the largest modal energy change value of the audio frame is selected from the modes of the audio frame and recorded as the noise mode; the energy of each same mode of all neighboring frames is arranged in time sequence to form each energy sequence, and the lag frame number of the remaining modes is obtained by analyzing the cross-correlation of the energy sequence between the noise mode of the audio frame and the remaining modes; The spectral change characteristic value of each mode is obtained by the difference degree of the marginal spectrum of each same mode between the audio frame and its neighboring frames; the speech characteristic value of each mode is obtained by combining the proportion of the number of lagged frames of each mode in the neighboring frames with the spectral change characteristic value, and the mode with the largest speech characteristic value of the audio frame is selected from the modes of the audio frame and recorded as the main speech mode; the audio modal characteristic values of other modes are obtained by the similarity of the Mel-frequency cepstral coefficients between the main speech mode and other modes; The discreteness of the audio modal feature values of all other modes of each audio frame is obtained, and the noise frame and the new noise spectrum are obtained through the discreteness of each audio frame and a preset number of adjacent audio frames thereafter.

2. The audio optimization method for speech recognition according to claim 1, wherein: The evaluation of whether to reselect the noise spectrum when enhancing the audio data includes: Calculate the average energy of all neighboring frames of each audio frame; Recording the absolute value of the slope of the fitted straight line of the mean value of all neighboring frames of each audio frame as the first absolute value of each audio frame; Recording the absolute value of the slope of the fitted straight line of the mean values of all comparison frames of each audio frame as the second absolute value of each audio frame; Extracting the upper quartile of the first absolute value of all comparison frames of each audio frame; When the mean of the first absolute value and the second absolute value is greater than the upper quartile, the noise spectrum is reselected; otherwise, the noise spectrum is not reselected.

3. The audio optimization method for speech recognition according to claim 1, wherein: The obtaining of the modal energy change value of each mode includes: The absolute value of the slope of the fitted straight line of the energy of each mode of the same mode in all neighboring frames is taken as the modal energy change value of each mode.

4. The audio optimization method for speech recognition according to claim 1, wherein: The process of obtaining the delayed frame number is as follows: Obtaining a cross-correlation sequence of energy sequences between the noise mode of the audio frame and the remaining modes; The absolute value of the difference between the sequence number corresponding to the maximum value in the cross-correlation sequence and the center sequence number of the cross-correlation sequence is used as the lag frame number of the remaining modes.

5. The audio optimization method for speech recognition according to claim 1, wherein: The calculation process of the spectrum change characteristic value is: Each marginal spectrum sequence is normalized, and the DTW distance between the marginal spectrum sequences of the same mode of any two adjacent audio frames in all neighboring frames is calculated. The spectrum change characteristic value is the average value of the DTW distance between all any two adjacent audio frames in all neighboring frames.

6. The audio optimization method for speech recognition according to claim 1, wherein: The method for obtaining the speech feature value is: For each mode, the expression of the speech feature value of the mode is: ; Where F represents the speech feature value of the modality; m represents the number of delayed frames of the modality; M represents the number of neighboring frames; ε represents a preset positive integer; Indicates the spectral variation eigenvalue of the mode.

7. The audio optimization method for speech recognition according to claim 1, wherein: The method for obtaining the audio modal eigenvalue is: The first 12 coefficients of the Mel-frequency cepstral coefficients are used to form a speech feature vector. The audio modality feature value is a normalized value of the cosine similarity of the speech feature vectors between the main speech modality and other modalities.

8. The audio optimization method for speech recognition according to claim 1, wherein: The method for obtaining the noise frame and the new noise spectrum is: Each audio frame and the audio frame with the smallest discreteness among a preset number of adjacent audio frames thereafter are used as noise frames, and the frequency spectrum of the noise frame is used as a new noise spectrum.

9. An audio optimization system for speech recognition, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the audio optimization method applied to speech recognition as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Quick estimation method for abrupt change noise

    CN111933165A

  • Method for enhancing field bird buzzing audio data

    CN118173106A