Audio optimization method and system applied to speech recognition

By evaluating the energy changes and spectrum characteristics of audio frames in real time, and dynamically selecting the noise spectrum, the problem of poor noise adaptability in outdoor environments is solved, and the accuracy of speech recognition and the accuracy of noise frame selection is improved.

CN120299470AActive Publication Date: 2025-07-11FUZHOU UNIV ZHICHENG COLLEGE

Patent Information

Application Number
CN202510787147.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

In outdoor environments, the spectrum subtraction of fixed noise spectrum is poor in adaptability to complex noise signals, resulting in a decrease in speech recognition accuracy.

Method used

By collecting audio data in real time, evaluating the energy changes of neighboring frames and contrast frames, reselecting the noise spectrum, using the energy lag characteristics of the speech signal and the differences in the spectrum energy distribution, extracting the speech characteristics, selecting the main speech mode and updating the noise spectrum.

Benefits of technology

It improves the accuracy of speech recognition, enhances the adaptability of spectrum subtraction in complex noise environments outdoors, and reduces the interference of noise time-degeneration on speech feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299470A_ABST
    Figure CN120299470A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech enhancement, in particular to an audio optimization method and system applied to speech recognition, and the method comprises the steps: collecting audio data in real time, and uniformly dividing the audio data into audio frames; for each audio frame and each neighbor frame of the preset audio frame, evaluating whether a noise spectrum is re-selected when the audio data is enhanced by adopting a spectral subtraction method; if reselection is carried out, each mode of the audio frame is acquired; selecting a noise mode from the modes of the audio frame, and obtaining the number of lagging frames of other modes; obtaining a frequency spectrum change characteristic value and a voice characteristic value of each mode, and selecting a main voice mode; acquiring audio mode characteristic values of other modes except the main voice mode; and a noise frame and a new noise spectrum are obtained. The invention aims to improve the accuracy of voice recognition by improving the accuracy of noise frame selection and enhancing the audio enhancement effect of the spectral subtraction for the voice features of the audio frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech enhancement technology, and particularly to an audio optimization method and system applied to speech recognition. Background Art

[0002] Speech recognition is a technology that converts speech signals into text. In speech recognition, noise signals will interfere with speech signals and reduce the recognition accuracy. As a commonly used audio optimization method, in speech recognition, the spectral subtraction method uses technologies such as Fourier transform to remove background noise through frequency domain processing, retain the speech phase information, and thus enhance the speech signal. The principle of spectral subtraction is simple and easy to implement, has a wide range of applications, and causes little damage to speech. It can effectively improve the signal-to-noise ratio of audio data, provide a clearer input signal for the speech recognition system, reduce noise interference, and thus improve the accuracy and reliability of speech recognition.

[0003] In practical applications, the selection of the noise spectrum directly affects the optimization effect of the spectral subtraction method on audio data. In audio data in outdoor environments, there are many outdoor noise sources, resulting in strong time-variability of outdoor noise signals. When using the spectral subtraction method with a fixed noise spectrum to enhance audio data in outdoor environments, the spectral subtraction method has poor adaptability to complex outdoor noise signals. When the characteristics of the noise signal change, the fixed noise spectrum directly causes useful speech signals to be wrongly filtered out, resulting in a decrease in the accuracy of speech recognition. Summary of the Invention

[0004] In view of the above, it is necessary to provide an audio optimization method and system applied to speech recognition. Compared with traditional audio optimization methods and systems applied to speech recognition, by improving the accuracy of noise frame selection, enhancing the audio enhancement effect of the spectral subtraction method for the speech features of audio frames, and further improving the accuracy of speech recognition: In a first aspect, an embodiment of this application provides an audio optimization method applied to speech recognition. The method includes the following steps: Real-time collect audio data and evenly divide it into audio frames of a preset length; use the spectral subtraction method to enhance the audio data; For each audio frame, preset each neighboring frame and each comparison frame of the audio frame before the audio frame, and evaluate whether to reselect the noise spectrum when enhancing the audio data through the degree of change in the energy of all neighboring frames and all comparison frames; If reselected, obtain each modality of the audio frame; obtain the modality energy change value of each modality through the degree of change in the energy of the same modality of all neighboring frames, and select the noise modality from the modalities of the audio frame; arrange the energy of the same modality of all neighboring frames in time sequence to form each energy sequence, and obtain the lag frames of the remaining modalities by analyzing the cross-correlation between the energy sequences of the noise modality and the remaining modalities of the audio frame; Obtain the spectral change eigenvalue of each modality based on the degree of difference in the marginal spectra of the same modality between the audio frame and its neighboring frames; obtain the voice eigenvalue of each modality by combining the proportion of the number of lag frames of each modality in the neighboring frames with the spectral change eigenvalue, and select the main voice modality from the modalities of the audio frame; obtain the audio modality eigenvalue of each of the other modalities based on the similarity of the Mel Frequency Cepstral Coefficients (MFCCs) between the main voice modality and each of the other modalities; Obtain the dispersion degree of the audio modality eigenvalues of all other modalities of each audio frame, and obtain the noise frame and the new noise spectrum based on the dispersion degree of each audio frame and a preset number of neighboring audio frames after it.

[0005] In one embodiment, when evaluating whether to reselect the noise spectrum during audio data enhancement, it includes: Calculate the mean value of the energies of all neighboring frames of each audio frame; Denote the absolute value of the slope of the fitting line of the mean values of all neighboring frames of each audio frame as the first absolute value of each audio frame; Denote the absolute value of the slope of the fitting line of the mean values of all comparison frames of each audio frame as the second absolute value of each audio frame; Extract the upper quartile of the first absolute values of all comparison frames of each audio frame; When the mean value of the first absolute value and the second absolute value is greater than the upper quartile, reselect the noise spectrum; otherwise, do not reselect the noise spectrum.

[0006] In one embodiment, the process of obtaining the modal energy change value of each modality and selecting the noise modality from the modalities of the audio frame includes: Take the absolute value of the slope of the fitting line of the energies of the same modality of all neighboring frames as the modal energy change value of each modality, and take the modality with the largest modal energy change value of the audio frame as the noise modality.

[0007] In one embodiment, the process of obtaining the number of lag frames is as follows: Obtain the cross-correlation sequence of the energy sequences between the noise modality and the other modalities of the audio frame; Take the absolute value of the difference between the sequence number corresponding to the maximum value in the cross-correlation sequence and the central sequence number of the cross-correlation sequence as the number of lag frames of the other modalities.

[0008] In one embodiment, the calculation process of the spectral change eigenvalue is as follows: Normalize each marginal spectrum sequence, calculate the DTW distance between the marginal spectrum sequences of the same modality of any two adjacent audio frames within all neighboring frames, and the spectral change eigenvalue is the average value of the DTW distances between all any two adjacent audio frames within all neighboring frames.

[0009] In one of the embodiments, the method for obtaining the voice feature value is as follows: For each modality, the expression of the voice feature value of the modality is: ; where F represents the voice feature value of the modality; m represents the number of lag frames of the modality; M represents the number of neighboring frames; ε represents a preset positive integer; represents the spectral change feature value of the modality.

[0010] In one of the embodiments, the main voice modality is the modality with the largest voice feature value of the audio frame.

[0011] In one of the embodiments, the method for obtaining the audio modality feature value is as follows: The first 12 coefficients of the Mel-frequency cepstral coefficients are formed into a voice feature vector, and the audio modality feature value is the normalized value of the cosine similarity of the voice feature vectors between the main voice modality and other modalities.

[0012] In one of the embodiments, the method for obtaining the noise frame and the new noise spectrum is as follows: The audio frame with the smallest dispersion among each audio frame and a preset number of neighboring audio frames after it is used as the noise frame, and the spectrum of the noise frame is used as the new noise spectrum.

[0013] In a second aspect, an audio optimization system for speech recognition provided by an embodiment of the present application includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the audio optimization method for speech recognition described in any one of the above are implemented.

[0014] The present application has at least the following beneficial effects: The present application utilizes the lag characteristic of the voice signal in energy change and its difference characteristic from the noise signal in spectral energy distribution, adopts a time-frequency analysis method to extract voice features, reduces the interference of the time-varying outdoor noise on the voice feature extraction, improves the accuracy of the voice signal feature extraction, and further improves the accuracy of the subsequent selection of noise frames, enhancing the adaptability to audio enhancement using the spectral subtraction method in a complex outdoor noise environment; Furthermore, according to the characteristic that the noise signal and the voice signal in the outdoor environmental audio data are independent of each other, different modalities of the alternative noise frames are estimated, and the distribution differences of different modality features in the voice frames and non-voice frames are utilized to avoid the influence of the time-varying characteristics of the outdoor noise on the selection of noise frames, improve the accuracy of the noise frame selection, enhance the audio enhancement effect of the spectral subtraction method on the voice features of the audio frames, so as to improve the accuracy of speech recognition. Description of the Drawings

[0015] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] Figure 1 It is a flowchart of the steps of an audio optimization method for speech recognition provided in an embodiment of the present application; Figure 2 It is a schematic flowchart for evaluating whether to reselect the noise spectrum; Figure 3 It is a schematic flowchart for obtaining noise frames. Specific Embodiments

[0017] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example", etc. are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, the use of words such as "exemplary", "or", "for example" is intended to present relevant concepts in a specific manner.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. It should be understood that unless otherwise stated in this application, " / " means "or".

[0019] In addition, it should be noted that the terms "first" and "second" in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0020] The following specifically describes the specific solutions of the audio optimization method and system for speech recognition provided by this application in conjunction with the drawings.

[0021] Please refer to Figure 1 , which shows a flowchart of the steps of an audio optimization method for speech recognition provided in an embodiment of the present application. The method includes the following steps: Step 1, collect audio data in real time and evenly divide it into audio frames of a preset length.

[0022] In this embodiment, the sampling rate of the audio data is 22.05 kHz. For the convenience of data processing, the audio data is framed, and the frame length of each audio frame is set to 25 milliseconds. To minimize the distortion rate as much as possible, this embodiment sets that there is an overlap between two adjacent audio frames, where the frame length of the overlapping part is 1 / 4 of the frame length of the audio frame. The sampling rate, the frame length of the audio frame, and the frame length of the overlapping part are preset manually, and the implementer can set them by himself / herself, and this application does not make special restrictions.

[0023] Meanwhile, to avoid spectrum leakage and information loss caused by the framing process of the audio data, this application performs windowing processing on each audio frame respectively.

[0024] In this embodiment, the Hamming window function is used to perform windowing processing on each audio frame respectively.

[0025] Step 2, for each audio frame, several neighboring frames and several comparison frames of the audio frame are preset before the audio frame, and whether to reselect the noise spectrum when enhancing the audio data is evaluated through the degree of change in the energy of all neighboring frames and all comparison frames.

[0026] The energy of the speech signal is relatively fixed, and the energy of the outdoor noise signal changes randomly. When the energy of the noise signal changes, the fixed noise spectrum directly causes the useful speech signal to be filtered out incorrectly, resulting in inaccurate speech recognition.

[0027] For each audio frame, N - 1 audio frames are taken forward, and the N audio frames including itself are denoted as the comparison frames of each audio frame, and the energy of each comparison frame is obtained. When the number of audio frames before each audio frame is less than N - 1, no subsequent calculation is performed.

[0028] The change of the outdoor noise signal directly causes the change of its energy. Considering the instantaneous change and the complex situation of the steady change of the energy of the outdoor noise signal, it is judged whether each audio frame is in the enhanced state of the time-varying nature of the noise signal.

[0029] Considering the strong time-varying characteristics of the speech signal in a short time, the energy of a single audio frame changes quickly. Therefore, taking M - 1 audio frames forward with each audio frame as the end frame, the M audio frames including each audio frame are denoted as the neighboring frames of each audio frame. The mean value of the energy of all neighboring frames of each audio frame is calculated.

[0030] The fitting straight line of the mean values of all neighboring frames of each audio frame is obtained, and the absolute value of the slope of the fitting straight line is denoted as the first absolute value, which is used to reflect the degree of instantaneous energy change near each audio frame. Among them, the calculation of the slope is a well-known technology, and this application will not elaborate further.

[0031] Obtain the fitting line of the means of all comparison frames of each audio frame, and denote the absolute value of the slope of the fitting line as the second absolute value, which is used to reflect the degree of smooth energy change near each audio frame.

[0032] In this embodiment, to simultaneously analyze the degree of instantaneous energy change and the degree of smooth energy change near each audio frame, set the value of N to 120 and the value of M to 10. When the value of N is at least 10 times the value of M, the values of N and M can be preset manually, and the implementer can set them according to the actual situation.

[0033] In this embodiment, the least squares method is used to obtain the fitting line. The least squares method is a well-known technology and will not be elaborated in this application. As other implementation manners, on the basis of being able to obtain the fitting line of the means of all neighboring frames of each audio frame and the fitting line of the means of all comparison frames of each audio frame, the implementer can adopt other existing technologies, such as the weighted least squares method, etc., and this application does not make special restrictions.

[0034] Extract the upper quartile of the first absolute value of all comparison frames of each audio frame, and denote it as the energy change threshold. When the mean of the first absolute value and the second absolute value of each audio frame is greater than the energy change threshold, it indicates that the time-varying property of the noise signal is enhanced. When using spectral subtraction to enhance audio data, reselect the noise spectrum; otherwise, do not reselect the noise spectrum. The flow schematic diagram for evaluating whether to reselect the noise spectrum is as Figure 2 shown.

[0035] Step 3, if reselected, obtain each modality of the audio frame; obtain the modality energy change value of each modality through the degree of change of the energy of the same modality of all neighboring frames, and select the noise modality from the modalities of the audio frame; arrange the energies of the same modality of all neighboring frames in time sequence to form each energy sequence, and obtain the lag frames of the remaining modalities by analyzing the cross-correlation of the energy sequences between the noise modality and the remaining modalities of the audio frame.

[0036] For each audio frame, take N - 1 more audio frames backward, and denote the N audio frames including itself as the alternative noise frames of each audio frame. Considering that within a short period of time, the number of noise sources and their vibration modes change relatively fixed, obtain each modality of the alternative noise frames and calculate the energy of each modality. At the same time, obtain the marginal spectrum sequence of each modality, which is used to reflect the spectral energy distribution of each modality. The calculation process of the marginal spectrum sequence is a well-known technology and will not be elaborated in this application. When the number of audio frames after each audio frame is less than N - 1, no subsequent calculation is performed.

[0037] In this embodiment, the empirical mode decomposition method is used to obtain each mode of each alternative noise frame. The ensemble empirical mode decomposition algorithm is a well-known technology and will not be elaborated in this application. As other implementation manners, on the basis of being able to obtain each mode of each alternative noise frame, implementers can adopt other existing technologies, such as the ensemble empirical mode decomposition algorithm, the complementary ensemble empirical mode decomposition method, etc. This application does not make special restrictions.

[0038] In an outdoor environment, compared with the energy change of the noise signal, the energy change of the speech signal has a certain lag. For example, during an outdoor report, after the energy of the noise signal in the environment increases, the speaker will increase their speaking volume, and the energy of the speech signal will also increase accordingly. However, compared with the environmental noise, its energy change lags behind.

[0039] At the same time, as the pitch and volume increase, the characteristics of the speaker's speech signal also change slightly. And when the speaker adapts to the new noise signal environment, it takes a certain amount of time, and during this period, the energy change of the speech signal will continue to lag behind the energy change of the noise signal.

[0040] To extract the lag characteristics of the speech signal in the audio data, for each alternative noise frame, taking the M - 1 audio frames forward with each alternative noise frame as the end frame, the M audio frames including each alternative noise frame are recorded as the corresponding frames of each alternative noise frame.

[0041] Obtain the fitting line of the energy of any same mode of all corresponding frames of each alternative noise frame, and take the absolute value of the slope of the fitting line as the modal energy change value of the any same mode. Among them, when using the mode decomposition algorithm to obtain each mode of each alternative noise frame, the modes are numbered according to the acquisition order, and the modes with the same number of any two alternative noise frames are used as the same mode.

[0042] In this embodiment, the least squares method is used to obtain the fitting line. The least squares method is a well-known technology and will not be elaborated in this application. As other implementation manners, on the basis of being able to obtain the fitting line of the energy of any same mode of all corresponding frames of each alternative noise frame, implementers can adopt other existing technologies, such as the weighted least squares method, etc. This application does not make special restrictions.

[0043] For each alternative noise frame, the mode with the largest modal energy change value of the alternative noise frame is used as the noise mode, which is used to represent the vibration mode of the noise source where the main energy change occurs. In addition, affected by the resonance of the noise source system, the energy changes of the remaining vibration modes of the same noise source also have a lag, but the lag is weaker than that of the speech signal.

[0044] Arrange the energies of the same modality of all reference frames of each alternative noise source in sequence according to time series to form an energy sequence for each modality; obtain the cross-correlation sequence of the energy sequences between the noise modality and the other modalities of each alternative noise source, and take the absolute value of the difference between the serial number corresponding to the maximum value in the cross-correlation sequence and the central serial number of the cross-correlation sequence as the lag frames of the other modalities of each alternative noise source, which is used to reflect the lag degree of the energy change of the other modalities compared with the noise modality. If the lag frames of a certain modality are larger, the lag of the energy change of the certain modality is stronger, and the probability of containing a speech signal in the certain modality is greater.

[0045] Step 4: Obtain the spectral change characteristic values of each modality through the difference degree of the marginal spectra of the same modality between an audio frame and its neighboring frames; obtain the speech characteristic values of each modality by combining the spectral change characteristic values with the proportion of the lag frames of each modality in the neighboring frames, and select the main speech modality from the modalities of the audio frame; obtain the audio modality characteristic values of the other modalities through the similarity of the Mel Frequency Cepstral Coefficients between the main speech modality and the other modalities.

[0046] The phonemes of the speech signal and the frequency distribution of the outdoor noise signal both have a certain degree of time variability. However, the time-varying characteristics of the speech signal are concentrated in a specific frequency range, and the energy change degrees of different frequencies are relatively large, with strong time variability at a small scale; while the frequency change range of the outdoor noise signal is relatively large, but the energy change degrees of different frequencies are relatively small, making it possible to be treated as a stationary signal in a short time.

[0047] Normalize the marginal spectrum sequence to control the calculation range. For all reference frames of each alternative noise frame, calculate the DTW (Dynamic Time Warping) distance between the marginal spectrum sequences of the same modality of any two adjacent audio frames within all reference frames, and take the average value of the DTW distances between all any two adjacent audio frames within all reference frames as the spectral change characteristic value of each modality of each alternative noise frame, which is used to reflect the change degree of the spectral energy of each modality in all reference frames of the alternative noise frame. The larger the spectral change characteristic value, the stronger the time variability of the short-time spectral energy, and the greater the probability of containing a speech signal in the modality.

[0048] In this embodiment, the Min-Max normalization method is used to normalize the marginal spectrum sequence.

[0049] Furthermore, obtain the speech characteristic values of each modality by combining the proportion of the lag frames of each modality of each alternative noise frame in the reference frame with the spectral change characteristic value of each modality, which is used to reflect the probability of containing a speech signal in each modality of the alternative noise frame. For each modality, the expression of the speech characteristic value of the modality is: ; where \(F\) represents the voice feature value of the mode; \(m\) represents the number of lag frames of the mode; \(M\) represents the number of neighboring frames; \(\varepsilon\) represents a preset positive integer used to avoid affecting the calculation result of the voice feature value when the numerator is 0. In this embodiment, the value of \(\varepsilon\) is 1, and the value of \(\varepsilon\) can be preset manually, and this application does not make special restrictions; represents the spectral change feature value of the mode.

[0050] Furthermore, the mode with the largest voice feature value among the alternative noise frames is used as the main voice mode, representing the mode with the highest probability of containing a voice signal.

[0051] Step 5: Obtain the dispersion of the audio mode feature values of all other modes of each audio frame. Through the dispersion of each audio frame and a preset number of neighboring audio frames after it, obtain the noise frames and new noise spectra.

[0052] Considering the inherent voiceprint characteristics of the voice signal, this application calculates the Mel-frequency cepstral coefficients of each mode of each alternative noise frame. The Mel-frequency cepstral coefficient is a vector containing 13 coefficients, and the first 12 coefficients are used to form the voice feature vector of each mode. The Mel-frequency cepstral coefficient is a well-known technology, and this application will not elaborate.

[0053] Obtain the audio mode feature values of all other modes through the similarity of the voice feature vectors between the main voice mode and all other modes of each alternative noise frame, which is used to reflect the difference degree of the voice features between all other modes and the main voice mode; for each alternative noise frame, the expression of the audio mode feature value of all other modes of the alternative noise frame is: ; where represents the audio mode feature value of the \(k\)th mode of the alternative noise frame; \(norm[\ ]\) represents the arctangent normalization function; \(sim(\ )\) represents the cosine similarity function; 、 respectively represent the voice feature vectors of the main voice mode and the \(k\)th mode of the alternative noise frame.

[0054] Furthermore, calculate the dispersion of the audio mode feature values of all other modes of each alternative noise frame, which is used to reflect the distribution of voice features within different modes. The less voice components contained in the alternative noise frame, the closer the distribution of voice features between each mode, the smaller the calculated dispersion, and the more accurately it can reflect the noise situation in the time period where each alternative noise frame is located.

[0055] In this embodiment, the dispersion of the audio mode feature value is the standard deviation. As other implementation manners, on the basis of being able to measure the uneven degree of the distribution of the audio mode feature value, the implementer can adopt other existing technologies, such as variance, coefficient of variation, etc., and this application does not make special restrictions.

[0056] Further, the alternative noise frame with the smallest dispersion among all the alternative noise frames of each audio frame is used as the noise frame, and the spectrum of the noise frame is used as the new noise spectrum. The schematic diagram of the acquisition process of the noise frame is as shown in Figure 3 shown.

[0057] The audio signal is enhanced by using the new noise spectrum through spectral subtraction to complete audio optimization. It should be noted that when there is no need to reselect the noise spectrum, the audio signal is directly enhanced by spectral subtraction. Among them, spectral subtraction is a well-known existing technology and will not be elaborated in this application.

[0058] Based on the same inventive concept as the above method, the embodiment of the present application also provides an audio optimization system applied to speech recognition, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods for audio optimization applied to speech recognition.

[0059] In summary, the present application utilizes the lag characteristic of the speech signal in energy change and its difference characteristic from the noise signal in spectral energy distribution, adopts the time-frequency analysis method to extract speech features, reduces the interference of the time-varying outdoor noise on the extraction of speech features, improves the accuracy of speech signal feature extraction, and further improves the accuracy of subsequent selection of noise frames, enhancing the adaptability of spectral subtraction for audio enhancement in complex outdoor noise environments; Further, according to the characteristic that the noise signal and the speech signal in the outdoor environmental audio data are independent of each other, different modes of the alternative noise frames are estimated, and the distribution differences of different mode features in the speech frames and non-speech frames are utilized to avoid the influence of the time-varying features of outdoor noise on the selection of noise frames, improve the accuracy of noise frame selection, enhance the audio enhancement effect of spectral subtraction on the speech features of audio frames, so as to improve the accuracy of speech recognition.

[0060] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. In the description corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur out of the order disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0061] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and without departing from the basic characteristics of the present application, the present application can be implemented in other specific forms. Therefore, from any point of view, the above-described embodiments of the present application should be regarded as exemplary and non-limiting.

Claims

1. An audio optimization method applied to speech recognition, characterized in that, The method includes the following steps: Collect audio data in real time and evenly divide it into audio frames of a preset length; use spectral subtraction to enhance the audio data; For each audio frame, preset each neighboring frame and each comparison frame of the audio frame before the audio frame, and evaluate whether to re-select the noise spectrum when enhancing the audio data through the degree of change in the energy of all neighboring frames and all comparison frames; If re-selection is required, obtain each modality of the audio frame; obtain the modality energy change value of each modality through the degree of change in the energy of the same modality of all neighboring frames, and select the noise modality from the modalities of the audio frame; arrange the energy of the same modality of all neighboring frames in time sequence to form each energy sequence, and obtain the lag frames of the remaining modalities by analyzing the cross-correlation of the energy sequences between the noise modality and the remaining modalities of the audio frame; Obtain the spectral change eigenvalue of each modality through the degree of difference in the marginal spectra of the same modality between the audio frame and its neighboring frames; obtain the speech eigenvalue of each modality by combining the proportion of the lag frames of each modality in the neighboring frames and the spectral change eigenvalue, and select the main speech modality from the modalities of the audio frame; obtain the audio modality eigenvalue of the other modalities by the similarity of the Mel-frequency cepstral coefficients between the main speech modality and the other modalities; Obtain the dispersion of the audio modality eigenvalues of all other modalities of each audio frame, and obtain the noise frames and the new noise spectrum through the dispersion of each audio frame and the preset number of neighboring audio frames after it.

2. The audio optimization method applied to speech recognition according to claim 1, characterized in that, When evaluating whether to re-select the noise spectrum when enhancing the audio data, it includes: Calculate the mean value of the energy of all neighboring frames of each audio frame; Denote the absolute value of the slope of the fitting line of the mean value of all neighboring frames of each audio frame as the first absolute value of each audio frame; Denote the absolute value of the slope of the fitting line of the mean value of all comparison frames of each audio frame as the second absolute value of each audio frame; Extract the upper quartile of the first absolute value of all comparison frames of each audio frame; When the mean value of the first absolute value and the second absolute value is greater than the upper quartile, re-select the noise spectrum, otherwise, do not re-select the noise spectrum.

3. The audio optimization method applied to speech recognition according to claim 1, wherein, The process of obtaining the modality energy change value of each modality and selecting the noise modality from the modalities of the audio frame includes: Take the absolute value of the slope of the fitting line of the energy of the same modality of all neighboring frames as the modality energy change value of each modality, and take the modality with the largest modality energy change value of the audio frame as the noise modality.

4. The audio optimization method applied to speech recognition according to claim 1, wherein The process of obtaining the lag frames is: Obtain the cross-correlation sequence of the energy sequences between the noise modality and the remaining modalities of the audio frame; Take the absolute value of the difference between the serial number corresponding to the maximum value in the cross-correlation sequence and the central serial number of the cross-correlation sequence as the lag frames of the remaining modalities.

5. The audio optimization method applied to speech recognition according to claim 1, characterized in that, The calculation process of the spectral change eigenvalue is: Normalize each marginal spectrum sequence, calculate the DTW distance between the marginal spectrum sequences of the same modality of any two adjacent audio frames within all neighboring frames, and the spectral change eigenvalue is the average value of the DTW distances between all any two adjacent audio frames within all neighboring frames.

6. The audio optimization method applied to speech recognition according to claim 1, wherein The method for obtaining the speech eigenvalue is: For each modality, the expression of the voice feature value of the modality is as follows: ; where F represents the voice feature value of the mode; m represents the number of lag frames of the mode; M represents the number of neighboring frames; ε represents a preset positive integer; represents the spectral change feature value of the mode.

7. The audio optimization method applied to speech recognition according to claim 1, characterized in that The main voice modality is the modality with the largest voice feature value of the audio frame.

8. The audio optimization method applied to speech recognition according to claim 1, characterized in that, The method for obtaining the audio modality feature value is as follows: The first 12 coefficients of the Mel-frequency cepstral coefficients are used to form a voice feature vector, and the audio modality feature value is the normalized value of the cosine similarity of the voice feature vector between the main voice modality and other modalities.

9. The audio optimization method applied to speech recognition according to claim 1, characterized in that The method for obtaining the noise frame and the new noise spectrum is as follows: The audio frame with the smallest dispersion among each audio frame and a preset number of adjacent audio frames after it is used as the noise frame, and the spectrum of the noise frame is used as the new noise spectrum.

10. An audio optimization system applied to speech recognition, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the audio optimization method for speech recognition according to any one of claims 1-9.

Citation Information

Patent Citations

  • Quick estimation method for abrupt change noise

    CN111933165A

  • Speech signal denoising method based on improved variational mode decomposition and principal component analysis

    CN113851144A

  • Method for enhancing field bird buzzing audio data

    CN118173106A

  • Data enhancement optimization method in bird chirp recognition in complex field environment

    CN118212928A

  • Method for spectral subtraction in speech enhancement

    US20050071156A1

Cited By

  • Large model-based external call center voice recognition method and system

    CN121811863A

  • A Large-Model-Based Speech Recognition Method and System for Outbound Call Centers

    CN121811863B