Equalizer-based speech recognition methods, speech translation devices, storage media, and computer program products
By constructing an audio sample library for multi-dimensional acoustic scenarios and performing equalizer spectrum shaping, and selecting the optimal parameter set, the problem of insufficient translation quality in complex scenarios for speech recognition is solved, and the speech recognition engine achieves high accuracy and robustness in different environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN TIMEKETTLE TECH CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, speech recognition has limited effect on improving translation quality in complex scenarios because the front-end input speech signal processing cannot dynamically adapt to complex and ever-changing noise environments and device frequency response characteristics.
By constructing an audio sample library covering multiple acoustic scenarios, setting multiple candidate parameter sets, using an equalizer to perform spectral shaping on the audio signal, and selecting the parameter set that meets the recognition criteria as the adjustment parameter set, the speech recognition effect is optimized.
It significantly improves the recognition accuracy and robustness of the speech recognition engine in different or specific acoustic environments, and optimizes the spectral characteristics of the input signal.
Smart Images

Figure CN122493848A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of AI translation, and more specifically, to a speech recognition method, speech translation device, storage medium, and computer program product based on equalizer adjustment. Background Technology
[0002] In AI translation, the accuracy of speech recognition directly determines the final translation quality. Related technologies improve speech recognition accuracy through iterative optimization of backend algorithm models. However, the processing of the front-end input speech signal often follows fixed preprocessing procedures (e.g., general noise reduction, fixed gain adjustment), making it impossible to dynamically adapt to complex scenarios (e.g., complex and variable noise environments, device frequency response characteristics, etc.). Therefore, the effect of improving translation quality in complex scenarios is limited. Summary of the Invention
[0003] In view of the above problems, this application proposes a speech recognition method, speech translation device, storage medium and computer program product based on equalizer adjustment, which can effectively improve the translation quality.
[0004] In a first aspect, embodiments of this application provide a speech recognition method based on equalizer adjustment. The method includes: acquiring an original audio signal; performing spectral shaping processing on the original audio signal using an equalizer according to an adjustment parameter set to generate a target audio signal; wherein the adjustment parameter set is determined for a multi-dimensional acoustic scene or the current acoustic scene; and inputting the target audio signal into a speech recognition engine for recognition to obtain target text information.
[0005] Furthermore, the method also includes: constructing an audio sample library covering multi-dimensional acoustic scenarios; setting multiple candidate parameter sets for an equalizer applicable to multi-dimensional acoustic scenarios; for each candidate parameter set in the multiple candidate parameter sets, performing spectral shaping on each sample audio signal in the audio sample library according to the candidate parameter set using the equalizer to generate a set of test audio signals; evaluating the recognition index corresponding to the candidate parameter set based on the recognition result of the speech recognition engine on the current set of test audio signals; after obtaining the recognition index corresponding to multiple candidate parameter sets, selecting the candidate parameter set whose recognition index meets the value conditions from the multiple candidate parameter sets as the adjustment parameter set.
[0006] Furthermore, the multi-dimensional acoustic scene includes multiple pre-selected acoustic scenes; multiple candidate parameter sets of equalizers are set for use in the multi-dimensional acoustic scene, including: setting multiple candidate parameter sets by combining the equalizer parameter setting strategies corresponding to various pre-selected acoustic scenes in the multi-dimensional acoustic scene.
[0007] Furthermore, the pre-selected acoustic scenarios include a first scenario, a second scenario, a third scenario, and a fourth scenario. In the first scenario, the speech is mixed with background noise, and the speech volume level is the first volume level. In the second scenario, the speech is clean, and the speech volume level is the first volume level. In the third scenario, the speech is mixed with background noise, and the speech volume level is the second volume level. In the fourth scenario, the speech is clean, and the speech volume level is the second volume level. The second volume level is higher than the first volume level. The equalizer parameter setting strategy for the first scenario is: to perform attenuation processing on the first frequency band of each sample audio signal within a first attenuation amplitude range, and to perform attenuation processing on the first frequency band of each sample audio signal within a first attenuation amplitude range. The equalizer parameter setting strategy for the second scenario is to perform gain processing on the second frequency band within the first gain amplitude range for each sample audio signal; the equalizer parameter setting strategy for the third scenario is to perform attenuation processing on the first frequency band within the second attenuation amplitude range for each sample audio signal, and to perform gain processing on the third frequency band within the third gain amplitude range for each sample audio signal; the equalizer parameters for the fourth scenario are: to perform attenuation processing on the fourth frequency band within the third attenuation amplitude range for each sample audio signal, and to perform gain processing on the fifth frequency band within the fourth gain amplitude range for each sample audio signal. Among them, the first frequency band is the noise frequency band for identification, the second frequency band contains speech consonant information and speech formants, the minimum frequency value of the third frequency band is greater than the maximum frequency value of the first frequency band, the maximum frequency value of the fourth frequency band is less than the minimum frequency value of the fifth frequency band, the minimum amplitude value of the first gain amplitude range is greater than the maximum amplitude value of the second gain amplitude range, the minimum amplitude value of the second gain amplitude range is greater than the maximum amplitude value of the third gain amplitude range and is also greater than the maximum amplitude value of the fourth gain amplitude range, the third gain amplitude range and the fourth gain amplitude range are the same or different, the first attenuation amplitude range and the second attenuation amplitude range are the same or different, and the maximum amplitude value of the third attenuation amplitude range is less than the minimum amplitude value of the first attenuation amplitude range and is also less than the minimum amplitude value of the second attenuation amplitude range.
[0008] Furthermore, if the recognition metric includes average recognition accuracy, then the value condition includes the highest average recognition accuracy; if the recognition metric includes average word error rate, then the value condition includes the lowest average word error rate.
[0009] Furthermore, the method also includes: determining the scene category to which the current acoustic scene belongs; and determining the preset parameter set corresponding to the scene category as the adjustment parameter set.
[0010] Furthermore, the adjustment parameter set is solidified as configuration data and stored in the digital signal processing chip of the audio acquisition hardware, or in the audio preprocessing module of the software application; the original audio signal is spectrally shaped by the equalizer according to the adjustment parameter set to generate the target audio signal, including: calling the solidified adjustment parameter set from the digital signal processing chip or the audio preprocessing module; and equalizing and filtering the original audio signal by the equalizer according to the adjustment parameter set to generate the target audio signal.
[0011] Secondly, embodiments of this application also provide a speech translation device, which includes: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the equalizer-based speech recognition method described above.
[0012] Thirdly, embodiments of this application also provide a computer-readable storage medium storing program code, which can be called by a processor to execute the above-described equalizer-based speech recognition method.
[0013] Fourthly, embodiments of this application also provide a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the above-described equalizer-based speech recognition method.
[0014] The technical solution provided in this application includes the following method: acquiring an original audio signal; performing spectral shaping on the original audio signal using an equalizer according to an adjustment parameter set to generate a target audio signal; wherein the adjustment parameter set is determined for a multi-dimensional acoustic scene or the current acoustic scene; and inputting the target audio signal into a speech recognition engine for recognition to obtain target text information. Therefore, by utilizing an equalizer adjustment parameter set determined for a multi-dimensional acoustic scene or the current acoustic scene to adaptively shape the original audio signal's spectrum, the spectral characteristics of the input signal are optimized, thereby significantly improving the recognition accuracy and robustness of the speech recognition engine in different or specific acoustic environments. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0016] Figure 1 The diagram shows a flowchart of a speech recognition method based on equalizer adjustment provided in an embodiment of this application.
[0017] Figure 2 A schematic diagram of the structure of a speech recognition device based on equalizer adjustment provided in an embodiment of this application is shown.
[0018] Figure 3 A schematic diagram of the structure of a speech translation device provided in an embodiment of this application is shown.
[0019] Figure 4 This illustration shows a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of this application.
[0020] Figure 5 A schematic diagram of the structure of a computer program product provided in an embodiment of this application is shown. Detailed Implementation
[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0023] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0024] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0025] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0026] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0027] Please see Figure 1 , Figure 1 This diagram illustrates a flowchart of a speech recognition method based on equalizer adjustment provided in an embodiment of this application. Figure 1 As shown, the method may include steps 110 to 130.
[0028] In step 110, the original audio signal is acquired.
[0029] In some implementations, the raw audio signal to be processed is acquired via a Bluetooth headset system. The Bluetooth headset system may include a headset terminal for wearing on a user's ear (as an audio acquisition front end) and a processing terminal (e.g., a smartphone, a dedicated translator, or a cloud server) that is wirelessly connected to the headset via the Bluetooth protocol.
[0030] In one specific implementation, the headset terminal integrates at least one high-sensitivity microelectromechanical system (MEMS) microphone. When the user activates translation mode, the microphone senses sound wave vibrations in the air in real time and converts them into analog electrical signals. The audio codec or digital signal processor inside the headset terminal samples and quantizes the analog electrical signals, converting them into raw audio signals in digital format, thus acquiring the raw audio signal to be processed. In other words, the raw audio signal to be processed has not yet been optimized by an equalizer for a specific scenario, and it may contain ambient noise, wind noise, or frequency response loss due to distance.
[0031] After the headphone terminal completes the acquisition of the raw audio signal, it transmits the raw audio signal to be processed to the processing terminal in real time via a Bluetooth link (e.g., Bluetooth Classic A2DP or BLE Audio / LC3 encoding).
[0032] In step 120, the original audio signal is spectrally shaped by an equalizer according to the set of adjustment parameters to generate the target audio signal.
[0033] The set of adjustment parameters is determined for multi-dimensional acoustic scenarios or the current acoustic scenario.
[0034] In some implementations, the adjustment parameter set can be a pre-configured combination of parameters used to control the equalizer (EQ) to perform spectral shaping on the raw audio signal. The adjustment parameter set is not a single fixed value, but rather a library or preset set of parameters established for multi-dimensional acoustic scenarios common in speech recognition tasks (e.g., combinations of different volume levels with different noise environments) or the current acoustic scenario.
[0035] In some implementations, the set of adjustment parameters may include a target frequency band, adjustment amplitude, and / or quality factor. The target frequency band may be a specified range of audio frequencies to be processed; the adjustment amplitude may be a specified gain or attenuation intensity for the target frequency band; and the quality factor may be a specified filter bandwidth.
[0036] In some implementations, an acoustic scene can refer to the specific acoustic environment in which the voice acquisition device is located when acquiring signals. An acoustic scene is not just a physical spatial location, but a multidimensional acoustic state that affects the performance of voice recognition, composed of environmental noise characteristics and the intensity of the target voice signal.
[0037] In the embodiments of this application, the acoustic scene is the basis for determining the set of adjustment parameters, which can characterize the spectral defects present in the original audio signal. For example, low signal-to-noise ratio, uneven frequency response, or signal overload.
[0038] Having clarified the definition of the adjustment parameter set and the acoustic scenario, those skilled in the art will understand that, in order to ensure that the aforementioned adjustment parameter set can maximize the performance of the speech recognition engine, these parameters should not be arbitrarily set based solely on experience, but should be quantitatively selected based on actual recognition results. Therefore, in some embodiments, this speech recognition method based on equalizer adjustment further includes steps 120-1 to 120-5: In step 120-1, an audio sample library covering multi-dimensional acoustic scenarios is constructed; In step 120-2, the equalizer is set to be applicable to multiple candidate parameter sets for multi-dimensional acoustic scenarios; In step 120-3, for each candidate parameter set in multiple candidate parameter sets, the equalizer performs spectral shaping on each sample audio signal in the audio sample library according to the candidate parameter set to generate a set of test audio signals. In step 120-4, the recognition metrics corresponding to the candidate parameter set are evaluated based on the recognition results of the speech recognition engine on the current group of test audio signals. In step S120-5, after obtaining the identification indicators corresponding to multiple candidate parameter sets, the candidate parameter set whose identification indicators meet the value conditions is selected from the multiple candidate parameter sets as the adjustment parameter set.
[0039] The audio sample library is the foundation for parameter selection, and its construction process fully considers various complex situations that may occur in the real world. Specifically, in practice, the audio sample library can include not only clean speech but also impure speech formed by superimposing background noise, thereby achieving coverage of multi-dimensional acoustic scenarios.
[0040] For example, multiple testers are arranged in the same quiet environment, and each tester records a standard voice recording covering common words and with clear pronunciation (e.g., "Please open navigation" from a smart voice assistant) at three quantifiable (e.g., by sound pressure level measurement) volume levels of low, medium, and high, forming a clean voice set.
[0041] Multiple signal-to-noise ratios (SNRs) are pre-set (e.g., 0 dB, 5 dB, 10 dB, and 15 dB, etc.). For each SNR, in the mixing room, each clean speech in the clean speech set is mixed with selected background noise (e.g., traffic noise, office voices, and wind noise, etc.) according to that SNR to form a non-clean speech set, so as to simulate real interference through the non-clean speech set.
[0042] By combining clean and impure speech sets, we can identify clean speech at low volume levels as sample audio signals for the "quiet environment - low volume level" scenario, clean speech at medium volume levels as sample audio signals for the "quiet environment - medium volume level" scenario, clean speech at high volume levels as sample audio signals for the "quiet environment - high volume level" scenario, impure speech at low volume levels mixed with background noise as sample audio signals for the "noisy environment - low volume level" scenario, impure speech at medium volume levels mixed with background noise as sample audio signals for the "noisy environment - medium volume level" scenario, and impure speech at high volume levels mixed with background noise as sample audio signals for the "noisy environment - high volume level" scenario. This results in an audio sample library covering multiple acoustic scenarios, including sample audio signals from various two-dimensional acoustic scenarios: "ambient noise dimension" and "speech volume dimension".
[0043] By constructing an audio sample library covering multiple acoustic scenarios, we can ensure that the subsequently selected set of adjustment parameters has broad applicability and robustness.
[0044] In multi-dimensional acoustic scenarios, an equalizer is set to apply multiple candidate parameter sets to the multi-dimensional acoustic scenario. These multiple candidate parameter sets constitute a parameter pool to be selected.
[0045] In practice, a series of parameter combinations with different center frequencies, gain / attenuation amplitudes, and bandwidths (Q values) can be set according to the equalizer's hardware capabilities. For example, one set of parameters can be preset for low-frequency attenuation, another set for high-frequency boost, or multiple sets of complex parameters combining different frequency bands. These candidate parameter sets cover various spectrum shaping strategies that may improve recognition performance.
[0046] Furthermore, in some implementations, the multi-dimensional acoustic scene includes multiple pre-selected acoustic scenes; the step of setting the equalizer to multiple candidate parameter sets applicable to the multi-dimensional acoustic scene may include: setting multiple candidate parameter sets by combining the equalizer parameter setting strategies corresponding to each type of pre-selected acoustic scene in the multiple pre-selected acoustic scene.
[0047] In practical applications, based on actual business needs, multiple common or typical scenarios can be selected in advance from the above-mentioned multiple two-dimensional acoustic scenarios, and the resulting multiple pre-selected acoustic scenarios can be used as multi-dimensional acoustic scenarios.
[0048] For each of the multiple pre-selected acoustic scenarios, considering the audio signal characteristics under that scenario, an equalizer parameter setting strategy is defined for that scenario. Furthermore, multiple candidate parameter sets are set by combining the equalizer parameter setting strategies for each of the multiple pre-selected acoustic scenarios.
[0049] Specifically, in some implementations, the multiple pre-selected acoustic scenarios include a first scenario, a second scenario, a third scenario, and a fourth scenario. In the first scenario, the speech is mixed with background noise, and the speech volume level is a first volume level; in the second scenario, the speech is clean speech, and the speech volume level is a first volume level; in the third scenario, the speech is mixed with background noise, and the speech volume level is a second volume level; in the fourth scenario, the speech is clean speech, and the speech volume level is a second volume level; the second volume level is higher than the first volume level.
[0050] In one specific implementation, the equalizer parameter setting strategy corresponding to the first scenario is as follows: perform attenuation processing on the first frequency band of each sample audio signal within a first attenuation amplitude range, and perform gain processing on the second frequency band of each sample audio signal within a first gain amplitude range.
[0051] In one specific implementation, the equalizer parameter setting strategy for the second scenario is to perform gain processing on the second frequency band of each sample audio signal within the second gain amplitude range.
[0052] In one specific implementation, the equalizer parameter setting strategy for the third scenario is as follows: the first frequency band in each sample audio information is subjected to attenuation processing within the second attenuation amplitude range, and the third frequency band in each sample audio signal is subjected to gain processing within the third gain amplitude range.
[0053] In one specific implementation, the equalizer parameter setting strategy for the fourth scenario is as follows: the fourth frequency band in each sample audio signal is subjected to attenuation processing within the third attenuation amplitude range, and the fifth frequency band in each sample audio signal is subjected to gain processing within the fourth gain amplitude range.
[0054] Among them, the first frequency band is the noise frequency band for identification, the second frequency band contains speech consonant information and speech formants, the minimum frequency value of the third frequency band is greater than the maximum frequency value of the first frequency band, the maximum frequency value of the fourth frequency band is less than the minimum frequency value of the fifth frequency band, the minimum amplitude value of the first gain amplitude range is greater than the maximum amplitude value of the second gain amplitude range, the minimum amplitude value of the second gain amplitude range is greater than the maximum amplitude value of the third gain amplitude range and is also greater than the maximum amplitude value of the fourth gain amplitude range, the third gain amplitude range and the fourth gain amplitude range are the same or different, the first attenuation amplitude range and the second attenuation amplitude range are the same or different, and the maximum amplitude value of the third attenuation amplitude range is less than the minimum amplitude value of the first attenuation amplitude range and is also less than the minimum amplitude value of the second attenuation amplitude range.
[0055] In practical applications, we can first define the equalizer parameter setting strategy corresponding to a single-dimensional acoustic scene, and then define the equalizer parameter setting strategy corresponding to each of the multiple pre-selected acoustic scenes based on the equalizer parameter setting strategy corresponding to the relevant dimensions of that pre-selected acoustic scene.
[0056] For example, two single dimensions are defined as "ambient noise dimension" and "voice volume dimension". Assuming the first volume level is low and the second volume level is high, the multiple pre-selected acoustic scenarios include the first scenario (i.e. "noisy environment - low volume level" scenario), the second scenario (i.e. "quiet environment - low volume level" scenario), the third scenario (i.e. "noisy environment - high volume level" scenario), and the fourth scenario (i.e. "quiet environment - high volume level" scenario).
[0057] The equalizer parameter setting strategy corresponding to the acoustic scene of "low volume level" is predefined as Strategy 1: perform gain processing on the mid-to-high frequency band (e.g., 1kHz~4kHz) of the audio signal that contains speech consonant information and speech formants within a gain amplitude range (e.g., 3dB~6dB).
[0058] By performing gain processing on the mid-to-high frequency bands that contain speech consonant information and speech formants, the distinguishability of phonemes can be directly improved, effectively enhancing speech clarity.
[0059] The equalizer parameter setting strategy for the acoustic scene of "noisy environment" is predefined as Strategy 2: Attenuate the noise frequency band in the audio signal (e.g., traffic noise is concentrated in 80Hz~300Hz) within a certain attenuation range (e.g. 4dB~8dB).
[0060] Among these methods, spectral analysis can be performed on mixed background noise to determine the frequency bands where its energy is concentrated, thereby identifying the noise frequency bands in the audio signal.
[0061] Alternatively, strategy 1 above can be combined to enhance the mid-to-high frequency band while attenuating noise.
[0062] The equalizer parameter setting strategy corresponding to the acoustic scene of "high volume level" is predefined as Strategy 3: perform attenuation processing on the low frequency band (e.g., below 150Hz) of the audio signal within a range of attenuation amplitude (e.g., 2dB~4dB), and perform gain processing on the mid-high frequency band (e.g., 2kHz~5kHz) of the audio signal within a range of gain amplitude (e.g., 1dB~2dB).
[0063] By attenuating the low-frequency band, we can prevent excessive low frequencies from causing overload or masking effects. By slightly increasing the gain of the mid-to-high frequency band, we can maintain the brightness and clarity of speech.
[0064] Therefore, for the first scenario (i.e., the "noisy environment - low volume level" scenario), according to the above strategy 1 and strategy 2, the corresponding equalizer parameter setting strategy is defined as follows: the first frequency band identified as noise in each sample audio signal is subjected to attenuation processing within the first attenuation amplitude range, and the second frequency band containing speech consonant information and speech formants in each sample audio signal is subjected to gain processing within the first gain amplitude range.
[0065] For the second scenario (i.e., the "quiet environment - low volume level" scenario), according to strategy 1 above, the corresponding equalizer parameter setting strategy is defined as follows: perform gain processing on the second frequency band containing speech consonant information and speech formants in each sample audio signal within the second gain amplitude range.
[0066] For the third scenario (i.e., the "noisy environment - high volume level" scenario), according to the above strategy 2 and strategy 3, the corresponding equalizer parameter setting strategy is defined as follows: the first frequency band identified as noise in each sample audio information is subjected to attenuation processing within the second attenuation amplitude range, and the third frequency band in each sample audio signal, which is mid-high frequency compared to the first frequency band, is subjected to gain processing within the third gain amplitude range.
[0067] For the fourth scenario (i.e., the "quiet environment - high volume level" scenario), according to strategy 3 above, the corresponding equalizer parameter setting strategy is defined as follows: perform attenuation processing on the lower frequency fourth frequency band in each sample audio signal within the third attenuation amplitude range, and perform gain processing on the higher frequency fifth frequency band in each sample audio signal within the fourth gain amplitude range.
[0068] Users set the amplitude range according to the following principles: the minimum amplitude value of the first gain amplitude range is greater than the maximum amplitude value of the second gain amplitude range, the minimum amplitude value of the second gain amplitude range is greater than the maximum amplitude value of the third gain amplitude range and is also greater than the maximum amplitude value of the fourth gain amplitude range, and the third gain amplitude range and the fourth gain amplitude range are the same or different; the first attenuation amplitude range is the same as or different from the second attenuation amplitude range, and the maximum amplitude value of the third attenuation amplitude range is less than the minimum amplitude value of the first attenuation amplitude range and is also less than the minimum amplitude value of the second attenuation amplitude range.
[0069] By allowing users to set the amplitude range according to the above principles, the following can be achieved: in a "noisy environment - low volume level" scenario, significantly (or moderately) attenuate low-frequency noise and significantly enhance mid-to-high frequencies; in a "quiet environment - low volume level" scenario, moderately enhance mid-to-high frequencies; in a "noisy environment - high volume level" scenario, significantly (or moderately) attenuate low-frequency noise and slightly enhance mid-to-high frequencies; and in a "quiet environment - high volume level" scenario, slightly attenuate low frequencies and slightly enhance high frequencies. This can better optimize audio.
[0070] After obtaining the equalizer parameter setting strategies corresponding to various pre-selected acoustic scenarios, multiple candidate parameter sets are set in combination with the equalizer parameter setting strategies corresponding to various pre-selected acoustic scenarios.
[0071] In practical applications, since the selected frequency bands and amplitude ranges may differ in the equalizer parameter setting strategies corresponding to various pre-selected acoustic scenarios, a candidate parameter set can be set by superimposing the frequency bands from each strategy and processing each frequency band with a selected amplitude value. This results in multiple candidate parameter sets. It is understandable that the amplitude values selected for each frequency band in any candidate parameter set are not exactly the same as the amplitude values selected for each frequency band in the other candidate parameter sets.
[0072] For each candidate parameter set in the multiple candidate parameter sets, the system uses it as the current processing configuration and performs spectral shaping on each sample audio signal in the audio sample library according to the candidate parameter set through the equalizer to generate a set of test audio signals.
[0073] This step simulates the front-end processing of a speech recognition device in actual operation. After equalization, the individual sample audio signals in the audio sample library are converted into a set of test audio signals with specific spectral characteristics.
[0074] The system inputs the current group of test audio signals into a standard speech recognition engine, obtains the recognition results output by the speech recognition engine, and evaluates the recognition metrics corresponding to the candidate parameter set based on the recognition results.
[0075] In practice, the speech recognition engine outputs the recognized text of each test audio signal in the current group of test audio signals. The recognized text of each test audio signal is then compared with the real text of the corresponding sample audio signal. Based on the comparison results, the recognition metric corresponding to that candidate parameter set is evaluated. For example, the recognition metric could be the average word error rate, average character error rate, or average recognition accuracy. Through the feedback from machine hearing, the actual contribution of each candidate parameter set to the recognition performance can be objectively quantified.
[0076] After obtaining the identification indicators corresponding to multiple candidate parameter sets, the preset value conditions are determined, and the candidate parameter set whose identification indicators meet the value conditions is selected from the multiple candidate parameter sets as the adjustment parameter set.
[0077] In some implementations, if the recognition metric includes average recognition accuracy, then the value condition includes the highest average recognition accuracy; if the recognition metric includes average word error rate, then the value condition includes the lowest average word error rate.
[0078] For example, if the recognition metric is the average recognition accuracy, the system will select the candidate parameter set with the highest average recognition accuracy from multiple candidate parameter sets as the adjustment parameter set. As another example, if the recognition metric is the average word error rate, the system will select the candidate parameter set with the lowest average word error rate from multiple candidate parameter sets as the adjustment parameter set.
[0079] The set of adjustment parameters represents the optimal solution for multi-dimensional acoustic scenarios. This rigorous selection mechanism ensures that the final set of adjustment parameters optimizes speech recognition performance to the greatest extent possible.
[0080] In some implementations, the set of adjustment parameters can be embedded as configuration data and stored in the system's digital signal processing chip or software module to guide audio signal processing in the actual product.
[0081] In some implementations, the adjustment parameter set is fixed as configuration data and stored in the digital signal processing chip of the audio acquisition hardware, or in the audio preprocessing module of the software application; the original audio signal is spectrally shaped by an equalizer according to the adjustment parameter set to generate the target audio signal, including steps 120-6 to 120-7: In step 120-6, the pre-set set of adjustment parameters is retrieved from the digital signal processing chip or audio preprocessing module; In step 120-7, the original audio signal is subjected to equalization filtering processing by an equalizer according to the adjustment parameter set to generate the target audio signal.
[0082] Low-latency response can be achieved by calling pre-defined adjustment parameter sets from digital signal processing chips or audio preprocessing modules. Since the selected adjustment parameter set is pre-stored in the device's non-volatile memory or registers, the central processing unit or DSP chip can directly read the parameter configuration at the corresponding address when processing of the raw audio signal is required. This pre-defined approach avoids complex real-time calculations or network requests during runtime, ensuring the real-time performance and stability of parameter loading, making it particularly suitable for resource-constrained embedded speech recognition devices.
[0083] Subsequently, the equalizer performs a convolution operation on the original audio signal based on the set of adjustment parameters called. This process essentially redistributes the spectral energy distribution of the original signal. Specifically, the equalizer enhances key speech segments masked by ambient noise (e.g., boosting high-frequency gain by 1kHz-4kHz in low-frequency noise environments to restore consonant clarity) based on the characteristics of the current scene, while moderately attenuating useless noise bands (e.g., cutting off low-frequency wind noise or power hum).
[0084] Through the aforementioned targeted peak shaving and valley filling or energy compensation, the original audio signal, which was distorted due to environmental interference, is converted into a target audio signal with spectral equalization and improved signal-to-noise ratio. This target audio signal retains more complete speech feature information, thereby significantly reducing the decoding difficulty of the subsequent speech recognition engine and improving the final recognition accuracy.
[0085] Thus, the optimal set of adjustment parameters has been determined and solidified offline for multi-dimensional acoustic scenarios. However, in practical applications, speech recognition devices often need to deal with diverse acoustic scenarios. To further improve the device's scenario adaptability and the robustness of speech recognition, in some implementations, this equalizer-based speech recognition method further includes steps 140 to 150: In step 140, the scene category to which the current acoustic scene belongs is determined; In step 150, the preset parameter set corresponding to the scene category is determined as the adjustment parameter set.
[0086] In practical applications, after acquiring the raw audio signal in the current acoustic scene, the system can first extract acoustic parameters that characterize the scene features (e.g., spectral centroid, zero-crossing rate, Mel frequency cepstral coefficients, or frequency domain entropy). Then, using a pre-trained acoustic scene classification model (e.g., a classifier based on convolutional neural networks or bidirectional gated recurrent units), the system compares and analyzes these real-time extracted acoustic parameters with the acoustic parameters of various predefined scenes in the database.
[0087] Through the above analysis, the system can accurately classify the current acoustic scene into a specific scene category. Examples include "high-noise, low-volume scene," "quiet office scene," or "vehicle driving scene." This classification not only identifies the presence of noise but also clarifies its nature and energy distribution characteristics, providing a basis for subsequent parameter selection.
[0088] In some implementations, the system's internal memory maintains a mapping table of "scene category - preset parameter set". Each preset parameter set stored in this table is the optimal combination of parameters (including specific gain values, center frequency, and bandwidth, etc.) that has been trained and selected offline for a specific scene category.
[0089] Once the scene category of the current acoustic scene corresponding to the original audio signal is determined, the central processing unit or digital signal processing chip will immediately use the scene category as an index to quickly find and lock the corresponding preset parameter set in the mapping table, thereby determining the preset parameter set as the adjustment parameter set, so that the equalizer can perform spectrum shaping processing on the original audio signal according to the adjustment parameter set to generate the target audio signal.
[0090] For example, once the system identifies a "high noise, low volume" category, it will automatically invoke a parameter set specifically optimized for that category. This parameter set typically includes a strong suppression strategy for low-frequency noise and a significant boosting strategy for key speech frequency bands. Through this dynamic matching mechanism, the system can ensure that the processing strategy best suited to the current acoustic scenario is used at any given time, thereby maximizing the performance of the speech recognition engine.
[0091] In step 130, the target audio signal is input into the speech recognition engine for recognition to obtain the target text information.
[0092] The system uses the target audio signal generated by the aforementioned spectrum shaping process as a high signal-to-noise ratio input source and feeds it to the speech recognition engine in real time to recognize the text information in the original audio signal, effectively improving the translation quality.
[0093] Please see Figure 2, Figure 2 This illustration shows a schematic diagram of a speech recognition device based on equalizer adjustment according to an embodiment of this application. The speech recognition device 200 based on equalizer adjustment includes: an acquisition module 210, a generation module 220, and a recognition module 230. Specifically: Acquisition module 210 is used to acquire the raw audio signal.
[0094] The generation module 220 is used to perform spectral shaping processing on the original audio signal according to the adjustment parameter set by the equalizer to generate the target audio signal; wherein the adjustment parameter set is determined for the multi-dimensional acoustic scene or the current acoustic scene.
[0095] The recognition module 230 is used to input the target audio signal into the speech recognition engine for recognition to obtain target text information.
[0096] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0097] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.
[0098] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0099] Please see Figure 3 , Figure 3 The diagram shows a structural schematic of a speech translation device provided in an embodiment of this application. The speech translation device 300 in this application may include one or more of the following components: a processor 310, a memory 320, and one or more application programs. The one or more application programs may be stored in the memory 320 and configured to be executed by one or more processors 310. The one or more programs are configured to perform the equalizer-based speech recognition method as described in the foregoing method embodiments.
[0100] Processor 310 may include one or more processing cores. Processor 310 connects to various parts within the speech translation device 300 using various interfaces and lines, and performs various functions and processes data of the speech translation device 300 by running or executing instructions, programs, code sets, or instruction sets stored in memory 320, and by calling data stored in memory 320. Optionally, processor 310 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 310 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 310 and may be implemented separately through a communication chip.
[0101] The memory 320 may include random access memory (RAM) or read-only memory (ROM). The memory 320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 320 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created by the voice translation device 300 during use.
[0102] Please see Figure 4 , Figure 4 The diagram shows a computer-readable storage medium 400 provided in an embodiment of this application. The computer-readable storage medium 400 stores program code, which can be called by a processor to execute the equalizer-based speech recognition method described in the above method embodiment.
[0103] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code 410 that performs any of the method steps described above. This program code can be read from or written to one or more computer program devices. The program code 410 may be compressed, for example, in a suitable form.
[0104] Please see Figure 5 , Figure 5 The diagram illustrates the structure of a computer program product according to an embodiment of this application. The computer program product 500 includes a computer program 510 stored on a computer-readable storage medium. The computer program 510 includes program instructions, which, when executed by a computer, cause the computer to perform the aforementioned speech recognition method based on equalizer adjustment.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A speech recognition method based on equalizer adjustment, characterized in that, The method includes: Acquire the raw audio signal; The original audio signal is spectrally shaped using an equalizer according to a set of adjustment parameters to generate the target audio signal; wherein the set of adjustment parameters is determined for a multi-dimensional acoustic scene or the current acoustic scene. The target audio signal is input into a speech recognition engine for recognition to obtain the target text information.
2. The speech recognition method based on equalizer adjustment according to claim 1, characterized in that, The method further includes: Construct an audio sample library covering the aforementioned multi-dimensional acoustic scenarios; The equalizer is configured to apply multiple candidate parameter sets to the multi-dimensional acoustic scene; For each candidate parameter set in the plurality of candidate parameter sets, the equalizer performs spectral shaping on each sample audio signal in the audio sample library according to the candidate parameter set to generate a set of test audio signals. Based on the recognition results of the speech recognition engine on the current group of test audio signals, evaluate the recognition metrics corresponding to the candidate parameter set; After obtaining the identification indicators corresponding to the multiple candidate parameter sets, the candidate parameter set whose identification indicators meet the value conditions is selected from the multiple candidate parameter sets as the adjustment parameter set.
3. The speech recognition method based on equalizer adjustment according to claim 2, characterized in that, The multi-dimensional acoustic scene includes multiple pre-selected acoustic scenes; The setting of the equalizer to be applicable to multiple candidate parameter sets in the multi-dimensional acoustic scene includes: Based on the equalizer parameter setting strategies corresponding to various pre-selected acoustic scenarios, the multiple candidate parameter sets are set.
4. The speech recognition method based on equalizer adjustment according to claim 3, characterized in that, The multiple pre-selected acoustic scenarios include a first scenario, a second scenario, a third scenario, and a fourth scenario; in the first scenario, the speech is mixed with background noise, and the speech volume level is a first volume level; in the second scenario, the speech is clean speech, and the speech volume level is the first volume level; in the third scenario, the speech is mixed with background noise, and the speech volume level is a second volume level; in the fourth scenario, the speech is clean speech, and the speech volume level is the second volume level; the second volume level is higher than the first volume level. The equalizer parameter setting strategy corresponding to the first scenario is: to perform attenuation processing on the first frequency band of each sample audio signal within a first attenuation amplitude range, and to perform gain processing on the second frequency band of each sample audio signal within a first gain amplitude range. The equalizer parameter setting strategy corresponding to the second scenario is: to perform gain processing on the second frequency band in each sample audio signal within the second gain amplitude range; The equalizer parameter setting strategy corresponding to the third scenario is as follows: perform attenuation processing on the first frequency band in each sample audio information within the second attenuation amplitude range, and perform gain processing on the third frequency band in each sample audio signal within the third gain amplitude range. The equalizer parameters corresponding to the fourth scene are: attenuation processing of the fourth frequency band in each sample audio signal within the third attenuation amplitude range, and gain processing of the fifth frequency band in each sample audio signal within the fourth gain amplitude range. Wherein, the first frequency band is the identified noise frequency band, the second frequency band contains speech consonant information and speech formants, the minimum frequency value of the third frequency band is greater than the maximum frequency value of the first frequency band, the maximum frequency value of the fourth frequency band is less than the minimum frequency value of the fifth frequency band, the minimum amplitude value of the first gain amplitude range is greater than the maximum amplitude value of the second gain amplitude range, the minimum amplitude value of the second gain amplitude range is greater than the maximum amplitude value of the third gain amplitude range and greater than the maximum amplitude value of the fourth gain amplitude range, the third gain amplitude range and the fourth gain amplitude range are the same or different, the first attenuation amplitude range and the second attenuation amplitude range are the same or different, and the maximum amplitude value of the third attenuation amplitude range is less than the minimum amplitude value of the first attenuation amplitude range and less than the minimum amplitude value of the second attenuation amplitude range.
5. The speech recognition method based on equalizer adjustment according to claim 2, characterized in that, If the identification metric includes average identification accuracy, then the value selection condition includes the highest average identification accuracy. If the identification metric includes the average word error rate, then the value selection condition includes the lowest average word error rate.
6. The speech recognition method based on equalizer adjustment according to claim 1, characterized in that, The method further includes: Determine the scene category to which the current acoustic scene belongs; The preset parameter set corresponding to the scene category is determined as the adjustment parameter set.
7. The speech recognition method based on equalizer adjustment according to any one of claims 1 to 6, characterized in that, The set of adjustment parameters is solidified into configuration data and stored in the digital signal processing chip of the audio acquisition hardware, or in the audio preprocessing module of the software application. The step of performing spectral shaping on the original audio signal using an equalizer according to a set of adjustment parameters to generate the target audio signal includes: The preset set of adjustment parameters is retrieved from the digital signal processing chip or the audio preprocessing module; The target audio signal is generated by performing equalization filtering on the original audio signal using the equalizer according to the set of adjustment parameters.
8. A voice translation device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the equalizer-based speech recognition method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code, which can be called by a processor to execute the speech recognition method based on equalizer adjustment as described in any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform the equalizer-based speech recognition method according to any one of claims 1-7.