Voice processing method and system

By extracting multimodal feature information and dynamically adjusting the module states and parameters in the speech processing link, the problem of resource waste in existing speech processing systems under different acoustic scenarios is solved, achieving more efficient speech enhancement and resource utilization.

CN121884841APending Publication Date: 2026-04-17MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MOORE THREADS TECH CO LTD
Filing Date
2025-12-19
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing speech processing systems are difficult to adapt to different acoustic scenarios, resulting in high computational resource consumption and a lack of flexibility. Existing technologies are also unable to coordinate and adjust multiple speech processing modules in conjunction with factors such as environmental noise, far-end signals, and speaker direction.

Method used

By extracting multimodal feature information, the start/stop status and processing parameters of modules in the speech processing link are dynamically adjusted, and the speech processing system is adaptively configured according to the acoustic processing mode, including the start/stop and parameter adjustment of beamforming, echo cancellation and noise reduction modules.

Benefits of technology

It improves speech enhancement and resource utilization efficiency, reduces unnecessary computational overhead and power consumption, and enhances the operating efficiency and robustness of the speech processing system in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884841A_ABST
    Figure CN121884841A_ABST
Patent Text Reader

Abstract

The invention provides a voice processing method and system, the system comprises an initial voice processing link composed of a plurality of voice processing modules, and the method comprises the steps: obtaining a to-be-processed audio signal, and extracting multi-modal feature information for representing a current acoustic scene based on the audio signal; determining an acoustic processing mode corresponding to the current acoustic scene based on the multi-modal feature information; adjusting the start-stop state and / or processing parameters of a voice processing module in the initial voice processing link based on the acoustic processing mode to obtain a target voice processing link; and performing voice enhancement processing on the audio signal by using the target voice processing link to obtain an enhanced voice signal. According to the invention, flexible speech enhancement processing for different acoustic scenes can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and more particularly to speech processing methods and systems. Background Technology

[0002] With the widespread use of devices such as smartphones, conferencing terminals, in-vehicle voice systems, smart speakers, and voice assistants, users' demand for voice calls and interactions in noisy environments is constantly increasing. To improve the intelligibility and listening quality of voice signals, existing voice processing systems typically configure multiple voice processing modules at the front end to perform voice enhancement processing on the acquired audio signals. For example, they use beamforming to suppress spatial interference, echo cancellation to weaken far-end echoes, and noise suppression to reduce environmental noise.

[0003] In existing technologies, common speech enhancement systems typically employ a fixed processing sequence: for each frame of input audio, it sequentially passes through modules such as beamforming, acoustic echo cancellation (AEC), and denoising. These modules are cascaded in a preset order to form a fixed processing chain, widely used in embedded devices, conference terminals, and voice assistants. While this approach is simple to implement and stable, it uses the same processing chain and roughly the same processing intensity regardless of whether the acoustic environment is quiet, noisy, or involves only near-end speech or two-way communication between near and far ends. This makes it difficult to adapt to changes in the environment, leading to high computational resource consumption and a lack of flexibility.

[0004] To reduce unnecessary processing overhead, some technical solutions attempt to introduce Voice Activity Detection (VAD) to provide simple control over certain modules. For example, when a silence segment is detected, the noise reduction intensity may be reduced or certain processing steps may be skipped to reduce computational load. However, most of these solutions rely on a single feature (such as energy threshold or spectral features) to make a coarse-grained judgment of whether speech is present or absent. They can only control the switching on and off of individual modules in a limited dimension, making it difficult to combine multiple factors such as environmental noise, far-end signals, and speaker direction to coordinate and adjust multiple speech processing modules. Summary of the Invention

[0005] To achieve flexible speech enhancement processing for different acoustic scenarios, this invention provides a speech processing method and system.

[0006] According to a first aspect of the present invention, a speech processing method is provided, applied to a speech processing system, the speech processing system including an initial speech processing link composed of multiple speech processing modules, the speech processing method comprising: Acquire the audio signal to be processed, and extract multimodal feature information to characterize the current acoustic scene based on the audio signal; Based on the multimodal feature information, an acoustic processing mode corresponding to the current acoustic scene is determined; Based on the acoustic processing mode, the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link are adjusted to obtain the target speech processing link. The audio signal is enhanced using the target speech processing link to obtain an enhanced speech signal.

[0007] In some exemplary embodiments of the present invention, based on the foregoing scheme, the multimodal feature information includes speech activity features, or the multimodal feature information includes speech activity features and proximal energy features; The initial speech processing link includes a beamforming module, an echo cancellation module, and a noise reduction module; The acoustic processing mode includes an idle mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: If the speech activity feature indicates that there is no speech in the proximal end, or if the speech activity feature indicates that there is speech in the proximal end and the proximal energy feature is less than a first threshold, the acoustic processing mode corresponding to the current acoustic scene is determined to be idle mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the idle mode, the beamforming module, the noise reduction module, and the echo cancellation module are all adjusted to be in a closed or transparent state.

[0008] In some exemplary embodiments of the present invention, based on the foregoing scheme, the multimodal feature information includes speech activity features, main sound source directional stability features, and far-end energy features; or, the multimodal feature information includes speech activity features, main sound source directional stability features, near-end energy features, and far-end energy features; The multiple speech processing modules include a beamforming module, an echo cancellation module, and a noise reduction module; The acoustic processing mode includes the normal processing mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech near the end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, or when the speech activity feature indicates the presence of speech near the end, the near-end energy feature is greater than a first threshold and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the normal processing mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the normal processing mode, the beamforming module is controlled to update the beam at a first frequency, the noise reduction module is turned on, and the echo cancellation module is enabled or disabled according to the far-end energy characteristics.

[0009] In some exemplary embodiments of the present invention, based on the foregoing scheme, the multimodal feature information includes speech activity features, main sound source direction features, and far-end energy features; or, the multimodal feature information includes speech activity features, main sound source direction stability features, near-end energy features, and far-end energy features; The multiple speech processing modules include a beamforming module, an echo cancellation module, and a noise reduction module; The acoustic processing mode includes a direction tracking mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech near the end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, or when the speech activity feature indicates the presence of speech near the end, the near-end energy feature is greater than a first threshold and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the direction tracking mode; Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the direction tracking mode, the beamforming module is controlled to update the beam at a second frequency, and the noise reduction module and the echo cancellation module are enabled or disabled according to the far-end energy characteristics.

[0010] In some exemplary embodiments of the present invention, based on the foregoing scheme, the acoustic processing mode further includes a normal processing mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech near the end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, or when the speech activity feature indicates the presence of speech near the end, the near-end energy feature is greater than a first threshold and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the normal processing mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the normal processing mode, the beamforming module is controlled to update the beam at a first frequency, the noise reduction module is turned on, and the echo cancellation module is enabled or disabled according to the far-end energy characteristics. Wherein, the first frequency is less than the second frequency.

[0011] In some exemplary embodiments of the present invention, based on the foregoing scheme, the acoustic processing mode includes a normal processing mode, a direction tracking mode, and a dual-talk protection mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: With the current acoustic processing mode set to direction tracking mode... When the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, the direction tracking mode remains unchanged. When the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode is switched from the direction tracking mode to the normal processing mode. Furthermore, when the voice activity feature indicates the presence of voice at the near end and the far end energy feature is greater than a third threshold, the acoustic processing mode is switched to the dual-talk protection mode.

[0012] In some exemplary embodiments of the present invention, based on the foregoing scheme, the main sound source direction stability feature is used to characterize the stability of the main sound source direction estimate within a preset frame window, wherein, Within a preset frame window, calculate the difference between the estimated direction of the main sound source in the current frame and the estimated direction of the main sound source in the previous frame; When the maximum value of the difference is less than the second threshold, it is determined that the estimated value of the main sound source direction is stable within a preset frame window; When the maximum value of the difference is greater than or equal to the second threshold, it is determined that the estimated value of the main sound source direction is unstable within a preset frame window.

[0013] In some exemplary embodiments of the present invention, based on the foregoing scheme, the audio signal includes a multi-channel microphone signal, and the beamforming module includes the following when performing beamforming: Based on the estimated direction of the main sound source, delay compensation is performed on the microphone signals of each channel; The microphone signals from each channel, after time delay compensation, are weighted and superimposed according to preset or adaptive weighting values ​​to obtain the beam output signal.

[0014] In some exemplary embodiments of the present invention, based on the foregoing scheme, the feature extraction module includes a sound source localization submodule, and the beamforming module further includes the following when performing beamforming: When the control signal indicates that the sound source localization submodule is turned off, the microphone signals of each channel are delayed and weighted by using the main sound source directional characteristics and / or beamforming parameters of the previous moment to obtain the beam output signal.

[0015] In some exemplary embodiments of the present invention, based on the foregoing scheme, the multimodal feature information includes speech activity features and far-end energy features; or, the multimodal feature information includes speech activity features, near-end energy features, and far-end energy features; The multiple speech processing modules include a beamforming module, an echo cancellation module, and a noise reduction module; The acoustic processing mode includes a dual-talk protection mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the voice activity feature indicates that there is voice at the near end and the energy feature at the far end is greater than a third threshold, or when the voice activity feature indicates that there is voice at the near end, the energy feature at the near end is greater than a first threshold and the energy feature at the far end is greater than a third threshold, the acoustic processing mode corresponding to the current acoustic scene is determined to be the dual-talk protection mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the dual-talk protection mode, the beamforming module is controlled to update the beam at a first frequency, the noise reduction module is turned on, and the echo cancellation module is paused from updating the filter coefficients.

[0016] According to a second aspect of the present invention, a voice processing system is provided, comprising: The feature extraction module is used to acquire the audio signal and, in response to the audio signal, extract multimodal feature information to characterize the current acoustic scene. The acoustic processing mode decision module determines the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information. An adjustment module is used to adjust the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode. The execution module uses multiple adjusted speech processing modules to perform speech enhancement processing on the audio signal to obtain an enhanced speech signal.

[0017] In some exemplary embodiments of the present invention, based on the foregoing scheme, a plurality of the speech processing modules include a speech acquisition module, which is used to acquire audio signals of the current acoustic scene at a preset frequency and send the audio signals to the feature extraction module.

[0018] According to a third aspect of the present invention, an electronic device is provided, comprising: a processor; and a memory storing computer-readable instructions that, when executed by the processor, implement the method of the first aspect.

[0019] According to a fourth aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method of the first aspect.

[0020] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: By extracting multimodal feature information to characterize the current acoustic scene in response to the audio signal, and determining the corresponding acoustic processing mode based on the multimodal feature information, and then adjusting the start / stop status and / or processing parameters of multiple speech processing modules according to the acoustic processing mode, the speech processing system no longer uses a fixed speech enhancement processing link, but can adaptively configure the speech processing modules in the initial speech processing link according to the current acoustic scene. Therefore, on the one hand, more suitable module combinations and processing intensities can be selected in different acoustic scenes such as quiet scenes, high-noise scenes, near-end single-talk, and near-end and far-end dual-talk, improving the speech enhancement effect and the intelligibility of the speech signal; on the other hand, some speech processing modules can be shut down or the processing intensity reduced in scenarios where complex processing is not required, reducing unnecessary computational overhead and power consumption, and improving the operating efficiency and overall robustness of the speech processing system in resource-constrained environments.

[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0022] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the specification, serve to explain the principles of the invention.

[0023] Figure 1 A schematic diagram of the structure of a speech processing system to which embodiments of the present invention can be applied is shown; Figure 2 The schematic diagram illustrates the structure of an initial speech processing link according to some embodiments of the present invention; Figure 3 A flowchart illustrating a speech processing method according to some embodiments of the present invention is shown schematically; Figure 4 The schematic diagram illustrates the structure of a feature extraction module according to some embodiments of the present invention; Figure 5 The schematic diagram illustrates the processing flow of the acoustic processing mode decision module according to some embodiments of the present invention; Figure 6 The schematic diagram illustrates an electronic device according to some embodiments of the present invention; Figure 7 A schematic diagram of a computer-readable storage medium according to some embodiments of the present invention is shown. Detailed Implementation

[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0025] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0027] This invention provides a speech processing system that can apply the speech processing method provided in this invention. In some embodiments, the speech processing system 100 may include an audio device, which includes a speech acquisition module 110, a feature extraction module 120, an acoustic processing mode decision module 130, and multiple speech processing modules. The multiple speech processing modules are cascaded in a preset order to form an initial speech processing link 200.

[0028] refer to Figure 2 As shown, the initial speech processing link 200 may include a beamforming module 210, an echo cancellation module 220, and a noise reduction module 230, which are cascaded in a preset order to form the initial speech processing link 200. The beamforming module 210 is used to spatially filter the signals of each channel when using a multi-channel microphone array to enhance the signal from the main sound source direction and suppress interference from other directions; the echo cancellation module 220 is used to suppress acoustic echoes leaking from the speaker to the microphone in scenarios where a far-end reference signal exists; the noise reduction module 230 is used to suppress background noise and improve the signal-to-noise ratio of the speech signal. Of course, the initial speech processing link 200 may also include other speech processing modules such as echo cancellation and automatic gain control; this invention does not limit this.

[0029] The embodiments of the present invention will now be described in detail.

[0030] like Figure 3 As shown, Figure 3 This is a flowchart illustrating a speech processing method according to an exemplary embodiment of the present invention, comprising the following steps: S310: Acquire the audio signal to be processed, and extract multimodal feature information to characterize the current acoustic scene based on the audio signal; S320: Determine the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information; S330: Based on the acoustic processing mode, adjust the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link 200 to obtain the target speech processing link; S340: The audio signal is enhanced using the target speech processing link to obtain an enhanced speech signal.

[0031] By extracting multimodal feature information to characterize the current acoustic scene in response to audio signals, and determining the corresponding acoustic processing mode based on the multimodal feature information, the system then adjusts the start / stop status and / or processing parameters of multiple speech processing modules according to the acoustic processing mode. This allows the speech processing system 100 to no longer use a fixed speech processing link, but rather to adaptively configure the target speech processing link based on the current acoustic scene, selecting a more suitable processing method in different scenarios such as quiet, high noise, near-end single-speaker, and two-speaker scenarios. For example, in scenarios where complex processing is not required, the acoustic processing mode can control the shutdown of some speech processing modules or reduce the processing intensity, reducing unnecessary computational overhead and power consumption. In scenarios requiring enhanced processing, key modules can be activated and the processing intensity or update frequency can be appropriately increased, thereby improving the utilization efficiency of computing resources while ensuring speech quality and listening experience. This is particularly suitable for embedded devices and scenarios with limited computing power.

[0032] Secondly, since the configuration of the feature extraction module 120, the acoustic processing mode decision module 130, and the target speech processing link are all executed in real time based on the current audio signal, and the reconstruction of the processing link is completed by adjusting the start / stop status and processing parameters of the speech processing module in the initial speech processing link 200, there is no need to change the hardware structure or introduce additional large-scale model inference. Therefore, the closed-loop operation of feature acquisition, mode determination, link adjustment, and speech enhancement can be completed in each frame or a short time window, which significantly shortens the response time between changes in the acoustic scene and adjustments in the processing strategy. While ensuring the speech enhancement effect, it is beneficial to reduce end-to-end processing latency, thereby improving the response speed and overall real-time performance of the speech processing system 100 to scene changes. It is suitable for application scenarios with high real-time requirements, such as conference terminals, embedded devices, and voice assistants.

[0033] Because this invention abstracts the speech enhancement process using multimodal feature information, acoustic processing modes, and target speech processing links, without limiting specific module implementation algorithms or combinations, it can flexibly adapt to different hardware platforms and different speech processing architectures. For high-performance platforms, more speech processing modules can be configured and finely controlled through mode decision-making; for low-power platforms, simplified module combinations can be used, similarly achieving basic adaptive processing through mode decision-making.

[0034] In S310, the audio signal to be processed is acquired, and multimodal feature information for characterizing the current acoustic scene is extracted based on the audio signal.

[0035] The audio signal to be processed can be a pre-processed audio signal or the original audio signal; this invention does not impose any specific limitations.

[0036] In some implementations, when the audio signal to be processed is a preprocessed audio signal, the audio preprocessing module can preprocess the acquired raw audio signal, such as performing sampling, DC removal, framing, windowing, etc., to obtain a frame-level audio signal suitable for subsequent processing.

[0037] In other implementations, when the audio signal to be processed is the original audio signal, the feature extraction module can preprocess the original audio signal and use the preprocessed audio signal as the audio signal to be processed.

[0038] The raw audio signal is acquired by the voice acquisition module 110, which can be configured to include at least one microphone for acquiring the audio signal to be processed. The microphone can be a single-channel microphone or a microphone array composed of multiple microphones. The multi-channel microphone can be a linear array, a circular array, or other array forms. In some embodiments, a multi-channel microphone array is preferred in order to acquire acoustic information in the spatial dimension, but the present invention is not limited thereto.

[0039] In this embodiment of the invention, the following description will be given as an example of an audio signal including a multi-channel microphone signal acquired by a multi-channel microphone.

[0040] Subsequently, the feature extraction module 120 extracts multimodal feature information to characterize the current acoustic scene based on the audio signal.

[0041] Multimodal feature information may include, but is not limited to: Speech activity features used to characterize the presence of near-end speech in the current frame, such as speech activity detection results obtained based on short-time energy, spectral entropy, or deep neural network models; Noise characteristics used to characterize the intensity or type of background noise in the current environment, such as noise energy, signal-to-noise ratio estimate, and frequency band energy distribution; Far-end energy characteristics used to characterize the presence and intensity of far-end signals, such as energy statistics calculated based on echo reference signals; Spatial features used to characterize the spatial location or direction changes of the main sound source, such as the estimated direction of the main sound source calculated based on the microphone array and the degree of change within a preset frame window.

[0042] The aforementioned multimodal feature information can be used individually or by combining multiple features to form a multimodal feature vector, which can be used to comprehensively represent the current acoustic scene.

[0043] By extracting multimodal feature information to characterize the current acoustic scene in response to the audio signal, and determining the corresponding acoustic processing mode based on the multimodal feature information, the speech processing system can dynamically select the appropriate processing method according to different acoustic scenes (such as quiet environment, strong noise environment, near-end only speaking, near-end and far-end dual speaking, etc.), thereby avoiding the use of a single and fixed processing flow in multiple scenarios and improving the adaptability and robustness of speech enhancement algorithms to complex acoustic environments.

[0044] In S320, an acoustic processing mode corresponding to the current acoustic scene is determined based on the multimodal feature information.

[0045] The acoustic processing mode decision module 130 receives the multimodal feature information output by the feature extraction module, judges the current acoustic scene according to the pre-set decision rules or pre-trained model, and determines the corresponding acoustic processing mode.

[0046] In one possible implementation, several acoustic processing modes can be predefined, for example: Idle mode: When the speech activity features indicate that there is no speech in the near end and the energy in the far end is low, it is considered that the current situation is a silent or idle scene with low background noise. Normal processing mode: When the speech activity features indicate that there is near-end speech and the far-end energy is low, it is considered that the current situation is a near-end solo speech scenario; Direction tracking mode: If the direction of the main sound source changes drastically within a preset frame window, it is considered that the beamforming module 210 needs to be updated more frequently in order to track the moving speaker.

[0047] Dual-talk protection mode: When the voice activity characteristics indicate that there is near-end speech and far-end energy is high, it is considered that the current situation is a dual-talk scenario in which there is speech at both the near end and the far end. Of course, in practical applications, other acoustic processing modes can be set as needed, such as strong noise suppression mode, remote dominance mode, etc., and this implementation method is not limited to these modes.

[0048] The acoustic processing mode decision module 130 can be implemented based on a rule table, that is, different feature combinations are mapped to different acoustic processing modes; or it can be implemented through a machine learning model, such as a classification network or decision tree model, to automatically classify the acoustic scene and output the acoustic processing mode identifier corresponding to the classification result.

[0049] Instead of always activating all modules sequentially in a fixed manner, the voice processing system 100 adjusts the start / stop status and / or processing parameters of multiple voice processing modules based on acoustic processing modes. This allows the system to activate only the necessary processing modules in the current acoustic scenario and configure parameters such as processing intensity and update frequency for each module. This reduces unnecessary computation, lowering computational resource and power consumption, while also enhancing the role of key processing modules when needed, thus achieving a better balance between voice enhancement and resource utilization.

[0050] In S330, the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link 200 are adjusted based on the acoustic processing mode to obtain the target speech processing link.

[0051] After determining the current acoustic processing mode, the speech processing system 100 performs unified control on multiple speech processing modules in the initial speech processing link 200 according to the acoustic processing mode, and adjusts the start / stop status and / or processing parameters of each module to obtain the target speech processing link.

[0052] For example, in one specific implementation, adjustments can be made as follows: When the acoustic processing mode is idle mode, the beamforming module 210, echo cancellation module 220 and noise reduction module 230 are controlled to be off or bypassed, and only basic signal monitoring and low power monitoring functions are retained, thereby significantly reducing the amount of computation and power consumption. When the acoustic processing mode is normal processing mode, the beamforming module 210 and the noise reduction module 230 are enabled, the echo cancellation module 220 is selectively enabled or disabled according to the far-end energy characteristics, the update frequency of the beamforming module 210 is set to the first update frequency, and the suppression intensity of the noise reduction module 230 is set to a medium level, so as to control the computational complexity while ensuring the voice quality. When the acoustic processing mode is dual-talk protection mode, the beamforming module 210 and the noise reduction module 230 are enabled, and the filter update strategy of the echo cancellation module 220 is adjusted to protection mode, such as pausing the update or reducing the update step, to avoid filter divergence in dual-talk scenarios. At the same time, the noise reduction module 230's ability to suppress background noise and residual echo can be enhanced. When the acoustic processing mode is direction tracking mode, the beamforming module 210 remains enabled, and the beamforming parameter update frequency of the beamforming module 210 is increased to a second update frequency, which is higher than the first update frequency, thereby tracking the directional changes of the main sound source more quickly. The parameters or operating states of the noise reduction module 230 and the echo cancellation module 220 can be adjusted as needed. For example, the first update frequency is once every N frames, preferably every ten frames, and the second update frequency is updated every frame or every two frames, preferably every frame.

[0053] In the above way, a corresponding target speech processing link can be generated according to different acoustic processing modes: in the target speech processing link, which speech processing modules participate in the processing and what parameter configurations are used are all determined by the acoustic processing mode, thereby realizing the linkage control of multiple speech processing modules.

[0054] After jointly adjusting multiple speech processing modules, the adjusted speech processing modules are used to perform speech enhancement processing on the audio signal. Compared with the existing technology based on fixed processing links and fixed parameters, it can configure noise suppression, echo cancellation, beamforming and other processing processes more closely to the current acoustic scenario, reduce speech distortion and processing artifacts, improve the clarity and intelligibility of the target speech, and thus improve the subjective listening experience of end users.

[0055] In S340, the audio signal is enhanced using the target speech processing link to obtain an enhanced speech signal.

[0056] After obtaining the target speech processing link, the speech processing system 100 performs speech enhancement processing on the received audio signal according to the start / stop status and processing parameters of each module in the target speech processing link. Specifically, in each frame or each processing period, the audio signal is processed sequentially according to the order of each module in the target speech processing link. Modules in the enabled state execute the corresponding speech processing algorithm, while modules in the disabled or bypassed state do not participate in the signal processing of this frame, thereby obtaining an enhanced speech signal.

[0057] In one implementation, the determination of the acoustic processing mode and the configuration of the target speech processing link can be updated periodically or event-driven. For example, the acoustic scene can be re-evaluated after processing several frames of audio, or a new decision can be triggered when a significant change in acoustic features is detected, so as to ensure that the speech enhancement processing can respond to changes in the environment and scene in a timely manner.

[0058] By abstracting the speech enhancement process through acoustic processing modes and module start / stop / parameter adjustment, without limiting specific module combinations and implementation algorithms, this invention can be easily deployed on different hardware platforms and different speech processing architectures. For high-performance platforms, more complex speech processing modules can be configured and finely controlled through mode selection; for platforms with limited computing power, the processing chain can be simplified through mode adjustment, thereby improving the reusability and scalability of the solution in various application scenarios.

[0059] In some exemplary embodiments of the present invention, the multimodal feature information includes speech activity features; The plurality of the aforementioned voice processing modules include a beamforming module 210, an echo cancellation module 220, and a noise reduction module 230; The acoustic processing mode includes an idle mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity features indicate that there is no speech in the near end, the acoustic processing mode corresponding to the current acoustic scene is determined to be idle mode; Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link 200 based on the acoustic processing mode includes: In the idle mode, the beamforming module 210, the noise reduction module 230, and the echo cancellation module 220 are all adjusted to be in a closed or transparent state.

[0060] Here, for reference Figure 4 As shown, the feature extraction module 120 may further include a speech activity detection submodule 410.

[0061] The voice activity detection submodule 410 can be configured to perform a voice activity detection (VAD) algorithm based on frame-level audio data to determine whether the current frame is a voice frame.

[0062] Specifically, for each frame of preprocessed audio signal, at least one frame-level feature related to speech activity is calculated. The feature may include, but is not limited to: short-time energy; short-time zero-crossing rate; band energy, spectral entropy, spectral centroid; and the speech presence probability output by a trained classification model (e.g., deep neural network, convolutional neural network, or recurrent neural network).

[0063] Based on the aforementioned features, the speech activity detection submodule 410 can determine whether the current frame contains valid speech according to preset decision rules or a pre-trained model. For example, the features of the current frame can be compared with an adaptive noise threshold or a preset threshold, or the feature vector can be input into a pre-trained VAD classification model to obtain a decision result that the current frame belongs to "speech" or "non-speech".

[0064] For each frame of audio data, the speech activity detection submodule 410 can output a speech activity flag vad_flag, where: when the current frame is determined to be a speech frame, vad_flag=1 is set; when the current frame is determined to be a non-speech frame or only contains background noise, vad_flag=0 is set; that is, vad_flag∈{0,1}, which is used to identify the speech activity status of the current frame.

[0065] When vad_flag=0, it can be assumed that there is no near-end speech in the current frame. Therefore, the acoustic processing mode decision module 130 can determine the acoustic processing mode corresponding to the current acoustic scene based on the speech activity features: when the speech activity features output by the speech activity detection submodule 410 indicate that there is no near-end speech, for example, when vad_flag=0 in the current frame or the current time window, and optionally when the frames are determined to be non-speech frames for several consecutive frames, the acoustic processing mode decision module 130 determines that the current acoustic scene is in a state of near-end silence or only background noise, and determines the corresponding acoustic processing mode as idle mode.

[0066] Alternatively, in some implementations, the feature extraction module 120 may further include a speech activity detection submodule 410 and an energy calculation submodule 420.

[0067] The energy calculation submodule 420 can be configured to receive the audio signal from the microphone channel and calculate the energy characteristics of the corresponding frame. Specifically, the energy calculation submodule 420 can calculate the first root mean square (RMS) value of the frame-level signal of the microphone channel and use the first RMS value as the near-end energy characteristic mic_energy; mic_energy is used to characterize the overall energy level of speech and ambient noise received by the near-end microphone in the current frame.

[0068] When vad_flag=1, it can be assumed that there is near-end speech in the current frame. Based on this, the energy calculation submodule 420 calculates the near-end energy feature mic_energy of the current frame. If the near-end energy feature mic_energy is less than the first threshold, it indicates that the signal of the microphone channel is too small. The acoustic processing mode decision module 130 also determines that the acoustic processing mode corresponding to the current acoustic scene is the idle mode.

[0069] In idle mode, the control module can adjust the start / stop status of multiple speech processing modules based on the acoustic processing mode, setting the beamforming module 210, noise reduction module 230, and echo cancellation module 220 to the off state and / or pass-through state. Specifically, the beamforming module 210 can be controlled to stop performing beam weight calculation and multi-channel weighted superposition operation, and only output the input signal of any channel or the simple superimposed signal in a pass-through manner; the echo cancellation module 220 can be controlled to stop updating the adaptive filter coefficients and bypass its filtering operation, and directly output the input signal; the noise reduction module 230 can be controlled to stop noise spectrum estimation and gain calculation, and pass-through the input signal in a manner that basically does not change the amplitude and spectrum.

[0070] In the above manner, in idle mode, multiple speech processing modules no longer perform complex speech enhancement processing on the current audio signal, thereby significantly reducing unnecessary computation and storage access, reducing system power consumption and processing latency, while still retaining the basic monitoring capability of the input signal so as to switch to other acoustic processing modes in a timely manner when changes in speech activity characteristics are detected.

[0071] In some exemplary embodiments of the present invention, the multimodal feature information includes speech activity features, main sound source direction features, and far-end energy features; The plurality of the aforementioned voice processing modules include a beamforming module 210, an echo cancellation module 220, and a noise reduction module 230; The acoustic processing mode includes the normal processing mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech in the near end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the normal processing mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link 200 based on the acoustic processing mode includes: In the normal processing mode, the beamforming module 210 is controlled to update the beam at a first frequency, the noise reduction module 230 is turned on, and the echo cancellation module 220 is enabled or disabled according to the far-end energy characteristics.

[0072] Here, for reference Figure 4 As shown, the feature extraction module 120 may further include a speech activity detection submodule 410, an energy calculation submodule 420, and a sound source localization submodule 430.

[0073] The voice activity detection submodule 410 can be configured to have the same structure and function as the voice activity detection submodule 410 described above, and will not be described in detail here.

[0074] The energy calculation submodule 420 can be configured to receive audio signals from the microphone channel and audio signals from the far-end channel, and calculate the energy features of the corresponding frames respectively. Specifically, the energy calculation submodule 420 can calculate the second root mean square (RMS) value only for the frame-level signal of the far-end channel and use the second RMS value as the far-end energy feature farend_energy; or it can simultaneously calculate the first root mean square (RMS) value for the frame-level signal of the microphone channel and use the first RMS value as the near-end speech energy feature mic_energy; and calculate the second root mean square (RMS) value for the frame-level signal of the far-end channel and use the second RMS value as the far-end energy feature farend_energy.

[0075] Here, mic_energy is used to characterize the overall energy level of speech and ambient noise received by the near-end microphone in the current frame, and farend_energy is used to characterize the energy level of the far-end signal in the current frame. In some embodiments, the energy calculation submodule 420 can also normalize, average, or perform threshold comparison on mic_energy and / or farend_energy to generate energy features suitable for subsequent acoustic processing mode decisions.

[0076] The sound source localization submodule 430 can be configured to: when the voice acquisition module 110 uses a multi-channel microphone array, receive frame-level audio data from each channel, and use a Steered Response Power-Phase Transform (SRP-PHAT) algorithm and / or a Multiple Signal Classification (MUSIC) algorithm to estimate the direction of the main sound source in the current frame, obtaining a main sound source direction estimate value doa_value, and using this main sound source direction estimate value doa_value as the main sound source direction feature. doa_value can be an angle value used to represent the azimuth angle of the main sound source relative to the microphone array reference direction, for example, limited to the range of 0 to 180°.

[0077] To characterize the stability of the main sound source direction, the sound source localization submodule 430 can also perform statistical analysis on the doa_value of multiple consecutive frames within a preset frame window, calculate the difference Δdoa_value between the estimated main sound source direction doa_value(t) of the current frame and the estimated main sound source direction doa_value(t-1) of the previous frame. When the maximum value of the difference Δdoa_value is less than a preset second threshold, the current main sound source direction is determined to be stable, and a direction stability flag is output, for example, doa_stable_flag=1; when the maximum value of the difference Δdoa_value is greater than or equal to the preset second threshold, the current main sound source direction is determined to be unstable, and a direction instability flag is output, for example, doa_stable_flag=0.

[0078] The second threshold can be determined empirically, for example, by setting it based on the typical range of change in the main sound source direction when the speaker makes slight head movements in a real deployment environment; it can also be obtained through statistical analysis or simulation experiments on a large amount of offline collected speech and scene data, for example, by selecting a threshold range that can better distinguish between stable and unstable states of the main sound source direction on a dataset labeled "stable / unstable". This invention does not limit this, and the specific value of the second threshold can be adjusted according to different application scenarios and microphone array arrangements.

[0079] Therefore, the sound source localization submodule 430 can simultaneously output the main sound source direction estimate (doa_value) and a direction stability feature characterizing the stability of the main sound source direction. On the one hand, compared to instantaneous judgment relying solely on the single-frame main sound source direction estimate (doa_value), this invention utilizes multi-frame statistics to suppress direction estimation glitches caused by noise, echoes, or short-term jitter, enabling a more accurate distinction between slight direction jitter and true direction migration, thereby improving the reliability and robustness of direction stability / instability determination. On the other hand, this stability feature provides a clear decision basis for the adaptive scheduling of subsequent acoustic processing modes (e.g., normal processing mode and direction tracking mode) and beamforming parameter update frequency: when determined to be stable, the beam update frequency can be reduced to save computational resources and power consumption; when determined to be unstable, the update frequency can be increased in a timely manner to enhance the tracking ability of moving speakers. This ensures speech enhancement performance while improving the system's real-time performance, resource utilization efficiency, and overall operational stability, avoiding performance degradation caused by frequent mode jitter or erroneous switching.

[0080] The vad_flag output by the speech activity detection submodule 410, the doa_value and its directional stability features output by the sound source localization submodule 430, and the farend_energy, or mic_energy and farend_energy output by the energy calculation submodule 420 can be provided as part of the multimodal feature information characterizing the current acoustic scene to the upper-layer acoustic processing mode decision module 130 to determine the acoustic processing mode corresponding to the current acoustic scene, and thereby drive the start / stop status and / or adjustment of processing parameters of multiple speech processing modules.

[0081] In some implementations, the acoustic processing mode decision module 130 can determine the acoustic processing mode corresponding to the current acoustic scene as the normal processing mode based on the speech activity detection submodule 410 outputting the speech activity feature vad_flag=1; the sound source localization submodule 430 outputting the main sound source direction stability doa_stable_flag=1; and the far end energy feature farend_energy calculated by the energy calculation submodule 420. That is, the current scene is considered to be a normal processing scene with near-stable speech and a basically fixed main sound source direction.

[0082] In other implementations, the acoustic processing mode decision module 130 can determine the acoustic processing mode corresponding to the current acoustic scene as the normal processing mode based on the speech activity detection submodule 410 outputting the speech activity feature vad_flag=1, the near-end energy feature mic_energy calculated by the energy calculation submodule 420 being greater than a first threshold, the main sound source direction stability doa_stable_flag=1 output by the sound source localization submodule 430, and the far-end energy feature farend_energy calculated by the energy calculation submodule 420. That is, the current scene is considered to be a normal processing scene with near-end stable speech and a basically fixed main sound source direction.

[0083] In this embodiment, the control module can adjust the start / stop status and / or processing parameters of multiple speech processing modules based on a determined normal processing mode, thereby obtaining a target speech processing link suitable for normal speech enhancement scenarios. Specifically, adjusting the start / stop status and / or processing parameters of the speech processing modules in the initial speech processing link 200 based on the acoustic processing mode includes: a. Set the beamforming module 210 to the enabled state and control it to update the beam at the first frequency.

[0084] Specifically, the beamforming module 210 can calculate and / or update the delay compensation and complex weights of the microphone signals of each channel based on the doa_value of the current frame or several frames. In normal processing mode, in order to reduce computational overhead, the update frequency of the beamforming parameters can be set to a first frequency f1, for example, updating the beamforming weights once every K frames, or performing a beam update once at a fixed time interval (e.g., the time length corresponding to several frames), thereby achieving a lower beam update overhead under the premise that the direction of the main sound source is stable.

[0085] b. Set the noise reduction module 230 to the on state.

[0086] The noise reduction module 230 can calculate the noise reduction gain and perform spectral subtraction or other forms of noise suppression processing on the audio signal based on information such as noise estimation, frequency band signal-to-noise ratio, or power spectrum, in order to improve the signal-to-noise ratio and intelligibility of the speech signal. In this mode, the noise reduction module 230 can use a medium or preset suppression intensity to balance noise suppression effect and speech naturalness.

[0087] c. Control the activation or deactivation of the echo cancellation module 220 based on the far-end energy characteristic farend_energy.

[0088] For example, a far-end energy threshold E0 can be preset: when farend_energy is greater than E0, it is determined that there is a significant far-end signal, and the echo cancellation module 220 is set to the enabled state, so that it performs echo path modeling and echo suppression processing based on the far-end reference signal and microphone signal, and allows its adaptive filter coefficients to be updated normally; when farend_energy is less than or equal to E0, it is determined that there is no far-end signal or it is weak, and the echo cancellation module 220 can be set to the disabled or bypassed state, such as stopping the filtering operation and keeping the filter coefficients from being updated, or directly bypassing the echo cancellation module 220 to output the original signal, so as to reduce unnecessary computation and avoid introducing additional processing artifacts to near-end speech.

[0089] With the above configuration, in normal processing mode, the target speech processing link includes a beamforming module 210 that is enabled and updates the beam at a first frequency, a noise reduction module 230 that is enabled, and an echo cancellation module 220 that is enabled or disabled based on the far-end energy characteristics. Thus, in typical call scenarios where there is stable speech at the near end and the direction of the main sound source is basically fixed, the system can ensure speech enhancement while avoiding excessively frequent or unnecessary calculations on the echo cancellation module 220 and the beamforming module 210, thereby improving speech quality while balancing computational resources and real-time processing performance.

[0090] In some exemplary embodiments of the present invention, the multimodal feature information includes speech activity features, main sound source direction features, and far-end energy features; or, the multimodal feature information includes speech activity features, main sound source direction stability features, near-end energy features, and far-end energy features; The plurality of the aforementioned voice processing modules include a beamforming module 210, an echo cancellation module 220, and a noise reduction module 230; The acoustic processing mode includes a direction tracking mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech near the end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, or when the speech activity feature indicates the presence of speech near the end, the near-end energy feature is greater than a first threshold and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the direction tracking mode; Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link 200 based on the acoustic processing mode includes: In the direction tracking mode, the beamforming module 210 is controlled to update the beam at a second frequency, and the noise reduction module 230 and the echo cancellation module 220 are enabled or disabled according to the far-end energy characteristics.

[0091] Here, for reference Figure 4 As shown, the feature extraction module 120 may further include a speech activity detection submodule 410, an energy calculation submodule 420, and a sound source localization submodule 430.

[0092] The structure and function of the voice activity detection submodule 410, the energy calculation submodule 420, and the sound source localization submodule 430 are the same as those described above, and will not be repeated here.

[0093] Based on this, in some implementations, the acoustic processing mode decision module 130 can determine the acoustic processing mode corresponding to the current acoustic scene as the direction tracking mode based on the voice activity detection submodule 410 outputting the voice activity feature vad_flag=1; the sound source localization submodule 430 outputting the main sound source direction stability doa_stable_flag=0; and the far end energy feature farend_energy calculated by the energy calculation submodule 420. That is, the current acoustic scene is considered to be a direction tracking scene.

[0094] In other implementations, the acoustic processing mode decision module 130 can determine the acoustic processing mode corresponding to the current acoustic scene as a direction tracking mode based on the speech activity feature vad_flag=1 output by the speech activity detection submodule 410, the near-end energy feature mic_energy calculated by the energy calculation submodule 420 being greater than a first threshold, the main sound source direction stability doa_stable_flag=0 output by the sound source localization submodule 430, and the far-end energy feature farend_energy calculated by the energy calculation submodule 420. That is, the current acoustic scene is considered to be a direction tracking scene.

[0095] In this embodiment, the control module can adjust the start / stop status and / or processing parameters of multiple speech processing modules based on a determined direction tracking mode, thereby obtaining a target speech processing link suitable for normal speech enhancement scenarios. Specifically, adjusting the start / stop status and / or processing parameters of the speech processing modules in the initial speech processing link 200 based on the acoustic processing mode includes: a. Set the beamforming module 210 to the enabled state and control it to update the beam at the second frequency.

[0096] Specifically, the beamforming module 210 can calculate or update the delay compensation and weighting values ​​of each microphone channel based on the estimated direction of the main sound source (doa_value) in the current frame or several frames. In normal processing mode, the beamforming parameters can be updated according to a first update frequency f1, for example, the beam weights are updated once every M1 frames. In direction tracking mode, in order to track changes in the direction of the main sound source more quickly, the update frequency of the beamforming module 210 is increased to a second update frequency f2, so that the beamforming module 210 recalculates the beamforming parameters in a shorter frame interval, for example, the beam weights are updated once every M2 frames, where M2 is less than M1, or f2 is greater than f1. In this way, the beam pointing speed can be improved in scenarios where the direction of the main sound source is constantly changing, and the deviation between the actual location of the main sound source and the beam pointing can be reduced.

[0097] b. The noise reduction module 230 and the echo cancellation module 220 are enabled or disabled according to the remote energy characteristics.

[0098] The energy calculation submodule 420 can calculate the root mean square (RMS) value of the frame-level signal of the far-end channel to obtain the far-end energy characteristic, farend_energy. One or more far-end energy thresholds can be preset, such as a first far-end energy threshold E1 and a second far-end energy threshold E2, where E1 ≤ E2. Based on the comparison result between farend_energy and the threshold, the noise reduction module 230 and the echo cancellation module 220 are controlled to start and stop. For example: When farend_energy is less than the first far-end energy threshold E1, it is determined that the current far-end signal is basically non-existent or has very low energy. The echo cancellation module 220 can be set to a disabled or bypassed state to avoid performing unnecessary echo path modeling and filtering operations. At the same time, according to system requirements, the noise reduction module 230 can be set to a low-intensity working state or temporarily disabled to further reduce the computational burden. When farend_energy is between the first farend energy threshold E1 and the second farend energy threshold E2, it is determined that there is a certain degree of background noise and / or the farend signal is not obvious. The noise reduction module 230 can be activated first to suppress the background noise, while a conservative strategy is adopted for the echo cancellation module 220, such as maintaining only the filter coefficients, reducing the update frequency, or keeping it in standby state. When farend_energy is greater than or equal to the second far-end energy threshold E2, it is determined that the current far-end signal energy is high and may produce obvious echo. The echo cancellation module 220 can be set to the enabled state, so that it performs echo suppression based on the far-end reference signal and allows the adaptive filter to be updated normally. At the same time, the noise reduction module 230 can be enabled or enhanced to suppress background noise and residual echo components.

[0099] With the above configuration, in normal processing mode and direction tracking mode, the target speech processing link includes a beamforming module 210 that updates the beam at a second update frequency, and a noise reduction module 230 and an echo cancellation module 220 that are enabled or disabled as needed based on the far-end energy characteristics. This ensures that the beamforming module 210 has high direction tracking capability when the direction of the main sound source changes rapidly, and also allows for reasonable configuration of noise reduction and echo cancellation processing based on the far-end signal energy, avoiding unnecessary long-term activation of all complex modules. Thus, it balances speech enhancement effect, real-time performance, and computational resource utilization efficiency in complex acoustic scenarios.

[0100] Based on this, acoustic processing modes also include normal processing modes; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech near the end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, or when the speech activity feature indicates the presence of speech near the end, the near-end energy feature is greater than a first threshold and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the normal processing mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link 200 based on the acoustic processing mode includes: In the normal processing mode, the beamforming module 210 is controlled to update the beam at a first frequency, the noise reduction module 230 is turned on, and the echo cancellation module 220 is enabled or disabled according to the far-end energy characteristics. Wherein, the first frequency is less than the second frequency.

[0101] In this embodiment, the normal processing mode is the same as the normal processing mode described above. Based on this, the start / stop status and / or processing parameters of the voice processing module in the initial voice processing link 200 in the normal processing mode are adjusted. That is, the adjustment of the start / stop status and / or processing parameters of the voice processing module in the initial voice processing link 200 in the normal processing mode is described above. This invention will not elaborate on this.

[0102] With the above configuration, in normal processing mode, the target speech processing link includes a beamforming module 210 that updates the beam at a first frequency f1, a noise reduction module 230 that is in the enabled state, and an echo cancellation module 220 that is enabled or disabled as needed based on the far-end energy characteristics. Compared to the configuration of updating the beam at a higher second frequency f2 in direction tracking mode, this embodiment reduces the update frequency of the beamforming parameters when the main sound source direction is stable, thereby reducing computational complexity and power consumption while ensuring speech enhancement effect, and improving system real-time performance and overall resource utilization efficiency.

[0103] Based on this, in some exemplary embodiments of the present invention, the multimodal feature information includes speech activity features, main sound source directional stability features, and far-end energy features; the acoustic processing modes include normal processing mode, directional tracking mode, and dual-talk protection mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: With the current acoustic processing mode set to direction tracking mode... When the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, the direction tracking mode remains unchanged. When the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode is switched from the direction tracking mode to the normal processing mode. Furthermore, when the voice activity feature indicates the presence of voice at the near end and the far end energy feature is greater than a third threshold, the acoustic processing mode is switched to the dual-talk protection mode.

[0104] The specific structure and algorithm implementation of each of the above sub-modules can be found in the aforementioned implementation methods, and will not be repeated here.

[0105] Based on this, in one specific implementation, the acoustic processing mode decision module 130 can dynamically adjust the acoustic processing mode based on doa_stable_flag and farend_energy, provided that the current acoustic processing mode is the direction tracking mode. Specifically, when the current acoustic processing mode is direction tracking mode, the acoustic processing mode decision module 130 reads the main sound source direction stability feature doa_stable_flag in each processing cycle: if doa_stable_flag indicates that the estimated value of the main sound source direction is still unstable within a preset frame window, for example, doa_stable_flag=0, then it is determined that the main sound source direction is still in a state of significant change. In order to maintain the ability to quickly track the speaker's direction, the current acoustic processing mode is kept unchanged as direction tracking mode, so that the beamforming module continues to update the beam at the second frequency; if doa_stable_flag indicates that the estimated value of the main sound source direction has stabilized within a preset frame window, for example, doa_stable_flag=1, and the dual-talk condition has not been triggered, then the acoustic processing mode decision module 130 can switch the current acoustic processing mode from direction tracking mode to normal processing mode, so that the subsequent beamforming module resumes updating the beam at the first frequency, thereby reducing the computational complexity and power consumption in the scenario where the direction is restored to stability.

[0106] Meanwhile, during the operation of the direction tracking mode, the acoustic processing mode decision module 130 detects the dual-talk scenario based on voice activity features and far-end energy features: when the voice activity feature indicates the presence of voice at the near end (vad_flag=1) and the far-end energy feature (farend_energy) is greater than a preset third threshold, it determines that the current scenario is a dual-talk scenario where voices exist simultaneously at the near and far ends. At this time, regardless of whether the direction of the main sound source is stable, the acoustic processing mode decision module 130 can prioritize switching the current acoustic processing mode from the direction tracking mode to the dual-talk protection mode and trigger the corresponding module control strategy, such as keeping the beamforming module 210 and the noise reduction module 230 on, while performing dual-talk protection operations such as freezing the filter coefficients of the echo cancellation module 220.

[0107] In this way, when the current acoustic processing mode is directional tracking mode, the system can adaptively switch between three states: continuing directional tracking, resuming normal processing, and entering dual-talk protection, based on the directional stability characteristics of the main sound source and the energy characteristics of the far end. This allows the system to balance speech enhancement, real-time performance, and system stability in complex acoustic scenarios such as moving speakers and dual-talk.

[0108] In some exemplary embodiments of the present invention, the multimodal feature information includes speech activity features and far-end energy features; or, the multimodal feature information includes speech activity features, near-end energy features, and far-end energy features. The plurality of the aforementioned voice processing modules include a beamforming module 210, an echo cancellation module 220, and a noise reduction module 230; The acoustic processing mode includes a dual-talk protection mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the voice activity feature indicates that there is voice at the near end and the energy feature at the far end is greater than a third threshold, or when the voice activity feature indicates that there is voice at the near end, the energy feature at the near end is greater than a first threshold and the energy feature at the far end is greater than a third threshold, the acoustic processing mode corresponding to the current acoustic scene is determined to be the dual-talk protection mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link 200 based on the acoustic processing mode includes: In the dual-talk protection mode, the beamforming module 210 is controlled to update the beam at a first frequency, the noise reduction module 230 is turned on, and the echo cancellation module 220 is paused from updating the filter coefficients.

[0109] In other words, based on the aforementioned feature extraction module 120 and multi-speech processing module, a dual-talk protection mode is further set. At this time, the multimodal feature information used to characterize the current acoustic scene includes the aforementioned speech activity features and far-end energy features, and optionally also includes near-end energy features; the multiple speech processing modules include the aforementioned beamforming module 210, echo cancellation module 220, and noise reduction module 230.

[0110] Here, as Figure 4 As shown, the feature extraction module 120 may further include a speech activity detection submodule 410 and an energy calculation submodule 420. The structure and function of the speech activity detection submodule 410 and the energy calculation submodule 420 can be referred to the description of speech activity features and energy features in the foregoing embodiments, and will not be repeated here.

[0111] Based on this, in some implementations, the acoustic processing mode decision module 130 can determine the acoustic processing mode corresponding to the current acoustic scenario as a dual-talk protection mode based on the voice activity features output by the voice activity detection submodule 410 and the far-end energy features output by the energy calculation submodule 420. Specifically: When the voice activity feature output by the voice activity detection submodule 410 satisfies vad_flag=1, indicating that there is currently near-end voice, and the far-end energy feature farend_energy calculated by the energy calculation submodule 420 is greater than the preset third threshold, the acoustic processing mode decision module 130 can determine that the current situation is a dual-talk scenario where near-end and far-end voices exist simultaneously, and set the corresponding acoustic processing mode to dual-talk protection mode. In other implementations, when the voice activity detection submodule 410 outputs vad_flag = 1, the near-end energy feature mic_energy calculated by the energy calculation submodule 420 is greater than a preset first threshold, and the far-end energy feature farend_energy is greater than a third threshold, the acoustic processing mode decision module 130 can also determine the corresponding acoustic processing mode as a dual-talk protection mode to cover the dual-talk scenario where "near-end voice energy is sufficient and far-end voice is significantly present".

[0112] The first threshold is the same as the aforementioned first threshold, which will not be elaborated upon here. The third threshold can be used to distinguish between scenarios with weak far-end signals and significant far-end speech. The specific value can be set through system calibration or statistical analysis of offline data, and this invention does not limit it in this regard.

[0113] In this embodiment, the control module can adjust the start / stop status and / or processing parameters of multiple voice processing modules based on a determined dual-talk protection mode, thereby obtaining a target voice processing link suitable for dual-talk scenarios. Specifically, adjusting the start / stop status and / or processing parameters of the voice processing modules in the initial voice processing link 200 based on the acoustic processing mode may include: a. The beamforming module 210 is controlled to update the beam at a first frequency.

[0114] In dual-talk protection mode, the beamforming module 210 can remain enabled, and the beamforming parameters can be updated using the first update frequency f1 in normal processing mode. For example, the delay compensation amount and weighting value related to the main sound source can be updated once every M1 frames processed, instead of using the higher second update frequency f2 in directional tracking mode. In this way, the complexity of beamforming-related calculations can be controlled while ensuring a certain spatial directivity enhancement capability, thus balancing the voice enhancement effect and system real-time performance in dual-talk scenarios.

[0115] b. Enable the noise reduction module 230.

[0116] In dual-talk protection mode, the control module can set the noise reduction module 230 to the on state. The noise reduction module 230 can perform noise suppression processing on the beam output signal based on the above, so as to reduce the interference of environmental noise and some residual echo on near-end speech, suppress background noise and echo residue, and improve the clarity and intelligibility of near-end speech.

[0117] c. Control the echo cancellation module 220 to pause updating the filter coefficients.

[0118] In dual-talk protection mode, the control module can set the echo cancellation module 220 to a protected state: the echo cancellation module 220 continues to filter the input signal using the adaptive filter coefficients from the most recent valid update to maintain the existing echo suppression effect, but suspends the update of the filter coefficients, that is, freezes the weight adjustment process of the adaptive algorithm during dual-talk. For example, when dual-talk protection mode is detected, the step size parameter of the adaptive filter can be set to zero, or the weight update logic based on the error signal can be temporarily turned off, retaining only the filtering operation of the existing filter structure.

[0119] With the above configuration, in dual-talk protection mode, the target speech processing link includes a beamforming module 210 that performs beamforming processing at a first update frequency f1, a noise reduction module 230 that is in the enabled state, and an echo cancellation module 220 that still performs echo filtering but pauses updating the filter coefficients. This maintains spatial enhancement and noise suppression capabilities for near-end speech in dual-talk scenarios, and by freezing the filter updates of the echo cancellation module 220, avoids echo path estimation being contaminated by near-end speech under dual-talk conditions, thus improving the stability of echo cancellation and the overall robustness of the system.

[0120] It should be noted that, based on the above embodiments, in some example embodiments, the acoustic processing mode decision module 130 can also dynamically switch the acoustic processing mode based on the current acoustic processing mode and the multimodal feature information of consecutive frames. Optionally, the acoustic processing mode can be regarded as a state machine that switches between normal processing mode, direction tracking mode and dual-talk protection mode.

[0121] When the current acoustic processing mode is direction tracking mode, the acoustic processing mode decision module 130 can read the main sound source direction stability feature doa_stable_flag output by the sound source localization submodule 430 in each processing cycle, and execute the following decision logic accordingly: When doa_stable_flag=0, it is determined that the estimated direction of the main sound source is still unstable within the preset frame window. The acoustic processing mode decision module 130 keeps the current direction tracking mode unchanged, that is, continues to process the current frame and subsequent frames in the direction tracking mode. When doa_stable_flag=1, it is determined that the estimated value of the main sound source direction has stabilized within the preset frame window, and no double talk is detected at present (e.g., the conditions such as the farend energy feature farend_energy being lower than the third threshold are met). The acoustic processing mode decision module 130 can switch the acoustic processing mode from the direction tracking mode back to the normal processing mode, so that the subsequent frames are restored to the beam update frequency and module start / stop configuration in the normal processing mode.

[0122] Furthermore, in some implementations, to improve the robustness of the system in complex call scenarios, the acoustic processing mode decision module 130 can prioritize the detection of dual-talk scenarios: when the voice activity feature output by the voice activity detection submodule 410 indicates the presence of voice vad_flag=1 at the near end, and the far end energy feature farend_energy output by the energy calculation submodule 420 is greater than a preset third threshold (optionally, the near end energy feature mic_energy can also be required to be greater than the first threshold), that is, when it is determined that there is a dual-talk scenario with simultaneous near and far-end speech, regardless of whether the current acoustic processing mode is normal processing mode or direction tracking mode, the acoustic processing mode decision module 130 can directly switch the acoustic processing mode to dual-talk protection mode. In this way, it can be ensured that the dual-talk protection mode is entered in time when dual-talk is detected, triggering the aforementioned protection strategy for updating the filter coefficients of the echo cancellation module 220, and avoiding divergence or false convergence of adaptive echo cancellation under dual-talk conditions.

[0123] Optionally, to avoid frequent switching between direction tracking mode and normal processing mode, in some embodiments, the acoustic processing mode decision module 130 may also, based on the statistical results of doa_stable_flag over several consecutive processing cycles, trigger the switch from direction tracking mode to normal processing mode only when the directional stability characteristics of the main sound source indicate stability for multiple consecutive frames; similarly, it may switch to dual-talk protection mode only when the dual-talk condition continuously meets a preset number of frames. This invention does not impose limitations on this, and the specific frame number threshold can be configured according to the actual application scenario.

[0124] In some exemplary embodiments of the present invention, the audio signal includes a multi-channel microphone signal, and the beamforming module 210 includes the following when performing beamforming: Based on the estimated direction of the main sound source, delay compensation is performed on the microphone signals of each channel; The microphone signals from each channel, after time delay compensation, are weighted and superimposed according to preset or adaptive weighting values ​​to obtain the beam output signal.

[0125] In other words, assuming the microphone array includes M microphone channels, the estimated direction of the main sound source output by the sound source localization submodule 430 is doa_value, expressed as an azimuth angle of 0 to 180°. In this embodiment, the beamforming module 210 can calculate the delay compensation amount corresponding to each microphone channel based on the array geometry, such as the microphone spacing, array shape, and doa_value. For example, using a reference microphone, the relative arrival time difference of the main sound source wavefront to other microphones is calculated, and this arrival time difference is converted into a delay amount at the sampling point level or a phase offset of the frequency domain signal. Subsequently, the beamforming module 210 applies corresponding delay compensation processing to each microphone signal, which can be achieved by using time-domain delay (including integer sampling delay or fractional sampling delay implemented through a fractional delay filter) or frequency-domain phase rotation, so that the effective speech components from the main sound source direction are aligned in time as much as possible after delay compensation for each channel.

[0126] After delay compensation is completed, beamforming module 210 can assign a complex weight value to each delayed microphone signal, i.e., a weight system that simultaneously has amplitude and phase, and perform weighted superposition of the signals from each channel to obtain the beam output signal. The complex weight value can be a directional weighting coefficient preset according to the array design, such as the equal amplitude weight coefficient used to implement delay-and-sum beamforming; or it can be a weight coefficient adaptively calculated based on the statistical characteristics of the target direction signal and interference / noise, such as the weight coefficient updated in real time using minimum variance distortionless response (MVDR) beamforming or other adaptive beamforming algorithms.

[0127] By performing weighted superposition of the signals after delay compensation for each channel, a main lobe with enhanced gain can be formed in the direction of the main sound source, while one or more low-gain side lobes or nulls can be formed in the non-target direction. This suppresses interference and noise from other directions, resulting in a beam output signal with a high signal-to-noise ratio, which serves as the input for subsequent echo cancellation, noise reduction, and other processing modules.

[0128] In some exemplary embodiments of the present invention, the feature extraction module 120 includes a sound source localization submodule 430, and the beamforming module 210 further includes the following when performing beamforming: When the control signal indicates that the sound source localization submodule 430 is turned off, the microphone signals of each channel are delayed and weighted by using the main sound source directional characteristics and / or beamforming parameters of the previous moment to obtain the beam output signal.

[0129] As mentioned earlier, when the sound source localization submodule 430 is in the active state, the speech processing system 100 can estimate the direction of the main sound source based on the multi-channel microphone signals using the Steered Response Power-Phase Transform (SRP-PHAT) algorithm and / or the Multiple Signal Classification (MUSIC) algorithm, obtaining the estimated direction value of the main sound source doa_value, and beamforming parameters such as the delay compensation amount and complex weighting value of each channel calculated based on doa_value. The beamforming module 210 can cache the main sound source direction characteristics (such as doa_value and its stability flag) and the corresponding beamforming parameters (such as the delay amount of each channel, weighting coefficient vector, etc.) in the storage unit at this moment.

[0130] When the control module issues a control signal instructing the sound source localization submodule 430 to be turned off, it indicates that there is no need to re-estimate the direction of the main sound source for a period of time. At this time, when the beamforming module 210 performs beamforming, it no longer triggers the sound source localization submodule 430 to estimate the direction of the main sound source in the current frame. Instead, it directly reads the main sound source direction features and / or corresponding beamforming parameters determined in the previous moment (e.g., the previous frame or the previous update cycle) from the storage unit, and processes the microphone signals of each channel in the current frame based on these cached parameters. On the one hand, the microphone signals of each channel are delayed according to the delay compensation amount of the previous moment, so that the speech components from the direction of the main sound source in the previous moment remain aligned in time. On the other hand, the signal of each channel after delay compensation is weighted and superimposed according to the weighted value of the previous moment to obtain the beam output signal of the current frame.

[0131] In some implementations, when only the main sound source direction estimate, doa_value, is stored in the cache, the beamforming module 210 can recalculate the delay compensation and weighting values ​​based on doa_value, but without calling the sound source localization submodule 430 to perform a new direction search, thus avoiding repeated execution of the highly complex direction estimation algorithm. In other implementations, when the complete beamforming parameters (including delay and weight vector) for the previous moment have been calculated and cached, these parameters can be directly reused to perform delay compensation and weighted superposition on the current frame signal without recalculation.

[0132] In this way, during the period when the sound source localization submodule 430 is turned off, the beamforming module 210 can still maintain the beam pointing and enhancement effect, ensuring the continuity of speech enhancement performance when the main sound source direction remains basically unchanged or changes little. At the same time, it avoids frequently performing sound source localization operations with high computational complexity, which helps to reduce the overall computational load and power consumption, and reduces system latency jitter caused by frequent parameter updates.

[0133] refer to Figure 5 As shown, Figure 5 The specific processing flow of the acoustic processing mode decision module 130 of the present invention is shown below: After powering on, the voice processing system 100 first enters idle mode. In this mode, only audio acquisition and feature extraction functions are maintained to continuously acquire voice activity features, main sound source direction-related features, and near / far-end energy features. The beamforming module 210, noise reduction module 230, and echo cancellation module 220 are all in a closed or pass-through state to reduce computational and power consumption overhead. When the voice activity features indicate the presence of near-end voice within a preset number of frames, the voice processing system 100 switches from idle mode to normal processing mode, turns on the beamforming module 210 and noise reduction module 230, and enables or disables the echo cancellation module 220 as needed based on the far-end energy features. At the same time, the beamforming parameters are updated at a first update frequency to achieve conventional voice enhancement processing.

[0134] In normal processing mode, the speech processing system 100 continuously monitors changes in the direction of the main sound source using the stability characteristics of the main sound source direction. When the characteristic indicates that the direction of the main sound source remains stable within a preset frame window, the normal processing mode is maintained, and the beam is updated at a lower first update frequency, thereby saving computing resources while ensuring the enhancement effect. When the characteristic indicates that the direction of the main sound source is unstable within a preset frame window, and is optionally determined to be a valid speech scene based on near-end energy characteristics, the speech processing system 100 switches the current acoustic processing mode from normal processing mode to direction tracking mode. In direction tracking mode, the beamforming module 210 remains on and the beam update frequency is increased to a second update frequency, updating the delay compensation and weights based on the estimated value of the main sound source direction at a higher frequency, thereby enhancing the tracking capability for rapid changes in the direction of the main sound source. The noise reduction module 230 remains on, and the echo cancellation module 220 is still enabled or disabled as needed based on the far-end energy characteristics. If the direction of the main sound source is detected to stabilize again and no double speech occurs during the direction tracking process, the speech processing system 100 can return from the direction tracking mode to the normal processing mode; if the near-end speech is detected to have ended, it returns to the idle mode.

[0135] In either the normal processing mode or the direction tracking mode, the speech processing system 100 will continuously detect the two-way conversation scenario: when the speech activity features indicate the presence of speech at the near end and the far-end energy features exceed a preset third threshold (optionally, also combined with the near-end energy features exceeding a first threshold), it is determined that two-way conversation is currently present, and the acoustic processing mode switches to the two-way conversation protection mode. In the two-way conversation protection mode, the beamforming module 210 continues to operate, generally using the first update frequency in the normal processing mode to maintain spatial directional enhancement, and the noise reduction module 230 remains on to suppress environmental noise and some residual echoes; the echo cancellation module 220 enters a protection state, that is, it continues to use the existing filter coefficients for echo filtering, but suspends the adaptive update of the filter coefficients to avoid the near-end speech components in the error signal interfering with echo path modeling under two-way conversation conditions. When the far-end energy features decrease and remain below the third threshold, and it is determined that the two-way conversation has ended, the speech processing system 100 can recover to the appropriate mode between the normal processing mode and the direction tracking mode based on the directional stability features of the main sound source; when the near-end speech is detected to have ended, it returns to the idle mode again. By implementing automatic mode switching and parameter scheduling based on voice activity, main sound source direction stability, and dual-talk scenarios, dynamic start / stop and frequency update control of beamforming, echo cancellation, and noise reduction modules 230 are achieved under different acoustic scenarios, thus balancing voice enhancement effect, real-time performance, and computational resource utilization efficiency.

[0136] In an exemplary embodiment of the present invention, an electronic device capable of implementing the above-described voice processing method is also provided. (See reference) Figure 6 As shown, the electronic device 600 includes a processor 601 and a memory 602. The memory 602 stores computer-readable instructions, which, when executed by the processor 601, implement the above-described method.

[0137] In an exemplary embodiment of the present invention, a computer-readable storage medium is also provided, on which computer program code instructions are stored, which, when invoked by a processor of a network communication device, cause the processor to execute the method described in the embodiments.

[0138] refer to Figure 7 As shown, a program product 700 for implementing the above-described method according to an embodiment of this application is described. It may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0139] Through the description of the above embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.

[0140] Finally, the above preferred embodiments are only used to illustrate the technical solutions of this application and are not restrictive. Although this application has been described in detail, those skilled in the art should understand that changes in form and detail can be made without departing from the scope defined by the claims of this application. The dimensions in the drawings are not related to the specific physical object, and the physical object dimensions can be arbitrarily changed.

Claims

1. A speech processing method applied to a speech processing system, the speech processing system comprising an initial speech processing link consisting of multiple speech processing modules, characterized in that, The speech processing method includes: Acquire the audio signal to be processed, and extract multimodal feature information to characterize the current acoustic scene based on the audio signal; Based on the multimodal feature information, an acoustic processing mode corresponding to the current acoustic scene is determined; Based on the acoustic processing mode, the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link are adjusted to obtain the target speech processing link. The audio signal is enhanced using the target speech processing link to obtain an enhanced speech signal.

2. The speech processing method according to claim 1, characterized in that, The multimodal feature information includes speech activity features, or the multimodal feature information includes speech activity features and proximal energy features; The initial speech processing link includes a beamforming module, an echo cancellation module, and a noise reduction module; The acoustic processing mode includes an idle mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: If the speech activity feature indicates that there is no speech in the proximal end, or if the speech activity feature indicates that there is speech in the proximal end and the proximal energy feature is less than a first threshold, the acoustic processing mode corresponding to the current acoustic scene is determined to be idle mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the idle mode, the beamforming module, the noise reduction module, and the echo cancellation module are all adjusted to be in a closed or transparent state.

3. The speech processing method according to claim 1, characterized in that, The multimodal feature information includes speech activity features, main sound source directional stability features, and far-end energy features; or, the multimodal feature information includes speech activity features, main sound source directional stability features, near-end energy features, and far-end energy features. The multiple speech processing modules include a beamforming module, an echo cancellation module, and a noise reduction module; The acoustic processing mode includes the normal processing mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech near the end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, or when the speech activity feature indicates the presence of speech near the end, the near-end energy feature is greater than a first threshold and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the normal processing mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the normal processing mode, the beamforming module is controlled to update the beam at a first frequency, the noise reduction module is turned on, and the echo cancellation module is enabled or disabled according to the far-end energy characteristics.

4. The speech processing method according to claim 1, characterized in that, The multimodal feature information includes speech activity features, main sound source direction features, and far-end energy features; or, the multimodal feature information includes speech activity features, main sound source direction stability features, near-end energy features, and far-end energy features. The multiple speech processing modules include a beamforming module, an echo cancellation module, and a noise reduction module; The acoustic processing mode includes a direction tracking mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech near the end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, or when the speech activity feature indicates the presence of speech near the end, the near-end energy feature is greater than a first threshold and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the direction tracking mode; Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the direction tracking mode, the beamforming module is controlled to update the beam at a second frequency, and the noise reduction module and the echo cancellation module are enabled or disabled according to the far-end energy characteristics.

5. The speech processing method according to claim 4, characterized in that, The acoustic processing mode also includes a normal processing mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the speech activity feature indicates the presence of speech near the end and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, or when the speech activity feature indicates the presence of speech near the end, the near-end energy feature is greater than a first threshold and the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode corresponding to the current acoustic scene is determined to be the normal processing mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the normal processing mode, the beamforming module is controlled to update the beam at a first frequency, the noise reduction module is turned on, and the echo cancellation module is enabled or disabled according to the far-end energy characteristics. Wherein, the first frequency is less than the second frequency.

6. The speech processing method according to claim 1, characterized in that, The multimodal feature information includes speech activity features, main sound source directional stability features, and far-end energy features; the acoustic processing modes include normal processing mode, direction tracking mode, and dual-talk protection mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: With the current acoustic processing mode set to direction tracking mode... When the main sound source direction stability feature indicates that the estimated value of the main sound source direction is unstable within a preset frame window, the direction tracking mode remains unchanged. When the main sound source direction stability feature indicates that the estimated value of the main sound source direction is stable within a preset frame window, the acoustic processing mode is switched from the direction tracking mode to the normal processing mode. Furthermore, when the voice activity feature indicates the presence of voice at the near end and the far end energy feature is greater than a third threshold, the acoustic processing mode is switched to the dual-talk protection mode.

7. The speech processing method according to any one of claims 3-6, characterized in that, The main sound source direction stability feature is used to characterize the stability of the estimated main sound source direction within a preset frame window, wherein, Within a preset frame window, calculate the difference between the estimated direction of the main sound source in the current frame and the estimated direction of the main sound source in the previous frame; When the maximum value of the difference is less than the second threshold, it is determined that the estimated value of the main sound source direction is stable within a preset frame window; When the maximum value of the difference is greater than or equal to the second threshold, it is determined that the estimated value of the main sound source direction is unstable within a preset frame window.

8. The speech processing method according to any one of claims 3-6, characterized in that, The audio signal includes a multi-channel microphone signal, and the beamforming module includes the following when performing beamforming: Based on the estimated direction of the main sound source, delay compensation is performed on the microphone signals of each channel; The microphone signals from each channel, after time delay compensation, are weighted and superimposed according to preset or adaptive weighting values ​​to obtain the beam output signal.

9. The speech processing method according to claim 8, characterized in that, The feature extraction module includes a sound source localization submodule, and the beamforming module further includes the following when performing beamforming: When the control signal indicates that the sound source localization submodule is turned off, the microphone signals of each channel are delayed and weighted by using the main sound source directional characteristics and / or beamforming parameters of the previous moment to obtain the beam output signal.

10. The speech processing method according to any one of claims 1-6, characterized in that, The multimodal feature information includes speech activity features and far-end energy features; or, the multimodal feature information includes speech activity features, near-end energy features, and far-end energy features. The multiple speech processing modules include a beamforming module, an echo cancellation module, and a noise reduction module; The acoustic processing mode includes a dual-talk protection mode; Determining the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information includes: When the voice activity feature indicates that there is voice at the near end and the energy feature at the far end is greater than a third threshold, or when the voice activity feature indicates that there is voice at the near end, the energy feature at the near end is greater than a first threshold and the energy feature at the far end is greater than a third threshold, the acoustic processing mode corresponding to the current acoustic scene is determined to be the dual-talk protection mode. Adjusting the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode includes: In the dual-talk protection mode, the beamforming module is controlled to update the beam at a first frequency, the noise reduction module is turned on, and the echo cancellation module is paused from updating the filter coefficients.

11. A voice processing system, characterized in that, Including the initial voice processing link consisting of multiple voice processing modules, it also includes: The feature extraction module is used to acquire audio signals and, in response to the audio signals, extract multimodal feature information to characterize the current acoustic scene. The acoustic processing mode decision module determines the acoustic processing mode corresponding to the current acoustic scene based on the multimodal feature information. An adjustment module is used to adjust the start / stop status and / or processing parameters of the speech processing module in the initial speech processing link based on the acoustic processing mode. The execution module uses multiple adjusted speech processing modules to perform speech enhancement processing on the audio signal to obtain an enhanced speech signal.

12. The speech processing system according to claim 11, characterized in that, The multiple speech processing modules include a speech acquisition module, which is used to acquire audio signals of the current acoustic scene at a preset frequency and send the audio signals to the feature extraction module.

13. An electronic device, characterized in that, include: processor; as well as A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that, It stores a computer program thereon, which, when executed by a processor, implements the method as described in any one of claims 1 to 10.