Echo cancellation method and device, electronic equipment, storage medium and program product
By sensing the user's emotional state and intensity in real time and dynamically adjusting the echo estimation filter parameters, the problem of deep learning-driven echo cancellation technology being unable to adapt to changes and emotional fluctuations in the home environment is solved, resulting in clearer voice interaction and a better user experience.
Patent Information
- Application Number
- CN202511475167.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-02-06
AI Technical Summary
Existing deep learning-driven echo cancellation technology cannot adapt to changes in the home environment and fluctuations in user emotions, resulting in echo residue or speech distortion.
By sensing users' emotional state and intensity in real time, the filter parameters of the echo estimation filter are dynamically adjusted, including adjusting the filter convergence direction and step size. Combined with the emotional influence coefficient and user interaction feedback data with a preset iteration period, the echo cancellation process is optimized.
It effectively adapts to environmental changes and user emotional fluctuations, avoids echo residue or echo cancellation of voice distortion, and improves the clarity of voice interaction and user experience of smart home devices.
Smart Images

Figure CN121483271A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and more particularly to an echo cancellation method, device, electronic device, storage medium, and program product. Background Technology
[0002] Echo cancellation of audio signals is a crucial step in enabling smart voice functionality in home smart devices.
[0003] Currently, the echo cancellation technology commonly used in smart home devices is deep learning-driven echo cancellation technology. This technology builds an echo cancellation model based on a deep neural network, learns the difference between the echo and the target speech using the echo cancellation model, and generates an echo suppression mask to achieve echo cancellation.
[0004] However, because deep learning-driven echo cancellation technology only focuses on the characteristics of the audio signal itself, it is prone to misclassifying the target speech under emotional fluctuations as echo or noise when there are sudden changes in the audio signal characteristics, such as a sudden increase in sound pressure level when the user is angry or a slow speech rate when the user is sad. This results in excessive speech suppression. In other words, deep learning-driven echo cancellation technology has the problem of not being able to adapt to environmental changes and user emotional fluctuations. Summary of the Invention
[0005] This application provides an echo cancellation method, apparatus, electronic device, storage medium, and program product to address the shortcomings of existing technologies that cannot adapt to environmental changes and user emotional fluctuations.
[0006] This application provides an echo cancellation method, including: Adjust the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal; Based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the echo-cancelled near-end audio signal is determined.
[0007] According to the echo cancellation method provided in this application, the filter parameters include the filter convergence direction and the filter step size; adjusting the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal includes: adjusting the filter convergence direction based on the emotional state; and adjusting the filter step size based on the emotional intensity value.
[0008] According to an echo cancellation method provided in this application, adjusting the filter step size based on the emotion intensity value includes: adjusting the filter step size based on the filter base step size, the emotion influence coefficient, and the emotion intensity value; the emotion influence coefficient is determined based on the emotion state.
[0009] According to the echo cancellation method provided in this application, after determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the adjusted filter parameters, the method further includes: adjusting the mapping relationship between the emotion influence coefficient and the emotion state according to the user interaction feedback data within a preset iteration period.
[0010] According to an echo cancellation method provided in this application, before adjusting the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal, the method further includes: determining a multimodal input vector based on the near-end mixed audio signal; inputting the multimodal input vector into a pre-trained emotion perception model to obtain the emotion probability distribution and the emotion intensity value output by the emotion perception model; the emotion probability distribution includes emotion categories and their probabilities; and determining the emotional state based on the emotion category with the highest probability.
[0011] According to an echo cancellation method provided in this application, determining a multimodal input vector based on the near-end mixed audio signal includes: determining the multimodal input vector based on the acoustic features of the near-end mixed audio signal, a timestamp, and a text sequence corresponding to the near-end mixed audio signal.
[0012] According to the echo cancellation method provided in this application, after determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the adjusted filter parameters, the method further includes: adjusting the emotion category classification weights of the emotion perception model according to the preset iteration period based on user interaction feedback data within a preset iteration period.
[0013] According to an echo cancellation method provided in this application, before determining the echo-cancelled near-end audio signal based on a reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the method further includes: adjusting the model weights of a depth mask generation model according to the emotional state; the depth mask generation model is constructed based on a convolutional neural network; determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters includes: determining an echo estimation signal based on the reference audio signal and the echo estimation filter with adjusted filter parameters; inputting the echo estimation signal and the near-end mixed audio signal into the depth mask generation model with adjusted model weights to obtain the echo-cancelled near-end audio signal output by the depth mask generation model.
[0014] This application also provides an echo cancellation device, comprising: The parameter adjustment module is used to adjust the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal. The signal echo cancellation module is used to determine the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the filter parameters adjusted.
[0015] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement any of the echo cancellation methods described above.
[0016] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the echo cancellation methods described above.
[0017] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the echo cancellation methods described above.
[0018] The echo cancellation method, apparatus, electronic device, storage medium, and program product provided in this application consider the influence of emotions on echo cancellation, perceive the user's emotional state and intensity in real time, and then dynamically adjust the filter parameters of the echo estimation filter based on the real-time emotional state and intensity. Echo cancellation is performed based on the echo estimation filter adjusted according to the convergence direction and step size. This can adapt to environmental changes and user emotional fluctuations, and avoid echo residue or echo cancellation speech distortion problems in emotional fluctuation scenarios. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is one of the flowcharts of the echo cancellation method provided in this application.
[0021] Figure 2 This is the second flowchart of the echo cancellation method provided in this application.
[0022] Figure 3 This is a schematic diagram of the echo cancellation device provided in this application.
[0023] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] It should be noted that, in the description of this application, the term "comprising" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Those skilled in the art will understand the specific meaning of the above terms in this application according to the specific circumstances.
[0026] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.
[0027] The following is combined with Figures 1-4 This application describes the echo cancellation method, apparatus, electronic device, storage medium, and program product provided.
[0028] Current echo cancellation technologies in smart home devices mainly include traditional adaptive filtering technology and deep learning-driven echo cancellation technology.
[0029] Traditional adaptive filtering techniques refer to the use of classic adaptive algorithms such as NLMS (Normalized Least Mean Squares) and RLS (Recursive Least Squares) to construct an adaptive filter to estimate the echo path, and then cancel the echo signal. Basic echo cancellation is achieved by iteratively updating the filter with a fixed step size.
[0030] Because traditional adaptive filtering technology uses filter parameters with fixed step size and fixed convergence direction, it is prone to echo residue or target speech distortion when the home environment changes dynamically, such as changes in echo paths caused by family members moving around, or differences in acoustic characteristics between different rooms (such as an open living room versus a closed bedroom).
[0031] Deep learning-driven echo cancellation technology refers to building an echo cancellation model based on deep neural networks. By extracting the time-domain or frequency-domain features of the audio signal, the echo cancellation model learns the difference between the echo and the target speech, generating an echo suppression mask to achieve echo cancellation and improve echo suppression capabilities in complex noise environments. The echo cancellation model can be implemented based on any of the deep neural networks such as DTLN-aec (Dual Time Domain Linear Network Echo Cancellation Model) and ConvTasNet (Convolutional Time Domain Audio Separation Network).
[0032] Because deep learning-driven echo cancellation technology only focuses on the characteristics of the audio signal itself, when the sound pressure level suddenly increases when the user is angry or the speech rate slows down when the user is sad, it will cause a sudden change in the characteristics of the audio signal. The echo cancellation model based on deep neural network is prone to misjudging the target speech under emotional fluctuation as an echo or noise, resulting in excessive speech suppression.
[0033] It is evident that current echo cancellation technology in smart home devices suffers from an inability to adapt to environmental changes and user emotional fluctuations. In view of this, this application provides an echo cancellation method, apparatus, electronic device, storage medium, and program product, which addresses at least one of the aforementioned problems.
[0034] Figure 1 This is one of the flowcharts illustrating the echo cancellation method provided in this application, such as... Figure 1 As shown, the echo cancellation method includes, but is not limited to, steps 101 to 102.
[0035] It should be noted that the entity performing the echo cancellation method provided in this application can be any of the following: smart home devices, servers, and computer devices, such as smart speakers, smart TVs, mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, Ultra-Mobile Personal Computers (UMPCs), netbooks, or Personal Digital Assistants (PDAs).
[0036] For example, the echo cancellation method provided in this application is implemented by a home smart device with a lightweight, large-scale model that adapts to edge computing power.
[0037] To better illustrate the echo cancellation method provided in this application, the echo cancellation methods provided in the following embodiments are all described with the execution subject being a home smart device.
[0038] Step 101: Adjust the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal.
[0039] Near-end mixed audio signals are audio signals picked up near the home smart device, which are a mixture of near-end speech signals (such as human voices), echo audio signals, and ambient noise. The echo audio signal is the audio signal after acoustic reflection of the reference audio signal played by the speaker. The reference audio signal is the original audio signal played by the speaker.
[0040] Emotional state is the category of user emotion identified by analyzing acoustic features (such as pitch, speech rate, volume, spectral energy distribution, formant changes, etc.) in near-end mixed audio signals, such as anger, sadness, happiness, calmness, etc.; emotional intensity value indicates the intensity of the emotional state, such as very happy, slightly sad, etc.
[0041] Echo estimation filters are filters used to estimate and simulate echo paths, providing a basis for echo cancellation.
[0042] It should be noted that the echo estimation filter can be implemented using any of the classic adaptive algorithms such as NLMS and RLS, and this application does not impose any restrictions on this.
[0043] The filter parameters of an echo estimation filter include, but are not limited to, any one of the following: filter convergence direction, filter step size, filter order, and filter coefficients.
[0044] Specifically, smart home devices acquire near-end mixed audio signals using built-in microphones or microphone arrays (two or more microphones) at a preset sampling rate (e.g., 16kHz, 16-bit quantization). A lightweight, large-scale model adapted to edge computing power is then used to analyze the acoustic features of the near-end mixed audio signals, obtaining their emotional state and intensity values. Furthermore, based on the emotional state and intensity values of the near-end mixed audio signals, the filter parameters of the echo estimation filter are adjusted.
[0045] Step 102: Determine the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the adjusted filter parameters.
[0046] Specifically, in addition to acquiring the near-end mixed audio signal, a reference audio signal played by the speaker is also acquired. At the same time, based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the near-end mixed audio signal is echo-cancelled using relevant echo cancellation technology to obtain the echo-cancelled near-end audio signal.
[0047] Optionally, determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters includes: Based on the echo estimation filter adjusted by the filter parameters and the reference audio signal, an echo estimation signal is determined; based on the echo estimation signal and the near-end mixed audio signal, a near-end audio signal after echo cancellation is determined.
[0048] More specifically, the reference audio signal is processed using an echo estimation filter with adjusted filter parameters to estimate the echo estimation signal representing the echo path. The echo estimation signal is then used to cancel the echoes in the near-end mixed audio signal, resulting in the echo-cancelled near-end audio signal.
[0049] The echo cancellation method provided in this application considers the influence of emotions on echo cancellation, perceives the user's emotional state and intensity in real time, and then dynamically adjusts the filter parameters of the echo estimation filter based on the real-time emotional state and intensity. Echo cancellation is performed based on the echo estimation filter adjusted according to the convergence direction and step size. This method can adapt to environmental changes and user emotional fluctuations, and avoid echo residue or echo cancellation speech distortion problems in emotional fluctuation scenarios.
[0050] Meanwhile, by adapting lightweight large-scale home smart devices to edge computing power, echo cancellation can be implemented without relying on cloud computing power and high-end dedicated echo cancellation chips, reducing hardware costs and improving the clarity of voice interaction, device reliability, and user experience in dynamic home environments.
[0051] Furthermore, by addressing the echo interference problem in scenarios involving emotional fluctuations and environmental changes, the clarity and fluency of voice interaction in smart home devices can be improved, reducing the rate of command misrecognition caused by echo issues, enhancing user stickiness, and improving the user experience of smart home devices.
[0052] Optionally, when multiple smart home devices are working simultaneously, each working smart home device can communicate with each other via a local area network to share reference audio signals and emotion perception results (emotional state and / or emotion intensity values) to build a distributed dynamic echo cancellation system, thereby achieving cross-device collaborative echo cancellation and avoiding echo interference between multiple devices.
[0053] Based on the above embodiments, as an optional embodiment, the filter parameters include the filter convergence direction and the filter step size; The step of adjusting the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal includes: Based on the emotional state, the convergence direction of the filter is adjusted; The filter step size is adjusted based on the emotion intensity value.
[0054] Here, the filter convergence direction is the direction of the optimization path followed by the filter coefficients of the echo estimation filter when updating; the filter step size is the adjustment range of the echo estimation filter for each parameter update.
[0055] Specifically, in the process of adjusting the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal, on the one hand, the convergence direction of the echo estimation filter is adjusted according to the emotional state of the near-end mixed audio signal; on the other hand, the step size of the echo estimation filter is adjusted according to the emotional intensity value of the near-end mixed audio signal.
[0056] Generally speaking, there is a positive correlation between the emotion intensity value and the filter step size; that is, the larger the emotion intensity value, the larger the filter step size; and the smaller the emotion intensity value, the smaller the filter step size.
[0057] Understandably, by pre-determining an emotion-convergence mapping table between emotional states and filter convergence directions, the corresponding filter convergence direction can be directly determined based on this table when adjusting the filter convergence direction according to the emotional state. The emotion-convergence mapping table represents the mapping relationship between emotional states and filter convergence directions.
[0058] For example, in the case of "anger", the sound pressure level of the near-end mixed audio signal is high and the speech energy fluctuates greatly. The convergence direction of the filter that prioritizes the suppression of high-intensity echo signals can be pre-associated so that when adjusting the filter parameters of the echo estimation filter, the convergence direction of the filter is adjusted to prioritize the suppression of high-intensity echo signals.
[0059] For example, in the case of "sadness", the speech rate of the near-end mixed audio signal is slow and the speech energy is low. The convergence direction of the filter that prioritizes the preservation of low-energy target speech can be pre-associated. When adjusting the filter parameters of the echo estimation filter, the convergence direction of the filter is adjusted to prioritize the preservation of low-energy target speech, so as to avoid excessive suppression.
[0060] Optionally, adjusting the filter convergence direction based on the emotional state includes: determining a target convergence direction associated with the emotional state from an emotion-convergence mapping table based on the emotional state; and adjusting the filter convergence direction based on the target convergence direction.
[0061] The echo cancellation method provided in this application perceives the user's emotional state and intensity in real time, and then adjusts the convergence direction of the echo estimation filter according to the real-time emotional state and dynamically adjusts the step size of the echo estimation filter according to the real-time emotional intensity value. Echo cancellation is performed based on the echo estimation filter adjusted according to the convergence direction and step size. This method can adapt to environmental changes and user emotional fluctuations, and avoid echo residue or echo cancellation speech distortion problems in emotional fluctuation scenarios.
[0062] Based on the above embodiments, as an optional embodiment, adjusting the filter step size based on the emotion intensity value includes: The filter step size is adjusted based on the filter base step size, the emotion influence coefficient, and the emotion intensity value. The emotion influence coefficient is determined based on the emotional state.
[0063] The emotional impact coefficient is a coefficient that represents the influence of emotional state on echo cancellation, and it is determined based on emotional state.
[0064] For example, when the emotional state is anger or joy, the emotional influence coefficient is determined as the first influence coefficient; when the emotional state is sadness or fear, the emotional influence coefficient is determined as the second influence coefficient; when the emotional state is calm, the emotional influence coefficient is determined as the third influence coefficient; the absolute value of the first influence coefficient is greater than the absolute value of the second influence coefficient, and the absolute value of the second influence coefficient is greater than the absolute value of the third influence coefficient.
[0065] More specifically, the first influence coefficient is 0.5, the second influence coefficient is -0.3, and the third influence coefficient is 0.
[0066] Understandably, by pre-determining an emotion-influence coefficient mapping table between emotional states and emotional influence coefficients, the target influence coefficient can be determined from the emotion-influence coefficient mapping table based on the emotional state when adjusting the filter step size. This target influence coefficient can then be used to adjust the filter step size. The emotion-influence coefficient mapping table represents the mapping relationship between emotional states and emotional influence coefficients.
[0067] Optionally, the calculation formula for adjusting the filter step size based on the filter base step size, the emotion influence coefficient, and the emotion intensity value is as follows: ; in, This is the filter step size; This is the basic step size of the filter; The coefficient representing the influence of emotions; This represents the intensity of emotion, ranging from 0 to 1.
[0068] Taking a filter with a default base step size of 0.01, an emotional state of "anger," and an emotional intensity value of 0.8 (strong) as an example, the emotional influence coefficient is determined to be 0.5 based on the emotional state. Therefore, the filter step size to be adjusted is calculated to be 0.01 × (1 + 0.5 × 0.8) = 0.014. Increasing the basic step size of the filter accelerates the convergence speed of the echo estimation filter.
[0069] Taking a filter with a default base step size of 0.01, an emotional state of "calm," and an emotional intensity value of 0.2 (moderate) as an example, the emotional influence coefficient is set to 0 based on the emotional state. Therefore, the filter step size to be adjusted is calculated to be 0.01 × (1 + 0 × 0.2) = 0.01. With the filter's base step size remaining constant, filter stability is guaranteed.
[0070] The echo cancellation method provided in this application establishes a correlation mechanism between emotion perception results and echo estimation filters, thereby enabling real-time perception of the user's emotional state and intensity. It dynamically adjusts the core parameters of the echo estimation filter, namely the convergence direction and step size, to perform echo cancellation based on the echo estimation filter adjusted by the convergence direction and step size. This method can adapt to environmental changes and user emotional fluctuations, avoiding echo residue or echo cancellation speech distortion problems in emotional fluctuation scenarios.
[0071] Based on the above embodiments, as an optional embodiment, after determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the method further includes: Based on user interaction feedback data within a preset iteration period, the mapping relationship between the emotion influence coefficient and the emotion state is adjusted according to the preset iteration period.
[0072] The preset iteration cycle is a pre-defined cycle that updates the mapping relationship between the emotion influence coefficient and the emotion state based on user interaction feedback data.
[0073] For example, the preset iteration cycle is 1 hour, 2 hours, 12 hours or 24 hours, etc.
[0074] User interaction feedback data is determined based on user interaction feedback regarding the near-end audio signal after echo cancellation.
[0075] Optionally, user interaction feedback data includes, but is not limited to, at least one of the following: false wake-up rate of smart home devices, voice command recognition accuracy, number of times the user manually adjusts the volume, and number of times the user repeats the pronunciation.
[0076] The false wake-up rate is the ratio between the number of times a smart home device is mistakenly woken from sleep / standby mode due to residual echoes or noise during device wake-up operations based on echo-cancelled near-end audio signals, and the total number of times the user actively tries to wake up the device. The false wake-up rate reflects the situation where the device is accidentally triggered when emotional fluctuations occur.
[0077] Voice command recognition accuracy is the ratio between the number of correctly recognized commands and the total number of user-pronounced commands when a smart home device recognizes voice commands based on echo-cancelled near-end audio signals. Voice command recognition accuracy reflects the correct recognition of voice commands after echo cancellation.
[0078] Specifically, Figure 2 This is the second flowchart of the echo cancellation method provided in this application, as shown below. Figure 2 As shown, after eliminating the echo in the near-end mixed audio signal and obtaining the near-end audio signal, the home smart device adjusts the mapping relationship between the emotion influence coefficient and the emotional state based on the user interaction feedback data of the echo-eliminated near-end audio signal within a preset iteration cycle.
[0079] By adjusting the mapping relationship between the emotion influence coefficient and the emotion state, the result of the step of "adjusting the filter step size based on the filter base step size, emotion influence coefficient and emotion intensity value" can be adjusted, so that the echo estimation filter after the filter parameters are adjusted can better adapt to environmental changes and user emotional fluctuations.
[0080] The echo cancellation method provided in this application improves the echo cancellation effect by constructing a closed-loop system of "emotion perception - filter parameter adjustment - interactive feedback - emotion influence coefficient and emotion state mapping adjustment".
[0081] Based on the above embodiments, as an optional embodiment, before adjusting the filter parameters of the echo estimation filter according to the emotional state and emotional intensity value of the near-end mixed audio signal, the method further includes: Based on the near-end mixed audio signal, determine the multimodal input vector; The multimodal input vector is input into a pre-trained emotion perception model to obtain the emotion probability distribution and the emotion intensity value output by the emotion perception model; the emotion probability distribution includes emotion categories and their probabilities. The emotional state is determined based on the emotion category with the highest probability.
[0082] The multimodal input vector is the input vector determined based on the multimodal input data obtained after time-frequency conversion and automatic speech recognition (ASR) conversion of the near-end mixed audio signal.
[0083] Optionally, the emotion categories adopt Ekman's six basic emotions classification, specifically including joy, anger, sadness, fear, surprise, and calmness.
[0084] Specifically, in combination Figure 2 As shown, before adjusting the filter parameters of the echo estimation filter based on the emotional state and emotional intensity values of the near-end mixed audio signal, the smart home device needs to perform acoustic feature analysis on the acquired near-end mixed audio signal and determine the emotional state and emotional intensity values based on the acoustic feature analysis results.
[0085] Specifically, an emotion perception model is pre-built based on a lightweight edge-side large model, and the pre-training of the emotion perception model is completed.
[0086] In actual inference, time-frequency conversion and automatic speech recognition are performed on the near-end mixed audio signal to obtain multimodal input data, including acoustic features of the near-end mixed audio signal and text sequence (text features). The multimodal input data is then encoded to obtain a multimodal input vector.
[0087] The multimodal input vector is fed into a pre-trained emotion perception model to obtain the emotion probability distribution and emotion intensity value output by the model. The emotion probability distribution includes each emotion category and its corresponding probability; the emotion intensity value ranges from 0 to 1, with higher values indicating stronger emotions. Finally, the emotion category with the highest probability is determined as the emotional state.
[0088] It is understandable that the emotion perception model is pre-trained based on multiple training samples; each training sample includes a multimodal training vector sample and its corresponding emotion category label and emotion intensity value label; the vector size and encoding method are the same between the multimodal training vector sample and the multimodal input vector.
[0089] Optionally, the emotion perception model on the smart home device side adopts a lightweight large model based on the MiniLLaMA-2B architecture, with parameters compressed to a preset size (e.g., 500M). The model backbone network is frozen through Low-Rank Adaptation (LoRA) technology, and only the emotion perception-related adapter layers with a parameter ratio less than a preset ratio (e.g., 5%) are fine-tuned to reduce the computational power requirement of the device side inference, so that the CPU utilization rate during inference is less than a preset CPU utilization rate threshold (e.g., 30%) and the memory utilization rate is less than a preset memory utilization rate threshold (e.g., 512MB).
[0090] Alternatively, the emotion perception model on the smart home device side can also be implemented based on Phi-2 or Qwen-1.8B.
[0091] Many solutions that incorporate emotion computing into smart home devices rely on large cloud models for emotion analysis, which presents data transmission delays (typically 100-300ms) and privacy risks. On the other hand, if large models are deployed directly on the device side, the computing power (such as CPU clock speeds of 1-2GHz) and memory (often 1-4GB) of smart home devices make it difficult to achieve real-time emotion perception and echo cancellation coordination.
[0092] By employing the aforementioned model pruning and LoRA fine-tuning methods to implement an emotion perception model on the edge of home smart devices, the edge inference latency can be reduced to less than 30ms, meeting the industry requirement of less than 100ms latency for real-time voice interaction. At the same time, the model's memory usage is less than 512MB, and the CPU utilization rate is less than 30%, making it basically compatible with mainstream mid-to-low-end home smart devices. It has better edge deployment performance and realizes an emotion perception solution that meets the computing power and memory constraints of edge devices.
[0093] The echo cancellation method provided in this application, by deploying a lightweight large-scale emotion perception model on the edge of a smart home device, can quickly and in real-time realize the real-time emotion perception of the user's voice, obtain the emotional state and emotional intensity value of the near-end mixed audio signal, and adjust the filter parameters of the echo estimation filter. Echo cancellation is then performed based on the parameter-adjusted echo estimation filter. The entire echo cancellation process is performed on the edge, without the need for data transmission to the cloud, thus avoiding data transmission delays and privacy leakage risks. It can also adapt to environmental changes and user emotional fluctuations, avoiding echo residue or echo cancellation speech distortion problems in emotional fluctuation scenarios.
[0094] Optionally, determining the multimodal input vector based on the near-end mixed audio signal includes: determining the multimodal input vector based on the near-end mixed audio signal and user visual data.
[0095] User visual data refers to user video or image data captured by cameras installed in smart home devices. Its types include, but are not limited to, facial expression data and user behavior data.
[0096] Optionally, determining the multimodal input vector based on the near-end mixed audio signal includes: determining the multimodal input vector based on the near-end mixed audio signal and user physiological data.
[0097] For example, user physiological data includes heart rate, blood pressure, and other data obtained using smart bracelets connected to smart home devices.
[0098] Optionally, determining the multimodal input vector based on the near-end mixed audio signal includes: determining the multimodal input vector based on the near-end mixed audio signal, user visual data, and user physiological data.
[0099] In addition to achieving emotion perception based on a single voice modality, further expanding and constructing a multimodal emotion perception system that integrates "voice + vision + physiology" can improve the accuracy of identifying emotional states and emotional intensity values.
[0100] Based on the above embodiments, as an optional embodiment, determining the multimodal input vector based on the near-end mixed audio signal includes: The multimodal input vector is determined based on the acoustic features, timestamps, and text sequences corresponding to the near-end mixed audio signals.
[0101] Optionally, the acoustic features include, but are not limited to, at least one of the following features of the near-end mixed audio signal: 13-dimensional Mel-Frequency Cepstral Coefficients (MFCC), spectral centroid, fundamental frequency (F0).
[0102] Specifically, in combination Figure 2 As shown, when determining the multimodal input vector, home smart devices extract acoustic features such as MFCC, spectral centroid, and fundamental frequency of the near-end mixed audio signal.
[0103] For example, the near-end mixed audio signal is pre-emphasized to enhance its high-frequency components; the pre-emphasized near-end mixed audio signal is divided into frames with a frame length of 20ms and a frame shift of 10ms; the framed near-end mixed audio signal is windowed according to the Hanning window; the windowed near-end mixed audio signal is subjected to a Short-Time Fourier Transform (STFT) to convert the time-domain signal into frequency-domain features, and acoustic features are extracted from the frequency-domain features.
[0104] On the other hand, smart home devices use their built-in ASR (Automatic Speech Recognition) module to convert near-end mixed audio signals into corresponding text sequences.
[0105] On the other hand, smart home devices record voice interaction timestamps through their built-in clock modules.
[0106] Finally, home smart devices simultaneously encode the acoustic features, timestamps, and text sequences of the near-end mixed audio signals, for example, by concatenating the acoustic features, timestamps, and text sequences together, to obtain a multimodal input vector.
[0107] As an optional embodiment, determining the multimodal input vector based on the acoustic features of the near-end mixed audio signal, the timestamp, and the text sequence corresponding to the near-end mixed audio signal includes: determining a dialogue context sequence based on the timestamp and the text sequence; and determining the multimodal input vector based on the dialogue context sequence and the acoustic features.
[0108] Optionally, determining the dialogue context sequence based on the timestamp and the text sequence includes: encoding the text sequence and the timestamp using a context semantic encoding module to determine the dialogue context sequence; the context semantic encoding module is built based on a Transformer encoder.
[0109] The Transformer encoder can capture the logic of the conversation history and reflect the emotional changes in the user's voice.
[0110] For example, a user's historical text sequence at a historical timestamp indicates "turn up the volume," while the current text sequence at the current timestamp indicates "too noisy." This means that the user previously said "turn up the volume" but now says "too noisy," reflecting changes in the user's emotions and needs.
[0111] Optionally, determining the multimodal input vector based on the dialogue context sequence and the acoustic features includes: fusing the dialogue context sequence and the acoustic features based on a cross-attention mechanism to obtain the multimodal input vector.
[0112] By fusing dialogue context sequences and acoustic features through a cross-attention mechanism, the connection between tone and semantics can be strengthened.
[0113] Optionally, determining the dialogue context sequence based on the timestamp and the text sequence includes: obtaining a historical dialogue context sequence based on historical timestamps and historical text sequences; obtaining a current dialogue context sequence based on the current timestamp and current text sequence; and determining the dialogue context sequence based on the historical dialogue context sequence and the current dialogue context sequence.
[0114] On the one hand, most deep learning echo cancellation technologies only focus on the features of the audio itself, without combining the contextual semantics and emotional state of the user's speech. On the other hand, although some smart home devices integrate basic affective computing modules and use affective computing combined with voice interaction technology to extract acoustic features such as fundamental frequency, speech rate, and energy of the speech signal, and combine them with simple rules to judge the user's emotional state (such as calm, happy, angry), and use the emotional state to optimize the voice interaction response (such as emotional voice broadcasting), they are not linked with the echo cancellation function, and the echo cancellation still uses a filtering scheme with fixed parameters.
[0115] The echo cancellation method provided in this application fully considers the contextual semantics and emotional state of the user's speech during emotion perception. It performs emotion perception based on a multimodal input vector consisting of "acoustic features + text sequence + timestamp" obtained by processing the near-end mixed audio signal. The result is composed of the emotional state and emotional intensity value of the near-end mixed audio signal. Based on this, the filter parameters of the echo estimation filter are dynamically adjusted, which improves the adaptability of echo cancellation to environmental changes and user emotional fluctuations.
[0116] Based on the above embodiments, as an optional embodiment, after determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the method further includes: Based on user interaction feedback data within a preset iteration period, the emotion category classification weights of the emotion perception model are adjusted according to the preset iteration period.
[0117] The preset iteration cycle is a pre-defined cycle for updating the emotion category classification weights of the emotion perception model based on user interaction feedback data.
[0118] For example, the preset iteration cycle is 1 hour, 2 hours, 12 hours or 24 hours, etc.
[0119] User interaction feedback data is determined based on user interaction feedback regarding the near-end audio signal after echo cancellation.
[0120] Optionally, user interaction feedback data includes, but is not limited to, at least one of the following: false wake-up rate of smart home devices, voice command recognition accuracy, number of times the user manually adjusts the volume, and number of times the user repeats the pronunciation.
[0121] Specifically, in combination Figure 2As shown, after eliminating the echo in the near-end mixed audio signal and obtaining the near-end audio signal, the home smart device adjusts the emotion category classification weights of the emotion perception model according to the user interaction feedback data of the echo-eliminated near-end audio signal within a preset iteration cycle. For example, the emotion classification weights of the emotion perception model (i.e., the edge-side large model) are fine-tuned using the mini-batch gradient descent method.
[0122] By adjusting the emotion category classification weights of the emotion perception model, the step of "inputting the multimodal input vector into the pre-trained emotion perception model to obtain the emotion probability distribution and emotion intensity value output by the emotion perception model" can be adjusted, so that the emotion perception results are more adaptable to environmental changes and user emotional fluctuations.
[0123] The echo cancellation method provided in this application improves the echo cancellation effect by constructing a closed-loop system of "emotion perception - filter parameter adjustment - interactive feedback - emotion category classification weight adjustment".
[0124] In one embodiment, after determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the method further includes: Based on user interaction feedback data within a preset iteration period, the mapping relationship between the emotion influence coefficient and the emotion state is adjusted according to the preset iteration period, and the emotion category classification weight of the emotion perception model is adjusted according to the preset iteration period.
[0125] Experimental data shows that the parameters of the echo estimation filter in relevant echo cancellation schemes are always fixed. For example, when people walk around or doors and windows are opened and closed in the home environment, the echo cancellation effect is reduced by about 20%. The echo cancellation scheme provided in this application adjusts the model and filter parameters in real time through a "feedback-iteration" mechanism, and the effect attenuation rate is controlled within 5%, which significantly improves stability and has stronger dynamic environmental adaptability.
[0126] Optionally, the step of adjusting the mapping relationship between the emotion influence coefficient and the emotion state according to the preset iteration period based on user interaction feedback data within the preset iteration period, and adjusting the emotion category classification weights of the emotion perception model according to the preset iteration period, includes: Based on user interaction feedback data and family scenarios within a preset iteration period, the mapping relationship between the emotion influence coefficient and the emotion state is adjusted according to the preset iteration period, and the emotion category classification weights of the emotion perception model are adjusted according to the preset iteration period.
[0127] Family settings include, but are not limited to, children's rooms, elderly people's rooms, and living rooms.
[0128] For example, in a family setting such as a children's room, the acoustic feature extraction weights and filter step size range of the model are adjusted to take into account the characteristics of children's speech, which have obvious high-frequency features and frequent emotional fluctuations.
[0129] For example, in a home setting such as an elderly person's room, a strategy for preserving low-frequency speech can be optimized to address the characteristics of elderly people who speak slowly and have low speech energy.
[0130] By considering different, segmented family scenarios when optimizing emotion perception and filter parameters, scenario-based customized optimization can be achieved.
[0131] Based on the above embodiments, as an optional embodiment, before determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the following further step is taken: The model weights of the deep mask generation model are adjusted according to the emotional state; the deep mask generation model is determined based on a convolutional neural network (CNN). The process of determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters includes: The echo estimation signal is determined based on the reference audio signal and the echo estimation filter after the filter parameters are adjusted. The echo estimation signal and the near-end mixed audio signal are input into the depth mask generation model after the model weight adjustment to obtain the echo-cancelled near-end audio signal output by the depth mask generation model.
[0132] Overall, the echo cancellation method provided in this embodiment adopts a hybrid architecture of "adaptive filtering + depth mask generation". It not only uses the emotion perception results to adjust the filter parameters of the echo estimation filter, but also uses the emotion perception results to adjust the depth mask generation steps.
[0133] Optionally, the depth mask generation model is implemented based on a lightweight convolutional neural network including three CNN layers.
[0134] Specifically, a deep mask generation model is pre-built based on a convolutional neural network, and the pre-training of the deep mask generation model is completed.
[0135] In actual echo cancellation, on the one hand, the filter parameters of the echo estimation filter in the hybrid architecture are adjusted based on the emotional state and emotional intensity values in the emotion perception results; on the other hand, the model weights of the depth mask generation model are adjusted based on the emotional state in the emotion perception results to optimize the depth mask.
[0136] For example, when the emotional state is anger, the masking suppression strength in the high-frequency range (2-4kHz) is increased to reduce echo spikes. Similarly, when the emotional state is sadness, the masking retention strength in the low-frequency range (200-500Hz) is increased to avoid voice distortion and a low-pitched tone.
[0137] After dynamically adjusting the filter parameters of the echo estimation filter and the model weights of the depth mask generation model, the reference audio signal is processed using the echo estimation filter with adjusted filter parameters to estimate the echo estimation signal representing the echo path. The echo estimation signal and the near-end mixed audio signal are then input into a pre-trained depth mask generation model whose model weights have been adjusted based on real-time emotion perception results. The echo-cancelled near-end audio signal output by the depth mask generation model completes the echo cancellation process.
[0138] It is understandable that the pre-training of the depth mask generation model is based on multiple training signal samples; each training signal sample includes an input signal training sample and its corresponding echo estimation signal label; the input signal training sample consists of an echo estimation signal sample and a near-end mixed audio signal sample.
[0139] The echo cancellation method provided in this application adopts a hybrid architecture of "adaptive filtering + depth mask generation" during echo cancellation. It not only dynamically and in real time adjusts the filter parameters of the echo estimation filter using the emotion perception results, but also dynamically and in real time adjusts the model weights of the depth mask generation using the emotion perception results. This enables the echo cancellation to adapt to environmental changes and user emotional fluctuations, avoiding echo residue or echo cancellation speech distortion problems in emotional fluctuation scenarios.
[0140] Combination Figure 2 As shown, the echo cancellation method provided in this application is a dynamic echo cancellation method for home smart devices based on edge-side large model emotion perception. It constructs a four-layer architecture within the home smart device, consisting of a "multimodal data acquisition layer, edge-side large model emotion perception layer, dynamic adjustment layer, and effect feedback layer." The layers interact with each other through a data bus, forming a closed-loop collaborative mechanism to achieve dynamic collaborative optimization of emotion perception and echo cancellation.
[0141] The multimodal data acquisition layer is responsible for acquiring voice signals and device status data. In the multimodal data acquisition layer, the microphone array built into the smart home device acquires the near-end mixed audio signal, acquires the reference audio signal played by the speaker, the clock module records the timestamp, extracts acoustic features from the near-end mixed audio signal, and the ASR module converts the near-end mixed audio signal into a text sequence.
[0142] The edge-side large-scale model's emotion perception layer achieves real-time emotion recognition and semantic understanding based on collected data. In this layer, a contextual semantic encoding module based on a Transformer encoder encodes the text sequence and timestamps to determine the dialogue context sequence. Then, an acoustic emotion feature fusion module based on a cross-attention mechanism fuses the dialogue context sequence and acoustic features to obtain a multimodal input vector. This multimodal input vector is fed into a pre-trained emotion perception model to obtain the emotion probability distribution and emotion intensity value output by the model. The emotion category with the highest probability is determined as the emotion state, resulting in an emotion perception result that includes both the emotion state and the emotion intensity value.
[0143] The dynamic adjustment layer optimizes the echo cancellation filter parameters and depth mask generation model weights based on emotion perception results. In this layer, the parameter adjustment unit adjusts the convergence direction of the echo estimation filter based on the emotion state, the filter step size based on the emotion intensity value, and the model weights of the depth mask generation model based on the emotion state. The echo estimation filter, adjusted by the convergence direction and step size, is then used to process the reference audio signal to estimate the echo estimation signal representing the echo path. Finally, the echo estimation signal and the near-end mixed audio signal are input into the depth mask generation model with adjusted model weights to obtain the echo-cancelled near-end audio signal output by the depth mask generation model, completing the echo cancellation process.
[0144] The effect feedback layer optimizes the performance of the depth mask generation model and filters based on user interaction data feedback. In this layer, the user interaction data acquisition unit collects user interaction feedback data such as false wake-up rate, voice command recognition accuracy, number of times the user manually adjusts the volume, and number of times the user repeats the pronunciation. The model and parameter iteration unit, based on the user interaction feedback data within a preset iteration period, uses mini-batch gradient descent to adjust the mapping relationship between the emotion influence coefficient and the emotional state according to the preset iteration period, and adjusts the emotion category classification weights of the emotion perception model according to the preset iteration period.
[0145] Compared to deep learning solutions that rely solely on audio signal features for echo cancellation, introducing emotion perception into a large edge model addresses the issue of abrupt changes in audio features caused by emotional fluctuations, achieving deep synergy between emotion perception and echo cancellation. Experimental data shows that in a home environment, when a user's emotion changes from calm to anger, the existing solution retains approximately 15% of the echo residual energy, while this solution reduces it to below 5%; when the emotion changes from calm to sadness, the existing solution achieves approximately 8% target speech distortion, while this solution reduces it to below 3%.
[0146] Figure 3 This is a schematic diagram of the echo cancellation device provided in this application, as shown below. Figure 3 As shown, the echo cancellation device includes, but is not limited to, a parameter adjustment module 301 and a signal echo cancellation module 302.
[0147] The parameter adjustment module 301 is used to adjust the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal.
[0148] The signal echo cancellation module 302 is used to determine the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the filter parameters adjusted.
[0149] It should be noted that the echo cancellation device provided in this application can perform the echo cancellation method described in any of the above embodiments during actual operation, which will not be elaborated in this embodiment.
[0150] The echo cancellation device provided in this application considers the influence of emotions on echo cancellation, perceives the user's emotional state and intensity in real time, and then dynamically adjusts the filter parameters of the echo estimation filter based on the real-time emotional state and intensity. Echo cancellation is performed based on the echo estimation filter adjusted by the convergence direction and step size. This device can adapt to environmental changes and user emotional fluctuations, and avoid echo residue or echo cancellation speech distortion problems in emotional fluctuation scenarios.
[0151] Optionally, the parameter adjustment module further includes a convergence direction adjustment module and a step size adjustment module.
[0152] The convergence direction adjustment module is used to adjust the convergence direction of the filter based on the emotional state.
[0153] The step size adjustment module is used to adjust the filter step size based on the emotion intensity value.
[0154] Optionally, the step size adjustment module is further configured to adjust the filter step size based on the filter base step size, the emotion influence coefficient, and the emotion intensity value; the emotion influence coefficient is determined based on the emotion state.
[0155] Optionally, the echo cancellation device further includes a mapping relationship adjustment module.
[0156] The mapping relationship adjustment module is used to adjust the mapping relationship between the emotion influence coefficient and the emotion state according to the preset iteration period based on user interaction feedback data within the preset iteration period.
[0157] Optionally, the echo cancellation device further includes an emotion perception module.
[0158] The emotion perception module further includes: a vector determination module, a model processing module, and a state determination module.
[0159] The vector determination module is used to determine the multimodal input vector based on the near-end mixed audio signal.
[0160] The model processing module is used to input the multimodal input vector into a pre-trained emotion perception model to obtain the emotion probability distribution and the emotion intensity value output by the emotion perception model; the emotion probability distribution includes emotion categories and their probabilities.
[0161] The state determination module is used to determine the emotional state based on the emotion category with the highest probability.
[0162] Optionally, the vector determination module is further configured to determine the multimodal input vector based on the acoustic features of the near-end mixed audio signal, the timestamp, and the text sequence corresponding to the near-end mixed audio signal.
[0163] Optionally, the echo cancellation device further includes a classification weight adjustment module.
[0164] The classification weight adjustment module is used to adjust the emotion category classification weights of the emotion perception model according to the preset iteration period based on user interaction feedback data within the preset iteration period.
[0165] Optionally, the echo cancellation device further includes a model weight adjustment module.
[0166] The model weight adjustment module is used to adjust the model weights of the depth mask generation model according to the emotional state; the depth mask generation model is built based on a convolutional neural network.
[0167] The signal echo cancellation module includes: The echo estimation module is used to determine the echo estimation signal based on the reference audio signal and the echo estimation filter after the filter parameters are adjusted. The model cancellation module is used to input the echo estimation signal and the near-end mixed audio signal into the depth mask generation model after the model weight adjustment, so as to obtain the echo-cancelled near-end audio signal output by the depth mask generation model.
[0168] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call logical instructions in the memory 430 to execute the echo cancellation method provided in any of the above embodiments. The echo cancellation method includes, but is not limited to, the following steps: adjusting the filter parameters of the echo estimation filter according to the emotional state and emotional intensity value of the near-end mixed audio signal; and determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the adjusted filter parameters.
[0169] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0170] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the echo cancellation method provided in any of the above embodiments. The echo cancellation method includes, but is not limited to, the following steps: adjusting the filter parameters of the echo estimation filter according to the emotional state and emotional intensity value of the near-end mixed audio signal; and determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the adjusted filter parameters.
[0171] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the echo cancellation method provided in any of the above embodiments. The echo cancellation method includes, but is not limited to, the following steps: adjusting the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal; and determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the adjusted filter parameters.
[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An echo cancellation method, characterized in that, include: Adjust the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal; Based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the echo-cancelled near-end audio signal is determined.
2. The echo cancellation method according to claim 1, characterized in that, The filter parameters include the filter convergence direction and the filter step size; The step of adjusting the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal includes: Based on the emotional state, the convergence direction of the filter is adjusted; The filter step size is adjusted based on the emotion intensity value.
3. The echo cancellation method according to claim 2, characterized in that, Adjusting the filter step size based on the emotion intensity value includes: The filter step size is adjusted based on the filter base step size, the emotion influence coefficient, and the emotion intensity value. The emotion influence coefficient is determined based on the emotional state.
4. The echo cancellation method according to claim 3, characterized in that, After determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the method further includes: Based on user interaction feedback data within a preset iteration period, the mapping relationship between the emotion influence coefficient and the emotion state is adjusted according to the preset iteration period.
5. The echo cancellation method according to claim 1, characterized in that, Before adjusting the filter parameters of the echo estimation filter based on the emotional state and emotional intensity values of the near-end mixed audio signal, the method further includes: Based on the near-end mixed audio signal, determine the multimodal input vector; The multimodal input vector is input into a pre-trained emotion perception model to obtain the emotion probability distribution and the emotion intensity value output by the emotion perception model; the emotion probability distribution includes emotion categories and their probabilities. The emotional state is determined based on the emotion category with the highest probability.
6. The echo cancellation method according to claim 5, characterized in that, The determination of the multimodal input vector based on the near-end mixed audio signal includes: The multimodal input vector is determined based on the acoustic features, timestamps, and text sequences corresponding to the near-end mixed audio signals.
7. The echo cancellation method according to claim 5, characterized in that, After determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the method further includes: Based on user interaction feedback data within a preset iteration period, the emotion category classification weights of the emotion perception model are adjusted according to the preset iteration period.
8. The echo cancellation method according to claim 1, characterized in that, Before determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters, the method further includes: The model weights of the depth mask generation model are adjusted according to the emotional state; the depth mask generation model is built based on a convolutional neural network. The process of determining the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with adjusted filter parameters includes: The echo estimation signal is determined based on the reference audio signal and the echo estimation filter after the filter parameters are adjusted. The echo estimation signal and the near-end mixed audio signal are input into the depth mask generation model after the model weight adjustment to obtain the echo-cancelled near-end audio signal output by the depth mask generation model.
9. An echo cancellation device, characterized in that, include: The parameter adjustment module is used to adjust the filter parameters of the echo estimation filter based on the emotional state and emotional intensity value of the near-end mixed audio signal. The signal echo cancellation module is used to determine the echo-cancelled near-end audio signal based on the reference audio signal, the near-end mixed audio signal, and the echo estimation filter with the filter parameters adjusted.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the echo cancellation method as described in any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the echo cancellation method as described in any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the echo cancellation method as described in any one of claims 1 to 8.