Voice response method, device and storage medium

By acquiring acoustic performance parameters from smart devices and determining the correspondence between frequency points and gain compensation values, gain compensation processing is performed on voice signals, solving the problem of multi-device wake-up decision errors, achieving accurate identification of target devices, and improving user experience.

CN116110382BActive Publication Date: 2026-06-05BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2021-11-09
Publication Date
2026-06-05

Smart Images

  • Figure CN116110382B_ABST
    Figure CN116110382B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a voice response method, device and storage medium, the method comprising: receiving a first voice signal; performing gain compensation processing on the first voice signal according to a correspondence relationship between frequency points of a voice signal and gain compensation values obtained in advance to obtain a second voice signal; wherein the correspondence relationship is determined according to acoustic performance parameters corresponding to the same voice signal received by a first device in a current sound field environment and a preset sound field environment; determining first energy information of the second voice signal; wherein the first energy information is used to determine a target device responding to the first voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of speech signal processing technology, and in particular to a speech response method, apparatus and storage medium. Background Technology

[0002] With the development of smart devices, users are encountering more and more smart devices in scenarios such as home, work, and leisure. In scenarios involving multiple smart devices, if there are multiple smart devices with the same wake word, when the user says the wake word, multiple smart devices may respond to the wake word and be woken up, and all of them will recognize and respond to the user's subsequent voice commands.

[0003] In related technologies, in order to ensure that only one device responds, the server or local device usually determines the smart device closest to the user based on the voice energy of the received voice signal, and that smart device responds to the user, while other smart devices do not respond to the wake word.

[0004] However, in scenarios with multiple smart devices, the positions of different smart devices may be different, and the sound field environment they are in may also be different. This may affect the voice energy of the voice signal received by the smart device, leading to wake-up decision errors. In the end, the smart device that responds to the user is not the smart device closest to the user, resulting in a very bad user experience. Summary of the Invention

[0005] This disclosure provides a voice response method, apparatus, and storage medium.

[0006] According to a first aspect of the present disclosure, a voice response method is provided, comprising:

[0007] Receive the first voice signal;

[0008] Based on the pre-obtained correspondence between the frequency points of the speech signal and the gain compensation value, the first speech signal is subjected to gain compensation processing to obtain the second speech signal; wherein, the correspondence is determined based on the acoustic performance parameters corresponding to the same speech signal received by the first device in the current sound field environment and the preset sound field environment.

[0009] First energy information of the second voice signal is determined; wherein the first energy information is used to determine the target device responding to the first voice signal.

[0010] Optionally, the acoustic performance parameters include at least the frequency response;

[0011] The correspondence is obtained by calculating the difference between the first frequency response and the second frequency response; wherein, the first frequency response is the frequency response generated by the first device receiving the test audio signal in the preset sound field environment; and the second frequency response is the frequency response generated by the first device receiving the test audio signal in the current sound field environment.

[0012] Optionally, the step of performing gain compensation processing on the first speech signal according to the pre-obtained correspondence between the frequency points of the speech signal and the gain compensation value to obtain the second speech signal includes:

[0013] Perform a Fourier transform on the first speech signal to obtain a first frequency domain signal;

[0014] Based on the correspondence, the amplitude values ​​corresponding to each frequency point in the first frequency domain signal are compensated to obtain the second frequency domain signal;

[0015] The second frequency domain signal is subjected to inverse Fourier transform to obtain the second speech signal.

[0016] Optionally, the first speech signal is subjected to gain compensation processing based on the pre-obtained correspondence between frequency points and gain compensation values ​​to obtain the second speech signal:

[0017] Based on the correspondence, determine the filtering parameters and the first filter corresponding to the filtering parameters;

[0018] The method involves inputting the first speech signal into the first filter to obtain a second speech signal output by the first filter. Optionally, the method further includes:

[0019] The system receives second energy information sent by a second device, wherein the second energy information is the energy information of the voice signal generated by the second device after receiving the first voice signal and performing gain compensation processing; and determines whether the first device is the target device responding to the first voice signal based on the first energy information and the second energy information.

[0020] or,

[0021] The first energy information is sent to the second device that determines the target device.

[0022] Optionally, determining the first energy information of the second speech signal includes:

[0023] Obtain the first energy of each frame of the second speech signal;

[0024] The first energy information of the second speech signal is determined based on the ratio between the energy of the first energy of each frame of the second speech signal and the number of signal frames in the second speech signal.

[0025] According to a second aspect of the present disclosure, a voice response device is provided, comprising:

[0026] The compensation module is used to receive a first voice signal; and to perform gain compensation processing on the first voice signal according to the pre-obtained correspondence between the frequency points of the voice signal and the gain compensation value to obtain a second voice signal; wherein, the correspondence is determined according to the acoustic performance parameters corresponding to the same voice signal received by the first device in the current sound field environment and the preset sound field environment.

[0027] A determining module is used to determine first energy information of the second voice signal; wherein the first energy information is used to determine the target device responding to the first voice signal.

[0028] Optionally, the acoustic performance parameters include at least the frequency response;

[0029] The correspondence is obtained by calculating the difference between the first frequency response and the second frequency response; wherein, the first frequency response is the frequency response generated by the first device receiving the test audio signal in the preset sound field environment; and the second frequency response is the frequency response generated by the first device receiving the test audio signal in the current sound field environment.

[0030] Optionally, the compensation module is further configured to:

[0031] Perform a Fourier transform on the first speech signal to obtain a first frequency domain signal;

[0032] Based on the correspondence, the amplitude values ​​corresponding to each frequency point in the first frequency domain signal are compensated to obtain the second frequency domain signal;

[0033] The second frequency domain signal is subjected to inverse Fourier transform to obtain the second speech signal.

[0034] Optionally, the determining module is further configured to:

[0035] Based on the correspondence, determine the filtering parameters and the first filter corresponding to the filtering parameters;

[0036] The first speech signal is input to the first filter to obtain the second speech signal output by the first filter.

[0037] Optionally, the determining module is further configured to:

[0038] Receive second energy information sent by the second device, wherein the second energy information is the energy information of the voice signal generated by the second device after receiving the first voice signal and performing gain compensation processing;

[0039] Based on the first energy information and the second energy information, determine whether the first device is the target device responding to the first voice signal;

[0040] or,

[0041] The first energy information is sent to the second device that determines the target device.

[0042] Optionally, the determining module is further configured to:

[0043] Obtain the first energy of each frame of the second speech signal;

[0044] The first energy information of the second speech signal is determined based on the ratio between the energy of the first energy of each frame of the second speech signal and the number of signal frames in the second speech signal.

[0045] According to a third aspect of the present disclosure, a voice response device is provided, comprising:

[0046] processor;

[0047] Memory used to store executable instructions;

[0048] The processor is configured to, when executing executable instructions stored in the memory, implement the steps of the voice response method according to the first aspect of the present disclosure.

[0049] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of a voice response device, the voice response device is enabled to perform steps in the voice response method as described in the first aspect of the present disclosure.

[0050] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:

[0051] This embodiment of the invention determines the correspondence between frequency points and gain compensation values ​​of a voice signal based on the acoustic performance parameters corresponding to the same voice signal received by the first device in the current sound field environment and the preset sound field environment. Based on the correspondence, the gain compensation value corresponding to each frequency point of the first voice signal is determined, and gain compensation processing is performed on the first voice signal received by the first device to reduce the influence of the current sound field environment of the first device on the acoustic performance parameters of the first voice signal. Energy information is determined using the second voice signal obtained by the gain compensation processing, and the target device responding to the first voice signal is accurately determined based on the energy information. This reduces the situation where multiple smart devices respond to the first voice signal at the same time, and also reduces the judgment error caused by different sound field environments of smart devices. It accurately determines the smart device closest to the user (i.e., the user's true intention to wake it up), improving the user's voice interaction experience.

[0052] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0053] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0054] Figure 1 This is a flowchart illustrating a voice response method according to an exemplary embodiment. Figure 1 .

[0055] Figure 2 This is a flowchart illustrating a voice response method according to an exemplary embodiment. Figure 2 .

[0056] Figure 3 This is a schematic diagram of the structure of a voice response device according to an exemplary embodiment.

[0057] Figure 4 This is a block diagram illustrating a voice response device according to an exemplary embodiment. Detailed Implementation

[0058] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses consistent with some aspects of this disclosure as detailed in the appended claims.

[0059] With the development of artificial intelligence and the increasing maturity of 5G technology, more and more smart devices are being deployed in home environments, and scenarios with multiple smart devices in the home are now very common. The wake words for different smart devices can be the same or different. If the wake words are the same, when a user says the wake word, multiple smart devices with the same wake word may simultaneously respond to the user's voice command and interact with the user, which can easily cause confusion and create a very poor user experience.

[0060] To ensure that only one device responds, the various devices in the home environment first interconnect within a local area network via Wi-Fi. When the user issues a wake-up voice command, the microphone of each device picks up the user's voice signal and extracts voice features (such as voice energy) from the voice signal. These voice features are then shared via Wi-Fi connection.

[0061] Each device within the same local area network (LAN) can acquire the voice characteristics of other devices within the LAN and determine whether to respond to the wake-up voice command based on the voice characteristics; alternatively, the gateway device acquires the voice characteristics of each device within the LAN, determines the responding device from among the multiple devices within the LAN based on the voice characteristics, and controls the responding device to respond to the wake-up voice command.

[0062] However, different smart devices are in different sound field environments, which may affect the voice characteristics of the voice signals picked up by the smart devices, leading to wake-up decision errors.

[0063] For example, consider a first smart device placed on a coffee table (open sound field environment) and a second smart device placed in a recessed compartment of a TV cabinet (semi-closed sound field environment), where the distance between the first smart device and the user is less than the distance between the second smart device and the user.

[0064] Generally, the voice signal picked up by the first smart device has stronger voice energy than that picked up by the second smart device, and the first smart device is the responding device. However, because the second smart device is in a semi-enclosed sound field environment and has more transmission circuits, the voice signal picked up by the second smart device has even greater voice energy. During the wake-up decision, based on the corresponding voice energy of the first and second smart devices, the second smart device is identified as the responding device, leading to a decision error.

[0065] Based on this, the present disclosure provides a voice response method. Figure 1 This is a flowchart illustrating a voice response method according to an exemplary embodiment. Figure 1 ,like Figure 1 As shown, the method includes:

[0066] Step S101: Receive the first voice signal;

[0067] Step S102: According to the pre-obtained correspondence between the frequency points of the speech signal and the gain compensation value, the first speech signal is subjected to gain compensation processing to obtain the second speech signal; wherein, the correspondence is determined based on the acoustic performance parameters corresponding to the same speech signal received by the first device in the current sound field environment and the preset sound field environment.

[0068] Step S103: Determine the first energy information of the second voice signal; wherein the first energy information is used to determine the target device responding to the first voice signal.

[0069] In this embodiment of the disclosure, the voice response method can be applied to a first device, wherein the first device can be any electronic device with voice interaction function in the current local area network, and the first device may specifically include: smart speakers, smart TVs, smartphones or smart refrigerators and other electronic devices with voice interaction function.

[0070] In step S101, the first device acquires the user's first voice signal through the audio acquisition module.

[0071] Here, the first device includes at least: an audio acquisition module and an audio output module; the audio acquisition module may be a microphone with the function of acquiring or picking up sound; the audio output module is a speaker.

[0072] The first device can collect ambient sounds in real time through the audio acquisition module, or periodically collect ambient sounds through the audio acquisition module.

[0073] In step S102, based on the pre-obtained correspondence between the frequency points of the voice signal and the gain compensation value, the gain compensation value corresponding to each frequency point of the first voice signal received by the first device is determined. Then, based on the gain compensation value, the amplitude corresponding to each frequency point in the first voice signal is compensated to obtain the compensated second voice signal.

[0074] It's important to note that when multiple smart devices are present, they are typically placed in different locations. For example, in a smart home environment, multiple smart devices can be stored in different areas such as the living room, kitchen, and bedroom. Generally, users wake up the smart devices based on their location; that is, users usually wake up the smart device closest to them. For instance, when a user is in the living room, they will usually want to wake up the smart device located in the living room, not the one in the kitchen or bedroom.

[0075] In related technologies, the target device responding to the voice signal is typically determined from multiple smart devices within the same local area network based on the acoustic performance parameters (e.g., voice intensity) of the same voice signal received by those devices. However, considering that different smart devices may be placed in different locations, the current sound field environment in which the different smart devices are located may affect the acoustic performance parameters of the same voice signal received by the different smart devices, leading to inaccurate determination of the target device.

[0076] For example, when the first device is in a semi-enclosed sound field environment, due to the large number of reflection loops of the speech signal in the semi-enclosed sound field environment, the signal strength of the speech signal received by the first device in the semi-enclosed sound field environment is greater than the signal strength of the same speech signal received by the first device in the open sound field environment.

[0077] Therefore, in order to reduce the impact of different sound field environments on the acoustic performance parameters of the voice signal received by the smart device, the embodiments of this disclosure perform gain compensation processing on the voice signal received by the smart device.

[0078] Based on the acoustic performance parameters corresponding to the same speech signal received by the first device in the current sound field environment and the preset sound field environment, the correspondence between the frequency points of the speech signal received by the first device and the gain compensation value is determined; thereby, based on the correspondence, the gain compensation value corresponding to the first speech signal is determined, and the first speech signal is subjected to gain compensation processing.

[0079] Here, the preset sound field environment can be set according to actual needs. For example, the preset sound field environment can be an open sound field environment.

[0080] It should be noted that the gain compensation value of the first device is related to the current sound field environment. Understandably, different sound field environments correspond to different gain compensation values.

[0081] When the position of the first device changes, the acoustic performance parameters of the first device receiving voice signals in the current sound field environment can be obtained; and the gain compensation value corresponding to the current sound field environment can be determined based on the acoustic performance parameters of the first device receiving the same voice signals in the preset sound field environment and the current sound field environment.

[0082] After determining the gain compensation value corresponding to the current sound field environment, the gain compensation value can be stored in the memory of the first device so that after the first device receives the voice signal, it can retrieve the gain compensation value from the memory and perform gain compensation processing on the voice signal.

[0083] In some embodiments, prior to performing gain compensation processing on the first speech signal, the method further includes:

[0084] The first voice signal is identified to determine whether it contains the wake-up word of the first device. If the first voice signal contains the wake-up word of the first device, the first voice signal is subjected to gain compensation processing according to the pre-obtained correspondence between the frequency points of the voice signal and the gain compensation value to obtain the second voice signal.

[0085] Here, the wake word can be a word pre-set by the user in the smart device, or it can be a word set before the smart device leaves the factory. For example, the wake word can be a specific word such as Xiao Ai or Xiao Q.

[0086] In step S103, the first energy information may include, but is not limited to, the average energy value and total energy value of the second voice signal.

[0087] Here, the first energy information can be used to reflect the distance between the first device and the first user who emitted the first voice signal; it is understood that, since the voice signal will have energy attenuation during transmission, the longer the transmission distance, the lower the first energy information of the voice signal picked up by the first device.

[0088] The average energy value of the second speech signal can be obtained by sampling the second speech signal, summing the values ​​at each sampling point, and dividing by the number of sampling points. Alternatively, the total energy value of the second speech signal can be obtained by sampling the second speech signal, summing the values ​​at each sampling point, and then summing the sums.

[0089] In some embodiments, the frequency domain energy spectrum of the second speech signal can be obtained, and the frequency domain energy value of the second speech signal can be determined based on the frequency domain energy spectrum of the second speech signal; the frequency domain energy value of the second speech signal can be determined as the first energy information of the second speech signal.

[0090] In other embodiments, the first energy information of the second speech signal can be determined by obtaining the root mean square statistical energy value of the second speech signal.

[0091] Here, the root mean square statistical energy value of the second speech signal can be obtained by sampling the second speech signal, accumulating the squares of the values ​​of each sampling point of the second speech signal, dividing by the number of sampling points, and taking the square root.

[0092] In this embodiment of the disclosure, after the first device determines the first energy information based on the second voice signal obtained through gain compensation processing, it can send the first energy information to the server.

[0093] The server receives the energy information of the voice signal generated after multiple smart devices in the same local area network receive the same voice signal and perform gain compensation processing, determines the maximum energy information, and identifies the device corresponding to the maximum energy information as the target device responding to the voice signal.

[0094] In this embodiment of the disclosure, when determining the target device, the energy information of the voice signal generated after gain compensation processing of multiple smart devices is used for judgment. On the one hand, this can avoid the situation where multiple smart devices respond to the first voice signal at the same time. On the other hand, it can reduce the judgment error caused by the different sound field environment of the smart devices. It can accurately determine the smart device closest to the user (i.e., the user's real intention to wake it up) and improve the user's voice interaction experience.

[0095] Optionally, the acoustic performance parameters include at least the frequency response;

[0096] The correspondence is obtained by calculating the difference between the first frequency response and the second frequency response; wherein, the first frequency response is the frequency response generated by the first device receiving the test audio signal in the preset sound field environment; and the second frequency response is the frequency response generated by the first device receiving the test audio signal in the current sound field environment.

[0097] In this embodiment of the disclosure, the gain compensation processing is used to reduce the impact of different sound field environments on the acoustic performance parameters of the speech signals received by different smart devices. Through gain compensation processing, the acoustic performance parameters of the speech signals generated by multiple smart devices at different locations after gain compensation processing are approximately the same as the acoustic performance parameters of the speech signals received by multiple smart devices in the same sound field environment.

[0098] Therefore, the gain compensation value for gain compensation processing should at least indicate the difference between the acoustic performance parameters corresponding to the same speech signal received by the smart device in the preset sound field environment and the current sound field environment.

[0099] The step of calculating the difference between the first frequency response and the second frequency response to obtain the correspondence includes:

[0100] Based on the first frequency response and the second frequency response, a first difference between the amplitudes of the first frequency response and the second frequency response at the same frequency point is determined.

[0101] The correspondence is determined based on the plurality of frequency points and the first difference corresponding to the plurality of frequency points.

[0102] In this embodiment of the disclosure, a first frequency response generated by the first device receiving a test audio signal in a preset sound field environment and a second frequency response generated by the first device receiving a test audio signal in the current sound field environment can be obtained. Based on the first frequency response and the second frequency response, a first difference between the amplitudes corresponding to each frequency point of the test audio signal can be determined. It can be understood that the first difference corresponding to each frequency point is the gain compensation value corresponding to each frequency point. Therefore, the first difference corresponding to each frequency point is determined as the gain compensation value based on the first difference corresponding to each frequency point.

[0103] It is understood that the gain compensation value can be used to characterize the difference between the first frequency response and the second frequency response, and the second frequency response is related to the current sound field environment in which the first device is located; that is, the second frequency response is different for different sound field environments, and thus the gain compensation value is also different for different sound field environments.

[0104] If a change in the location of the first device is detected, that is, the acoustic environment in which the first device is located may also change, it is necessary to redetermine the gain compensation value of the first device after the change in the acoustic environment.

[0105] Optionally, step S102, which involves performing gain compensation processing on the first speech signal based on the pre-obtained correspondence between the frequency points of the speech signal and the gain compensation value, to obtain the second speech signal, includes:

[0106] Perform a Fourier transform on the first speech signal to obtain a first frequency domain signal;

[0107] Based on the correspondence, the amplitude values ​​corresponding to each frequency point in the first frequency domain signal are compensated to obtain the second frequency domain signal;

[0108] The second frequency domain signal is subjected to inverse Fourier transform to obtain the second speech signal.

[0109] In this embodiment of the disclosure, the correspondence is used to describe the relationship between the frequency point of the speech signal and the gain compensation value corresponding to the frequency point; since the gain compensation value is the amplitude difference at the same frequency value determined based on the first frequency response generated by the first device receiving the test audio signal in the preset sound field environment and the second frequency response generated by the first device receiving the test audio signal in the current sound field environment; and the first speech signal received by the first device is a time domain signal.

[0110] Therefore, it is necessary to perform a Fourier transform on the first speech signal to convert the first speech signal in the time domain into a first frequency domain signal in the frequency domain; according to the correspondence, determine the gain compensation value corresponding to each frequency point in the first frequency domain signal, and compensate the amplitude value of each frequency point in the first frequency domain signal.

[0111] An amplitude compensation value corresponding to each frequency point in the first frequency domain signal can be determined based on the frequency value. The amplitude value corresponding to each frequency point in the first frequency domain signal is compensated based on the amplitude compensation value to obtain a second frequency domain signal. An inverse Fourier transform is then performed on the second frequency domain signal to convert the second frequency domain signal in the frequency domain into a second speech signal in the time domain.

[0112] It should be noted that, in order to eliminate the influence of different sound field environments on the acoustic performance parameters of the voice signal received by the smart device, gain compensation processing can be performed based on the gain compensation value. The frequency response of the second voice signal obtained after gain compensation processing of the smart device in different sound field environments can be approximately the frequency response of the first voice signal received by the smart device in the preset sound field environment.

[0113] Optionally, step S102, which involves performing gain compensation processing on the first speech signal based on the pre-obtained correspondence between the frequency points of the speech signal and the gain compensation value, to obtain the second speech signal, includes:

[0114] Based on the correspondence, determine the filtering parameters and the first filter corresponding to the filtering parameters;

[0115] The first speech signal is input to the first filter to obtain the second speech signal output by the first filter. In this embodiment, the correspondence between the frequency points of the speech signal and the gain compensation value in the current sound field environment can be determined based on the current sound field environment in which the first device is located; based on the correspondence, filtering parameters and the filters corresponding to the filtering parameters are determined. Here, the frequency response compensation curve of the filter corresponding to the filtering parameters is the same as the correspondence between the frequency points of the speech signal and the gain compensation value in the current sound field environment.

[0116] After receiving a first voice signal containing a wake-up word, the first device can directly input the first voice signal into the filter and use the filter to perform gain compensation processing on the first voice signal; and determine the first energy information of the second voice signal based on the second voice signal output by the filter, thereby determining the target device responding to the first voice signal based on the first energy information.

[0117] In some embodiments, the filter may store multiple sets of filtering parameters; the correspondence between the frequency points of the speech signal and the gain compensation value is different for different filtering parameters.

[0118] Based on the current sound field environment in which the first device is located, the correspondence between the frequency points of the speech signal and the gain compensation value in the current sound field environment can be determined; the filter parameters that match the correspondence can be determined; based on the filter parameters that match the correspondence, the filter parameters can be updated; and the input first speech signal can be processed by the filter with updated parameters.

[0119] Optionally, the method further includes:

[0120] The system receives second energy information sent by a second device, wherein the second energy information is the energy information of the voice signal generated by the second device after receiving the first voice signal and performing gain compensation processing; and determines whether the first device is the target device responding to the first voice signal based on the first energy information and the second energy information.

[0121] or,

[0122] The first energy information is sent to the second device that determines the target device.

[0123] In this embodiment of the disclosure, the second device can be any device with voice interaction function in the current local area network other than the first device. Specifically, there can be one or more second devices with voice interaction function belonging to the same local area network as the first device; and the wake-up word of the second device is the same as that of the first device.

[0124] The first device can receive second energy information sent by a second device within the same local area network, and determine whether the first device responds to the target device of the first voice signal based on the comparison result between the first energy information and the second energy information.

[0125] If the first energy information includes an average energy value, the first device can be determined to be the target device responding to the first voice signal if the average energy value indicated by the first energy information is greater than the average energy value indicated by the second energy information; if the average energy value indicated by the first energy information is less than the average energy value indicated by the second energy information, the first device is determined not to be responding to the target device, and the first device remains silent.

[0126] The first device can also send the determined first energy information to a second device on the same local area network, so that the second device can determine the target device based on the first energy information of the first device and the second energy information of the second device.

[0127] It should be noted that smart devices can obtain the energy information of voice signals generated by other smart devices within the same local area network after gain compensation processing through local decision-making; and then determine whether to respond to the first voice signal based on the energy information.

[0128] In some embodiments, determining the first energy information of the second speech signal includes:

[0129] Obtain the first energy of each frame of the second speech signal;

[0130] The first energy information of the second speech signal is determined based on the ratio between the energy of the first energy of each frame of the second speech signal and the number of signal frames in the second speech signal.

[0131] In this embodiment of the disclosure, the first energy is the short-time energy of each frame of the first speech signal.

[0132] It should be noted that speech signals are non-stationary random signals that vary over time. Although speech signals are time-varying, they have short-term correlation, meaning that the characteristics of speech signals remain basically unchanged over a short period of time. Therefore, the analysis of speech signals is usually a short-time analysis.

[0133] The second speech signal can be divided into frames based on a preset duration to obtain multiple speech frame signals of the same length; and the short-time energy of each frame signal can be extracted based on a preset window function; the short-time average energy of the second speech signal can be determined based on the short-time energy of each frame signal in the second speech signal, and the short-time average energy of the second speech signal can be determined as the first energy information of the second speech signal.

[0134] Here, the preset duration can be a pre-set duration, for example, any duration within 10-30ms. It should be noted that within a short time range (generally considered to be 10-30ms), the characteristics of the speech signal remain essentially unchanged, i.e., relatively stable; therefore, when performing feature analysis on the speech signal, the feature parameters of each frame can be analyzed. Thus, for the entire speech signal, the analyzed result is a time series of feature parameters composed of the feature parameters of each frame. This disclosure also provides the following embodiments:

[0135] Figure 2 This is a flowchart illustrating a voice response method according to an exemplary embodiment. Figure 2 The method includes:

[0136] Step S201: Receive the first voice signal;

[0137] In this example, when the user emits a first voice signal, the first device uses a microphone to capture the first voice signal.

[0138] Here, the first device is any electronic device with voice interaction capabilities in the current local area network. For example, the first device could be a smart speaker, a smart refrigerator, or a smart TV.

[0139] Step S202: If the first voice signal contains a wake-up word of the first device, based on the pre-obtained correspondence between frequency points and gain compensation values, determine the gain compensation value corresponding to each frequency point of the first voice signal; wherein, the correspondence is obtained by calculating the difference between the first frequency response and the second frequency response; the first frequency response is: the frequency response generated by the first device receiving the test audio signal in the preset sound field environment; the second frequency response is: the frequency response generated by the first device receiving the ranging audio signal in the current generation environment;

[0140] In this example, the preset sound field environment is an open sound field environment.

[0141] The correspondence between the frequency points and the gain compensation values ​​is data pre-stored in the memory of the first device. It can be understood that the correspondence is determined based on the difference between the acoustic performance parameters corresponding to the same speech signal received by the first device in the current sound field environment and the preset sound field environment. Therefore, the correspondence between the frequency points and the gain compensation values ​​of the corresponding speech signals will also be different depending on the current sound field environment in which the first device is located.

[0142] For example, before the first device leaves the factory, in an anechoic chamber environment (i.e., an open sound field environment), a test audio signal (e.g., a sweep frequency signal) can be output using the speaker of the first device, and the test audio signal output by the speaker can be acquired using the microphone of the first device; based on the acquired test audio signal, a first frequency response of the ranging audio signal acquired by the first device in the open sound field environment can be determined; and the first spatial frequency response can be stored in the memory of the first device.

[0143] Each time the first device is powered on, a test audio signal is played through the speaker of the first device, and the test audio signal output by the speaker is collected through the microphone of the first device; based on the collected test frequency audio, the second frequency response of the ranging audio signal collected by the first device in the current sound field environment is determined;

[0144] Based on the first frequency response and the second frequency response, the amplitude difference at the same frequency value is determined; based on the amplitude difference corresponding to multiple frequency values, the gain compensation value of the first device in the current sound field environment is determined.

[0145] Step S203: Based on the correspondence, perform gain compensation processing on the first speech signal to obtain the second speech signal;

[0146] In this example, based on the correspondence, the amplitude difference that needs to be compensated for at each frequency point in the first speech signal is determined, and the first speech signal is compensated to obtain the second speech signal.

[0147] In some embodiments, the step of performing gain compensation processing on the first speech signal based on the gain compensation value of the first device to obtain a second speech signal includes:

[0148] Perform a Fourier transform on the first speech signal to obtain a first frequency domain signal;

[0149] Based on the correspondence, the amplitude values ​​corresponding to each frequency point in the first frequency domain signal are compensated to obtain the second frequency domain signal;

[0150] The second frequency signal is subjected to an inverse Fourier transform to obtain the second speech signal.

[0151] In this example, the first voice signal x1(t) acquired by the first device is a time-domain signal;

[0152] The first frequency domain signal X1(k) is obtained by performing a Fourier transform on the first speech signal x1(t);

[0153] Based on the gain compensation value delta(k), the amplitude values ​​at each frequency point in the first frequency domain signal are updated to obtain the second frequency domain signal X′1(k):

[0154] X′1(k)=X1(k)+delta(k);

[0155] Finally, the second frequency domain signal X′1(k) is subjected to inverse Fourier transform to obtain the second speech signal x′1(t) in the time domain.

[0156] In other embodiments, the method further includes:

[0157] Based on the correspondence, determine the filtering parameters and the first filter corresponding to the filtering parameters;

[0158] The first speech signal is input to the first filter to obtain the second speech signal output by the first filter. In this example, after determining the gain compensation value of the first device in the current sound field environment, the filter parameters can be designed to make the frequency response of the filter approximate the gain compensation value of the first device in the current sound field environment.

[0159] After the first device receives the first voice signal, a filter can be used to filter and increase the input first voice signal x1(t) to obtain the second voice signal x′1(t) output by the filter.

[0160] This example demonstrates how gain compensation of the first speech signal makes the first device appear to be placed in an open sound field environment, thus eliminating the influence of different sound field environments on the acoustic performance parameters of the speech signal acquired by the first device.

[0161] Step S204: Determine the first energy information of the second speech signal;

[0162] In this example, the first energy information may be the average energy information of the second speech signal.

[0163] Here, the length of the wake word in the second speech signal can be determined by a speech activity detection algorithm; based on the length of the wake word in the second speech signal, the average energy information of the wake word can be determined.

[0164] For example, the average energy information of the second speech signal can be determined by the following formula:

[0165]

[0166] Where E1 is the average energy value of the second speech signal; and T is the length of the wake-up word.

[0167] Step S205: Receive second energy information sent by the second device, wherein the second energy information is the energy information of the voice signal generated by the second device after receiving the first voice signal and performing gain compensation processing; determine whether the first device is the target device responding to the first voice signal based on the first energy information and the second energy information; or, send the first energy information to the second device that determined the target device.

[0168] In this example, the second device can be any device with voice interaction function in the current local area network other than the first device, and the wake word of the second device is the same as the wake word of the first device.

[0169] After receiving the first voice signal, each device in the same local area network will share the energy information of the voice signal generated after gain compensation processing of the first voice signal with other devices in the same local area network via WIFI. After obtaining the energy information shared by other devices via WIFI, each device will determine the maximum energy information based on its own energy information and the energy information of other devices; and determine whether to respond to the first voice signal based on the maximum energy information.

[0170] Here, if the device's own energy information is at its maximum energy level, then the device responds to the first voice signal; if the device's own energy information is not at its maximum energy level, the device remains silent.

[0171] In other embodiments, the first device may send first energy information to the server so that the server can determine the target device responding to the first voice signal based on the energy information of different devices.

[0172] In this example, after receiving the first voice signal, each device within the same local area network sends the energy information of the voice signal generated after gain compensation processing of the first voice signal to the server via WIFI; the server determines the maximum energy information based on the energy information of each device; and determines the device corresponding to the maximum energy information as the target device responding to the first voice signal.

[0173] This disclosure also provides a voice response device. Figure 3 This is a schematic diagram illustrating the structure of a voice response device according to an exemplary embodiment, such as... Figure 3 As shown, the voice response device 100 includes:

[0174] The compensation module 101 is used to receive a first voice signal; and to perform gain compensation processing on the first voice signal according to the pre-obtained correspondence between the frequency points of the voice signal and the gain compensation value to obtain a second voice signal; wherein, the correspondence is determined according to the acoustic performance parameters corresponding to the same voice signal received by the first device in the current sound field environment and the preset sound field environment.

[0175] The determining module 102 is used to determine the first energy information of the second voice signal; wherein the first energy information is used to determine the target device responding to the first voice signal.

[0176] Optionally, the acoustic performance parameters include at least the frequency response;

[0177] The correspondence is obtained by calculating the difference between the first frequency response and the second frequency response; wherein, the first frequency response is the frequency response generated by the first device receiving the test audio signal in the preset sound field environment; and the second frequency response is the frequency response generated by the first device receiving the test audio signal in the current sound field environment.

[0178] Optionally, the compensation module 101 is further configured to:

[0179] Perform a Fourier transform on the first speech signal to obtain a first frequency domain signal;

[0180] Based on the correspondence, the amplitude values ​​corresponding to each frequency point in the first frequency domain signal are compensated to obtain the second frequency domain signal;

[0181] The second frequency domain signal is subjected to inverse Fourier transform to obtain the second speech signal.

[0182] Optionally, the determining module 102 is configured to:

[0183] Based on the correspondence, determine the filtering parameters and the first filter corresponding to the filtering parameters;

[0184] The first speech signal is input to the first filter to obtain the second speech signal output by the first filter.

[0185] Optionally, the determining module 102 is further configured to:

[0186] Receive second energy information sent by the second device, wherein the second energy information is the energy information of the voice signal generated by the second device after receiving the first voice signal and performing gain compensation processing;

[0187] Based on the first energy information and the second energy information, determine whether the first device is the target device responding to the first voice signal;

[0188] or,

[0189] The first energy information is sent to the second device that determines the target device.

[0190] Optionally, the determining module 102 is further configured to:

[0191] Obtain the first energy of each frame of the second speech signal;

[0192] The first energy information of the second speech signal is determined based on the ratio between the energy of the first energy of each frame of the second speech signal and the number of signal frames in the second speech signal.

[0193] Figure 4 This is a block diagram illustrating a voice response device according to an exemplary embodiment. For example, device 800 may be a mobile phone, mobile computer, etc.

[0194] Reference Figure 4 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0195] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0196] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of this data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0197] Power supply component 806 provides power to various components of device 800. Power supply component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to device 800.

[0198] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0199] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0200] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0201] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in the position of device 800 or a component of device 800, the presence or absence of user contact with device 800, the orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0202] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as Wi-Fi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0203] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0204] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of the device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0205] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0206] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A voice response method, characterized in that, Applied to a first device, the method includes: Receive the first voice signal; Based on the pre-obtained correspondence between the frequency points of the speech signal and the gain compensation value, the first speech signal is subjected to gain compensation processing to obtain the second speech signal; wherein, the correspondence is determined based on the acoustic performance parameters corresponding to the same speech signal received by the first device in the current sound field environment and the preset sound field environment. Determine the first energy information of the second speech signal; The first energy information is sent to the server so that the server can compare the first energy information with the energy information obtained by multiple devices in the same local area network after performing gain compensation processing based on the first voice signal, and determine the target device responding to the first voice signal. If a change in the position of the first device is detected, the gain compensation value of the first device after the change in the sound field environment is re-determined, and the correspondence is updated based on the re-determined gain compensation value.

2. The method according to claim 1, characterized in that, The acoustic performance parameters include at least the frequency response; The correspondence is obtained by calculating the difference between the first frequency response and the second frequency response; wherein, the first frequency response is the frequency response generated by the first device receiving the test audio signal in the preset sound field environment; and the second frequency response is the frequency response generated by the first device receiving the test audio signal in the current sound field environment.

3. The method according to claim 1, characterized in that, The step of performing gain compensation processing on the first speech signal based on the pre-obtained correspondence between the frequency points of the speech signal and the gain compensation value to obtain the second speech signal includes: Perform a Fourier transform on the first speech signal to obtain a first frequency domain signal; Based on the correspondence, the amplitude values ​​corresponding to each frequency point in the first frequency domain signal are compensated to obtain the second frequency domain signal; The second frequency domain signal is subjected to inverse Fourier transform to obtain the second speech signal.

4. The method according to claim 1, characterized in that, The step of performing gain compensation processing on the first speech signal to obtain the second speech signal based on the pre-obtained correspondence between the frequency points of the speech signal and the gain compensation value includes: Based on the correspondence, determine the filtering parameters and the first filter corresponding to the filtering parameters; The first speech signal is input to the first filter to obtain the second speech signal output by the first filter.

5. The method according to claim 1, characterized in that, The method further includes: The system receives second energy information sent by a second device, wherein the second energy information is the energy information of the voice signal generated by the second device after receiving the first voice signal and performing gain compensation processing; and determines whether the first device is the target device responding to the first voice signal based on the first energy information and the second energy information. or, The first energy information is sent to the second device that determines the target device.

6. The method according to claim 1, characterized in that, The determination of the first energy information of the second speech signal includes: Obtain the first energy of each frame of the second speech signal; The first energy information of the second speech signal is determined based on the ratio between the energy of the first energy of each frame of the second speech signal and the number of signal frames in the second speech signal.

7. A voice response device, characterized in that, Applied to the first device, including: The compensation module is used to receive a first voice signal; and to perform gain compensation processing on the first voice signal according to the pre-obtained correspondence between the frequency points of the voice signal and the gain compensation value to obtain a second voice signal; wherein, the correspondence is determined according to the acoustic performance parameters corresponding to the same voice signal received by the first device in the current sound field environment and the preset sound field environment. The determining module is used to determine the first energy information of the second speech signal; The sending module is used to send the first energy information to the server so that the server can compare the first energy information with the energy information obtained by multiple devices in the same local area network after performing gain compensation processing based on the first voice signal, and determine the target device responding to the first voice signal. The compensation module is further configured to, if it detects a change in the position of the first device, redetermine the gain compensation value of the first device after the change in the sound field environment, and update the correspondence based on the redetermined gain compensation value.

8. The apparatus according to claim 7, characterized in that, The acoustic performance parameters include at least the frequency response; The correspondence is obtained by calculating the difference between the first frequency response and the second frequency response; wherein, the first frequency response is the frequency response generated by the first device receiving the test audio signal in the preset sound field environment; and the second frequency response is the frequency response generated by the first device receiving the test audio signal in the current sound field environment.

9. The apparatus according to claim 7, characterized in that, The compensation module is also used for: Perform a Fourier transform on the first speech signal to obtain a first frequency domain signal; Based on the correspondence, the amplitude values ​​corresponding to each frequency point in the first frequency domain signal are compensated to obtain the second frequency domain signal; The second frequency domain signal is subjected to inverse Fourier transform to obtain the second speech signal.

10. The apparatus according to claim 7, characterized in that, The determining module is used for: Based on the correspondence, determine the filtering parameters and the first filter corresponding to the filtering parameters; The first speech signal is input to the first filter to obtain the second speech signal output by the first filter.

11. The apparatus according to claim 7, characterized in that, The determining module is also used for: Receive second energy information sent by the second device, wherein the second energy information is the energy information of the voice signal generated by the second device after receiving the first voice signal and performing gain compensation processing; Based on the first energy information and the second energy information, determine whether the first device is the target device responding to the first voice signal; or, The first energy information is sent to the second device that determines the target device.

12. The apparatus according to claim 7, characterized in that, The determining module is used for: Obtain the first energy of each frame of the second speech signal; The first energy information of the second speech signal is determined based on the ratio between the energy of the first energy of each frame of the second speech signal and the number of signal frames in the second speech signal.

13. A voice response device, characterized in that, include: processor; Memory used to store executable instructions; The processor is configured to implement the voice response method according to any one of claims 1-6 when executing executable instructions stored in the memory.

14. A non-transitory computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of a voice response device, enable the voice response device to perform the voice response method of any one of claims 1-6.