Smart speaker and audio signal processing method

By processing the volume characteristics and voiceprint characteristics of the microphone array, the target frame length is screened and the user voice signal is weighted and separated, which solves the problem of speaker audio interference and improves the voice recognition and interaction performance of the smart speaker.

CN120452482BActive Publication Date: 2025-09-12SHENZHEN ZHUOYUE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510956121.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-12
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

When a smart speaker plays audio, the audio played by the speaker and the user's voice act on the microphone array at the same time, causing the user's voice signal to be severely aliased and interfered with. The noise reduction process cannot obtain a pure user voice signal, affecting the accuracy of voice recognition.

Method used

The azimuth angle and audio signal of each microphone are obtained through the microphone array. Combined with the user's pure voice signal, the volume eigenvalue and voiceprint eigenvalue are determined, the target frame length is screened, the audio signal matrix is ​​constructed and weighted processing is performed, and the user voice signal is separated using the FastICA algorithm.

Benefits of technology

It improves the extraction sensitivity and accuracy of user voice signals, enhances the voice interaction performance of smart speakers, and improves the real-time and accuracy of voice recognition and interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452482B_ABST
    Figure CN120452482B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of voice processing technology, and specifically to a smart speaker and audio signal processing method, comprising: in a microphone array of the smart speaker, based on the audio signals of microphones at different azimuth angles under each frame length and the energy values ​​at different frequencies in the audio signals, combined with the fundamental frequency in the user's pure voice signal, determining the volume characteristic value of each microphone under each frame length, thereby screening out a target frame length, and then determining the voiceprint characteristic value of each microphone under the target frame length, thereby determining a voice signal weight vector, constructing an audio signal matrix for the audio signals of all microphones under the target frame length, and weighting them to determine the user's voice signal under the target frame length. The present invention enhances the role of the audio signal collected by the microphone facing the user's sound source direction by extracting the volume characteristics and voiceprint characteristics of the audio signal of the microphone array, thereby improving the ability to capture the user's voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice processing technology, and in particular to an intelligent speaker and an audio signal processing method. Background Art

[0002] Smart speakers are devices that combine voice interaction, audio playback, and smart home control functions, and are an important part of the modern smart home ecosystem. Through built-in voice assistants, they allow users to interact with the device through natural language, enabling a variety of functions and services.

[0003] The voice interaction function of smart speakers places high demands on the quality of user voice signals to improve the accuracy of subsequent voice recognition and enhance voice interaction performance. An existing smart speaker and its control method uses a dual-microphone array to collect the user's first and second audio signals. After processing through multiple noise reduction sub-units, the noise-reduced signals are integrated to remove redundant audio signals and improve the audio noise reduction effect, thereby enhancing the smart speaker's voice interaction performance.

[0004] Existing Problem: When a user interacts with a smart speaker while it's playing audio, the audio from the speaker and the user's voice simultaneously interact with the microphone array, causing severe aliasing interference. Furthermore, the speaker's audio volume is high, and the frequencies of the audio and the user's voice often overlap significantly. This makes it difficult to obtain a pure user voice signal through noise reduction alone, resulting in a decrease in the accuracy of subsequent voice recognition and poor voice interaction performance on the smart speaker. Summary of the Invention

[0005] The present invention provides an intelligent speaker and an audio signal processing method to solve the existing problems.

[0006] The present invention provides a smart speaker and an audio signal processing method using the following technical solutions:

[0007] An embodiment of the present invention provides a method for processing an audio signal of an intelligent speaker, the method comprising the following steps:

[0008] In the microphone array of the smart speaker, obtain the azimuth angle, number, audio signal of each microphone at each frame length, and a historical audio signal set consisting of the audio signals of all microphones at several historical frame lengths, as well as the user's pure voice signal;

[0009] Determine the volume characteristic value of the audio signal of each microphone in each frame length based on the audio signals of the microphones at different azimuth angles in each frame length and the energy values ​​at different frequencies in the audio signals, combined with the fundamental frequency of the user's pure voice signal;

[0010] Filter out the target frame length based on the comparison results of the volume feature values ​​of the audio signals of all microphones at each frame length in the historical audio signal set;

[0011] Determine the voiceprint feature value of the audio signal of each microphone at the target frame length based on the similarity between the audio signal of each microphone at the target frame length and the voiceprint feature in the user's pure voice signal;

[0012] According to the numbering order of the microphones, an audio signal matrix is ​​constructed using the audio signals of all microphones at the target frame length; based on the voiceprint eigenvalues ​​and volume eigenvalues ​​of the audio signals of all microphones at the target frame length, a speech signal weight vector is determined; based on the speech signal weight vector and the audio signal matrix at the target frame length, the user voice signal at the target frame length is determined.

[0013] Furthermore, the specific steps of determining the volume characteristic value of the audio signal of each microphone at each frame length include the following:

[0014] Determine the azimuth angle interval energy of each microphone in the current frame length according to the audio signals of the microphones at different azimuth angles in the current frame length;

[0015] Using the user's pure voice signal as input, the autocorrelation algorithm is used to obtain the user's fundamental frequency;

[0016] The preset frequency range size A is centered on the user's fundamental frequency and a target frequency range of size A is constructed;

[0017] For the audio signal of each microphone in the current frame length, a fast Fourier transform algorithm is used to obtain energy values ​​at several different frequencies. The sum of the energy values ​​at all frequencies within the target frequency range is used as the user voice energy of each microphone in the current frame length.

[0018] The volume characteristic value of the audio signal of each microphone in the current frame length is determined based on the difference between the azimuth angle interval energies of the microphones in the current frame length and the difference between the user voice energies.

[0019] Furthermore, the step of determining the azimuth angle interval energy of each microphone at the current frame length according to the audio signals of the microphones at different azimuth angles at the current frame length includes the following specific steps:

[0020] The MUSIC algorithm uses the azimuth angles of all microphones and the audio signals of all microphones at the current frame length as input, and outputs the spatial spectrum energy at the current frame length; the spatial spectrum energy is composed of energy values ​​at different azimuth angles;

[0021] The size of the angle interval θ is preset, and an azimuth angle interval of size θ is constructed with the azimuth angle of any microphone as the center;

[0022] In the spatial spectrum energy under the current frame length, the sum of the energy values ​​under all azimuth angles in the azimuth angle interval corresponding to each microphone is used as the azimuth angle interval energy of each microphone under the current frame length.

[0023] Furthermore, the method of determining the volume characteristic value of the audio signal of each microphone at the current frame length according to the difference between the azimuth angle interval energies of the microphones at the current frame length and the difference between the user voice energies includes the following specific steps:

[0024] At the current frame length, calculate the difference between the azimuth angle interval energy of each microphone and the average azimuth angle interval energy of all microphones, then calculate the ratio of the user voice energy of each microphone to the average user voice energy of all microphones, and input the product of the difference and the ratio into an exponential function with a natural constant as the base to obtain the output value, which is recorded as the volume characteristic value of the audio signal of each microphone at the current frame length.

[0025] Furthermore, the specific steps of screening out the target frame length include the following:

[0026] Obtain the range of the volume characteristic values ​​of the audio signals of all microphones in the current frame length, which is recorded as the volume characteristic range size R in the current frame length;

[0027] In the historical audio signal set, the volume feature range size under each historical frame length is obtained according to the method for obtaining the volume feature range size under the current frame length;

[0028] A first constant C1 is preset, and the mean R_avg and standard deviation V of the volume feature range sizes under all historical frame lengths are obtained. If the difference between the volume feature range size R under the current frame length and the mean R_avg of the volume feature range sizes under all historical frame lengths is greater than C1 times the standard deviation V, then the current frame length is recorded as the target frame length.

[0029] Furthermore, the step of determining the voiceprint feature value of the audio signal of each microphone at the target frame length includes the following specific steps:

[0030] Preset the second constant C2, and obtain the sequence of Mel-frequency cepstral coefficients of the first C2 orders of the audio signal of each microphone under the target frame length, and record it as the voiceprint sequence of each microphone under the target frame length;

[0031] Obtain a voiceprint sequence of the user's pure voice signal according to the method for obtaining the voiceprint sequence of each microphone at the target frame length;

[0032] The cosine similarity between the voiceprint sequence of each microphone and the voiceprint sequence of the user's clean speech signal at the target frame length is obtained, and the sum of the cosine similarity and a preset third constant is recorded as the voiceprint feature value of the audio signal of each microphone at the target frame length.

[0033] Furthermore, the audio signal matrix is ​​constructed using the audio signals of all microphones at the target frame length according to the numbering order of the microphones, and the specific steps include the following:

[0034] At the target frame length, the audio signals of each microphone are arranged row by row from top to bottom according to the numbers of all microphones from small to large to form an audio signal matrix.

[0035] Furthermore, the determining of the speech signal weight vector includes the following specific steps:

[0036] At the target frame length, obtain the sum of the voiceprint eigenvalue and volume eigenvalue of the audio signal of each microphone, record it as the comprehensive eigenvalue of each microphone, calculate the ratio of the comprehensive eigenvalue of each microphone to the mean of the comprehensive eigenvalues ​​of all microphones, and record it as the weight of each microphone at the target frame length;

[0037] Under the target frame length, the weights of each microphone are arranged in a row from left to right according to the numbers of all microphones from small to large, forming a weight vector;

[0038] The FastICA algorithm is used to iterate the weight vector at the target frame length. When the Euclidean distance between the weight vectors of two consecutive update iterations is less than the preset convergence threshold, the update iteration is stopped and the weight vector of the last update iteration is recorded as the speech signal weight vector at the target frame length.

[0039] Furthermore, the step of determining the user voice signal at the target frame length includes the following specific steps:

[0040] The product of the speech signal weight vector at the target frame length and the audio signal matrix is ​​recorded as the user speech signal at the target frame length.

[0041] The present invention also proposes a smart speaker, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. The processor executes the computer program stored in the memory to implement the steps of the aforementioned smart speaker audio signal processing method.

[0042] The beneficial effects of the technical solution of the present invention are:

[0043] In an embodiment of the present invention, in the microphone array of the smart speaker, the azimuth angle of each microphone and the audio signal at each frame length are obtained. According to the audio signals of the microphones at different azimuth angles at each frame length and the energy values ​​at different frequencies in the audio signals, combined with the fundamental frequency in the user's pure voice signal, the volume characteristic value of the audio signal of each microphone at each frame length is determined, thereby screening out the target frame length. Based on the energy changes of the audio signals collected by the microphone array and the difference in the sound field generated by the user's sound source and the speaker's sound source, it is determined whether the current audio signal contains the user's voice signal, thereby improving the sensitivity and accuracy of user voice extraction and improving the real-time performance of the smart speaker voice interaction. The voiceprint characteristic value of each microphone at the target frame length is determined, and the voice signal weight vector is determined based on the voiceprint characteristic value and volume characteristic value of all microphones at the target frame length. The audio signal matrix of all microphones at the target frame length is constructed and weighted to determine the user's voice signal at the target frame length. Thus, the present invention enhances the role of the audio signal collected by the microphone facing the user's sound source direction by extracting the volume characteristics and voiceprint characteristics of the audio signal of the microphone array, thereby improving the ability to capture the user's voice signal. At the same time, the direction of the user's sound source is used as prior knowledge to adjust the initialization process of the weight vector in the sound source separation algorithm, which not only speeds up the convergence of the algorithm, but also improves the separation accuracy of the user's voice signal, thereby improving the performance of smart speaker voice recognition and voice interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 This is a flowchart of the steps of a smart speaker audio signal processing method of the present invention;

[0046] Figure 2 The following is a flowchart for obtaining user voice signals. DETAILED DESCRIPTION

[0047] To further illustrate the technical means and effects of the present invention to achieve the intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effects of a smart speaker and audio signal processing method proposed in accordance with the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable form.

[0048] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0049] The following describes in detail a smart speaker and an audio signal processing method provided by the present invention with reference to the accompanying drawings.

[0050] See also Figure 1 , which shows a flowchart of a method for processing an audio signal of an intelligent speaker provided by one embodiment of the present invention, the method comprising the following steps:

[0051] Step S001: In the microphone array of the smart speaker, obtain the azimuth angle, number, audio signal of each microphone at each frame length, and a historical audio signal set consisting of the audio signals of all microphones at several historical frame lengths and the user's pure voice signal.

[0052] It should be noted that the smart speaker proposed in this embodiment is composed of a microphone array, a speaker, a processor, a memory, a Wi-Fi module, a touch screen, a power module, and a plastic shell. The microphone array is installed at the outer ring of the bottom of the smart speaker and uses a microphone array to collect audio data; the speaker is installed in the center of the top of the smart speaker and is used to play audio; the processor is used to run the operating system and perform tasks such as audio processing, voice recognition, and network communication; the memory is used to store the operating system, applications, user data, and audio data collected by the microphone array; the Wi-Fi module is used to connect to the Internet and interact with other devices; the touch screen is installed on the side of the smart speaker to display information and interactive interfaces, and to realize manual operation of the smart speaker; the power module is used to power the smart speaker and manage energy consumption; the plastic shell is used to protect internal components and improve the aesthetics of the smart speaker. Taking into account the situation where users perform voice interaction during the audio playback of the smart speaker, this embodiment proposes an audio processing method based on the smart speaker to separate the user voice signal from the audio signal collected by the microphone array.

[0053] In the microphone array of the smart speaker, the azimuth angle, number, audio signal of each microphone at each frame length, and the historical audio signal set consisting of the audio signals of all microphones at several historical frame lengths and the user's pure voice signal are obtained.

[0054] It should be noted that during voice interaction between the user and the smart speaker, when the smart speaker is not playing audio, the user voice signal collected by the microphone array is relatively pure and can be directly used for voice recognition after noise reduction processing. Therefore, the user's pure voice signal is stored in the smart speaker memory for subsequent extraction of the voiceprint features of the user's voice signal. However, when the user performs voice interaction while the smart speaker is playing audio, the audio played by the speaker and the user's voice act on the microphone array at the same time, and the user's voice signal is subject to severe aliasing interference. It is difficult to obtain a pure user voice signal through noise reduction processing alone, and it cannot be used as input for subsequent voice recognition.

[0055] It should be further explained that the microphone array in this embodiment is a circular array of six microphones, with each microphone numbered, for example, {1, 2, 3, 4, 5, 6}. A two-dimensional coordinate system is constructed on the two-dimensional plane where the microphone array resides, with the center point of the microphone array as the origin, the horizontal axis pointing to the right as the horizontal square, and the vertical axis pointing upward as the vertical square. The position coordinates of each microphone are determined in this two-dimensional coordinate system. The azimuth angle of each microphone is then obtained by rotating the system counterclockwise, with the horizontal axis pointing to the right as 0 degrees. This is described as an example. During audio playback by the smart speaker, each microphone in the array collects audio signals in real time at a sampling frequency of 16 kHz. Simultaneously, the audio signals collected by each microphone are framed over the total duration, with the total duration divided into 25 millisecond (ms) frames with a frame shift of 10 milliseconds (ms). This is described as an example. Framing is a well-known technique, and framing involves dividing the audio signal into short frames of fixed length. The purpose of framing is to convert a non-stationary audio signal into a series of approximately stationary short-term signals. The framing steps include: frame length selection (selecting the length of each frame) and frame shift selection (selecting the overlap between frames, with the frame shift typically being smaller than the frame length to ensure sufficient overlap between adjacent frames). Furthermore, the smart speaker's memory stores 5,000 consecutive historical frames of audio signals collected by six microphones at 30,000 historical frames. This constitutes a historical audio signal set, which is used to subsequently determine whether the audio signal contains a user voice signal. This example will be used to illustrate this.

[0056] Step S002: Determine the volume characteristic value of the audio signal of each microphone in each frame length based on the audio signals of the microphones at different azimuth angles in each frame length and the energy values ​​at different frequencies in the audio signals, combined with the fundamental frequency in the user's pure voice signal.

[0057] It should be noted that for the audio signals collected by the six microphones at any given frame length, since the smart speaker is located above the center of the microphone array, the sound field it produces is the same at each microphone. When the user is not engaging in voice interaction, the energy of the audio signals collected by different microphones is the same. However, the distances between the user's sound source and the different microphones vary, and due to the differences in the sound collection directions of the different microphones, the energy of the audio signals collected by different microphones varies when the user is engaging in voice interaction with the smart speaker. The closer the microphone's sound collection direction is to the direction the smart speaker is currently pointing towards the user, the greater the proportion of the voice signal in the audio signal it collects. Therefore, the spatial spectrum energy at the current frame length can be obtained based on the audio signals collected by the six microphones, and the azimuth angle interval energy of each microphone can be calculated accordingly.

[0058] Preferably, in one embodiment of the present invention, the method for obtaining the volume characteristic value of the audio signal of each microphone at each frame length includes:

[0059] The MUSIC algorithm uses the azimuth angles of all microphones and the audio signals of all microphones in the current frame length as input, and outputs the spatial spectrum energy in the current frame length. The spatial spectrum energy is composed of the energy values ​​at different azimuth angles.

[0060] It should be noted that the MUSIC algorithm (Multiple Signal Classification) is a well-known technology and the specific method will not be introduced here.

[0061] Preset angle interval size The preset frequency range A is 15 degrees and 100 Hz, which is used as an example for description.

[0062] Taking the azimuth angle of any microphone as the center, the size of the construction is azimuth angle range.

[0063] For example, if the azimuth angle of a microphone is 10 degrees, then its azimuth angle range is [2.5, 17.5].

[0064] In the spatial spectrum energy under the current frame length, the sum of the energy values ​​under all azimuth angles in the azimuth angle interval corresponding to each microphone is used as the azimuth angle interval energy of each microphone under the current frame length.

[0065] It should be noted that: considering that the volume of the audio played by the speaker is relatively high and the user voice is subject to strong aliasing interference, this embodiment calculates the volume characteristic value of each microphone based on the difference in frequency distribution between the user voice and the audio played by the speaker, as well as the azimuth angle interval energy obtained above.

[0066] The user's pure voice signal is used as input and the autocorrelation algorithm is used to obtain the user's fundamental frequency.

[0067] Centered on the user's fundamental frequency, a target frequency range of size A is constructed.

[0068] For example, if the user's fundamental frequency is 100 Hz, the target frequency range is [50,150].

[0069] For the audio signal of each microphone in the current frame length, the fast Fourier transform algorithm is used to obtain the energy values ​​at several different frequencies. The sum of the energy values ​​at all frequencies in the target frequency range is used as the user voice energy of each microphone in the current frame length.

[0070] It should be noted that the autocorrelation algorithm and the fast Fourier transform algorithm are both well-known technologies, and the specific methods are not introduced here. The target frequency range is the frequency range in which the user's voice energy is concentrated. According to Parseval's theorem (well-known technology), the amount of energy within the target frequency range reflects the energy of the user's voice signal near the fundamental frequency, and thus indicates the probability that the audio signal collected by the corresponding microphone contains the user's voice signal. Within the frequency range in which the user's voice energy is concentrated, the greater the microphone's signal energy, the greater the probability that its sound collection direction is toward the user during voice interaction, and the greater the proportion of the user's voice signal contained in the audio signal.

[0071] At the current frame length, calculate the difference between the azimuth angle interval energy of each microphone and the average azimuth angle interval energy of all microphones, then calculate the ratio of the user speech energy of each microphone to the average user speech energy of all microphones, and input the product of the difference and the ratio into an exponential function with a natural constant as the base to obtain the output value, which is recorded as the volume characteristic value of the audio signal of each microphone at the current frame length.

[0072] It should be noted that in this embodiment, an exponential function with a natural constant as the base is used to amplify the differences between different microphones caused by the user voice signal.

[0073] Step S003: Filter out a target frame length based on a comparison result of the volume feature values ​​of the audio signals of all microphones in a historical audio signal set at each frame length.

[0074] Preferably, in one embodiment of the present invention, the method for obtaining the target frame length includes:

[0075] Get the range of the volume characteristic values ​​of the audio signals of all microphones under the current frame length, and record it as the volume characteristic range under the current frame length .

[0076] In the historical audio signal set, the volume feature range size under each historical frame length is obtained according to the method for obtaining the volume feature range size under the current frame length.

[0077] The first constant C1 is preset to 3, and this is used as an example for description.

[0078] Get the mean value of the volume feature range size under all historical frame lengths and standard deviation , if the volume feature range size under the current frame length Subtract the mean of the volume feature range size under all historical frame lengths The difference is greater than C1 times the standard deviation , then the current frame length is recorded as the target frame length.

[0079] It should be noted that in this embodiment, if the user interacts with the smart speaker during the time period corresponding to the target frame length, user voice signal extraction is required. Otherwise, it is assumed that no user voice is present in the audio signal at the current frame length, and user voice signal extraction is not required. This allows real-time determination of whether user voice is present in the audio signal at each frame length.

[0080] Step S004: determining the voiceprint feature value of the audio signal of each microphone at the target frame length according to the similarity between the audio signal of each microphone at the target frame length and the voiceprint feature in the pure speech signal of the user.

[0081] It should be noted that, given the differences in sound characteristics between audio played by a smart speaker and the user's voice, the degree of match between the audio signal collected by the microphone and the user's voice signal in this embodiment reflects the proportion of the audio signal containing the user's voice signal. The higher the degree of sound characteristic match between the two, the greater the proportion of the audio signal containing the user's voice signal, indicating a greater probability that the corresponding microphone is collecting sound towards the user.

[0082] Preferably, in one embodiment of the present invention, the method for obtaining the voiceprint feature value of each microphone at the target frame length includes:

[0083] The second constant C2 is preset to 12, and the third constant C3 is preset to 1.5, and this is used as an example for description.

[0084] At the target frame length, a sequence consisting of the first C2 orders of the Mel-frequency cepstral coefficients of the audio signal of each microphone is obtained, and recorded as the voiceprint sequence of each microphone at the target frame length.

[0085] The voiceprint sequence of the user's pure voice signal is obtained according to the method of obtaining the voiceprint sequence of each microphone under the target frame length.

[0086] Obtain the cosine similarity between the voiceprint sequence of each microphone and the voiceprint sequence of the user's clean speech signal at the target frame length, and record the sum of the cosine similarity and C3 as the voiceprint feature value of the audio signal of each microphone at the target frame length.

[0087] It should be noted that both Mel-cepstral coefficients and cosine similarity are well-known technologies, and the specific methods will not be introduced here. Mel-cepstral coefficients compress the spectrum of the speech signal to a logarithmic scale to characterize the voiceprint characteristics of the speech signal, avoiding interference from high-frequency noise and secondary information. In this embodiment, only Mel-cepstral coefficients with orders from 1 to 12 are retained, and the voiceprint sequence is constructed from small to large orders, which contains most of the speech information. Cosine similarity is used to measure the similarity between two sequences, and its value range is between -1 and 1. In this embodiment, C3 is added to the cosine similarity to ensure that the voiceprint feature value is non-negative, so the setting of C3 needs to be no less than 1. The larger the voiceprint feature value, the higher the match between the audio signal of the corresponding microphone and the pure voice signal of the user in terms of sound characteristics, the greater the proportion of the user's voice signal in the audio signal, and the greater the probability that the corresponding microphone's sound collection direction is facing the user.

[0088] Step S005: Construct an audio signal matrix based on the numbering order of the microphones and the audio signals of all microphones at the target frame length; determine the speech signal weight vector based on the voiceprint eigenvalues ​​and volume eigenvalues ​​of the audio signals of all microphones at the target frame length; determine the user speech signal at the target frame length based on the speech signal weight vector and the audio signal matrix at the target frame length.

[0089] It should be noted that: in this embodiment, the FastICA algorithm is used for sound source separation. Considering that the algorithm is sensitive to the selection of the initial value of the projection matrix, the random selection method of the initial value of the projection matrix, on the one hand, slows down the convergence speed, affecting the real-time performance of the smart speaker voice interaction; on the other hand, it causes the convergence direction to fall into the local optimal solution of the algorithm objective function, resulting in a decrease in the separation accuracy of the user voice signal. Among them, the FastICA algorithm (Fast Independent Component Analysis algorithm) is a well-known technology, and the specific method is not introduced here. Therefore, the volume characteristics and voiceprint characteristics of the user voice signal and the speaker audio signal are used as prior knowledge to adjust the initialization process of the weight vector in the sound source separation algorithm, thereby accelerating the convergence speed of the algorithm while improving the separation accuracy of the user voice signal, thereby improving the performance of the smart speaker voice recognition and voice interaction.

[0090] Preferably, in one embodiment of the present invention, the method for acquiring a user voice signal at a target frame length includes:

[0091] At the target frame length, the audio signals of each microphone are arranged row by row from top to bottom according to the numbers of all microphones from small to large to form an audio signal matrix.

[0092] It should be noted that the audio signal matrix undergoes zero-mean centering and pre-whitening processing, which is a well-known technology and the specific method will not be introduced here. Considering that the voice recognition of the smart speaker only requires the user voice signal, and the FastICA algorithm can independently separate the signal of a single sound source, in this embodiment, only the user voice signal is separated from the audio signal collected by the microphone array. Therefore, the weight vector of the voice signal can be initialized according to the volume characteristics and voiceprint characteristics of the user voice signal and the speaker audio signal to enhance the role of the audio signal collected by the microphone facing the direction of the user sound source in user voice extraction.

[0093] At the target frame length, obtain the sum of the voiceprint eigenvalue and volume eigenvalue of the audio signal of each microphone, record it as the comprehensive eigenvalue of each microphone, calculate the ratio of the comprehensive eigenvalue of each microphone to the mean of the comprehensive eigenvalues ​​of all microphones, and record it as the weight of each microphone at the target frame length.

[0094] Under the target frame length, the weights of each microphone are arranged in a row from left to right according to the numbers of all microphones from small to large, forming a weight vector.

[0095] The preset convergence threshold is 0.001, which is used as an example for description.

[0096] The FastICA algorithm is used to iterate the weight vector at the target frame length. When the Euclidean distance between the weight vectors of two consecutive update iterations is less than the preset convergence threshold, the update iteration is stopped and the weight vector of the last update iteration is recorded as the speech signal weight vector at the target frame length.

[0097] The product of the speech signal weight vector at the target frame length and the audio signal matrix is ​​recorded as the user speech signal at the target frame length.

[0098] What needs to be explained is that the voice signal weight vector is a matrix with 1 row and 6 columns, while the audio signal matrix is ​​a matrix with 6 rows and each row of audio signal has a duration of 25 milliseconds. That is, the number of columns of the voice signal weight vector is equal to the number of rows of the audio signal matrix. They can be multiplied to obtain a row of audio signal with a duration of 25 milliseconds, that is, the user voice signal under the target frame length. In this way, the user voice signal under each frame length can be obtained in real time to complete the separation of the user voice signal. The flow chart for obtaining the user voice signal is as follows: Figure 2 shown.

[0099] The present invention also provides a smart speaker, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. The processor executes the computer program stored in the memory to implement the steps of the aforementioned smart speaker audio signal processing method.

[0100] So far, the present invention is completed.

[0101] In summary, in an embodiment of the present invention, in the microphone array of the smart speaker, the azimuth angle of each microphone and the audio signal at each frame length are obtained, and the volume characteristic value of the audio signal of each microphone at different azimuth angles under each frame length and the energy value at different frequencies in the audio signal are combined with the fundamental frequency in the user's pure voice signal to determine the volume characteristic value of the audio signal of each microphone under each frame length, thereby screening out the target frame length, and then determining the voiceprint characteristic value of the audio signal of each microphone under the target frame length. According to the voiceprint characteristic value and volume characteristic value of all microphones under the target frame length, the voice signal weight vector is determined, and the audio signal matrix of all microphones under the target frame length is constructed for weighting, thereby determining the user voice signal under the target frame length. The present invention enhances the role of the audio signal collected by the microphone facing the user's sound source direction by extracting the volume characteristics and voiceprint characteristics of the audio signal of the microphone array, and improves the ability to capture the user's voice signal.

[0102] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for processing audio signals of a smart speaker, characterized in that: The method comprises the following steps: In the microphone array of the smart speaker, obtain the azimuth angle, number, audio signal of each microphone at each frame length, and a historical audio signal set consisting of the audio signals of all microphones at several historical frame lengths, as well as the user's pure voice signal; Determine the volume characteristic value of the audio signal of each microphone in each frame length based on the audio signals of the microphones at different azimuth angles in each frame length and the energy values ​​at different frequencies in the audio signals, combined with the fundamental frequency of the user's pure voice signal; Filter out the target frame length based on the comparison results of the volume feature values ​​of the audio signals of all microphones at each frame length in the historical audio signal set; Determine the voiceprint feature value of the audio signal of each microphone at the target frame length based on the similarity between the audio signal of each microphone at the target frame length and the voiceprint feature in the user's pure voice signal; According to the numbering order of the microphones, an audio signal matrix is ​​constructed using the audio signals of all microphones at the target frame length; based on the voiceprint eigenvalues ​​and volume eigenvalues ​​of the audio signals of all microphones at the target frame length, a speech signal weight vector is determined; based on the speech signal weight vector and the audio signal matrix at the target frame length, the user voice signal at the target frame length is determined.

2. The method for processing audio signals of a smart speaker according to claim 1, wherein: The specific steps of determining the volume characteristic value of the audio signal of each microphone at each frame length include the following: Determine the azimuth angle interval energy of each microphone in the current frame length according to the audio signals of the microphones at different azimuth angles in the current frame length; Using the user's pure voice signal as input, the autocorrelation algorithm is used to obtain the user's fundamental frequency; The preset frequency range size A is centered on the user's fundamental frequency and a target frequency range of size A is constructed; For the audio signal of each microphone in the current frame length, a fast Fourier transform algorithm is used to obtain energy values ​​at several different frequencies. The sum of the energy values ​​at all frequencies within the target frequency range is used as the user voice energy of each microphone in the current frame length. The volume characteristic value of the audio signal of each microphone in the current frame length is determined based on the difference between the azimuth angle interval energies of the microphones in the current frame length and the difference between the user voice energies.

3. The method for processing audio signals of a smart speaker according to claim 2, wherein: The specific steps of determining the azimuth angle interval energy of each microphone under the current frame length according to the audio signals of the microphones at different azimuth angles under the current frame length are as follows: The MUSIC algorithm uses the azimuth angles of all microphones and the audio signals of all microphones at the current frame length as input, and outputs the spatial spectrum energy at the current frame length; The spatial spectrum energy is composed of energy values ​​at different azimuth angles; Preset angle interval size , with the azimuth angle of any microphone as the center, the size of the construction is Azimuth angle interval; In the spatial spectrum energy under the current frame length, the sum of the energy values ​​under all azimuth angles in the azimuth angle interval corresponding to each microphone is used as the azimuth angle interval energy of each microphone under the current frame length.

4. The method for processing audio signals of a smart speaker according to claim 2, wherein: The method of determining the volume characteristic value of the audio signal of each microphone at the current frame length according to the difference between the azimuth angle interval energies of the microphones at the current frame length and the difference between the user voice energies includes the following specific steps: At the current frame length, calculate the difference between the azimuth angle interval energy of each microphone and the average azimuth angle interval energy of all microphones, then calculate the ratio of the user voice energy of each microphone to the average user voice energy of all microphones, and input the product of the difference and the ratio into an exponential function with a natural constant as the base to obtain the output value, which is recorded as the volume characteristic value of the audio signal of each microphone at the current frame length.

5. The method for processing audio signals of a smart speaker according to claim 1, wherein: The specific steps of filtering out the target frame length are as follows: Get the range of the volume characteristic values ​​of the audio signals of all microphones under the current frame length, and record it as the volume characteristic range under the current frame length ; In the historical audio signal set, the volume feature range size under each historical frame length is obtained according to the method for obtaining the volume feature range size under the current frame length; Preset the first constant C1 to obtain the mean value of the volume feature range size under all historical frame lengths and standard deviation , if the volume feature range size under the current frame length Subtract the mean of the volume feature range size under all historical frame lengths The difference is greater than C1 times the standard deviation , then the current frame length is recorded as the target frame length.

6. The method for processing audio signals of a smart speaker according to claim 1, wherein: The specific steps of determining the voiceprint feature value of the audio signal of each microphone at the target frame length include the following: Preset the second constant C2, and obtain the sequence of Mel-frequency cepstral coefficients of the first C2 orders of the audio signal of each microphone under the target frame length, and record it as the voiceprint sequence of each microphone under the target frame length; Obtain a voiceprint sequence of the user's pure voice signal according to the method for obtaining the voiceprint sequence of each microphone at the target frame length; The cosine similarity between the voiceprint sequence of each microphone and the voiceprint sequence of the user's clean speech signal at the target frame length is obtained, and the sum of the cosine similarity and a preset third constant is recorded as the voiceprint feature value of the audio signal of each microphone at the target frame length.

7. The method for processing audio signals of a smart speaker according to claim 1, wherein: The method of constructing an audio signal matrix using the audio signals of all microphones at a target frame length according to the numbering order of the microphones includes the following specific steps: At the target frame length, the audio signals of each microphone are arranged row by row from top to bottom according to the numbers of all microphones from small to large to form an audio signal matrix.

8. The method for processing audio signals of a smart speaker according to claim 1, wherein: The specific steps of determining the speech signal weight vector are as follows: At the target frame length, obtain the sum of the voiceprint eigenvalue and volume eigenvalue of the audio signal of each microphone, record it as the comprehensive eigenvalue of each microphone, calculate the ratio of the comprehensive eigenvalue of each microphone to the mean of the comprehensive eigenvalues ​​of all microphones, and record it as the weight of each microphone at the target frame length; Under the target frame length, the weights of each microphone are arranged in a row from left to right according to the numbers of all microphones from small to large, forming a weight vector; The FastICA algorithm is used to iterate the weight vector at the target frame length. When the Euclidean distance between the weight vectors of two consecutive update iterations is less than the preset convergence threshold, the update iteration is stopped and the weight vector of the last update iteration is recorded as the speech signal weight vector at the target frame length.

9. The method for processing audio signals of a smart speaker according to claim 1, wherein: The specific steps of determining the user voice signal at the target frame length include the following: The product of the speech signal weight vector at the target frame length and the audio signal matrix is ​​recorded as the user speech signal at the target frame length.

10. A smart speaker comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is executed by a processor, the steps of the smart speaker audio signal processing method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Microphone equipment anti-interference method and microphone equipment

    CN107682786A

  • Semantic recognition device and method for tracking target person

    CN107862060A