Sound signal processing system, method and smart speaker
By using FM technology to return speaker audio data in smart speakers and voice interaction devices, the problem of interference between devices is solved, and low-latency voice interaction and echo cancellation is achieved, improving the ease of use of the device.
Patent Information
- Application Number
- CN202010704870.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-07-21
AI Technical Summary
Smart speakers and other voice interactive devices may interfere with each other when they work at the same time, affecting their normal work.
Frequency modulation (FM) technology is used to pass the acquired speaker audio data back and forth, ensuring the low delay of back-passing and no high bandwidth and low delay circuitry is required, thereby achieving the fast response of voice interactive devices to user voice.
Through FM technology, the low-latency and reliable return of the associated speaker speaker signal is achieved, and accurate voice interaction and echo cancellation operations are supported, improving the overall ease of use of multiple voice interaction devices.
Smart Images

Figure CN113963711B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of signal processing, and in particular, to a sound signal processing system, method, and smart speaker. Background Art
[0002] With the development of voice interaction and Internet technologies, smart speakers that can use voice functions for information acquisition, music playback, and Internet of Things device management have become popular. In a home usage environment, users usually purchase a smart speaker and place it in the area where daily activities occur most frequently, such as the living room. Further, in order to provide a full-range voice interaction function, other voice interaction devices can also be arranged in addition to the living room smart speaker serving as the central node, such as other smart speakers, or simpler voice interaction devices such as smart voice stickers or smart voice switches.
[0003] These voice interaction devices can respectively provide sound playback and voice interaction functions to users in their respective corresponding areas. However, when these devices work simultaneously, they may interfere with other voice interaction devices, which is not conducive to the normal operation of these devices.
[0004] Therefore, an improved sound signal processing solution is needed. Summary of the Invention
[0005] One technical problem to be solved by the present disclosure is to provide an improved sound signal processing solution that uses FM to backhaul the collected speaker audio data, ensuring low latency in backhaul while avoiding the need for conventional high-bandwidth low-latency circuits, thereby ensuring the fast response of voice interaction devices to user voices and relatively low costs.
[0006] According to a first aspect of the present disclosure, there is provided a sound signal processing system, including a main speaker and a secondary speaker, wherein the secondary speaker is configured to: collect a secondary speaker signal; perform FM modulation on the collected secondary speaker signal and transmit it, and the main speaker is configured to: receive the FM signal sent by the secondary speaker; and perform an echo cancellation operation based on the FM signal.
[0007] According to a second aspect of the present disclosure, there is provided a sound signal processing method, including: the secondary speaker collects a secondary speaker signal; the secondary speaker performs FM modulation on the collected secondary speaker signal and transmits it; the main speaker receives the FM signal sent by the secondary speaker; the main speaker performs an echo cancellation operation based on the FM signal.
[0008] According to a third aspect of the present disclosure, there is provided a smart speaker, comprising: a speaker for playing an audio signal; a speaker signal acquisition unit for acquiring a speaker signal; an FM receiving unit for receiving an FM signal containing an associated speaker signal of the speaker; a microphone unit for acquiring a microphone signal; and a processor for performing the echo cancellation operation on the microphone signal based on the speaker signal and the associated speaker signal of the speaker.
[0009] According to a fourth aspect of the present disclosure, there is provided a playback device, comprising: a communication unit for receiving an audio signal; a speaker for playing the received audio signal; a speaker signal acquisition unit for acquiring a speaker signal; and an FM transmitting unit for transmitting an FM signal containing the speaker signal.
[0010] According to a fifth aspect of the present disclosure, there is provided an FM transceiver component, comprising: an FM transmitting unit for acquiring and transmitting a speaker signal of a first device; and an FM receiving unit for receiving the speaker signal of the first device and transmitting the speaker signal to a second device for performing an echo cancellation operation. The above FM transmitting unit and FM receiving unit can be externally connected to the first device and the second device (e.g., a smart speaker) respectively, or be built-in units of the first device or the second device, and perform corresponding FM signal transmission and reception when performing the echo cancellation operation.
[0011] According to a sixth aspect of the present disclosure, there is provided a smart speaker, comprising: a WiFi unit for acquiring an audio signal including at least two channels; a speaker for playing a first channel signal of the audio signal; a Bluetooth unit for transmitting other channel signals of the audio signal to an associated speaker; a speaker signal acquisition unit for acquiring a speaker signal; an FM receiving unit for receiving an FM signal containing an associated speaker signal of the speaker; a microphone unit for acquiring a microphone signal; and a processing unit for performing the echo cancellation operation on the microphone signal based on the speaker signal and the associated speaker signal of the speaker.
[0012] According to a seventh aspect of the present disclosure, there is provided a method for processing a sound signal, comprising: acquiring a user's instruction to play music; searching for a corresponding audio signal; allocating at least two channels; playing an audio signal of a first channel, and transmitting at least one second channel so that a sub-speaker simultaneously plays the audio signal of the second channel; acquiring a speaker signal of the audio signal of the first channel and receiving an FM signal via FM, the FM signal containing a speaker signal of the audio signal of the second channel; acquiring a microphone signal; and performing the echo cancellation operation on the microphone signal based on the speaker signal of the audio signal of the first channel and the speaker signal of the audio signal of the second channel.
[0013] According to an eighth aspect of the present disclosure, there is provided a computing device, including: a processor; and a memory storing executable code thereon, which when executed by the processor, causes the processor to execute the method as described in the seventh aspect above.
[0014] According to a ninth aspect of the present disclosure, there is provided a non-transitory machine-readable storage medium storing executable code thereon, which when executed by a processor of an electronic device, causes the processor to execute the method as described in the seventh aspect above.
[0015] Thus, the voice signal processing solution of the present invention can use a relatively low-cost FM circuit to achieve reliable low-latency feedback of associated speaker signals, facilitating echo cancellation operations required for accurate voice interaction and improving the overall usability of multiple voice interaction devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0017] Figure 1 The composition schematic diagram of a voice signal processing system according to an embodiment of the present invention is shown.
[0018] Figure 2 An example of the composition of the FM transmission unit is shown.
[0019] Figure 3 An example of the composition of a voice signal processing system including multiple sub-speakers is shown.
[0020] Figure 4 The schematic flowchart of a voice signal processing method according to an embodiment of the present invention is shown.
[0021] Figure 5 The block diagram of the composition of a smart speaker according to an embodiment of the present invention is shown.
[0022] Figure 6 An example of the internal composition of a smart speaker is shown.
[0023] Figure 7 An example of the internal composition of a playback device according to an embodiment of the present invention is shown.
[0024] Figure 8 An example of the method flow for monitoring whether a user has a voice command input while playing audio is shown.
[0025] Figure 9The figure shows a schematic structural diagram of a computing device that can be used to implement the above-mentioned voice signal processing method according to an embodiment of the present invention. Detailed implementation manners
[0026] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0027] In order to provide a voice interaction function in the entire home domain, in addition to the smart speaker in the living room as the central node, the user can arrange other voice interaction devices, for example, other smart speakers, or voice interaction devices with simpler functions, such as smart voice stickers or smart voice switches. These voice interaction devices can respectively provide sound playback and voice interaction functions to the user in their respective corresponding areas. However, when these devices work simultaneously, they may interfere with other voice interaction devices, which is not conducive to the normal operation of these devices.
[0028] For this reason, an improved voice signal processing scheme is needed. The present invention uses FM to backhaul the collected speaker audio data, ensuring low latency in the backhaul while avoiding the need for conventional high-bandwidth low-latency circuits, thereby ensuring the quick response of the voice interaction device to the user's voice and relatively low cost.
[0029] Figure 1 The figure shows a schematic composition diagram of a voice signal processing system according to an embodiment of the present invention. As shown in the figure, the system 100 includes a main speaker 110 and a secondary speaker 120. Both the main speaker and the secondary speaker include speakers (loudspeakers) 111 and 121 for playing sounds (including voices and music).
[0030] Among them, the secondary speaker 120 can collect the speaker signal emitted by the secondary speaker 121, and perform FM modulation on the collected secondary speaker signal and send it. For example, it is sent to the FM transmission unit (FMTX) 122 for transmission. Correspondingly, the main speaker 110 can receive the FM signal sent by the secondary speaker 110 via, for example, the FM reception unit (FMRX) 112, and operate based on the FM signal, such as an echo cancellation operation.
[0031] "Echo cancellation" can specifically refer to Acoustic Echo Cancellation (AEC) here. As the name implies, it is a technology used to eliminate acoustic echoes. Acoustic echoes are caused by the sound of the speaker being fed back to the microphone. Figure 1In the system 100 shown, the sound being played by the speaker 121 of the secondary speaker 120 is fed back to the microphone (mic) 113 of the primary speaker. For this purpose, it is necessary to collect the sound signal of the secondary speaker 121, transmit it back to the primary speaker, and process it via the processing circuit (e.g., CPU 114) of the primary speaker to cancel the sound information of the secondary speaker 121 contained in the data collected by the microphone.
[0032] In the present invention, FM (Frequency Modulation), that is, frequency modulation, also known as "frequency modulation" technology, is adopted to transmit the secondary speaker signal. As is well known, frequency modulation is a modulation method in which the instantaneous frequency of a carrier changes according to the variation law of the signal to be transmitted. In the present invention, the secondary speaker signal can be directly input into the FMTX122 through a voltage-dividing filter circuit for transmission.
[0033] The FM transmitting unit 122 can have a very simple structure. Figure 2 A composition example of the FM transmitting unit is shown. As shown in the figure, the FMTX222 can be composed of a crystal oscillator (e.g., a 24 MHz crystal oscillator), an FM chip, and an FM antenna. Figure 1 The FM RX112 shown can have a corresponding simple receiving structure. As Figure 2 Although the FM transmitting unit shown has a simple composition and low configuration cost, it can meet the requirements of low-latency and high-bandwidth data transmission in combination with the FM receiving unit, usually has a transmission distance of more than 10 meters, and the configuration of the FM unit also has less impact on the software systems of the primary and secondary speakers.
[0034] In contrast, the costs of UWG and 5G circuits that can meet the requirements of high-bandwidth and low-latency network backhaul are relatively high. For example, the near-field transmission of LBRT has too short a transmission distance (usually not exceeding 0.5 meters), and WiFi backhaul cannot achieve sufficient stability while ensuring low latency, while Bluetooth and Zigbee cannot meet the low-latency requirements. Therefore, by using FM technology for the backhaul of speaker data in the present invention, the low-latency requirements in AEC processing can be achieved at extremely low cost and with less modification to the system.
[0035] The scenarios that require acoustic echo cancellation are usually those where the speaker is playing (e.g., playing music), so it is necessary to cancel the sound with the speaker sound signal. Thus, in a common application scenario, the primary speaker 110 and the secondary speaker 120 can be used to obtain audio signals and play the audio signals of their respective channels according to the configuration. In different embodiments, the primary speaker 110 and the secondary speaker 120 can obtain the audio signals to be played based on different mechanisms. It should be understood that in the present invention, there is no wired connection between the primary speaker 110 and the secondary speaker 120, but the signal is transmitted wirelessly, for example, via FM for speaker signal transmission.
[0036] In one embodiment, the main speaker 110 and the secondary speaker 120 may be a set of kits, and the default pairing between the two speakers may be completed at the time of shipment, such as Bluetooth pairing. In other embodiments, the main speaker 110 and the secondary speaker 120 may be speakers purchased separately by the user, for example, and the user may pair the two speakers, such as Bluetooth pairing. In one embodiment, the main speaker 110 and the secondary speaker 120 may be devices with the same software and hardware configuration, and the identities of the main speaker and the secondary speaker may be determined based on settings or dynamically. In another embodiment, the main speaker 110 and the secondary speaker 120 may have different software and hardware configurations. For example, the main speaker 110 may have more powerful functions, such as a smart speaker equipped with a WiFi module and more powerful processing functions, while the secondary speaker 120 may be another intelligent voice interaction device with a simpler structure and function and without a WiFi module, such as a smart voice sticker.
[0037] The main speaker 110 and the secondary speaker 120 may determine their respective channel playback configurations during pairing, or may dynamically determine their respective channel configurations based on, for example, the user's current location. In a preferred embodiment, the main speaker 110 and the secondary speaker 120 may operate similar to TWS (True Wireless Stereo) headphones. To this end, the main speaker 110 may obtain the audio signal to be played from the server via the WiFi component 115, for example. The audio signal may include the audio signals of the left and right channels, and the audio signal of the corresponding channel may be sent to the secondary speaker 120 via the Bluetooth (BT) component 115. The secondary speaker 120 may obtain the audio signal of the secondary speaker sent by the main speaker via its own equipped Bluetooth (BT) component 123 and play it, and at the same time may collect the played audio signal of the secondary speaker as the speaker signal of the secondary speaker. In other embodiments, the secondary speaker 120 may also obtain the music information of the channel to be played by itself from the server (in this case, the secondary speaker 120 also needs to be equipped with a WiFi component, which will greatly increase the cost, so it is not a preferred solution).
[0038] Here, the audio signal may refer to the audio signal to be played by the secondary speaker 120, or the audio signal played simultaneously by the main speaker 110 and the secondary speaker 120 according to channel allocation. The audio signal may be an audio signal stored locally in the main speaker or the secondary speaker, or an audio signal obtained from the server or other local storage institutions. The audio signal may be a simple audio signal, such as a music signal, and may include human voices or pure music, or may be the audio part of an audio-visual signal, such as the multi-channel audio information in a stereo movie resource.
[0039] When the main speaker 110 also plays audio, in order to perform echo cancellation operations, it is also necessary to collect the speaker signals of the main speaker. At this time, the main speaker 110 is used to: obtain and play the main speaker audio signal, for example, the audio signal corresponding to the channel of the main speaker; collect the played main speaker audio signal as the main speaker speaker signal; and perform the echo cancellation operation based on the main speaker speaker signal.
[0040] Further, the main speaker 110 can use the microphone unit 113 to collect the main speaker microphone signal; and perform the echo cancellation operation on the main speaker microphone signal based on the secondary speaker speaker signal and the main speaker speaker signal.
[0041] It should be understood that for a speaker with a microphone, the signal collected by its microphone is the ambient sound including the audio played by the speaker, while its speaker signal collection unit directly obtains the audio signal played by the speaker. In a system as Figure 1 shown, if the main speaker 110 and the secondary speaker 120 are playing music signals synchronously by channels, and at the same time the user utters a voice command (for example, a wake-up word), then the signal collected by the microphone unit 113 at this time includes the signal played by the speaker 111 of the main speaker 110, the signal played by the speaker 121 of the secondary speaker 120, and the user's voice signal. In this way, based on the directly collected secondary speaker speaker signal and the main speaker speaker signal, the signal played by the speaker 111 of the main speaker 110 and the signal played by the speaker 121 of the secondary speaker 120 included in the microphone signal can be eliminated.
[0042] Thus, the main speaker 110 can perform a voice recognition operation on the microphone signal after the echo cancellation operation, for example, determine whether the microphone signal after the echo cancellation operation includes a wake-up word, or perform further voice recognition after waking up.
[0043] In some embodiments, the secondary speaker 120 can also be equipped with a microphone to implement a voice interaction function. Further, the secondary speaker 120 can use the microphone unit to collect the secondary speaker microphone signal. Subsequently, the main speaker or the secondary speaker performs the echo cancellation operation on the secondary speaker microphone signal based on the secondary speaker speaker signal and the main speaker speaker signal. The above echo cancellation and voice recognition operations for the secondary speaker microphone signal can be a supplement or replacement for the echo cancellation and voice recognition operations for the main speaker microphone signal. For example, if the user issues a voice command at a position closer to the secondary speaker, it is more likely to detect a meaningful signal by performing voice recognition using the secondary speaker microphone signal.
[0044] As described above, the main speaker 110 and the secondary speaker 120 can be a kit that has completed network configuration at the time of leaving the factory, or independent devices sold separately. In either case, when the main speaker 110 is used for the first time, it needs to be powered on for network configuration. For example, the authentication and network configuration with the cloud server are completed using the WiFi component 115. The main speaker 110 described above can be, for example, a smart speaker used as a home voice intelligent interaction center. When the secondary speaker needs to be network-configured, the main speaker can receive the network configuration request of the secondary speaker and send channel configuration information to the network-configured secondary speaker.
[0045] Further, after the main speaker is network-configured or connected to the secondary speaker, the main speaker 110 can scan available FM transmission frequencies and send the selected FM transmission frequency information to the secondary speaker. Thus, the main speaker 110 and the secondary speaker 120 can agree on the working frequency of the FM for speaker signal transmission.
[0046] After the network configuration is completed, the main speaker can receive a voice command from the user to play audio, obtain the audio signal based on the above command, and forward the audio signal of the corresponding channel to the secondary speaker. At this time, the main speaker 110 and the secondary speaker 120 can synchronously play the audio information of their respective corresponding channels, and keep collecting their respective speaker signals and the main speaker microphone signal, as well as performing AEC operations, so as to determine in real time whether the user has issued a voice command when playing the audio signal, for example, uttered a wake-up word.
[0047] In one embodiment, the system can include multiple secondary speakers. These secondary speakers can be network-configured simultaneously or one by one, and the channel allocation can be dynamically determined according to the number and position of the current secondary speakers. At this time, each secondary speaker can be used to: play the audio signal of the corresponding channel; collect its own secondary speaker signal; perform FM modulation on the collected secondary speaker signal of its own and send it. The main speaker can be used to: receive the FM signals from each secondary speaker simultaneously; and perform echo cancellation operations based on the FM signals.
[0048] When there are multiple secondary speakers, the main speaker needs to allocate different FM transmission frequencies to each secondary speaker and use different FM receivers to receive the FM signals sent by each secondary speaker. Figure 3 Shows a composition example of a sound signal processing system including multiple secondary speakers. As shown in the figure, in addition to including the main speaker 310 and the secondary speaker 320, the system 300 can also include a secondary speaker 330, and the secondary speaker 330 can form a 3-channel stereo playback system with the main speaker 310 and the secondary speaker 320. For this reason, in order to perform echo cancellation on the signal from the secondary speaker 330, the main speaker 310 also needs to additionally include a set of FM receivers 312', and ensure that the two FMRXs 312 and 312' work at different frequencies respectively.
[0049] Further, in addition to the horn 321 and the Bluetooth component 323 which are the same as those shown in Figure 1 , the slave speaker 320 may further include an audio power amplifier (AUPA) 324, which is used to amplify the audio signal obtained from the Bluetooth component 323 and send it to the horn 321 for playing. The slave speaker horn signal can be directly connected to the input end of the FM transmitting chip as shown in Figure 2 through a voltage dividing and filtering circuit, and sent to the main speaker 310 through an FM antenna.
[0050] Similarly, the main speaker 310 may also be equipped with an audio power amplifier (AUPA) 316, and may also be equipped with an ADC chip 317. When the FM receiving chip of the main speaker 310 receives an FM signal, the audio signal decoded by the receiving chip can be sent into the ADC chip 317 through an audio output path. The ADC 317 can also obtain the horn signal fed back by the main speaker, the microphone signal collected by the microphone 313, and optionally the audio signal from the slave speaker 330. These signals can be sent into the processor (CPU) 314 together for voice recognition. When the main speaker 310 is idle, it can scan the FM frequency in the background to update the frequency suitable for FM transmission of the horn signal. For example, find one or more clean FM frequencies under the current interference frequency.
[0051] As described above in combination with Figures 1 - 3 , a sound signal processing system suitable for implementing the present invention and its preferred embodiments are described. The solution of the present invention can also be implemented as a sound signal processing method. Figure 4 FIG. shows a schematic flowchart of a sound signal processing method according to an embodiment of the present invention.
[0052] As shown in the figure, in step S410, the slave speaker collects the slave speaker horn signal. In step S420, the slave speaker performs FM modulation on the collected slave speaker horn signal and sends it. In step S430, the main speaker receives the FM signal sent by the slave speaker. In step S440, the main speaker performs an echo cancellation operation based on the FM signal.
[0053] Here, the main speaker and the slave speaker can obtain audio signals and play the audio signals of their respective channels according to the configuration. At this time, in step S410, the slave speaker can collect the slave speaker audio signal as the slave speaker horn signal.
[0054] When the main speaker is also playing audio, the AEC operation also needs to collect the signal of the main speaker's speaker. To this end, the method may further include: collecting the played main speaker audio signal as the main speaker's speaker signal; collecting the main speaker microphone signal; and performing the echo cancellation operation on the main speaker microphone signal based on the secondary speaker's speaker signal and the main speaker's speaker signal for speech recognition.
[0055] Further, when the secondary speaker also has the function of voice interaction, the secondary speaker can also collect the microphone signal and perform speech recognition. To this end, the method may further include: collecting the secondary speaker microphone signal; and performing the echo cancellation operation on the secondary speaker microphone signal based on the secondary speaker's speaker signal and the main speaker's speaker signal for speech recognition.
[0056] Before audio playback, the main speaker can perform a network configuration operation. For example, a cloud network configuration operation via a WiFi component. Further, the main speaker can also receive a network configuration request from the secondary speaker and send channel configuration information and / or FM transmission frequency information to the network-configured secondary speaker.
[0057] In the sound signal processing system of the present invention, the "secondary speaker" for FM modulation and transmission and the "main speaker" for FM reception can be determined from multiple speakers according to different strategies. In one embodiment, the main speaker can be the default main speaker connected to the WiFi network. The main speaker can obtain a multi-channel audio signal and send its corresponding audio signal to the secondary speaker (e.g., via Bluetooth). To this end, when performing the echo cancellation operation, the secondary speaker in the audio playback is also defaulted to be the secondary speaker for FM transmission.
[0058] In another embodiment, the main speaker can be the main speaker set by the factory settings or specified by the user. For example, one of a pair of Bluetooth speakers has been set as the main speaker at the factory, or is set as the main speaker by the user when powered on / network configured, and then is also regarded as the main speaker for calculation operations when performing the echo cancellation operation.
[0059] In yet another embodiment, the main speaker and the secondary speaker can be dynamically determined by the current scenario. For example, the speaker closest to the user is set as the main speaker, and the speaker farther from the user is set as the secondary speaker, and the main speaker receives the FM signal from the secondary speaker for echo cancellation operation.
[0060] The present invention can also be implemented as a smart speaker. Figure 5The block diagram of the composition of a smart speaker according to an embodiment of the present invention is shown. This smart speaker is particularly suitable for being implemented as the main speaker as described above. As shown in the figure, the smart speaker 510 may include a speaker 511, a microphone 513, a speaker signal acquisition unit 517, an FM receiving unit 512, and a processing unit 514. It should be understood that in different drawings of the present invention, corresponding devices or components have similar numbers.
[0061] Here, the speaker 511 can be used to play an audio signal. This audio signal is preferably an audio signal such as a song broadcast based on a user instruction, but it can also be a voice feedback signal for the user instruction. The speaker signal acquisition unit 517 can be used to acquire the speaker signal, for example Figure 3 the ADC (analog-to-digital converter) shown. The ADC 517 can directly obtain the analog signal of the speaker 511 through a voltage division circuit and convert it into a digital signal as the acquired speaker signal.
[0062] At the same time, the FM receiving unit 512 can receive an FM signal containing the associated speaker signal of the speaker, for example, an FM signal containing the speaker signal of the slave speaker sent by the slave speaker at a predetermined frequency. The microphone unit 513, which is usually configured as a microphone array, can acquire the microphone signal. The processing unit 514 can then perform the echo cancellation operation on the microphone signal based on the speaker signal and the associated speaker signal of the speaker.
[0063] To obtain the audio signal to be played, the smart speaker may further include a first networking unit. The first networking unit can be a networking module capable of accessing resources in the cloud, for example, a WiFi module. To synchronize audio playback with the slave speaker, the smart speaker may further include a second networking unit. The second networking unit can send at least a part of the acquired audio signal to the associated speaker, for example, send the audio signal of the corresponding channel of the song resource acquired by the first networking unit from the cloud to the slave speaker. The second networking unit communicates with the second networking unit of the associated speaker based on a short-range communication protocol and can particularly be a Bluetooth module.
[0064] Furthermore, the processing unit 514 can scan available FM transmission frequencies (for example, clean frequencies with less interference) and send the FM transmission frequency information to the associated speaker via the second networking unit. Thus, it is convenient for the subsequent two devices to perform FM signal transmission and reception at this frequency.
[0065] The processing unit 514 can also be used for: performing a speech recognition operation on the microphone signal after the echo cancellation operation, and the speaker is also used for: performing voice feedback based on the result of the speech recognition.
[0066] When collecting the microphone signal of the slave speaker for speech recognition, the second networking unit can also be used to receive the microphone signal of the associated speaker to perform the echo cancellation operation.
[0067] In addition, when free identity switching is required, the smart speaker can also include an FM transmitting unit for sending an FM signal containing the speaker signal to the associated speaker. At this time, the smart speaker can switch its identity to act as a slave speaker.
[0068] In addition, the master speaker can also interact with multiple slave speakers. At this time, the smart speaker can include multiple FM receiving units for simultaneously receiving the associated speaker horn information from different associated speakers at different frequencies respectively.
[0069] Specifically, Figure 6 An internal composition example of the smart speaker is shown. The internal composition of the smart speaker 610 is the same as that of Figure 3 the exemplified master speaker 310. As shown in the figure, the first networking unit and the second networking unit can be respectively implemented as WiFi and BT components, and can jointly form a networking unit 615. This unit can utilize the WiFi component to obtain cloud resources and utilize the BT component to communicate with local devices, such as other IoT devices including slave speakers. In a scenario such as playing a song, after the WiFi component obtains the audio information from the cloud, the CPU 614 can perform channel allocation, send the audio information played by the local device into the audio power amplifier (AU PA) 616 for the convenience of playing by the speaker 611. At the same time, the BT component transmits the audio that the slave speaker needs to play to the slave speaker for playing. While playing the audio, the slave speaker collects the speaker signal and sends an FM signal through its own FX TX. At the same time, the smart speaker 610 acting as the master speaker also uses the speaker signal collection unit 617 implemented as an ADC to collect the speaker signal of the local device, and uses the ADC 617 to perform analog-to-digital conversion on the signal received via FM RX (for example, extracted as an analog signal by the FM receiving chip). The microphone 613 can be kept on or activated periodically to collect ambient sound. The analog signal collected by the microphone 613 can also be subjected to analog-to-digital conversion by the ADC617. Thus, the CPU614 can process the three signals obtained through ADC conversion: the smart speaker speaker signal, the associated speaker speaker signal, and the microphone signal, so as to cancel the speaker signal component contained in the microphone signal through the smart speaker speaker signal and the associated speaker speaker signal, thereby being able to more accurately determine whether the microphone signal contains a voice command such as a wake word.
[0070] To this end, the present invention can also be implemented as a smart speaker, including: a WiFi unit for obtaining an audio signal including at least two channels; a speaker for playing the first channel signal in the audio signal; a Bluetooth unit for sending the other channel signals in the audio signal to an associated speaker; a speaker signal acquisition unit for acquiring a speaker signal; an FM receiving unit for receiving an FM signal including the associated speaker signal; a microphone unit for acquiring a microphone signal; and a processing unit for performing the echo cancellation operation on the microphone signal based on the speaker signal and the associated speaker signal.
[0071] Figure 7 FIG. shows an internal composition example of a playback device according to an embodiment of the present invention. The playback device 720 can be implemented as the slave speaker as described above. As shown in the figure, the playback device 710 may include a speaker 721, an audio power amplifier (AUPA) 724, an FM transmission unit 722, and a communication unit 723.
[0072] The communication unit can be implemented as a Bluetooth (BT) module 723 for receiving an audio signal, such as an audio signal sent by the master speaker Bluetooth module. The audio power amplifier (AUPA) 724 can amplify the received signal in power, and the amplified audio signal is played by the speaker 721. The FMTX 722 can be used to send an FM signal including the speaker signal.
[0073] In addition to receiving an audio signal, such as an audio signal corresponding to a channel, the BT module 723 can also be used for network configuration operation with the master speaker, and can receive FM transmission frequency information, so that under the control of the processing unit of the playback device, the FM transmission unit can send the FM signal at the above-mentioned agreed FM transmission frequency. In a playback device with relatively simple functions and structures, the processing unit can be implemented as an MCU (micro control unit).
[0074] Similarly, the playback device can also include a microphone for acquiring a microphone signal, and the BT module 723 can be used to send the microphone signal, for example, to the master speaker for AEC operation or voice recognition operation for the microphone signal of the playback device.
[0075] In one embodiment, the present invention can also be implemented as an FM transceiver component, including: an FM transmission unit for obtaining and sending a speaker signal of a first device; and an FM receiving unit for receiving the speaker signal of the first device and transmitting the speaker signal to a second device for performing an echo cancellation operation. The above FM transmission unit and FM receiving unit can be externally connected to the first device and the second device (for example, a smart speaker) respectively, or be built-in units of the first device or the second device, and perform corresponding FM signal transmission and reception during the echo cancellation operation.
[0076] The present invention is particularly applicable to be implemented as a solution for monitoring whether a user has a voice command input while playing audio. Thus, the present invention can be implemented as a sound signal processing method, which is suitable for being executed by a main speaker, such as the intelligent speaker described above, and includes: obtaining a user's instruction to play audio; searching for a corresponding audio signal; allocating at least two channels; playing the audio signal of the first channel and sending at least one second channel so that the slave speaker plays the audio signal of the second channel simultaneously; collecting the speaker signal of the audio signal of the first channel and receiving an FM signal via FM, where the FM signal includes the speaker signal of the audio signal of the second channel; collecting a microphone signal; and performing the echo cancellation operation on the microphone signal based on the speaker signal of the audio signal of the first channel and the speaker signal of the audio signal of the second channel. Subsequently, a voice recognition operation can be performed on the microphone signal that has undergone the echo cancellation operation; and feedback can be performed based on the result of the voice recognition.
[0077] Specifically, obtaining a user's instruction to play music may include: collecting a microphone signal; performing voice recognition on the music play instruction included in the microphone signal, and performing feedback based on the result of the voice recognition includes: performing voice feedback based on the result of the voice recognition.
[0078] Before performing the above playback and monitoring operations, WiFi network configuration can also be performed and available FM frequencies can be searched; and the selected FM frequency can be sent to the associated speaker from which the FM signal is to be received.
[0079] Figure 8 A method flow example of monitoring whether a user has a voice command input while playing audio is shown.
[0080] This method can be particularly implemented by Figure 3 the system shown. This system can be implemented as a multi-channel intelligent speaker and includes a main speaker and at least one slave speaker. The main speaker can include a CPU, a BT module, a WiFi module, an ADC, an audio power amplifier (AUPA), an FMRX, a speaker (loudspeaker), and a microphone unit (MIC) as shown in Figure 3 shown. The slave speaker can include a micro control unit (MCU), a BT module, an audio power amplifier (AUPA), a speaker (loudspeaker), and an FMTX.
[0081] In step S810, the main speaker performs WiFi network configuration and can search for clean FM frequencies. For example, when the user gets the intelligent speaker, network configuration can be performed at the first power-on, for example, performing authentication operations with a cloud server, etc. At this time, the intelligent speaker acting as the main speaker can perform an FM reception scan in the background and select one or more frequencies with clean FM signals.
[0082] If the slave speaker needs to be networked, perform the networking operation. If the networking has been completed in a previous stage, for example, the networking of the master speaker and the slave speaker as a kit has been completed at the factory, then in step S820, the master speaker can synchronize the selected FM frequency to the slave speaker, for example, through Bluetooth to one or more slave speakers, as the FM transmission frequency of the slave speaker. In the case of including multiple slave speakers, different slave speakers use different FM transmission frequencies, and the master speaker includes FM receiving units corresponding to their respective frequencies.
[0083] After successful networking, in step S830, the user wakes up the master speaker and makes an audio playback request, such as a song playback request. At this time, the user can call the smart speaker to play a song, for example, "****(wake word), play XXXX (song name) by XXX (singer name)".
[0084] The master speaker can recognize and respond to the above voice command of the user. Then, in step S840, the master speaker downloads the song from the cloud through WiFi, configures the sound channels, synchronizes them to the slave speaker through Bluetooth, and performs sound synchronization (similar to TWS earphones), and the master and slave speakers play the song together. For example, the master speaker can obtain the audio information of XXXX (song name), for example, the information of the left and right channels. In the case where the master speaker is configured to play the left channel and the slave speaker is configured to play the right channel, the master speaker can synchronize the audio information of the right channel to the slave speaker via the Bluetooth module, so that the master speaker and the slave speaker play the audio of the left and right channels synchronously, thus realizing stereo playback.
[0085] In step S850, while playing the song, the slave speaker collects the speaker signal and transmits it through FM. Specifically, the speaker signal of the slave speaker can be connected to the input end of the FM transmitting chip through a voltage dividing and filtering circuit and transmitted to the master speaker through FM. At this time, the master speaker can also keep collecting its own speaker signal, for example, through the ADC circuit 317 as described above.
[0086] In step S860, the FM receiving chip of the master speaker receives the FM signal, gives the audio signal to the ADC chip through the audio output path, and in step S870, it is input to the CPU together with the sound collected by the master speaker and the microphone sound connected to the ADC. In step S880, the CPU processes the above signals for voice recognition, for example, to determine whether it contains the wake word. In addition, the speaker can also scan the FM frequency in the background when not playing a song to adapt to the change of interference frequency.
[0087] When the multi-channel smart speaker implemented according to the present invention plays a song in response to a voice command from the user, such as "**** (wake word), play XXXX (song name) by XXX (singer name)", it can obtain the speaker signals of the main speaker and the sub-speaker and use them to cancel the speaker signals in the microphone signal, so as to monitor in real time whether the user has a new voice command (such as a wake word) input. When detecting, for example, that the user inputs a wake word, the speaker can stop the currently playing music and switch to providing a voice response to the wake word, such as "Yes" or "How can I help you"; or it can provide a voice response such as "Yes" while continuing to play the music.
[0088] Figure 9 FIG. shows a schematic structural diagram of a computing device that can be used to implement the above-mentioned sound signal processing method according to an embodiment of the present invention. For example, the computing device can be implemented as a smart speaker and includes a speaker, a loudspeaker, and an FM transceiver unit.
[0089] See Figure 9 , the computing device 900 includes a memory 910 and a processor 920.
[0090] The processor 920 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 920 can include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a digital signal processor (DSP), and so on. In some embodiments, the processor 920 can be implemented using custom circuits, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).
[0091] The memory 910 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 920 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 910 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 910 can include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and instantaneous electronic signals transmitted wirelessly or wiredly.
[0092] Executable code is stored on the memory 910, and when the executable code is processed by the processor 920, it can cause the processor 920 to execute the sound signal processing method described above.
[0093] The sound signal processing solution according to the present invention has been described in detail above with reference to the accompanying drawings. The sound signal processing solution of the present invention can use a relatively low-cost FM circuit to achieve low-latency and reliable feedback of the associated speaker signals, thereby facilitating the echo cancellation operation required for accurate voice interaction and improving the overall usability of multiple voice interaction devices.
[0094] In addition, the method according to the present invention can also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present invention.
[0095] Alternatively, the present invention can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) having executable code (or computer program, or computer instruction code) stored thereon. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), the processor is caused to execute each step of the above-described method according to the present invention.
[0096] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both.
[0097] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a part thereof that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0098] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A voice signal processing system, including a main speaker and a secondary speaker, wherein the secondary speaker is used for: Collecting the secondary speaker signal; Performing FM modulation on the collected secondary speaker signal and transmitting it; The main speaker is used for: Receiving the FM signal sent by the secondary speaker; and Performing an echo cancellation operation based on the FM signal; wherein, The secondary speaker signal is an audio signal played by the secondary speaker.
2. The system according to claim 1, wherein, The main speaker and the secondary speaker are used for: Obtaining an audio signal and playing the audio signal of each channel according to the configuration.
3. The system according to claim 2, wherein, The secondary speaker is used for: Obtaining and playing the secondary speaker audio signal sent by the main speaker; and Collecting the played secondary speaker audio signal as the secondary speaker signal.
4. The system according to claim 3, wherein, The main speaker is used for: Obtaining and playing the main speaker audio signal; Collecting the played main speaker audio signal as the main speaker signal; and Performing the echo cancellation operation based on the main speaker signal.
5. The system according to claim 3, wherein, The main speaker is used for: Collecting the main speaker microphone signal by using the microphone unit; and Performing the echo cancellation operation on the main speaker microphone signal based on the secondary speaker signal and the main speaker signal.
6. The system according to claim 5, wherein, The main speaker is used for: Performing a speech recognition operation on the microphone signal after the echo cancellation operation.
7. The system according to claim 6, wherein, The main speaker is used for: Determining whether the microphone signal after the echo cancellation operation includes a wake word.
8. The system according to claim 5, wherein, The secondary speaker is used for: Collecting the secondary speaker microphone signal by using the microphone unit; and The main speaker or the secondary speaker performs the echo cancellation operation on the secondary speaker microphone signal based on the secondary speaker signal and the main speaker signal.
9. The system according to claim 2, wherein, The main speaker is used for: Receiving the network configuration request of the secondary speaker; and Sending channel configuration information to the network-configured secondary speaker.
10. The system according to claim 2, wherein, The main speaker is used for: Scanning available FM transmission frequencies; and Sending the selected FM transmission frequency information to the secondary speaker.
11. The system according to claim 2, wherein, The main speaker is used for: Obtaining the audio signal based on the received voice command to play the audio; and Forwarding the audio signal of the corresponding channel to the secondary speaker.
12. The system according to claim 2, wherein, The system includes multiple secondary speakers, and each secondary speaker is used for: Playing the audio signal of the corresponding channel; Collecting the respective secondary speaker signals; Performing FM modulation on the respectively collected secondary speaker signals and transmitting them; The main speaker is used for: Receiving the FM signals from each secondary speaker simultaneously; and Performing an echo cancellation operation based on the FM signals.
13. The system according to claim 1, wherein The main speaker is the default main speaker connected to the WiFi network; The main speaker is the main speaker set by the factory settings or specified by the user; and / or The main speaker and the secondary speaker are dynamically determined by the current scenario.
14. A method for processing sound signals, including: The secondary speaker collects the secondary speaker's speaker signal; The secondary speaker performs FM modulation on the collected secondary speaker's speaker signal and sends it; The main speaker receives the FM signal sent by the secondary speaker; The main speaker performs an echo cancellation operation based on the FM signal, wherein the secondary speaker's speaker signal is the audio signal played by the secondary speaker's speaker.
15. The method according to claim 14, further including: The main speaker and the secondary speaker obtain the audio signal and play the audio signals of their respective channels according to the configuration, wherein, The secondary speaker collects the secondary speaker's audio signal as the secondary speaker's speaker signal.
16. The method according to claim 15, further including: Collect the main speaker's audio signal played as the main speaker's speaker signal; Collect the main speaker's microphone signal; and Based on the secondary speaker's speaker signal and the main speaker's speaker signal, perform the echo cancellation operation on the main speaker's microphone signal for speech recognition.
17. The method according to claim 16, further including: Collect the secondary speaker's microphone signal; and Based on the secondary speaker's speaker signal and the main speaker's speaker signal, perform the echo cancellation operation on the secondary speaker's microphone signal for speech recognition.
18. The method according to claim 14, further including: Receive the network configuration request of the secondary speaker; and Send the channel configuration information and / or FM transmission frequency information to the network-configured secondary speaker.
19. An intelligent speaker, including: A speaker for playing the audio signal; A speaker signal collection unit for collecting the speaker signal; An FM receiving unit for receiving the FM signal containing the associated speaker's speaker signal; A microphone unit for collecting the microphone signal; and A processing unit for performing an echo cancellation operation on the microphone signal based on the speaker signal and the associated speaker's speaker signal, wherein the speaker signal is the audio signal played by the speaker.
20. The intelligent speaker according to claim 19, further including: A first networking unit for obtaining the audio signal to be played.
21. The intelligent speaker according to claim 20, further including: A second networking unit for sending at least a part of the obtained audio signal to the associated speaker.
22. The intelligent speaker according to claim 21, wherein, The first networking unit is used to obtain the audio signal from the server, and the second networking unit communicates with the second networking unit of the associated speaker based on the short-range communication protocol.
23. The intelligent speaker according to claim 21, wherein, The processing unit is further used for: Scanning the available FM transmission frequencies and sending the FM transmission frequency information to the associated speaker via the second networking unit.
24. The intelligent speaker according to claim 21, wherein, The second networking unit is further used to receive the microphone signal of the associated speaker for performing the echo cancellation operation.
25. The smart speaker according to claim 19, wherein, the processing unit is further configured to: perform a speech recognition operation on the microphone signal that has undergone the echo cancellation operation, the speaker is further configured to: perform a voice feedback based on the result of the speech recognition.
26. The smart speaker according to claim 19, further comprising: an FM transmitting unit configured to transmit an FM signal containing the speaker signal to the associated speaker.
27. The smart speaker according to claim 19, wherein, the smart speaker includes a plurality of FM receiving units configured to simultaneously receive associated speaker information from different associated speakers at different frequencies respectively.
28. An FM transceiver component, comprising: an FM transmitting unit configured to obtain a first device speaker signal and transmit it; and an FM receiving unit configured to receive the first device speaker signal and transmit the speaker signal to a second device for performing an echo cancellation operation, wherein the first device speaker signal is an audio signal played by a first device speaker.
29. A smart speaker, comprising: a WiFi unit configured to obtain an audio signal including at least two channels; a speaker configured to play a first channel signal of the audio signal; a Bluetooth unit configured to transmit other channel signals of the audio signal to an associated speaker; a speaker signal acquisition unit configured to acquire a speaker signal; an FM receiving unit configured to receive an FM signal containing an associated speaker's speaker signal; a microphone unit configured to acquire a microphone signal; and a processing unit configured to perform an echo cancellation operation on the microphone signal based on the speaker signal and the associated speaker's speaker signal, wherein the speaker signal is an audio signal played by the speaker.
30. A method for processing a sound signal, comprising: obtaining a user's instruction to play an audio; searching for a corresponding audio signal; allocating at least two channels; playing an audio signal of a first channel and transmitting at least one second channel so that a slave speaker simultaneously plays the audio signal of the second channel; acquiring a speaker signal of the audio signal of the first channel and receiving an FM signal via FM, the FM signal containing a speaker signal of the audio signal of the second channel; acquiring a microphone signal; and performing an echo cancellation operation on the microphone signal based on the speaker signal of the audio signal of the first channel and the speaker signal of the audio signal of the second channel, wherein the speaker signal is an audio signal played by a speaker.
31. The method according to claim 30, further comprising: performing a speech recognition operation on the microphone signal that has undergone the echo cancellation operation; performing a feedback based on the result of the speech recognition.
32. The method according to claim 31, wherein, obtaining a user's instruction to play music includes: acquiring a microphone signal; performing a speech recognition on a music playing instruction contained in the microphone signal, and performing a feedback based on the result of the speech recognition includes: performing a voice feedback based on the result of the speech recognition.
33. The method according to claim 30, further comprising: performing a WiFi network configuration and searching for available FM frequencies; and Send the selected FM frequency to the associated speaker from which the FM signal is to be received.
34. A computing device, comprising: a processor; and a memory storing executable code that, when executed by the processor, causes the processor to perform the method according to any one of claims 30-33.
35. A non-transitory machine-readable storage medium storing executable code that, when executed by a processor of an electronic device, causes the processor to perform the method according to any one of claims 30-33.
Citation Information
Patent Citations
One-to-many wireless frequency modulation (FM) loudspeaker box system
CN102572670A