Radio system, wearable device and device radio method

By combining the signal of the bone conduction microphone and the acoustic microphone in the wearable device to combine sound pickup and adjust the audio mode according to the scene signal, the problem of insufficient noise suppression ability in high-noise and high-wind environments is solved, and high-quality voice recognition effect is achieved.

CN119946495APending Publication Date: 2025-05-06ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411997611.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In extreme scenarios such as high noise and strong winds, existing wearable devices have limited noise suppression capabilities, which affects the voice recognition effect.

Method used

The signal of the bone conduction microphone and the acoustic microphone are fused to pick up sounds, and the controller determines the adapted audio mode according to the scene signal, recognizes different audio scenes and adapts to the corresponding audio modes.

Benefits of technology

It improves the high-quality sound pickup capability of the equipment in a highly noisy environment, enhances the noise suppression capability, and thus improves the voice recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946495A_ABST
    Figure CN119946495A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a sound receiving system, wearable equipment and an equipment sound receiving method, the sound receiving system comprises at least two acoustic microphones, at least one bone conduction microphone and a controller, the controller is used for determining a corresponding sound receiving mode according to a scene signal, the scene signal is used for representing a current sound receiving scene, the sound receiving scene comprises a noise scene, a wind noise scene and / or a preset sound receiving type scene, and the sound receiving mode adopts a microphone at a corresponding position to execute a sound receiving operation. Therefore, different sound receiving scenes can be recognized and corresponding sound receiving modes can be adapted, and meanwhile, the signals of the bone conduction microphone and the sound wave microphone are adopted to fuse sound pickup, so that high-quality sound pickup of the equipment in environments such as strong noise is realized, the noise suppression capability is improved, and the voice recognition effect is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and more specifically, to a sound receiving system, a wearable device and a device sound receiving method. Background Art

[0002] For most existing wearable devices, multiple acoustic microphones are usually sampled for sound collection. For example, smart glasses can be equipped with multiple acoustic microphones on the frame or temples. However, this type of microphone layout has a certain limit to its ability to suppress noise in extreme scenarios such as high noise and strong wind. Summary of the invention

[0003] In view of this, the embodiments of the present invention provide a sound reception system, a wearable device and a device sound reception method to identify different sound reception scenarios and adapt to corresponding sound reception modes, and at the same time use the signal fusion pickup of the bone conduction microphone and the acoustic wave microphone to achieve high-quality sound pickup of the device in environments such as strong noise, thereby improving the noise suppression capability, and further improving the speech recognition effect.

[0004] In a first aspect, an embodiment of the present invention provides a sound receiving system, the sound receiving system comprising:

[0005] at least two acoustic microphones;

[0006] at least one bone conduction microphone; and

[0007] A controller is used to determine a corresponding sound reception mode according to a scene signal, wherein the scene signal is used to characterize a current sound reception scene, wherein the sound reception scene includes a noise scene, a wind noise scene and / or a predetermined sound reception type scene, and the sound reception mode uses a microphone at a corresponding position to perform a sound reception operation.

[0008] Furthermore, the at least two acoustic wave microphones are distributed on both sides of the wearable device, and the bone conduction microphone is arranged in a predetermined area of ​​the wearable device, wherein at least a part of the predetermined area contacts the wearer after the wearable device is worn, or transmits a bone conduction signal with a strength that meets the conditions in a non-contact state.

[0009] Furthermore, the controller is used to receive the sound signals collected by each of the sound wave microphones, detect noise parameters of each of the sound signals, and determine the scene signal according to each of the noise parameters.

[0010] Further, the controller is used to determine that the sound recording scene is a noise scene in response to the noise parameter being greater than a first predetermined value, and to determine that the sound recording scene is a wind noise scene in response to the noise parameter representing wind noise and the noise parameter being greater than a second predetermined value; or,

[0011] The controller is used to determine whether the sound collection scene is a noise scene or a wind noise scene according to a predetermined noise classification model.

[0012] Furthermore, the controller is used to characterize wind noise in response to the noise parameters, determine the wind direction according to each noise parameter, and select a sonic microphone for performing a sound collection operation according to the wind direction.

[0013] Furthermore, the sound receiving system also includes:

[0014] The main control is used to perform signal fusion processing on the received sound signal to obtain the processed audio signal.

[0015] Furthermore, the controller is built into the main control, or the controller is an independent integrated circuit.

[0016] In a second aspect, an embodiment of the present invention provides a wearable device, the wearable device comprising:

[0017] The device body; and

[0018] A radio system as described above.

[0019] Furthermore, the wearable device is smart glasses, the sound wave microphones in the sound receiving system are distributed on the temples and / or the frames, and the bone conduction microphones in the sound receiving system are arranged in the area where the nose pads are located or the area where the temples are located.

[0020] Furthermore, the acoustic wave microphones are symmetrically distributed on the temples and / or frames on both sides.

[0021] In a third aspect, an embodiment of the present invention provides a device sound receiving method, which is applied to a sound receiving system of a wearable device, wherein the sound receiving system includes at least two sound wave microphones and at least one bone conduction microphone, and the method includes:

[0022] Collecting corresponding sound signals through each of the acoustic wave microphones and the bone conduction microphone;

[0023] Selecting a corresponding sound signal for processing based on a predetermined sound receiving mode to obtain a processed audio signal;

[0024] Among them, the sound reception mode is determined according to the corresponding scene signal, the scene signal is used to characterize the current sound reception scene, the sound reception scene includes a noise scene, a wind noise scene and / or a predetermined sound reception type scene, and the sound reception mode uses a microphone at a corresponding position to perform a sound reception operation.

[0025] The sound receiving system of the embodiment of the present invention includes at least two acoustic microphones, at least one bone conduction microphone and a controller, wherein the controller is used to determine the corresponding sound receiving mode according to the scene signal, the scene signal is used to characterize the current sound receiving scene, the sound receiving scene includes a noise scene, a wind noise scene and / or a predetermined sound receiving type scene, and the sound receiving mode uses the microphone at the corresponding position to perform the sound receiving operation. Therefore, this embodiment can identify different sound receiving scenes and adapt to the corresponding sound receiving mode, and at the same time use the signal fusion pickup of the bone conduction microphone and the acoustic microphone to achieve high-quality sound pickup of the device in strong noise environments, improve the noise suppression capability, and thus improve the speech recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0027] Figure 1 is a schematic diagram of a sound receiving system according to an embodiment of the present invention;

[0028] Figure 2 is a schematic diagram of a sound signal according to an embodiment of the present invention;

[0029] Figure 3 is a schematic diagram of a wearable device according to an embodiment of the present invention;

[0030] Figure 4 It is a flow chart of a device sound receiving method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.

[0032] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.

[0033] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".

[0034] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.

[0035] The solutions described in this specification and in the examples, if they involve the processing of personal information, will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.

[0036] Figure 1 is a schematic diagram of a sound receiving system in an embodiment of the present invention. The sound receiving system in an embodiment of the present invention includes at least two acoustic wave microphones, at least one bone conduction microphone and a controller. In this embodiment, a sound receiving system including four acoustic wave microphones and one bone conduction microphone is taken as an example. It should be understood that this embodiment does not limit the number of acoustic wave microphones and bone conduction microphones in the sound receiving system, and the number of each type of microphone can be set according to the specific application scenario and / or the appearance and wearing method of the wearable device.

[0037] like Figure 1 As shown, the sound pickup system of this embodiment includes acoustic wave microphones Mic11-Mic14, a bone conduction microphone 15 and a controller 16. The acoustic wave microphones Mic11-Mic14 capture sound signals after being controlled to pick up sound. The bone conduction microphone 15 senses bone vibration to capture sound after being controlled to pick up sound.

[0038] The controller is used to determine the corresponding sound recording mode according to the scene signal. Among them, the scene signal is used to characterize the current sound recording scene, and the sound recording scene includes a noise scene, a wind noise scene and / or a predetermined sound recording type scene. The sound recording mode uses a microphone at a corresponding position to perform a sound recording operation. Among them, the noise scene is used to characterize a scene where the surrounding sound noise is greater than a first predetermined value, and the wind noise scene is used to characterize a scene where the surrounding wind noise is greater than a second predetermined value. The predetermined sound recording type scene is a scene with specific requirements such as a stereo recording scene. It should be understood that the sound recording scene of this embodiment is not limited to this, and other scenes can also be set based on actual application requirements and environmental characteristics, which will not be explained one by one here.

[0039] Since the noise of the sound signal collected by the microphones at different positions or of different types may be different in different sound recording scenarios, the controller of this embodiment can identify the current sound recording scenario and switch the corresponding sound recording mode based on the corresponding sound recording scenario to achieve sound pickup based on the effective microphone in the corresponding sound recording mode, thereby improving the noise suppression capability of the sound recording system and further improving the speech recognition effect.

[0040] In an optional implementation, the controller 16 is used to perform noise detection on the sound signals collected by each sound wave microphone Mic11 - Mic14 to determine whether the current sound collection scene is in a noise scene or a wind noise scene.

[0041] Further optionally, the controller 16 receives the sound signals collected by the acoustic microphones Mic11-Mic14, detects the noise parameters of each sound signal, and determines the scene signal according to each noise parameter. Further, the controller 16 is used to determine that the sound recording scene is a noise scene in response to the noise parameter being greater than a first predetermined value. Optionally, in this embodiment, the noise parameter in the sound signal collected by any one of the acoustic microphones Mic11-Mic14 or at least a predetermined number (or a predetermined proportion) of the acoustic microphones is greater than a first predetermined value, and the current sound recording scene is determined to be a noise scene. It should be understood that the first predetermined value is a threshold value configured based on the type of noise parameter.

[0042] Furthermore, the noise parameter of the present embodiment is used to characterize the noise level, which can be defined based on different noise detection methods. For example, the controller 16 can use signal-to-noise ratio estimation (SNR estimation) to determine the noise parameter. The signal-to-noise ratio is the ratio of signal strength to background noise strength. The controller 16 can determine whether there is high noise in the current environment by calculating the signal-to-noise ratio of the collected sound signal. Among them, a low signal-to-noise ratio indicates the presence of strong noise. At this time, the noise parameter can be represented by a numerical value that is negatively correlated with the signal-to-noise ratio, such as using the inverse of the signal-to-noise ratio.

[0043] Furthermore, since noise often has a high spectral entropy, the controller 16 may also perform spectral entropy analysis on the sound signal to determine whether the current scene is a noisy one.

[0044] Furthermore, since the bone conduction microphone 15 can only detect bone vibration frequency and is not affected by environmental noise, this embodiment can also determine whether the current scene is high in noise based on the sound difference between the corresponding acoustic microphone and the bone conduction microphone 15 .

[0045] In other optional implementations, the controller 16 may also use voice activity detection, noise modeling or deep learning methods to determine whether the current scene is in a noise scene based on the collected sound signal. This embodiment does not limit the detection method of the noise scene, and it only needs to be able to detect high noise.

[0046] Furthermore, wind noise is usually manifested as broadband noise, has certain spectral characteristics, and is relatively obvious under strong wind conditions. The controller 16 of this embodiment determines that the sound recording scene is a wind noise scene in response to the noise parameter of the collected sound signal representing the wind noise and the noise parameter is greater than the second predetermined value. It should be understood that the second predetermined value is a threshold value configured based on the type of noise parameter.

[0047] Furthermore, since wind noise is mainly concentrated in the low frequency band, and the spectrum is usually flat or tilted toward the low frequency. Therefore, the controller 16 can determine whether there is wind noise based on spectrum analysis. For example, the controller 16 can determine whether there is wind noise by spectrum shape analysis, and estimate the wind noise level by the energy ratio of the low frequency band in the sound signal relative to the entire frequency band. If the energy of the low frequency band is significantly higher than that of other frequency bands, high wind noise may exist. That is, at this time, the noise parameter can be characterized by the low frequency band energy ratio or a value positively correlated with the low frequency band energy ratio.

[0048] Furthermore, since wind noise often has a high zero-crossing rate (Zero-Crossing Rate, ZCR), this embodiment can use the zero-crossing rate of the sound signal to characterize the noise parameters. The controller 16 detects the zero-crossing rate of the sound signal and compares it with a second predetermined value to determine whether it is in a high wind noise scene.

[0049] In other optional implementations, the controller 16 may also use short-time Fourier transform (STFT) or wavelet transform to analyze based on the acquired time-frequency diagram of the sound signal, or use the decay rate of the autocorrelation function to assist in the analysis, or build a wind noise model, or use a deep learning method to determine whether the current scene is a high wind noise scene. This embodiment does not limit the detection method of the wind noise scene, and it only needs to be able to detect high wind noise.

[0050] Further, the scene signal of the predetermined sound recording type scene can be triggered by starting the application of the corresponding requirement, such as the recording scene. In this embodiment, the corresponding scene signal can be generated after the recording control of the device is triggered. After receiving the scene signal, the controller 16 determines that the current scene is in the recording scene.

[0051] Furthermore, it should be understood that the noise type scene and the predetermined sound recording type scene may exist at the same time. For example, if strong winds occur during the recording process, the current scene is both a recording scene and a high wind noise scene. Therefore, the sound recording mode of this embodiment may include modes corresponding to different sound recording scenes and sound recording modes in the case of combining multiple sound recording scenes.

[0052] Figure 2 is a schematic diagram of the sound signal of an embodiment of the present invention. In an optional implementation, the controller 16 further characterizes the wind noise in response to the noise parameters of the sound signal, determines the wind direction according to each noise parameter, and selects the sound wave microphone to perform the sound reception operation according to the wind direction. Furthermore, the controller 16 can perform wind noise analysis on the sound signals collected by the sound wave microphones Mic11-Mic14 at different positions, respectively, to determine the wind direction based on the strength of the wind noise signals (i.e., noise parameters) at different positions, and select the sound wave microphone with low wind noise to perform the sound reception operation based on the wind direction. Figure 2 As shown, the sound signal 21 is collected by the sound wave microphone facing the wind direction, and the sound signal 22 is collected by the sound wave microphone facing away from the wind direction. The wind noise of the sound signal 22 is obviously much smaller than that of the sound signal 21. Therefore, this embodiment configures the corresponding sound collection mode in the wind noise scene so that the sound signal collected by the sound wave microphone facing away from the wind direction is used for voice recognition, thereby achieving the purpose of reducing wind noise, thereby improving the clarity of sound pickup.

[0053] In an optional implementation, at least two sound wave microphones in the sound receiving system are distributed on both sides of the wearable device, and the bone conduction microphone is arranged in a predetermined area of ​​the wearable device, wherein at least a part of the predetermined area contacts the wearer after the wearable device is worn.

[0054] Furthermore, if it is determined that the current scene is a noisy scene, the controller 16 can select a combination of part or all of the acoustic wave microphones and the bone conduction microphone 15 to pick up sound, so as to achieve better noise reduction while extracting the direction information of the target user and combining the characteristics of bone sound information conduction obtained by the bone conduction microphone 15.

[0055] If the current scene is a high wind noise scene, the controller 16 can select the acoustic microphone on the side with weaker wind noise in combination with the bone conduction microphone 15 to pick up sound, thereby achieving the effect of reducing wind noise and clearly picking up sound.

[0056] If the current scene is a video recording scene, the controller 16 may select acoustic microphones located in different directions to pick up sound to achieve a stereo effect.

[0057] It should be understood that for sound recording scenarios different from the above scenarios, corresponding sound recording modes can be configured according to specific needs, and this embodiment will not illustrate them one by one.

[0058] Furthermore, the sound receiving system of this embodiment further includes a main control 17. The main control 17 is used to perform signal fusion processing on the received sound signal, obtain the processed audio signal, and recognize the audio signal to realize corresponding functions such as call.

[0059] In an optional implementation, the controller 16 of this embodiment can be a controller with a signal processing module such as a DSP (Digital Signal Processing) or an MCU (Microcontroller Unit). Further, the controller 16 of this embodiment can be independently set outside the main control 17, or can be built into the main control 17. This embodiment is not limited to this, and it can be set according to specific application conditions.

[0060] Furthermore, after the controller 16 in the sound pickup system determines the current sound pickup scene to determine the target microphones for sound pickup, based on the sound signals collected by each target microphone, speech processing is performed on the collected sound signals to achieve sound pickup.

[0061] Taking the current sound collection scene as a wind noise scene, the acoustic wave microphones Mic11 and Mic12 and the bone conduction microphone 15 are selected as target microphones as an example. The controller 16 controls the acoustic wave microphones Mic11, Mic12 and the bone conduction microphone 15 to synchronously collect sound signals, and performs signal preprocessing on the sound signals captured by each target microphone, such as power amplification, filtering, and / or noise reduction and other signal processing.

[0062] Furthermore, feature extraction is performed on the preprocessed sound signals of the acoustic microphones Mic11 and Mic12, for example, MFCC (Mel-frequency cepstral coefficients), speech spectrogram, etc., to obtain sound features F11 and F12, and feature extraction is performed on the preprocessed sound signals of the bone conduction microphone 15, for example, time domain features (such as amplitude envelope), frequency domain features (such as spectral characteristics), and / or other advanced features (such as wavelet transform coefficients) are extracted to obtain bone conduction feature F15.

[0063] Furthermore, the sound features F11 and F12, and the bone conduction feature F15 are signal aligned and normalized. Signal alignment is used to solve the delay between signals, and time alignment can be performed by cross-correlation or other methods. Normalization is used to adjust the sound features F11 and F12, and the bone conduction feature F15 to the same scale for subsequent processing, for example, through normalization or standardization processing, the sound features F11 and F12, and the bone conduction feature F15 have a similar distribution range.

[0064] Further, in this embodiment, the aligned and normalized sound features F11 and F12, and the bone conduction feature F15 are subjected to multimodal fusion to obtain the fused audio signal or speech recognition result. Further, in this embodiment, a pre-trained multimodal fusion model may be used to input the processed sound features F11 and F12, and the bone conduction feature F15 into the multimodal fusion model to obtain the fused audio signal or speech recognition result. It should be understood that the multimodal fusion model of this embodiment can be trained based on specific required functions, such as training corresponding models based on required speech recognition, call, noise reduction and other functions, which will not be illustrated one by one here. Optionally, the multimodal fusion model of this embodiment can be obtained by training basic models such as support vector machine (SVM), random forest, convolutional neural network (CNN), or recurrent neural network (RNN) and its variant LSTM / GRU. It should be understood that the model structure of this embodiment is not limited thereto, and it can achieve the corresponding functions.

[0065] In this embodiment, the above-mentioned sound signal processing process can be deployed in the controller 16, or in the main control 17, or can be distributed in the controller 16 and the main control 17. For example, signal preprocessing, feature extraction, and signal alignment and normalization are deployed in the controller 16, and the multimodal fusion model is deployed in the main control 17. This embodiment does not limit the location where the specific signal processing operation module or device is deployed, and it can be deployed based on specific needs.

[0066] The sound receiving system of the embodiment of the present invention includes at least two acoustic microphones, at least one bone conduction microphone and a controller, wherein the controller is used to determine the corresponding sound receiving mode according to the scene signal, the scene signal is used to characterize the current sound receiving scene, the sound receiving scene includes a noise scene, a wind noise scene and / or a predetermined sound receiving type scene, and the sound receiving mode uses the microphone at the corresponding position to perform the sound receiving operation. Therefore, this embodiment can identify different sound receiving scenes and adapt to the corresponding sound receiving mode, and at the same time use the signal fusion pickup of the bone conduction microphone and the acoustic microphone to achieve high-quality sound pickup of the device in strong noise environments, improve the noise suppression capability, and thus improve the speech recognition effect.

[0067] Figure 3 is a schematic diagram of a wearable device according to an embodiment of the present invention. The wearable device according to the embodiment of the present invention includes a device body and a sound pickup system in any of the above implementations. This embodiment is described in detail by taking the wearable device as smart glasses as an example. It should be understood that this embodiment is not limited to this, and it can be any wearable device that needs to pick up sound. Figure 3 As shown, the smart glasses of this embodiment include a glasses body and a sound receiving system. The sound receiving system includes sound wave microphones Mic31-Mic34, a bone conduction microphone 35, and a controller and / or a main control (not shown in the figure).

[0068] Furthermore, if Figure 3 As shown, the acoustic wave microphones Mic31 and Mic32 are arranged in the temple 10 of the smart glasses, and the acoustic wave microphones Mic33 and Mic34 are arranged in the temple 20 of the smart glasses. Optionally, the arrangement positions of the acoustic wave microphones Mic31 and Mic32 are symmetrical with the arrangement positions of the acoustic wave microphones Mic33 and Mic34. However, it should be understood that the acoustic wave microphones arranged in the temple 10 and the temple 20 may also be arranged asymmetrically. In other optional implementations, some of the acoustic wave microphones may also be arranged on the frame, for example, the acoustic wave microphone Mic31 is arranged on the temple 10, the acoustic wave microphone Mic32 is arranged on the frame on the same side of the temple 10, the acoustic wave microphone Mic33 is arranged on the temple 20, and the acoustic wave microphone Mic34 is arranged on the frame on the same side of the temple 20. It should be understood that this embodiment does not limit the specific positions of the arrangement.

[0069] Further, the bone conduction microphone 35 is arranged in a predetermined area of ​​the smart glasses, that is, in the area 30 where the nose pads are located. It should be understood that in this embodiment, the size of the area where the nose pads are located can be determined based on experience or testing to ensure the sound signal capture effect when the wearer speaks. Further, the bone conduction microphone 35 is arranged in the nose pads, and it can also be arranged in the frame in the area 30 where the nose pads are located. In other optional implementations, the bone conduction microphone 35 can also be arranged in the contact area between the temple and the wearer. It should be understood that this embodiment does not limit the specific setting position of the bone conduction microphone 35, and it can capture the bone vibration when the wearer speaks. In some implementations, the bone conduction microphone 35 can also be arranged in a non-contact area, as long as a bone conduction signal with a strength that meets the conditions is transmitted in a non-contact state.

[0070] Furthermore, the controller or the main control determines the current sound reception scene. If the current sound reception scene is a noisy scene with high noise, the selected sound reception mode may be: the mode of bone conduction microphone 35+Mic31+Mic32, or the mode of bone conduction microphone 35+Mic33+Mic34, or the mode of bone conduction microphone 35+Mic31+Mic32+Mic33+Mic3, so that the sound wave microphone is combined with the bone conduction microphone to pick up sound, thereby achieving a better noise reduction effect.

[0071] If the current sound recording scene is a wind noise scene with high wind noise, and the temple 10 is facing the wind direction, that is, the acoustic microphones Mic31 and Mic32 face the wind direction, the wind noise of the sound signal captured by them is high. In this scene, the selected sound recording mode can be: bone conduction microphone 35+Mic33+Mic34 mode, to ensure the signal strength of the collected sound signal, and effectively avoid wind noise, and combine the bone conduction microphone for sound pickup, to achieve a better noise reduction effect, and ensure the clarity of sound pickup.

[0072] If the current sound recording scene is a video recording scene, the selected sound recording mode may be: Mic31+Mic32+Mic33+Mic3. If the current scene is still noisy, the bone conduction microphone 35 may be used to collect sound, thereby effectively distinguishing the voice of the wearer from the ambient sound, achieving stereo sound recording while ensuring the clarity of the wearer's sound collection.

[0073] After determining the sound reception mode, the controller or main control of the sound reception system in the smart glasses performs preprocessing, feature extraction, feature alignment and normalization, and feature multi-source fusion on the sound signals collected by the microphone corresponding to the determined sound reception mode, so as to realize sound pickup and corresponding functions such as calls, voice interaction, and recording. The processing of the corresponding sound signals will not be described in detail here.

[0074] The sound receiving system in the wearable device of the embodiment of the present invention includes at least two acoustic microphones, at least one bone conduction microphone and a controller, wherein the controller is used to determine the corresponding sound receiving mode according to the scene signal, the scene signal is used to characterize the current sound receiving scene, the sound receiving scene includes a noise scene, a wind noise scene and / or a predetermined sound receiving type scene, and the sound receiving mode uses the microphone at the corresponding position to perform the sound receiving operation. Therefore, this embodiment can identify different sound receiving scenes and adapt to the corresponding sound receiving mode, and at the same time use the signal fusion pickup of the bone conduction microphone and the acoustic microphone to achieve high-quality sound pickup of the device in strong noise environments, improve the noise suppression capability, and thus improve the speech recognition effect.

[0075] Figure 41 is a flow chart of a device sound receiving method according to an embodiment of the present invention. The device sound receiving method according to this embodiment is applied to a sound receiving system of a wearable device, and the sound receiving system includes at least two sound wave microphones and at least one bone conduction microphone. Figure 4 As shown, the device sound receiving method of the embodiment of the present invention includes the following steps:

[0076] Step S110, collecting corresponding sound signals through each acoustic microphone and the bone conduction microphone. After the wearable device starts the sound receiving system, each acoustic microphone and the bone conduction microphone in the sound receiving system are controlled to collect corresponding sound signals.

[0077] Step S120, based on the predetermined sound reception mode, the corresponding sound signal is selected for processing to obtain the processed audio signal. The sound reception mode is determined according to the corresponding scene signal, and the scene signal is used to characterize the current sound reception scene. The sound reception scene includes a noise scene, a wind noise scene and / or a predetermined sound reception type scene. The sound reception mode uses a microphone at a corresponding position to perform a sound reception operation. The noise scene is used to characterize a scene where the surrounding sound noise is greater than a first predetermined value, and the wind noise scene is used to characterize a scene where the surrounding wind noise is greater than a second predetermined value. Predetermined sound reception type scenes include scenes with specific requirements such as stereo recording scenes. It should be understood that the sound reception scenes of this embodiment are not limited to this, and other scenes can also be set based on actual application requirements and environmental characteristics, which will not be illustrated one by one here.

[0078] In an optional implementation, the present embodiment may also perform noise analysis and wind noise analysis based on the sound signals collected by each sound wave microphone in the sound receiving system to determine whether the current scene is a high-noise noise scene or a high-wind-noise wind noise scene.

[0079] Furthermore, in response to the noise parameter being greater than a first predetermined value, the present embodiment determines that the sound recording scene is a noise scene. The noise parameter of the present embodiment is used to characterize the noise level, which can be defined based on different noise detection methods. The four-loop itself can use signal-to-noise ratio estimation (SNR estimation), spectral entropy analysis, sound difference between acoustic microphone and bone conduction microphone 15, voice activity detection, noise modeling or deep learning methods to determine whether it is currently in a noise scene based on the collected sound signal. The present embodiment does not limit the detection method of the noise scene, and it can detect high noise. It should be understood that the first predetermined value is a threshold configured based on the type of noise parameter.

[0080] Furthermore, in response to the noise parameter of the collected sound signal characterizing the wind noise and the noise parameter being greater than a second predetermined value, the present embodiment determines that the sound recording scene is a wind noise scene. It should be understood that the second predetermined value is a threshold value configured based on the type of noise parameter. In the present embodiment, spectrum analysis, zero crossing rate analysis, short-time Fourier transform (STFT) or wavelet transform, attenuation rate of autocorrelation function, construction of a wind noise model, or deep learning methods may be used to determine whether the current scene is a high wind noise scene. The present embodiment does not limit the detection method of the wind noise scene, and it is sufficient that it can detect high wind noise.

[0081] Further, the scene signal of the predetermined sound recording type scene can be triggered by starting the application of the corresponding requirement, such as the recording scene. In this embodiment, the corresponding scene signal can be generated after the recording control of the device is triggered. After receiving the scene signal, the controller 16 determines that the current scene is in the recording scene.

[0082] Furthermore, it should be understood that the noise type scene and the predetermined sound recording type scene may exist at the same time. For example, if strong winds occur during the recording process, the current scene is both a recording scene and a high wind noise scene. Therefore, the sound recording mode of this embodiment may include modes corresponding to different sound recording scenes and sound recording modes in the case of combining multiple sound recording scenes.

[0083] Optionally, a noise classification model may be pre-trained to determine whether the sound recording scene is a noise scene or a wind noise scene.

[0084] After determining the sound reception mode, this embodiment also performs preprocessing, feature extraction, feature alignment and normalization, and feature multi-source fusion on the sound signals collected by the microphone corresponding to the determined sound reception mode to achieve sound pickup and corresponding functions such as calls, voice interaction, and recording. The processing of the corresponding sound signals will not be described in detail here.

[0085] The embodiment of the present invention determines the corresponding sound pickup mode according to the scene signal, the scene signal is used to characterize the current sound pickup scene, the sound pickup scene includes a noise scene, a wind noise scene and / or a predetermined sound pickup type scene, and the sound pickup mode uses a microphone at a corresponding position to perform a sound pickup operation. Therefore, the present embodiment can identify different sound pickup scenes and adapt to the corresponding sound pickup mode, and at the same time use the signal fusion pickup of the bone conduction microphone and the acoustic wave microphone to achieve high-quality sound pickup of the device in a strong noise environment, improve the noise suppression capability, and thus improve the speech recognition effect.

[0086] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, wherein the computer-readable program is used for a computer to execute part or all of the above method embodiments.

[0087] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0088] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A sound receiving system, characterized in that: The sound receiving system comprises: at least two acoustic microphones; at least one bone conduction microphone; and A controller is used to determine a corresponding sound reception mode according to a scene signal, wherein the scene signal is used to characterize a current sound reception scene, wherein the sound reception scene includes a noise scene, a wind noise scene and / or a predetermined sound reception type scene, and the sound reception mode uses a microphone at a corresponding position to perform a sound reception operation.

2. The sound receiving system according to claim 1, characterized in that: The at least two acoustic wave microphones are distributed on both sides of the wearable device, and the bone conduction microphone is arranged in a predetermined area of ​​the wearable device, wherein at least a part of the predetermined area contacts the wearer after the wearable device is worn, or transmits a bone conduction signal with a strength that meets the conditions in a non-contact state.

3. The sound receiving system according to claim 2, characterized in that: The controller is used to receive the sound signals collected by each of the sound wave microphones, detect the noise parameters of each of the sound signals, and determine the scene signal according to each of the noise parameters.

4. The sound receiving system according to claim 3, characterized in that: The controller is configured to determine that the sound recording scene is a noise scene in response to the noise parameter being greater than a first predetermined value, and to determine that the sound recording scene is a wind noise scene in response to the noise parameter representing wind noise and the noise parameter being greater than a second predetermined value; or The controller is used to determine whether the sound collection scene is a noise scene or a wind noise scene according to a predetermined noise classification model.

5. The sound receiving system according to claim 4, characterized in that: The controller is used to characterize wind noise in response to the noise parameters, determine the wind direction according to each noise parameter, and select a sound wave microphone to perform a sound collection operation according to the wind direction.

6. The sound receiving system according to claim 1, characterized in that: The sound receiving system also includes: The main control is used to perform signal fusion processing on the received sound signal to obtain the processed audio signal.

7. The sound receiving system according to claim 6, characterized in that: The controller is built into the main control, or the controller is an independent integrated circuit.

8. A wearable device, characterized in that: The wearable device comprises: The device body; and A sound receiving system as claimed in any one of claims 1 to 7.

9. The wearable device according to claim 8, characterized in that: The wearable device is smart glasses, the sound wave microphones in the sound receiving system are distributed on the temples and / or the frames, and the bone conduction microphones in the sound receiving system are arranged in the area where the nose pads are located or the area where the temples are located.

10. The wearable device according to claim 9, characterized in that: The acoustic microphones are symmetrically distributed on the temples and / or frames on both sides.

11. A device sound receiving method, applied to a sound receiving system of a wearable device, characterized in that: The sound pickup system includes at least two acoustic wave microphones and at least one bone conduction microphone, and the method includes: Collecting corresponding sound signals through each of the acoustic wave microphones and the bone conduction microphone; Selecting a corresponding sound signal for processing based on a predetermined sound receiving mode to obtain a processed audio signal; Among them, the sound reception mode is determined according to the corresponding scene signal, the scene signal is used to characterize the current sound reception scene, the sound reception scene includes a noise scene, a wind noise scene and / or a predetermined sound reception type scene, and the sound reception mode uses a microphone at a corresponding position to perform a sound reception operation.