Signal processing system, method, device and storage medium
By combining the signal processing of the microphone and vibration sensor, determining the relationship between noise components and performing noise reduction processing, the signal clarity problem of the vibration sensor under environmental noise interference is solved, and a higher signal-to-noise ratio is achieved.
Patent Information
- Application Number
- CN202180048143.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-19
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2041-03-19
AI Technical Summary
When collecting user voice signals, vibration sensors are easily interfered with by environmental noise, and existing technologies are difficult to effectively reduce noise interference.
By combining the sound signal collected by the microphone with the vibration signal collected by the vibration sensor, the relationship between the noise components is determined, and the vibration signal is subjected to noise reduction processing based on the relationship, including the use of voice activity detection and noise suppressor.
It effectively reduces the interference of environmental noise on the voice signal collected by the vibration sensor and improves the signal clarity and signal-to-noise ratio.
Smart Images

Figure CN115989681B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing, and more specifically, to a system, method, device and storage medium for processing vibration signals. Background Art
[0002] When people speak, their bones and skin vibrate simultaneously. These vibrations can be picked up by vibration sensors and converted into corresponding electrical signals or other types of signals. Because general environmental noise rarely causes bone or skin vibrations, vibration sensors can record cleaner speech signals compared to air conduction microphones, minimizing interference from ambient noise.
[0003] However, when the external environment is noisy, the noise can cause the human body's bones, skin, or the vibration sensor itself to vibrate, thereby interfering with the voice signal received by the vibration sensor. Therefore, it is necessary to provide a method for processing the voice signal collected by the vibration sensor to reduce the interference caused by external noise on the vibration sensor. Summary of the Invention
[0004] One aspect of an embodiment of the present application provides a signal processing system, comprising: at least one microphone, the at least one microphone being used to collect a sound signal, the sound signal including at least one of user voice and ambient noise; at least one vibration sensor, the at least one vibration sensor being used to collect a vibration signal, the vibration signal including at least one of the user voice and the ambient noise; and a processor being configured to: determine a relationship between a noise component in the sound signal and a noise component in the vibration signal; and perform noise reduction processing on the vibration signal based at least on the relationship to obtain a target vibration signal.
[0005] Another aspect of an embodiment of the present application provides a signal processing method, including: obtaining a sound signal collected by at least one microphone, the sound signal including at least one of user voice and environmental noise; obtaining a vibration signal collected by at least one vibration sensor, the vibration signal including at least one of the user voice and the environmental noise; determining a relationship between a noise component in the sound signal and a noise component in the vibration signal; and performing noise reduction processing on the vibration signal based at least on the relationship to obtain a target vibration signal.
[0006] Another aspect of an embodiment of the present application provides an electronic device, comprising at least one processor and at least one memory; the at least one memory is used to store computer instructions; and the at least one processor is used to execute at least part of the computer instructions to implement the operations described above.
[0007] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the method described above is executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The present application will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:
[0009] Figure 1 is a schematic diagram of an application scenario of a signal processing system provided in some embodiments of the present application;
[0010] Figure 2 is a flowchart of a signal processing method provided according to some embodiments of the present application;
[0011] Figure 3 is a schematic diagram of modules of a signal processing system provided according to some embodiments of the present application;
[0012] Figure 4 is a schematic diagram of the working principle of a vibration sensor noise suppressor in a signal processing system provided in some embodiments of the present application;
[0013] Figure 5 is a schematic diagram of a signal spectrum of a vibration sensor provided according to some embodiments of the present application;
[0014] Figure 6 is a schematic diagram of a signal spectrum received by a vibration sensor in a noisy environment according to some embodiments of the present application;
[0015] Figure 7 is a module diagram of a signal processing system provided according to some other embodiments of the present application;
[0016] Figure 8 is a schematic diagram of a signal spectrum obtained after processing according to some embodiments of the present application;
[0017] Figure 9 is a module diagram of a signal processing system provided according to some other embodiments of the present application;
[0018] Figure 10 is a module diagram of a signal processing system provided according to some other embodiments of the present application;
[0019] Figure 11 is a module diagram of a signal processing system provided according to some other embodiments of the present application; and
[0020] Figure 12Schematic diagram of a signal frequency-signal-to-noise ratio curve provided according to other embodiments of the present application. DETAILED DESCRIPTION
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0022] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0023] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0024] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0025] Vibration sensors detect skin or bone vibrations when people speak and convert them into electrical signals. However, when vibration sensors collect user speech, they are often accompanied by some noise signals, such as ambient noise, noise generated by chewing and walking, and noise caused by friction between the skin and the vibration sensor. Therefore, it is necessary to reduce the noise of the signals collected by vibration sensors to minimize the interference caused by these noise signals.
[0026] In response to the above problems, an embodiment of the present application provides a signal processing system and method, which combines the vibration signal collected by a vibration sensor with the sound signal collected by a microphone to determine the relationship between the vibration signal and the noise components in the sound signal, and denoises the vibration signal based on the relationship and the noise components in the sound signal, thereby reducing the interference caused by the noise.
[0027] The signal processing system and method provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0028] Figure 1 It is a schematic diagram of an application scenario of a signal processing system according to some embodiments of the present application.
[0029] like Figure 1 As shown, in some embodiments, the signal processing system 100 may include a microphone 110, a network 120, a vibration sensor 130, a processor 140, and a memory 150. In some embodiments, the various components in the system 100 can be interconnected via the network 120. For example, the microphone 110 and the processor 140 can be connected or communicated via the network 120, the microphone 110 and the memory 150 can be connected or communicated via the network 120, and the memory 150 and the processor 140 can be connected or communicated via the network 120. In some embodiments, the network 120 is not required. For example, the microphone 110, the vibration sensor 130, the processor 140, and the memory 150 can be integrated as different components in the same electronic device. The electronic device includes wearable devices such as headphones, glasses, and smart helmets. The different components of the electronic device can be connected and transmit data via metal wires.
[0030] In some embodiments, the signal processing system 100 may include one or more microphones 110, and one or more vibration sensors 130. The one or more microphones 110 can be used to collect user voice and environmental noise, and generate sound signals. The user voice and environmental noise can be transmitted to the microphone 110 by air conduction. The one or more vibration sensors 130 can be in contact with the user's body, such as the user's face or neck, and generate vibration signals by receiving physical vibrations of the contact part caused by the user's speech or environmental noise. In some embodiments, multiple microphones 110 can be arranged in an array to form a microphone array. The microphone array can identify air-conducted sounds from a specific direction, such as sounds from the user's mouth, sounds from directions other than the user's mouth, etc.
[0031] Network 120 may include any suitable network capable of facilitating information and / or data exchange within system 100. In some embodiments, at least one component of system 100 (e.g., microphone 110, vibration sensor 130, processor 140, memory 150) may exchange information and / or data with at least one other component of system 100 via network 120. For example, processor 140 may obtain a signal from microphone 110 or vibration sensor 130 via network 120. For another example, processor 140 may obtain preset processing instructions from memory 150 via network 120. Network 120 may include or include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN)), a wired network, a wireless network (e.g., an 802.11 network, a Wi-Fi network), a frame relay network, a virtual private network (VPN), a satellite network, a telephone network, a router, a hub, a switch, a server computer, and / or any combination thereof. For example, the network 120 may include a wired network, a cable network, an optical fiber network, a telecommunication network, an intranet, a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, or a wireless local area network (WLAN). TM network, near field communication (NFC) network, etc. or any combination thereof. In some embodiments, the network 120 may include at least one network access point. For example, the network 120 may include a wired and / or wireless network access point, such as a base station and / or an Internet exchange point, and at least one component of the system 100 may be connected to the network 120 through the access point to exchange data and / or information. In some embodiments, the microphone 110 and the vibration sensor 130 may be integrated into the same electronic device (such as a headset). The electronic device can communicate with other terminal devices through the network 120. For example, the electronic device can send the electrical signals generated by the microphone 110 and the vibration sensor 130 to a user terminal (such as a mobile phone) through the network 120, and the user terminal processes the received signals and then sends the processed signals back to the electronic device through the network 120. This approach can reduce the burden of signal processing on the electronic device, thereby effectively reducing the size of the signal processor (if any) and battery on the electronic device.
[0032] The processor 140 can process data and / or instructions obtained from the microphone 110, the vibration sensor 130, the memory 150, or other components of the system 100. For example, the processor 140 can obtain a sound signal from the microphone 110 and a vibration signal from the vibration sensor 130, and process the two to determine the relationship between the noise component in the sound signal and the noise component in the vibration signal. For another example, the processor 140 can retrieve pre-stored instructions from the memory 150 and execute the instructions to implement the signal processing method described below. By way of example only, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computer (RISC), a microprocessor, etc., or any combination thereof.
[0033] In some embodiments, processor 140 can be local or remote. For example, processor 140, microphone 110, and vibration sensor 130 can be integrated into the same electronic device, or distributed across different electronic devices. In some embodiments, processor 140 can be implemented on a cloud platform. For example, the cloud platform can include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an inter-cloud cloud, a multi-cloud, or any combination thereof.
[0034] The memory 150 can store data, instructions and / or any other information. In some embodiments, the memory 150 can store sound signals collected by the microphone 110 and / or vibration signals collected by the vibration sensor 130. In some embodiments, the memory 150 can store data and / or instructions used by the processor 140 to execute or use to complete the exemplary methods described in this application. In some embodiments, the memory 150 may include a large-capacity memory, a removable memory, a volatile read-write memory, a read-only memory (ROM), etc., or any combination thereof. Exemplary large-capacity memory may include a magnetic disk, an optical disk, a solid-state disk, etc. Exemplary removable memory may include a flash drive, a floppy disk, an optical disk, a memory card, a compressed disk, a magnetic tape, etc. Exemplary volatile read-write memory may include a random access memory (RAM). In some embodiments, the memory 150 may be implemented on a cloud platform.
[0035] In some embodiments, the memory 150 may be connected to the network 120 to communicate with at least one other component in the system 100 (e.g., the processor 140). At least one component in the system 100 may access data or instructions stored in the memory 150 or write data to the memory 150 through the network 120. In some embodiments, the memory 150 may be part of the processor 140.
[0036] It should be noted that the above description of the signal processing system 100 and its components is for ease of description only and does not limit this specification to the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the components or form subsystems connected to other modules without departing from these principles. In some embodiments, the components may share a single memory 150. In some embodiments, each component may have its own storage module. Such variations are within the scope of this specification.
[0037] In some embodiments, the signal processing system 100 can be applied to electronic devices, such as wearable electronic devices such as headphones, glasses, and smart helmets, to reduce noise interference with user voice signals collected by vibration sensors. It should be noted that the aforementioned devices or equipment are merely examples, and the signal processing system 100 provided in the embodiments of the present application can be applied to, but is not limited to, the aforementioned devices or electronic devices.
[0038] Figure 2 2 is a flow chart of a signal processing method according to some embodiments of the present application. In some embodiments, process 200 may be completed using one or more additional operations not described below, and / or not by one or more operations discussed below. Figure 2 The order of operations shown is not limiting. In some embodiments, process 200 may be applied to Figure 1 The signal processing system 100 is shown. In some embodiments, the process 200 may be executed by the processor 140.
[0039] like Figure 2 As shown, in some embodiments, process 200 may include the following steps:
[0040] Step 210: At least one microphone collects at least one of user voice and environmental noise to generate a sound signal.
[0041] In some embodiments, one or more microphones may capture user voice and / or ambient noise. User voice may refer to the sound produced by a user speaking or making a sound, such as the sound produced by normal user speech, as well as laughter, crying, shouting, etc., and ambient noise may refer to sounds other than user voice, such as the sound of wind, rain, traffic, machinery, and other objects. The term "user" herein refers to the person wearing the at least one microphone. When the user is speaking, the one or more microphones may simultaneously capture the user's voice and the ambient noise. The generated sound signal will then include both the user's voice component corresponding to the user's voice and the noise component corresponding to the ambient noise. When the user is silent, the one or more microphones only capture the ambient noise. The generated sound signal will then include only the noise component corresponding to the ambient noise. In some embodiments, the one or more microphones may be air conduction microphones. In some embodiments, the one or more microphones may comprise a single microphone or a microphone array. Different microphones in the microphone array may be located at different distances from the user's mouth.
[0042] In some embodiments, the processor 140 may acquire a sound signal generated by the one or more microphones. The sound signal may be an electrical signal or a signal in another form.
[0043] Step 220: At least one vibration sensor collects at least one of the user voice and the environmental noise to generate a vibration signal.
[0044] In some embodiments, while the aforementioned one or more microphones are collecting user voice and / or environmental noise, one or more vibration sensors may collect vibrations caused by the user voice and / or the environmental noise. At this time, the sound signal generated by the microphone and the vibration signal generated by the vibration sensor correspond to the same sound content. In some embodiments, the one or more vibration sensors may be in contact with the user's body, such as the face, neck, and other parts, to collect vibrations generated by the user's skin or bones when the user speaks. When there are multiple vibration sensors, the multiple vibration sensors may be located in different parts of the user's body, respectively collecting vibrations of different parts of the user and generating the vibration signal. For example, the vibration signal may be the electrical signal corresponding to the vibration sensor with the strongest signal strength among the multiple vibration sensors. For another example, the vibration signal may be formed by combining the electrical signals collected by each of the multiple vibration sensors.
[0045] In some embodiments, the processor 140 may obtain a vibration signal generated by the one or more vibration sensors. In some embodiments, the vibration signal may be an electrical signal or a signal in another form. In some embodiments, the vibration signal and the sound signal may be acquired at the same time or in the same time period. In some embodiments, the vibration signal and the sound signal may be synchronized based on the same clock signal.
[0046] Step 230: Determine the relationship between the noise component in the sound signal and the noise component in the vibration signal.
[0047] Since the noise component in the sound signal and the noise component in the vibration signal are both excited by environmental noise, there is a strong correlation between the two. Therefore, in some embodiments, the processor 140 can determine the relationship between the noise component in the sound signal and the noise component in the vibration signal based on the sound signal collected by at least one microphone and the vibration signal collected by at least one vibration sensor.
[0048] It should be noted that, in some embodiments, the sound signal may be collected by a single microphone or a microphone array (ie, multiple microphones).
[0049] In some embodiments, the processor 140 can identify a time interval in which the user does not utter a voice, and determine a first noise signal reflecting the ambient noise from the sound signal within the time interval, and determine the relationship between the first noise signal and the vibration signal within the time interval, and then use the relationship between the first noise signal and the vibration signal as the relationship between the noise component in the sound signal and the noise component in the vibration signal when the user utters a voice.
[0050] In some alternative embodiments, when the sound signal is collected by the microphone array, the processor 140 may identify the time interval during which the user uttered the voice, determine a second noise signal reflecting the ambient noise from the sound signal within that time interval, and simultaneously determine the correlation between different components of the vibration signal within that time interval and the second noise signal. For example, components of the vibration signal whose correlation with the second noise signal is higher than a preset threshold are considered noise, while components whose correlation with the second noise signal is lower than the preset threshold are considered user voice.
[0051] In some embodiments, when the sound signal is collected by a single microphone, the processor 140 can convert the sound signal and the vibration signal from time domain signals into frequency domain signals, and obtain the noise relationship between the noise component in the sound signal and the noise component in the vibration signal on at least one frequency domain subband. In some embodiments, the noise relationship between the noise component in the sound signal and the noise component in the vibration signal can be expressed as a power ratio or a signal spectrum ratio between the two. For more details on determining the noise relationship based on the sound signal collected by a single microphone, please refer to other places in this specification (for example Figure 4 Part and related discussions), which will not be described in detail here.
[0052] Step 240: Perform noise reduction processing on the vibration signal based at least on the relationship to obtain a target vibration signal.
[0053] In some embodiments, after obtaining the noise relationship between the noise component in the sound signal and the noise component in the vibration signal, the processor 140 can perform noise reduction processing on the vibration signal based on the noise relationship and the noise component in the sound signal to obtain a target vibration signal, that is, a clean vibration signal obtained after noise reduction processing.
[0054] For example, the processor 140 can determine the noise component in the vibration signal when the user speaks based on the noise relationship when the user does not speak and the noise component in the sound signal when the user speaks (for example, determined based on the sound signal obtained by the microphone array), and further remove the noise component from the vibration signal when the user speaks to obtain the target vibration signal. For another example, the processor 140 can obtain the noise relationship between the noise component in the sound signal and the noise component in the vibration signal on at least one frequency domain subband based on the noise relationship when the user does not speak, and further remove the noise component from the vibration signal when the user speaks based on the noise relationship corresponding to the specific frequency domain subband and the noise component of the specific frequency domain subband when the user speaks.
[0055] For more technical details on determining the relationship between the noise component in the sound signal and the noise component in the vibration signal, and reducing the noise of the vibration signal, please refer to other places in this specification (for example, Figure 4 、 Figure 9 、 Figure 10 Part and related discussions), which will not be described in detail here.
[0056] Figure 3 This is a module diagram of a signal processing system provided according to some embodiments of the present application.
[0057] Reference Figure 3In some embodiments, the signal processing system 300 may include a voice activity detector 341 and a vibration sensor noise suppressor 342 .
[0058] In some embodiments, the voice activity detector 341 and the vibration sensor noise suppressor 342 may be part of the processor 140. The voice activity detector 341 may be used to identify signal segments containing user speech within the sound signal collected by the microphone 310 and the vibration signal collected by the vibration sensor 330. In other words, the voice activity detector 341 may identify whether the user is speaking. The vibration sensor noise suppressor 342 may be used to determine the relationship between the noise components in the vibration signal and the noise components in the sound signal, and based on this relationship, perform noise reduction processing on the signal segments containing user speech within the vibration signal to obtain a target vibration signal.
[0059] In some embodiments, the voice activity detector 341 can use a machine learning model to recognize user voice in sound signals and vibration signals. In some embodiments, the machine learning model can be trained using data samples so that the machine learning model acquires the ability to recognize user voice features and distinguish user voice from sound signals or vibration signals. The data samples described herein may include positive data samples and negative data samples. Positive data samples may include a set of sound signal samples and vibration signal samples containing user voice, and negative data samples may include a set of sound signal samples and vibration signal samples that do not contain user voice.
[0060] In some embodiments, the voice activity detector 341 can determine whether the user is speaking based on the sound and / or vibration signals it receives. For example, considering that whether the user is speaking affects the strength of the signal generated by the vibration sensor, the voice activity detector 341 can determine whether the user is speaking based on the strength of the vibration signal. When the strength of the vibration signal exceeds a first threshold, the voice activity detector 341 determines that the user is speaking at that moment. Alternatively, when the change in the strength of the vibration signal exceeds a second threshold, the voice activity detector 341 determines that the user has begun speaking at that moment. For another example, the voice activity detector 341 can determine whether the user is speaking based on the ratio between the vibration signal and the sound signal. When the ratio between the strength of the vibration signal and the sound signal exceeds a third threshold, the voice activity detector 341 determines that the user is speaking at that moment. Optionally, before determining the ratio between the vibration signal and the sound signal, the voice activity detector 341 (or other similar components) can perform noise reduction processing on the vibration signal and / or the sound signal.
[0061] Figure 4 Schematic diagram of the structure of the vibration sensor noise suppressor in the signal processing system provided by some embodiments of the present application. Figure 4In some embodiments, the vibration sensor noise suppressor 342 may include a noise relationship calculator 4421 and an environmental noise suppressor 4422 .
[0062] In some embodiments, the output of the voice activity detector 341 can serve as input to the noise relationship calculator 4421 and the environmental noise suppressor 4422. Specifically, in some embodiments, the noise relationship calculator 4421 can determine the relationship between the noise component in the sound signal and the noise component in the vibration signal based on the signal segments in the sound signal and the vibration signal that do not contain the user's voice (i.e., the noise segments, represented by VAD=0). Since both the vibration signal and the sound signal contain only noise components during the time period that does not contain the user's voice, the relationship between the noise component in the sound signal and the noise component in the vibration signal is equivalent to the relationship between the sound signal and the vibration signal. Based on the relationship between the noise component in the sound signal and the noise component in the vibration signal, the environmental noise suppressor 4422 can perform noise reduction processing on the signal segments in the vibration signal that contain the user's voice (i.e., the voice segments, represented by VAD=1) to obtain the target vibration signal.
[0063] For ease of understanding, the following description uses the sound signal collected by a single microphone. When the user is silent (ie, VAD=0), the sound signal collected by the microphone can be expressed as:
[0064] y(t)=n y (t), (1)
[0065] The vibration signal collected by the vibration sensor at the same time can be expressed as:
[0066] x(t)=n x (t), (2)
[0067] At this time, the relationship between the noise component of the vibration signal and the noise component of the sound signal h(t) can be expressed as:
[0068] x(t)=h(t)*y(t), (3)
[0069] In some embodiments, when voice activity detector 341 does not detect user speech, noise relationship calculator 4421 may update h(t) in real time. When voice activity detector 341 detects that the current signal contains a user voice signal, noise relationship calculator 4421 stops updating the noise relationship between the vibration signal and the sound signal. In some embodiments, the frequency with which noise relationship calculator 4421 updates the noise relationship is related to the noise level. When the noise level is low, noise relationship h(t) updates more slowly or may stop updating.
[0070] The environmental noise suppressor 4422 can be used to suppress the environmental noise component in the vibration signal when the user is speaking. In some embodiments, the input signal of the environmental noise suppressor 4422 may include the vibration signal, the sound signal, the most recently updated noise relationship, and the output signal of the voice activity detector 341. In some embodiments, when the user's voice and environmental noise are present at the same time, the vibration signal can be represented as:
[0071] x(t)=s x (t)+n x (t), (4)
[0072] where s x (t) represents the user voice received by the vibration sensor, n x (t) represents the ambient noise received by the vibration sensor. Similarly, in the presence of both user voice and ambient noise, the sound signal in the noisy environment can be expressed as:
[0073] y(t)=s y (t)+n y (t), (5)
[0074] where s y (t) can represent the user voice received by the microphone, n y (t) can represent the ambient noise received by the microphone. The relationship between the vibration sensor and the ambient noise received by the microphone can be approximately expressed as:
[0075] n x (t)=h(t)*n y (t), (6)
[0076] In some embodiments, the above-mentioned sound signal and vibration signal may be converted into the frequency domain. Specifically, the converted vibration signal is expressed as:
[0077] X(ω)=S X (ω)+N X (ω), (7)
[0078] Among them S X (ω) represents the frequency domain distribution of the user’s voice received by the vibration sensor, N X (ω) represents the frequency domain distribution of the ambient noise signal received by the vibration sensor. The converted sound signal can be expressed as:
[0079] Y(ω)=S Y (ω)+N Y (ω), (8)
[0080] Among them S Y (ω) represents the frequency domain distribution of the user’s voice received by the microphone, NY (ω) represents the frequency domain distribution of the ambient noise signal received by the microphone. The relationship between the ambient noise signal received by the vibration sensor and the ambient noise received by the microphone can be expressed as:
[0081] N X (ω)=H(ω)*N Y (ω), (9)
[0082] Where H(ω) is the frequency domain expression of the noise relationship h(t) in formula (3), which represents the noise relationship between the noise component in the sound signal and the noise component in the vibration signal in the frequency domain.
[0083] In some embodiments, considering that below a certain frequency range, for example, below 3000 Hz, the signal-to-noise ratio of the sound signal received by the microphone is lower than the signal-to-noise ratio of the vibration signal received by the vibration sensor (for more information on the signal-to-noise ratio of sound signals and vibration signals, see Figure 12 ), then the sound signal collected by the microphone can be approximated as an estimate of the noise signal, that is:
[0084] Y(ω)≈N Y (ω), (10)
[0085] Furthermore, according to formula (7), formula (9) and formula (10), the frequency domain expression of the denoised vibration signal can be expressed as:
[0086] S(ω)=S X (ω)=X(ω)-N X (ω)=X(ω)-H(ω)*N Y (ω)≈X(ω)-H(ω)*Y(ω), (11)
[0087] The meaning of each parameter can be found in the previous text and will not be repeated here.
[0088] In some embodiments, the voice activity detector 341 can function as a start switch. When it is detected that the sound signal and the vibration signal do not contain user voice (i.e., when VAD = 0), the noise relationship calculator 4421 can be activated to update the noise relationship between the two, and the environmental noise suppressor 4422 can be disabled. When it is detected that the sound signal and the vibration signal contain user voice (i.e., when VAD = 1), updating the noise relationship between the two is stopped, and the environmental noise suppressor 4422 is activated to perform noise reduction on the vibration signal. By controlling the operating states of the noise relationship calculator 4421 and the environmental noise suppressor 4422 through this method, unnecessary processing resource usage by the noise relationship calculator 4421 and the environmental noise suppressor 4422 can be avoided, thereby reducing the computational load of the processor to a certain extent.
[0089] Continue to refer to Figure 4 In some embodiments, the vibration sensor noise suppressor 342 may further include a steady-state noise suppressor 4423. The steady-state noise suppressor 4423 may be used to eliminate steady-state noise (e.g., background noise, etc.) in the signal generated by the vibration sensor. In some embodiments, background noise (also known as background noise) may exist in the vibration signal collected by the vibration sensor. Within a specific frequency range, the background noise may seriously affect the voice signal. Specifically, when using a vibration sensor to collect user voice, since the skin and bones have a low-pass filtering effect on the transmission of voice, the vibration sensor can receive fewer high-frequency voice signals, and the high-frequency components of the voice signal in the vibration signal it generates are also relatively small. Figure 5 : is a schematic diagram of the spectrum of the vibration signal generated by the vibration sensor provided in some embodiments of the present application. Figure 5 , the frame 501 part can represent the time domain signal corresponding to the vibration signal generated by the vibration sensor, and the frame 502 part can represent the corresponding frequency domain signal. In the period corresponding to the voice signal (such as the part shown in frame 503), the signal strength of the frequency domain signal below 1kHz is strong, and the signal strength at higher frequencies (such as above 2kHz) is weak. Figure 5 It can be seen that the signal received by the vibration sensor when a person is speaking has more low-frequency components and fewer high-frequency components.
[0090] In the frequency band where the user voice signal is relatively low within the vibration signal, for example, within the range of 2kHz–8kHz, the signal-to-noise ratio of the user voice signal collected by the vibration sensor is relatively low compared to the background noise. In this case, the vibration signal collected by the vibration sensor can be processed by the steady-state noise suppressor 4423 to reduce the impact of the background noise on the user voice signal. In some embodiments, the steady-state noise suppressor 4423 can use methods or devices such as spectral subtraction, Wiener filtering, and adaptive filtering to eliminate background noise.
[0091] Figure 6 : is a schematic diagram of the spectrum of the vibration signal generated by the vibration sensor in a noisy environment according to some embodiments of the present application. Figure 6 As can be seen, the voice signal (i.e., the signal corresponding to the user's voice) is minimally affected by noise signals within 1000Hz, resulting in a relatively clear signal. The voice signal is relatively less affected by noise signals between 1000Hz and 1500Hz, but the signal-to-noise ratio is lower than that within 1000Hz. Above 1500Hz, the voice signal is significantly affected by noise and is essentially "swamped" by the noise. This is partly because the higher the frequency, the smaller the voice signal received by the vibration sensor; and partly because the vibration sensor is more susceptible to high-frequency ambient noise signals.
[0092] Figure 7 Schematic diagram of a module of a signal processing system according to some other embodiments of the present application. Figure 7 As shown, in some embodiments, the system 500 may include a microphone signal noise suppressor 543, which may be used to reduce noise on a sound signal collected by at least one microphone 510 to obtain a clean air-conducted speech signal. Figure 7 As shown, the output signal of voice activity detector 541 and the sound signal generated by microphone 510 can simultaneously serve as input signals to microphone signal noise suppressor 543. In some embodiments, microphone signal noise suppressor 543 can process only the signal segments of the sound signal collected by microphone 510 that contain the user's voice based on the recognition results of voice activity detector 541. For example, when voice activity detector 541 determines that the user is speaking, microphone signal noise suppressor 543 will reduce the noise of the sound signal output by microphone 510 to generate a target sound signal.
[0093] Continue to refer to Figure 7 In some embodiments, the system 500 may further include a spectrum mixer 544. The spectrum mixer 544 may be used to perform spectrum aliasing processing on the target vibration signal obtained by the vibration sensor noise suppressor 542 and the target sound signal obtained by the microphone signal noise suppressor 543. For example, the spectrum mixer 544 may alias some components (e.g., low-frequency components) in the target vibration signal with some components (e.g., high-frequency components) in the target sound signal to form a full-band target signal. In some embodiments, the frequency of the portion of the target vibration signal used for aliasing is lower than the frequency of the portion of the target sound signal used for aliasing. In some embodiments, the highest frequency of the portion of the target vibration signal used for aliasing is equal to or greater than the lowest frequency of the portion of the target sound signal used for aliasing.
[0094] In some embodiments, the frequency range of the target vibration signal and the frequency range of the target sound signal may overlap. For example, the frequency range of the target vibration signal may be between 0Hz-2000Hz, and the frequency range of the target sound signal may be between 1000Hz-8000Hz. For another example, the frequency range of the target vibration signal may be between 0Hz-2000Hz, and the frequency range of the target sound signal may be between 0Hz-10kHz. Optionally, the spectrum mixer 544 may include one or more filtering circuits for filtering the aliased parts of the target vibration signal and / or the target sound signal before mixing. It should be noted that the above data is only for illustrative purposes. In some embodiments, the frequency range of the target vibration signal and the target sound signal may be, but not limited to, the above numerical range.
[0095] It should be noted that Figure 7 The signal processing system shown is compared to Figure 3 A microphone signal noise suppressor 543 and a spectrum mixer 544 are added. The common parts can be referred to Figure 3 For example, more technical details about the voice activity detector 541 can be found in Figure 3 The voice activity detector 341 in will not be described in detail here.
[0096] Figure 8 According to the method provided in some embodiments of the present application Figure 6 Schematic diagram of signal spectrum obtained after processing the signal shown. Block 801 may represent a time domain signal obtained after processing the vibration signal generated by the vibration sensor, and block 802 may represent a frequency domain signal obtained after processing the vibration signal.
[0097] Compared to Figure 6 ,from Figure 8 As can be seen, the above processing method has a significant noise reduction effect on noise in the 1500Hz-4000Hz range. The target signal obtained through this processing method not only preserves the low-frequency (e.g., 0-1000Hz) user voice signal, but also reduces the noise of mid- and high-frequency (e.g., 1500-4000Hz) vibration signals, resulting in a target signal with a high signal-to-noise ratio.
[0098] Figure 9 Schematic diagram of a module of a signal processing system according to some other embodiments of the present application. Figure 9 As shown, in some embodiments, the system 600 may include a noise signal generator 643, which may be part of the processor. In some embodiments, since there are certain differences in the directions of the microphones in the microphone array 610 relative to the sound source, and this difference will cause certain differences in the amplitude and / or phase of the sound signals collected by different microphones in the microphone array 610, based on this principle, the noise signal generator 643 can determine the first noise signal from the sound signals collected by them according to the relative position relationship between the microphones in the microphone array 610. In some embodiments, the first noise signal may be a noise signal in a specific direction in the environment. For example, the first noise signal may be a noise signal synthesized from noises in all directions in the environment except the direction of the user's voice. It should be noted that, Figure 9 The signal processing system shown is Figure 3 The common parts of the systems shown can be referred to Figure 3 For example, more technical details about the voice activity detector 641 can be found in Figure 3The voice activity detector 341 in will not be described in detail here.
[0099] Furthermore, in some embodiments, the vibration sensor noise suppressor 642 may determine the relationship between the first noise signal and the vibration signal collected by the vibration sensor 630 according to the method described elsewhere in this specification, and perform noise reduction processing on the vibration signal based on the relationship.
[0100] In some embodiments, when the vibration sensor noise suppressor 642 determines the relationship between the first noise signal and the vibration signal collected by the vibration sensor 630 based on the first noise signal, if there is no user voice and only noise, the vibration signal can be expressed as x(t)=n x (t), the first noise signal can be expressed as n(t), and the relationship between the two can be expressed as:
[0101] x(t)=h(t)*n(t), (12)
[0102] Where h(t) is the calculated noise relationship.
[0103] In some embodiments, if user voice and noise are present simultaneously, the vibration signal in the noisy environment can be represented as:
[0104] x(t)=s(t)+n x (t), (13)
[0105] Where s(t) represents the user's voice, n x (t) represents the environmental noise received by the vibration sensor. x The relationship between (t) and the first noise signal is approximately:
[0106] n x (t) = h(t)*n(t), (14)
[0107] At this time, according to formulas (13) and (14), the environmental noise can be removed from the vibration signal to obtain a clean user voice signal.
[0108] In some alternative embodiments, the vibration sensor noise suppressor 642 may treat the components of the vibration signal whose correlation with the noise signal is higher than a preset threshold (e.g., 60%, 80%, 90%, etc.) as noise, and treat the components of the vibration signal whose correlation with the noise signal is lower than the preset threshold as user voice.
[0109] For example, the vibration sensor noise suppressor 642 can identify the time interval during which the user utters speech, and determine a second noise signal reflecting ambient noise from the sound signal within that time interval (for example, by identifying sound from a direction other than the user's mouth using the microphone array described above). Furthermore, the vibration sensor noise suppressor 642 can determine the correlation between different components of the vibration signal within that time interval and the second noise signal. For example, components of the vibration signal whose correlation with the second noise signal is higher than a preset threshold are considered noise, while components whose correlation with the second noise signal is lower than the preset threshold are considered user speech.
[0110] Figure 10 It is a module diagram of a signal processing system provided according to other embodiments of the present application.
[0111] like Figure 10 As shown, in some embodiments, the system 700 may include a noise signal generator 743 and a voice signal generator 744, which may be part of the processor 140. The noise signal generator 743 may determine a first noise signal from the sound signal collected by it based on the relative positional relationships between the microphones in the microphone array 710. Similarly, the voice signal generator 744 may determine a first voice signal from the sound signal collected by it based on the relative positional relationships between the microphones in the microphone array 710. In some embodiments, the first noise signal may represent the noise in a specific direction in the environment collected by the microphone array 710. For example, the first noise signal may be a noise signal synthesized from noise in all directions in the environment except the direction of the user's voice. The first voice signal may represent the sound from the direction of the user's mouth, i.e., the user's voice, in the sound signal collected by the microphone array 710.
[0112] In some embodiments, when the microphone array 710 is a beamforming microphone array, the first noise signal may be a signal of a noise beam; when the microphone array 710 is another type of array, the first noise signal may be noise calculated by other methods. Similarly, in some embodiments, when the microphone array 710 is a beamforming microphone array, the first speech signal may be a signal of a speech beam; when the microphone array 710 is another type of array, the first speech signal may be a speech signal calculated by other methods.
[0113] In some embodiments, the system 700 may further include a microphone signal noise suppressor 742, which may be part of the processor. In some embodiments, the microphone signal noise suppressor 742 may perform noise reduction processing on the voice signal collected by the microphone array 710 based on the first noise signal and the first voice signal to obtain a target voice signal. For example, the microphone signal noise suppressor 742 may further process the first voice signal to remove components that have the same characteristics as the first noise signal from the first voice signal, thereby obtaining the target voice signal. In some alternative embodiments, the microphone signal noise suppressor 742 may directly use the first voice signal as the target voice signal.
[0114] In some embodiments, the target voice signal processed by microphone signal noise suppressor 742 may be aliased with the target vibration signal processed by vibration sensor noise suppressor 642 to form a full-band target signal. In some embodiments, the frequency of the portion of the target vibration signal used for aliasing is lower than the frequency of the portion of the target sound signal used for aliasing. In some embodiments, the highest frequency of the portion of the target vibration signal used for aliasing is equal to or greater than the lowest frequency of the portion of the target sound signal used for aliasing.
[0115] In some embodiments, the output signal of the voice activity detector 741 can be used as the input signal of the microphone signal noise suppressor 742. The input signal of the voice activity detector 741 may include the sound signal collected by the microphone array 710 and the vibration signal collected by the vibration sensor 730. Specifically, the microphone signal noise suppressor 742 can perform noise reduction processing only on the signal segment containing the user's voice in the sound signal collected by the microphone array 710 based on the recognition result of the voice activity detector 741. It should be noted that Figure 10 The signal processing system shown is Figure 9 The common parts of the systems shown can be referred to Figure 9 For example, more technical details about the voice activity detector 741 can be found in Figure 9 The voice activity detector 641 in will not be described in detail here.
[0116] Considering that when using microphones to estimate noise, a microphone array can better estimate noise from directions other than the user's voice source (i.e., the user's mouth), but it is difficult to obtain noise close to or identical to the user's voice source direction. When using a single microphone signal for noise estimation, although the processed noise can include the user's mouth direction, it can only process frequency bands with a lower signal-to-noise ratio than the vibration sensor and cannot reduce noise in other frequency bands. Therefore, in some embodiments, microphone array noise reduction and single-microphone noise reduction can be combined to achieve better noise reduction effects.
[0117] Figure 11 It is a module diagram of a signal processing system provided according to other embodiments of the present application.
[0118] like Figure 11 As shown, in some embodiments, in order to combine the advantages of microphone array noise reduction and single microphone noise reduction, the system 800 can add a noise mixer 8424. The noise mixer 8424 can be part of the processor 140. In some embodiments, the input signal of the noise mixer 8424 can include a microphone signal collected by a microphone. For example, the noise signal can come from Figure 9 The first noise signal generated by the noise signal generator 643. The microphone signal may come from Figure 9 The output signal of one of the microphones in the microphone array 610, or Figure 7 In some embodiments, the noise mixer 8424 can mix the noise signal with the microphone signal to generate a sound signal. Figure 4 For the sound signal input into the noise relationship calculator, the noise characteristics can be reflected more accurately, thereby improving the accuracy of noise estimation.
[0119] Further, continue to refer to Figure 11 The noise relationship calculator 8421 can determine the noise relationship between the vibration signal collected by at least one vibration sensor and the signal segment that does not contain the user voice in the sound signal generated by the aforementioned noise mixer 8424 (that is, the noise segment with VAD=0).
[0120] It should be noted that by adding the noise mixer 8424, the mixed sound signal can be increased to include noise in the same direction as the user's voice compared to the first noise signal, while the user's voice signal can be reduced compared to the noise signal. The result is better than using the noise signal alone or the microphone signal alone, and a more reliable noise estimation can be obtained, thereby improving the accuracy of the noise estimation.
[0121] In some embodiments, the mixing method of the noise signal and the microphone signal can be a fixed ratio or other methods. In some embodiments, the noise mixer 8424 can obtain the noise level from the direction of the user's voice and determine the mixing ratio of the noise signal and the microphone signal based on the noise level. For example, the louder the noise sound from the same direction as the user's voice, the greater the mixing ratio of the microphone signal.
[0122] It should be noted that Figure 11 The signal processing system shown is Figure 4 The common parts of the systems shown can be referred to Figure 4 For example, more technical details about the ambient noise suppressor 8422 and the steady-state noise suppressor 8423 can be found in Figure 4 The environmental noise suppressor 4422 and steady-state noise suppressor 4423 are not described in detail here.
[0123] Figure 12 2 is a schematic diagram of a signal frequency-signal-to-noise ratio curve provided according to some embodiments of the present application.
[0124] It should be noted that the signal-to-noise ratio of the sound signal received by the microphone is different from the signal-to-noise ratio of the vibration signal received by the vibration sensor. Figure 12 As shown, within the frequency range of less than 3000 Hz, the signal-to-noise ratio of the vibration sensor is greater than that of the microphone; within the frequency range of 4000 Hz–8000 Hz, the signal-to-noise ratio of the vibration sensor is less than that of the microphone. The signal-to-noise ratios of the microphone and the vibration sensor overlap within the range of 3000 Hz–4000 Hz. In some embodiments, the sound signal collected by the microphone can be approximated as an estimate of the noise signal within a lower frequency range (e.g., less than 3000 Hz). Considering that the signal-to-noise ratio of the vibration signal decreases with increasing frequency, in some embodiments, when spectrally aliasing the target sound signal and the target vibration signal, the highest frequency of the portion of the target vibration signal used for aliasing can be set to no higher than 3000 Hz but no lower than 1000 Hz. Preferably, the highest frequency of the portion of the target vibration signal used for aliasing can be set to no higher than 2500 Hz but no lower than 1500 Hz. More preferably, the highest frequency of the portion of the target vibration signal used for aliasing can be set to no higher than 2000 Hz but no lower than 1000 Hz.
[0125] It should be noted that the above description of the signal-to-noise ratio of the vibration sensor and the microphone is for illustrative purposes only. In some embodiments, when the position of the vibration sensor or the microphone changes, the signal-to-noise ratio comparison between the two will be different, and the position where the signal-to-noise ratio overlaps will also change.
[0126] The embodiments of this specification also provide a computer-readable storage medium, which stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer implements the operations corresponding to the aforementioned signal processing method.
[0127] It should be noted that the above-mentioned storage medium may be included in the above-mentioned electronic device, processor or server; or it may exist independently without being assembled into the electronic device, processor or server.
[0128] The basic concepts have been described above. It will be apparent to those skilled in the art that the detailed disclosure above is merely illustrative and does not limit the present application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and amendments to the present application. Such modifications, improvements, and amendments are suggested in the present application and remain within the spirit and scope of the exemplary embodiments of the present application.
[0129] At the same time, this application uses specific terms to describe the embodiments of this application. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "one embodiment," "an embodiment," or "an alternative embodiment" mentioned twice or multiple times in different locations in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application may be appropriately combined.
[0130] In addition, it will be understood by those skilled in the art that various aspects of the present application can be illustrated and described by a number of patentable categories or situations, including any new and useful process, machine, product or combination of substances, or any new and useful improvements thereto. Accordingly, various aspects of the present application can be performed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software may all be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". In addition, various aspects of the present application may be represented as a computer product located in one or more computer-readable media, which includes computer-readable program code.
[0131] A computer storage medium may include a propagated data signal embodying the computer program code, for example, in baseband or as part of a carrier wave. The propagated signal may be in a variety of forms, including electromagnetic, optical, or any suitable combination thereof. A computer storage medium may be any computer-readable medium other than a computer-readable storage medium that can be connected to an instruction execution system, apparatus, or device to communicate, propagate, or transfer the program for use. The program code on the computer storage medium may be transmitted via any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of these.
[0132] The computer program code required for the operation of each part of the present application can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C language, Visual Basic, Fortran 2003, Perl, COBOL 2002, PHP, ABAP, dynamic programming languages such as Python, Ruby and Groovy, or other programming languages. The program code can be run entirely on the user's computer, or as a separate software package on the user's computer, or partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).
[0133] In addition, unless expressly stated in the claims, the order of the processing elements and sequences described in this application, the use of alphanumeric characters, or the use of other names are not intended to limit the order of the processes and methods of this application. Although the above disclosure discusses some of the invention embodiments currently considered useful through various examples, it should be understood that such details are only for illustrative purposes, and the attached claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the essence and scope of the embodiments of this application. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.
[0134] Similarly, it should be noted that, in order to simplify the presentation of this application and thus facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this application sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not mean that the subject matter of this application requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single embodiment disclosed above.
[0135] In some embodiments, numbers are used to describe the quantity of components and attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may change according to the required features of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of the present application are approximate values, in specific embodiments, the settings of such numerical values are as accurate as possible within the feasible range.
[0136] Each patent, patent application, patent application disclosure, and other materials, such as articles, books, specifications, publications, documents, etc., cited in this application is hereby incorporated by reference in its entirety. This includes application history documents that are inconsistent with or conflict with the content of this application, as well as documents (currently or subsequently attached to this application) that limit the broadest scope of the claims of this application. It should be noted that if the descriptions, definitions, and / or use of terms in the accompanying materials of this application are inconsistent or conflicting with the content of this application, the descriptions, definitions, and / or use of terms in this application shall prevail.
[0137] Finally, it should be understood that the embodiments described in this application are merely illustrative of the principles of the embodiments of this application. Other variations may also fall within the scope of this application. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this application may be considered consistent with the teachings of this application. Accordingly, the embodiments of this application are not limited to the embodiments explicitly introduced and described in this application.
Claims
1. A signal processing system, characterized in that: include: at least one microphone, wherein the at least one microphone is used to collect sound signals; at least one vibration sensor, wherein the at least one vibration sensor is used to collect vibration signals; as well as The processor is configured to: Determining a noise signal reflecting ambient noise from a sound signal within a time interval in which no user speech is emitted; determining a relationship between the noise signal and the vibration signal during the time interval when the user voice is not emitted as a relationship between the noise component in the sound signal and the noise component in the vibration signal when the user voice is emitted; as well as Noise reduction processing is performed on the vibration signal based at least on the relationship to obtain a target vibration signal.
2. The system according to claim 1, wherein Also included is a voice activity detector configured to: identifying a signal segment in the sound signal and the vibration signal that does not include the user voice; Determining the relationship between the noise signal and the vibration signal during the time interval in which no user voice is emitted includes: Based on signal segments of the sound signal and the vibration signal that do not include the user voice, a relationship between noise components in the sound signal and noise components in the vibration signal is determined.
3. The system according to claim 2, wherein: The processor is further configured to: The sound signal and the vibration signal contain a signal segment of the user's voice, and noise reduction processing is performed on the vibration signal based on the relationship to obtain the target vibration signal.
4. The system according to claim 3, wherein: The processor is further configured to suppress steady-state noise in the vibration signal to obtain the target vibration signal.
5. The system according to claim 2, wherein: The processor is further configured to: Converting the sound signal and the vibration signal from time domain signals to frequency domain signals; and A noise relationship between a noise component in the sound signal and a noise component in the vibration signal in at least one frequency domain subband is obtained.
6. The system according to claim 2, wherein: The processor is further configured to: The sound signal includes a signal segment of the user's voice, and the sound signal is subjected to noise reduction processing to obtain a target sound signal.
7. The system according to claim 6, wherein: The processor is further configured to: At least part of the components in the target vibration signal are mixed with at least part of the components in the target sound signal to obtain a target signal, wherein the frequency of at least part of the components in the target vibration signal is lower than the frequency of at least part of the components in the target sound signal.
8. The system according to claim 2, wherein: The at least one microphone includes a microphone array including a plurality of microphones, and determining, based on a signal segment of the sound signal and the vibration signal that does not include the user voice, a relationship between a noise component in the sound signal and a noise component in the vibration signal includes: determining a first noise signal from the sound signal based on a relative positional relationship between microphones in the microphone array in a signal segment of the sound signal and the vibration signal that does not include the user voice; and A relationship between the first noise signal and the vibration signal is determined.
9. The system according to claim 8, wherein The processor is further configured to: a signal segment containing the user's voice in the sound signal, and determining a first voice signal from the sound signal based on a relative position relationship between microphones in the microphone array; as well as The sound signal is subjected to noise reduction processing based on the first noise signal and the first speech signal to obtain a target sound signal, or the first speech signal is used as the target sound signal.
10. The system according to claim 1, wherein: The system includes a noise mixer and a plurality of microphones, and generating a sound signal includes: determining a first noise signal based on a relative position relationship between the plurality of microphones; Acquiring a microphone signal collected by at least one target microphone among the plurality of microphones; and The noise mixer mixes the first noise signal with the microphone signal to generate the sound signal.
11. The system according to claim 10, wherein: The noise mixer is configured to: A noise level from the direction of the user voice is acquired, and a mixing ratio of the first noise signal and the microphone signal is determined based on the noise level.
12. The system according to any one of claims 1 to 11, wherein: The signal-to-noise ratio of the at least one vibration sensor is greater than the signal-to-noise ratio of the at least one microphone in at least a portion of the frequency range.
13. A signal processing method, characterized in that: include: collecting a sound signal by at least one microphone; collecting a vibration signal by at least one vibration sensor; Determining a noise signal reflecting ambient noise from a sound signal within a time interval in which no user speech is emitted; determining a relationship between the noise signal and the vibration signal during the time interval when the user voice is not emitted as a relationship between the noise component in the sound signal and the noise component in the vibration signal when the user voice is emitted; as well as Noise reduction processing is performed on the vibration signal based at least on the relationship to obtain a target vibration signal.
14. The method according to claim 13, wherein The method further includes: identifying a signal segment in the sound signal and the vibration signal that does not contain the user's voice; Determining the relationship between the noise signal and the vibration signal during the time interval in which no user voice is emitted includes: Based on signal segments of the sound signal and the vibration signal that do not include the user voice, a relationship between noise components in the sound signal and noise components in the vibration signal is determined.
15. The method according to claim 14, wherein The performing noise reduction processing on the vibration signal based at least on the relationship to obtain a target vibration signal includes: The sound signal and the vibration signal contain a signal segment of the user's voice, and noise reduction processing is performed on the vibration signal based on the relationship to obtain the target vibration signal.
16. The method according to claim 15, wherein The method further includes suppressing steady-state noise in the vibration signal to obtain the target vibration signal.
17. The method according to claim 14, wherein The method further comprises: Converting the sound signal and the vibration signal from time domain signals to frequency domain signals; and A noise relationship between a noise component in the sound signal and a noise component in the vibration signal in at least one frequency domain subband is obtained.
18. The method according to claim 14, wherein The method further comprises: The sound signal includes a signal segment of the user's voice, and the sound signal is subjected to noise reduction processing to obtain a target sound signal.
19. The method according to claim 18, wherein The method further comprises: At least part of the components in the target vibration signal are mixed with at least part of the components in the target sound signal to obtain a target signal, wherein the frequency of at least part of the components in the target vibration signal is lower than the frequency of at least part of the components in the target sound signal.
20. The method of claim 14, wherein: The at least one microphone includes a microphone array including a plurality of microphones, and determining, based on signal segments in the sound signal and the vibration signal that do not include the user voice, a relationship between a noise component in the sound signal and a noise component in the vibration signal includes: determining a first noise signal from the sound signal based on a relative positional relationship between microphones in the microphone array in a signal segment of the sound signal and the vibration signal that does not include the user voice; and A relationship between the first noise signal and the vibration signal is determined.
21. The method according to claim 20, wherein The method further comprises: The sound signal includes a signal segment of the user's voice, and determining a first voice signal from the sound signal based on a relative position relationship between microphones in the microphone array; and The sound signal is subjected to noise reduction processing based on the first noise signal and the first speech signal to obtain a target sound signal, or the first speech signal is used as the target sound signal.
22. The method of claim 13, wherein: The at least one microphone includes a plurality of microphones, the method further comprising: determining a first noise signal based on a relative position relationship between the plurality of microphones; Acquiring a microphone signal collected by at least one target microphone among the plurality of microphones; and The first noise signal is mixed with the microphone signal to generate the sound signal.
23. The method according to claim 22, wherein The method further comprises: A noise level from the direction of the user voice is acquired, and a mixing ratio of the first noise signal and the microphone signal is determined based on the noise level.
24. The method according to any one of claims 13 to 23, wherein The signal-to-noise ratio of the at least one vibration sensor is greater than the signal-to-noise ratio of the at least one microphone in at least a portion of the frequency range.
25. An electronic device, characterized in that: comprising at least one processor and at least one memory; The at least one memory is for storing computer instructions; The at least one processor is configured to execute at least part of the computer instructions to implement the method according to any one of claims 13 to 24.
26. A computer-readable storage medium, characterized in that The storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the method according to any one of claims 13 to 24 is executed.
Citation Information
Patent Citations
Method and apparatus for reducing noise corruption from an alternative sensor signal during multi-sensory speech enhancement
US20060178880A1
System and method for authenticating voice commands for a voice assistant
US20180068671A1