Speech enhancement, interaction method, apparatus, program product, and device
By setting time intervals for noise feature extraction and beamformer updates on devices such as robotic vacuum cleaners and drones, and utilizing the differences in noise and speech signal characteristics, the low signal-to-noise ratio and multiple interference sources of mobile devices are addressed, achieving effective voice enhancement and interaction.
Patent Information
- Application Number
- CN202011508186.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-18
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2040-12-18
AI Technical Summary
Mobile devices such as robot vacuum cleaners and drones struggle to effectively enhance voice and facilitate interaction due to their own noise interference, low signal-to-noise ratio caused by movement, and multiple interference sources.
By setting a time interval between noise feature extraction and beamformer update, the characteristics of the device's own noise signal feature change being small while the speech signal feature change being large are taken advantage of to avoid the elimination of useful speech components. The beamformer is updated using noise covariance matrix and weight vector.
It effectively suppresses device noise under conditions of extremely low signal-to-noise ratio and multiple interference sources, achieving effective voice enhancement and interaction, and improving voice recognition accuracy.
Smart Images

Figure CN114648999B_ABST
Abstract
Description
Technical Field
[0001] This application relates to a speech enhancement, interaction method, apparatus, program product, and device, belonging to the field of computer technology. Background Technology
[0002] Speech enhancement refers to the technique of extracting useful speech signals from noisy backgrounds and suppressing or reducing noise interference when speech signals are interfered with or even drowned out by various kinds of noise. Speech enhancement is widely used in various human-computer interaction scenarios that require speech recognition.
[0003] As an important device in smart homes, robotic vacuum cleaners are gradually developing towards voice-activated and intelligent operation. However, direct voice interaction with robotic vacuum cleaners presents significant challenges: Firstly, the mechanical noise, motor noise, and vacuum cleaner noise generated during operation are considerable. Furthermore, the microphone is mounted on the vacuum cleaner itself, close to the noise source, resulting in extremely low signal-to-noise ratio. Additionally, multiple sources of noise can be generated on the vacuum cleaner, all located close to the microphone, creating a multi-source interference problem. Secondly, the vacuum cleaner moves during operation, causing the received signal to be dynamic and real-time, making it difficult to determine the direction of the voice source and thus hindering effective noise reduction. These same problems exist with other similar mobile smart devices requiring human-computer interaction, such as drones. Summary of the Invention
[0004] This invention provides a voice enhancement, interaction method, apparatus, program product, and device to improve the voice enhancement effect of mobile devices.
[0005] To achieve the above objectives, embodiments of the present invention provide a speech enhancement processing method, including:
[0006] During the first time period, microphone signals are acquired, and noise features are extracted based on the microphone signals;
[0007] After a second time interval, the beamformer is updated based on the noise characteristics;
[0008] The updated beamformer is used to enhance the speech of subsequent microphone signals.
[0009] This invention also provides a speech enhancement processing device, comprising:
[0010] The noise feature extraction module is used to collect microphone signals during a first time period and extract noise features based on the microphone signals.
[0011] A beamformer update module is used to update the beamformer according to the noise characteristics after a second time interval.
[0012] The voice enhancement processing module is used to enhance the voice of subsequent microphone signals using an updated beamformer.
[0013] This invention also provides an electronic device, comprising:
[0014] Memory, used to store programs;
[0015] A processor is configured to run the program stored in the memory to perform the above-described speech enhancement processing method.
[0016] This invention also provides a computer program product, including a computer program or instructions, which, when executed by a processor, cause the processor to implement the aforementioned speech enhancement processing method.
[0017] This invention also provides a voice interaction method, including:
[0018] Acquire sound signals;
[0019] The sound signal is subjected to speech enhancement processing. During the speech enhancement processing, noise features are extracted, and the beamformer used for speech enhancement processing is updated according to the noise features after a certain period of time.
[0020] The system performs voice command recognition on the audio signal after speech enhancement processing and executes corresponding processing based on the recognized voice.
[0021] This invention also provides a voice interaction device, comprising:
[0022] A sound signal acquisition module is used to acquire sound signals.
[0023] The speech enhancement processing module is used to perform speech enhancement processing on the sound signal. During the speech enhancement processing, noise features are extracted, and the beamformer used for speech enhancement processing is updated according to the noise features after a certain period of time.
[0024] The instruction recognition and processing module is used to recognize voice instructions from the voice-enhanced audio signal and perform corresponding processing based on the recognized voice.
[0025] This invention also provides a drone, wherein the drone includes the aforementioned voice interaction device.
[0026] This invention also provides a robotic vacuum cleaner, which includes the aforementioned voice interaction device.
[0027] This invention also provides an electronic device, comprising:
[0028] Memory, used to store programs;
[0029] A processor is configured to run the program stored in the memory to perform the aforementioned voice interaction method.
[0030] This invention also provides a computer program product, including a computer program or instructions, characterized in that, when the computer program or instructions are executed by a processor, the processor implements the aforementioned voice interaction method.
[0031] The speech enhancement, interaction method, apparatus, program product, and device of this invention utilize the characteristic that the device's own noise signal characteristics change little during device movement, while the external speech signal characteristics change significantly due to changes in the location of the sound source. A time interval is set between noise feature acquisition and beamformer update processing. During this time interval, the device moves, resulting in the numbering of the speech source location. By leveraging the characteristic that the device's own noise source remains unchanged despite the change in the speech source caused by this time interval, useful speech components are avoided from being eliminated, while simultaneously suppressing the device's own noise. The speech enhancement processing technology of this invention can effectively suppress the noise emitted by the device itself under conditions of extremely low signal-to-noise ratio, multiple interference sources, and moving sound sources, achieving effective speech enhancement.
[0032] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of one application scenario of the speech enhancement processing method according to an embodiment of the present invention;
[0034] Figure 2 This is a second schematic diagram illustrating an application scenario of the speech enhancement processing method according to an embodiment of the present invention;
[0035] Figure 3 This is a schematic flowchart of the speech enhancement processing method according to an embodiment of the present invention;
[0036] Figure 4 This is a schematic diagram of the speech enhancement processing device according to an embodiment of the present invention;
[0037] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0038] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0039] like Figure 1 and Figure 2 The diagram illustrates an application scenario of the voice enhancement processing method according to an embodiment of the present invention. This method can be applied to mobile devices, such as robotic vacuum cleaners and drones. The following uses a robotic vacuum cleaner as an example to explain the technical principle of the voice enhancement processing method. In the application scenario of a robotic vacuum cleaner, noise is mainly generated by the robot itself, including mechanical noise, motor noise, and vacuum cleaner noise. This self-generated noise is much louder than the user's voice. The microphone is located closer to the noise sources (motor, vacuum cleaner, etc.) than the voice source. Furthermore, the robotic vacuum cleaner moves during operation, causing the voice signal received by the microphone to be dynamic in real time. However, since the relative position of the noise source and the microphone is fixed, the noise signal received by the microphone is relatively stable.
[0040] This invention utilizes the difference in sound characteristics between noise and speech due to the different locations of the sound sources for noise reduction. Specifically, the time period for noise acquisition and the time period for updating the beamformer are spaced out, taking advantage of the different effects of the robot vacuum cleaner's movement on speech and noise sources to perform noise reduction.
[0041] The technical solution of the present invention will be further illustrated below through some specific embodiments.
[0042] Example 1
[0043] like Figure 1 As shown, the robotic vacuum cleaner moves continuously during operation. The diagram exemplifies the robot's movement trajectory over four time periods, from t1 to t5. The large circle represents the robotic vacuum cleaner, and the smaller circles with jagged edges outside the large circle represent noise sources such as the motor and vacuum cleaner itself. As shown, regardless of how the robot moves and / or rotates, the noise source remains stationary on the robotic vacuum cleaner. The lower part of the diagram shows the voice source, typically emitted by the user. Figure 1In the scenario shown, the user's position is assumed to be stationary while the robot vacuum is moving. As the robot vacuum's position changes, its distance and direction relative to the voice source also change. However, since the noise sources (motor, vacuum cleaner, etc.) move along with the microphone, their relative positions remain constant. Furthermore, if the robot vacuum rotates, its direction relative to the voice source changes, but its own noise sources (motor, vacuum cleaner) rotate synchronously with the microphone as a whole, so their relative positions remain unchanged.
[0044] Combination Figure 2 The upper horizontal time axis represents the noise feature extraction and beamformer update process. T2 to T5 represents a noise feature calibration cycle, during which the covariance matrix, representing the noise features, is calculated. Then, after waiting for T3 to T4, the beamformer is updated. After one noise feature calibration cycle is completed, the process can either begin the next cycle or end. If a new calibration cycle is needed, there can be a time interval between two calibration cycles, i.e., the time interval T1 to T2 in the diagram. The time interval T1 to T2 can be a preset time, and its length can be set as needed; its main purpose is to control the frequency of beamformer updates. The time interval T2 to T3 is also a preset value, depending on the noise feature extraction requirements. If it is desired to extract noise features based on microphone signals from more data frames, the time interval T2 to T3 will be correspondingly longer. The time interval T2 to T3 can be a preset value or dynamically determined based on the device's movement speed. The time interval from t3 to t4 is used for the beamformer update process. This time interval can be a preset value, depending on the processing speed of the device during the beamformer update process.
[0045] Figure 2 The lower timeline represents the execution process of speech enhancement processing. Speech enhancement processing and noise feature calibration are performed synchronously and do not affect each other. Before the beamformer is updated, the previous beamformer is used for speech enhancement processing.
[0046] like Figure 3 The diagram shown is a flowchart of the voice enhancement processing method according to an embodiment of the present invention. This method can be applied to mobile devices that require voice interaction, or it can be applied to the server side. The server obtains microphone signals from the device side through the network and performs voice enhancement processing. The specific processing procedure is as follows:
[0047] S101: During the first time period, microphone signals are acquired, and noise features are extracted based on the microphone signals. The first time period can correspond to... Figure 1During the time period from t2 to t3, the microphone signal is treated as noise signal for feature extraction. The extracted noise features can specifically be a noise covariance matrix. Specifically, this step can further include:
[0048] S1011: Acquire microphone signals from multiple consecutive data frames. The acquired microphone signals can be represented as x(τ), where τ is the data frame number. Here, the microphone signals are multi-channel audio signals acquired by a microphone array formed by multiple microphone units. x(τ) is a vector whose dimension is equal to the number of microphone units. Each element of the vector is the signal value acquired by each microphone unit at the time of data frame τ.
[0049] S1012: Based on a preset forgetting factor, calculate the noise covariance matrix of the current data frame according to the noise covariance matrix of the previous data frame and the microphone signal of the current data frame. The noise covariance matrix can be calculated using the following formula (1):
[0050] C(τ)=αC(τ-1)+(1-α)x(τ)x H (τ) Formula (1)
[0051] Wherein, α is the forgetting factor, which is used to adjust the influence of the noise covariance matrix corresponding to the previous data frame on the noise covariance matrix corresponding to the current data frame. The forgetting factor is used to adjust the influence of the previous iteration data on the current iteration calculation in the iterative averaging calculation. In the embodiment of the present invention, the noise covariance matrix is actually calculated by iterative averaging of multiple data frames in formula (1). In each iteration calculation, the noise covariance matrix obtained by the previous iteration calculation is used, and the forgetting factor is to adjust the influence of the noise covariance matrix of the previous iteration on the current iteration. The initial value of the noise covariance matrix can be a set value or the noise covariance matrix calculated in the previous noise feature period.
[0052] S1013: The noise covariance matrix corresponding to the last data frame in the first time period is used as the noise feature. Specifically, in the first time period, a preset number of data frames can be sampled as data for calculating the noise covariance matrix. Through the calculation of formula (1), the preset number of data frames are statistically processed. The noise covariance matrix corresponding to the current data frame is calculated based on the previous noise covariance matrix. When calculating the noise covariance matrix corresponding to the first data frame, the noise covariance matrix calculated in the first time period is used as C(τ-1) and included in the calculation. After multiple rounds of calculation using formula (1), the noise covariance matrix corresponding to the last data frame will be used as the noise feature extracted in the first time period for subsequent updates to the beamformer.
[0053] S102: After the second time interval, the beamformer is updated based on noise characteristics. The beamformer's function is to adjust the weighting of the signals received by each microphone channel by assigning weighting factors and then summing them. These weighting factors are the coefficients of the beamformer, specifically existing in the form of a weight vector, which will be explained in detail below. In this embodiment, the main function of the beamformer is to update this weight vector. The second time interval can correspond to... Figure 1 During the time interval t3 to t4, the position of the robot vacuum cleaner changed significantly. Waiting until this time interval to update the beamformer is primarily to avoid including useful speech components in the noise features collected during the t2 to t3 time interval. That is, the noise covariance matrix mentioned above contains speech component features. If the beamformer is updated immediately after updating the noise covariance matrix, then during the t3 to t4 time interval, since the speech features do not change much from those in the t2 to t3 time interval, the beamformer may filter out useful speech components during spatial filtering. However, if the interval is from t3 to t4, the spatial filtering process will be based on the directional characteristics of the microphone signal. Due to the movement characteristics of the robot vacuum cleaner, the location of the noise source will not change after the interval from t3 to t4, but the location of the speech source will change significantly. Even if there are speech feature components in the noise features during the interval from t2 to t3, the features of the new speech signal after the interval from t3 to t4 will be very different from the speech feature components during the interval from t2 to t3. Therefore, the beamformer will be updated based on the noise covariance matrix extracted as noise features during the interval from t2 to t3, without affecting the speech signal during the interval from t4 to t5 and beyond.
[0054] The length of the time interval t3 to t4 can be a preset value or dynamically set according to the actual situation. Specifically, the motion data of the device where the microphone is located can be detected, and the length of the second time interval can be determined based on the motion data. Taking a robot vacuum cleaner as an example, the moving speed of the robot vacuum cleaner can be detected, and the length of the time interval t3 to t4 can be set inversely proportional to the moving speed. The motion data mentioned here can include moving speed, moving direction, moving distance, and rotation angle, etc. In this embodiment of the invention, the purpose of setting the second time interval is to allow the relative position of the voice source to change significantly, so that the voice features change significantly, thereby ensuring that the beamformer updated based on the noise features (which may contain some voice features) collected in the time interval t2 to t3 will not affect the subsequent voice signal. Therefore, the purpose of waiting for the second time interval is to ensure that the relative position of the device relative to the voice source changes significantly. This significant change can be achieved by setting a threshold. In terms of design threshold, factors such as moving distance and / or moving direction can be considered. For example, when the robot moves far enough along a certain direction, the relative position of the voice source will definitely change significantly. In addition, although the robot does not move far, its moving direction or its own rotation angle changes significantly, which will also cause the relative position of the voice source relative to the microphone array on the device to change significantly. In practical applications, as a relatively practical strategy, the moving distance factor can be emphasized. The moving distance is proportional to the moving speed. Therefore, the second time period can be dynamically set based on the moving speed. The length of the second time period can be set to be inversely proportional to the moving speed, so that the second time period can meet the moving distance requirements of the device. The specific ratio coefficient can be determined according to the actual needs, and a trade-off is made between the update efficiency of the beamformer and the voice enhancement effect. Further, in this step, the weight vector of the beamformer is updated according to the preset steering vector and the noise covariance matrix as a noise feature. The update of the beamformer is mainly the update of its weight vector. Specifically, the weight vector can be updated using the following formula (2):
[0055]
[0056] Where w(τ) is the weight vector and a is the steering vector. In this embodiment of the invention, a fixed steering vector can be used to update the weight vector. Specifically, the value of the steering vector a can be: a = [1,…,1] T / M, where M is the number of microphone units in the microphone array. In this embodiment of the invention, the use of a fixed steering vector avoids the problem of steering vector estimation, improves the processing speed of speech enhancement, and reduces the computational resources used for speech enhancement processing.
[0057] S103: Use the updated beamformer to perform speech enhancement processing on the subsequent microphone signals. After the beamformer is updated, spatial filtering will continue to be performed using the weight vector determined in step S102 until it is updated again.
[0058] Furthermore, although noise features change less compared to speech features, the environment and operating status of the equipment are constantly changing. For example, the motor noise changes when the robot vacuum's speed changes, and the suction power of the vacuum cleaner may also change. Additionally, the noise generated by the robot vacuum will differ depending on the type of surface it traverses, such as carpet or tile. Therefore, noise feature extraction and beamformer updates can be performed continuously; that is, steps S101 to S103 can be continuously cyclically executed to effectively enhance speech. Therefore, the method can also include: entering a new noise feature calibration cycle after a third time interval. Steps S101 and S102 can be considered as one noise feature calibration cycle. After this cycle, new noise features are collected and the beamformer parameters are updated. A certain amount of time can be allowed between different noise feature calibration cycles, such as... Figure 1 and Figure 2 The time from t1 to t2 is shown in the figure.
[0059] Furthermore, the noise feature extraction and beamformer update processes described above can be performed band-wise. For different frequency bands of audio signals, the beamformer can employ different weight vectors. In step S101, extracting noise features based on the microphone signal can include: extracting noise features from multiple frequency bands based on the microphone signal; specifically, generating noise covariance matrices corresponding to multiple frequency bands. Correspondingly, in step S102, updating the beamformer based on the noise features can include: updating the beamformer corresponding to the multiple frequency bands based on the noise features of those multiple frequency bands; specifically, updating the weight vectors of the beamformers corresponding to those multiple frequency bands.
[0060] Furthermore, in this embodiment of the invention, the main computing power is concentrated in the matrix inversion part of the above formula (2). The number of frequency bands to be updated determines the number of times formula (2) will be executed. Since the noise characteristics are relatively stable or slowly changing in the scenario of this embodiment of the invention, in order to reduce the burden on computing power caused by the beamformer update process, noise characteristics of some frequency bands can be obtained from the noise characteristics of multiple frequency bands during the time period from t4 to t5, and then the beamformer corresponding to that part of the frequency band is updated. For example, the weight vector of the beamformer of one frequency band can be updated only once during the time period from t4 to t5. When the next round of execution reaches the time period from t4 to t5, the weight vector of the beamformer of another frequency band is updated. This can significantly reduce the computing power requirements of the robot vacuum cleaner and is well applicable to embedded systems with low computing resources.
[0061] This invention provides a speech enhancement processing method that leverages the characteristic that the noise signal characteristics of a device itself change relatively little during device movement, while the characteristics of external speech signals change significantly due to changes in the location of the sound source. By setting a time interval between noise feature acquisition and beamformer updates, useful speech components are avoided from being eliminated, thereby improving the performance of speech enhancement. This speech enhancement processing method can effectively suppress noise emitted by the device itself under conditions of extremely low signal-to-noise ratio, multiple interference sources, and moving sound sources, achieving effective speech enhancement.
[0062] Example 2
[0063] like Figure 4 The diagram shown is a structural schematic of the voice enhancement processing device according to an embodiment of the present invention. This device can be applied to mobile devices requiring voice interaction, or it can be applied to a server. The server obtains microphone signals from the device via a network and performs voice enhancement processing. The specific processing procedure is as follows:
[0064] The noise feature extraction module 11 is used to acquire microphone signals during a first time period and extract noise features based on the microphone signals. The first time period can correspond to... Figure 1In the time period from t2 to t3, the microphone signal is treated as a noise signal for feature extraction. The extracted noise feature can be specifically a noise covariance matrix. Specifically, this part of the processing can further include: acquiring microphone signals from multiple consecutive data frames; and calculating the noise covariance matrix corresponding to the current data frame based on a preset forgetting factor, according to the noise covariance matrix corresponding to the previous data frame and the microphone signal of the current data frame. The noise covariance matrix can be specifically calculated using the above formula (1). In the first time period, the noise covariance matrix corresponding to the last data frame is used as a noise feature for the beamformer update module 12 to update the beamformer.
[0065] The beamformer update module 12 is used to update the beamformer based on noise characteristics after a second time interval. The second time interval can correspond to... Figure 1 The time period from t3 to t4 is specified. The length of the t3 to t4 time period can be a preset value or dynamically set according to the actual situation. Specifically, the motion data of the device where the microphone is located can be detected, and the length of the second time period can be determined based on the motion data. Taking a robot vacuum cleaner as an example, the moving speed of the robot vacuum cleaner can be detected, and the length of the t3 to t4 time period can be set inversely proportional to the moving speed.
[0066] Furthermore, the processing of the beamformer update module 12 can specifically include: updating the weight vector of the beamformer based on the preset steering vector and the noise covariance matrix as a noise feature. The update of the beamformer mainly involves updating its weight vector, which can be done using the formula (2) above. A fixed steering vector can be used to update the weight vector, thereby avoiding the estimation problem of the steering vector and improving the processing speed of speech enhancement.
[0067] The speech enhancement processing module 13 is used to perform speech enhancement processing on subsequent microphone signals using the updated beamformer. After the beamformer is updated, spatial filtering will be performed using the already determined weight vector until it is updated again.
[0068] The specific details of the above processing procedure, the detailed explanation of the technical principles, and the detailed analysis of the technical effects have been described in the previous embodiments, and will not be repeated here.
[0069] This invention provides a speech enhancement processing device that utilizes the characteristic that the noise signal characteristics of the device itself change relatively little during device movement, while the external speech signal characteristics change significantly due to changes in the location of the sound source. By setting a time interval between noise feature acquisition and beamformer updates, useful speech components are avoided from being eliminated, thereby improving the performance of speech enhancement. The speech enhancement processing method can effectively suppress noise emitted by the device itself under conditions of extremely low signal-to-noise ratio, multiple interference sources, and moving sound sources, achieving effective speech enhancement.
[0070] Example 4
[0071] This invention provides a voice interaction method that can be applied to mobile devices requiring voice interaction or to a server. The server acquires microphone signals from the device via a network and returns the processing results to the device. Specifically, the method includes: S201: Acquiring sound signals. Specifically, sound signals can be acquired using a microphone array on the device.
[0072] S202: Perform speech enhancement processing on the sound signal. During the speech enhancement processing, noise features are extracted, and the beamformer used for speech enhancement processing is updated based on the noise features after a certain time interval. The speech enhancement processing in this step can employ the specific processing methods mentioned in the previous embodiments.
[0073] S203: The voice signal, after voice enhancement processing, is used to recognize voice commands and execute corresponding processing based on the recognized voice. The voice commands vary depending on the specific device; for example, for a drone, these could be commands to control its flight or take photos, while for a robotic vacuum cleaner, they could be commands to move, turn on the vacuum cleaner, or spray water.
[0074] By using the voice enhancement processing technology in the aforementioned embodiments during the above voice interaction process, the noise emitted by the device itself can be effectively suppressed under conditions of extremely low signal-to-noise ratio, multiple interference sources, and moving sound sources, thereby achieving effective voice enhancement. This enables the device to more accurately recognize the voice commands issued by the user and complete the corresponding processing actions.
[0075] Furthermore, embodiments of the present invention also provide a voice interaction device, including:
[0076] A sound signal acquisition module is used to acquire sound signals.
[0077] The speech enhancement processing module is used to perform speech enhancement processing on the sound signal. During the speech enhancement processing, noise features are extracted, and the beamformer used for speech enhancement processing is updated according to the noise features after a certain period of time.
[0078] The instruction recognition and processing module is used to recognize voice instructions from the voice-enhanced audio signal and perform corresponding processing based on the recognized voice.
[0079] The aforementioned voice interaction devices are installed on drones or robotic vacuum cleaners to enable voice interaction between users and the drones or robotic vacuum cleaners with voice enhancement, thereby suppressing the noise emitted by the devices themselves and improving the accuracy of voice command recognition.
[0080] Example 5
[0081] The preceding embodiments described the process flow and apparatus structure of the speech enhancement processing method and the speech interaction method. The functions of the above methods and apparatus can be implemented using an electronic device, such as... Figure 5 As shown, it is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention, specifically including: a memory 110 and a processor 120.
[0082] Memory 110 is used to store programs.
[0083] In addition to the procedures described above, memory 110 can also be configured to store various other data to support operation on the electronic device. Examples of such data include instructions for any application or method used to operate on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0084] The memory 110 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0085] The processor 120, coupled to the memory 110, is used to execute the program in the memory 110 to perform the operation steps of the voice enhancement processing method and / or voice interaction method described in the foregoing embodiments.
[0086] In addition, the processor 120 may also include the various modules described in the foregoing embodiments to perform voice enhancement and / or voice interaction processing, and the memory 110 may be used, for example, to store data required for these modules to perform operations and / or output data.
[0087] The specific details of the above processing procedure, the detailed explanation of the technical principles, and the detailed analysis of the technical effects have been described in the previous embodiments, and will not be repeated here.
[0088] Furthermore, as shown in the figure, the electronic device may also include other components such as a communication component 130, a power supply component 140, an audio component 150, and a display 160. The figure only schematically shows some components and does not imply that the electronic device includes only the components shown in the figure.
[0089] Communication component 130 is configured to facilitate wired or wireless communication between electronic devices and other devices. The electronic devices can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, and other mobile communication networks, or combinations thereof. In one exemplary embodiment, communication component 130 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 130 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0090] Power supply component 140 provides power to various components of an electronic device. Power supply component 140 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.
[0091] Audio component 150 is configured to output and / or input audio signals. For example, audio component 150 includes a microphone (MIC) configured to receive external audio signals when the electronic device is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 110 or transmitted via communication component 130. In some embodiments, audio component 150 also includes a speaker for outputting audio signals.
[0092] Display 160 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0093] Furthermore, embodiments of the present invention also provide a computer program product, including a computer program or instructions, which, when executed by a processor, cause the processor to implement the aforementioned speech enhancement processing method and / or speech interaction method.
[0094] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for speech enhancement processing, comprising: acquiring microphone signals in a first time period, and extracting noise features from the microphone signals; a noise source and a microphone have a fixed relative position; the microphone is arranged on a mobile device; after a second time period, updating a beamformer weight vector according to the noise features; the beamformer is used to adjust the proportion of signals received by each microphone channel by weighting factors and add them together, and the weighting factors exist in the form of the weight vector; wherein the relative position of the mobile device to a speech source after the second time period is different from the relative position of the mobile device to the speech source in the first time period, and the difference reaches a set threshold; using the updated beamformer to perform speech enhancement processing on subsequent microphone signals.
2. The method of claim 1, wherein, The noise features include a noise covariance matrix, and the acquiring microphone signals and extracting noise features from the microphone signals include: acquiring microphone signals of multiple continuous data frames; based on a preset forgetting factor, calculating a noise covariance matrix corresponding to a current data frame according to a noise covariance matrix corresponding to a previous data frame and the microphone signals of the current data frame; taking the noise covariance matrix corresponding to the last data frame in the first time period as the noise features.
3. The method of claim 1, wherein, The noise features include a noise covariance matrix, and updating the beamformer using the noise features includes: updating the weight vector of the beamformer according to a preset steering vector and the noise covariance matrix as the noise features.
4. The method of claim 3, wherein, The steering vector is a constant value.
5. The method of claim 1, wherein, Further comprising: after a third time period, entering a new noise feature calibration period.
6. The method of claim 1, wherein, The extracting noise features from the microphone signals includes extracting noise features of multiple frequency bands from the microphone signals; The updating the beamformer according to the noise features includes: updating the beamformer corresponding to the multiple frequency bands according to the noise features of the multiple frequency bands; or, obtaining noise features of part of the multiple frequency bands, and updating the beamformer corresponding to the part of the multiple frequency bands.
7. The method of claim 1, wherein, Further comprising: detecting motion data of a device where the microphone is arranged, and determining the length of the second time period according to the motion data. 8.A speech enhancement processing apparatus, comprising: a noise feature extraction module configured to acquire microphone signals in a first time period, and extract noise features from the microphone signals; a noise source and a microphone have a fixed relative position; the microphone is arranged on a mobile device; a beamformer updating module configured to update a beamformer weight vector according to the noise features after a second time period; the beamformer is used to adjust the proportion of signals received by each microphone channel by weighting factors and add them together, and the weighting factors exist in the form of the weight vector; wherein the relative position of the mobile device to a speech source after the second time period is different from the relative position of the mobile device to the speech source in the first time period, and the difference reaches a set threshold. The voice enhancement processing module is configured to perform voice enhancement processing on the subsequent microphone signals using the updated beamformer. 9.An electronic device comprising: a memory configured to store a program; a processor configured to execute the program stored in the memory to perform the voice enhancement processing method according to any one of claims 1 to 7.
10. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions, when executed by a processor, cause the processor to implement the voice enhancement processing method according to any one of claims 1 to 7. 11.A voice interaction method comprising: acquiring a sound signal; performing voice enhancement processing on the sound signal, wherein during the voice enhancement processing, noise features are extracted, and after a period of time, a beamformer weight vector used for voice enhancement processing is updated according to the noise features; the beamformer is configured to adjust the proportions of signals received by each microphone channel by assigning weight factors and then summing them up, and the weight factors exist in the form of the weight vector; the relative position between a noise sound source and a microphone is fixed; the microphone is arranged on a mobile device; wherein the relative position of the mobile device with respect to the noise sound source after the period of time is different from the relative position of the mobile device with respect to the noise sound source before the period of time, and the difference reaches a set threshold; performing voice command recognition on the sound signal after the voice enhancement processing, and performing corresponding processing according to the recognized voice. 12.A voice interaction device comprising: a sound signal acquisition module configured to acquire a sound signal; a voice enhancement processing module configured to perform voice enhancement processing on the sound signal, wherein during the voice enhancement processing, noise features are extracted, and after a period of time, a beamformer weight vector used for voice enhancement processing is updated according to the noise features; the beamformer is configured to adjust the proportions of signals received by each microphone channel by assigning weight factors and then summing them up, and the weight factors exist in the form of the weight vector; the relative position between a noise sound source and a microphone is fixed; the microphone is arranged on a mobile device; wherein the relative position of the mobile device with respect to the noise sound source after the period of time is different from the relative position of the mobile device with respect to the noise sound source before the period of time, and the difference reaches a set threshold; an instruction recognition and processing module configured to perform voice command recognition on the sound signal after the voice enhancement processing, and perform corresponding processing according to the recognized voice.
13. A drone, wherein, The unmanned aerial vehicle comprises the voice interaction device according to claim 12.
14. A robotic vacuum cleaner, wherein, The sweeping robot comprises the voice interaction device according to claim 12. 15.An electronic device comprising: a memory configured to store a program; a processor configured to execute the program stored in the memory to perform the voice interaction method according to claim 11.
16. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions, when executed by a processor, cause the processor to implement the voice interaction method according to claim 11.
Citation Information
Patent Citations
Microphone array signal processing system
CN108464015A
Voice pickup method and device, storage medium and mobile robot
CN110428850A