Voice processing method, audio-video communication device, and vehicle
By combining deep learning models with beamforming and notch filtering algorithms, the problem of voice quality degradation caused by noise and interference in AIoT voice interaction is solved, and clear pickup and recognition of voice information within the effective interaction area is achieved.
Patent Information
- Application Number
- CN202211020210.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-08-24
AI Technical Summary
In AIoT voice interaction scenarios, noise and interference from non-target sound sources cause a decline in voice quality within the effective interaction area, making it difficult to pick up clear voice information.
A deep learning model is used in conjunction with beamforming and notch filtering algorithms to enhance speech signals within and outside the effective interaction area, respectively. The deep learning model is then used to recover the signal and generate the target speech.
It effectively suppresses sound source interference and environmental noise from non-target directions, improves the extraction effect of voice information within the effective interaction area, and enhances voice recognition and interactive experience.
Smart Images

Figure CN115482829B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and more specifically, to a voice processing method, an audio and video communication device, and a vehicle. Background Technology
[0002] Currently, in the voice interaction scenario of AIoT (AI+IoT, Artificial Intelligence of Things), microphone arrays are used to pick up the voice of the target speaker and provide it to the subsequent speech recognition model for recognition. However, in the voice interaction environment, there is usually noise and interference from non-target sound sources, which will reduce the voice quality of the sound source in the effective interaction area and increase the difficulty of picking up voice information.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides a voice processing method, an audio and video communication device, and a vehicle to at least solve the technical problem in the related art of difficulty in picking up voice information of sound sources within an effective interaction area.
[0005] According to one aspect of the embodiments of this application, a speech processing method is provided, comprising: acquiring an original speech set collected by a sound pickup device, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within an effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, wherein the sound source located within the effective interaction area is a speech interaction object identified directionally by the sound pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; and using a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate target speech, wherein the target speech is speech information directionally picked up by the sound pickup device.
[0006] According to one aspect of the embodiments of this application, a speech processing method is provided, comprising: capturing an original speech set collected by a sound pickup device disposed on an audio-visual communication device, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within an effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is a speech interaction object identified directionally by the sound pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; using a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate target speech; and controlling the audio-visual communication device to output the target speech.
[0007] According to one aspect of the embodiments of this application, a voice processing method is provided, comprising: capturing an original voice set collected by a sound pickup device installed on a target vehicle, wherein the original voice set includes a first voice and a second voice, wherein the first voice is a voice signal emitted from a sound source located within an effective interaction area, and the second voice is a voice signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is a voice interaction object identified directionally by the sound pickup device; performing enhancement processing on the first voice and the second voice respectively to obtain enhanced first voice and second voice; using a deep learning model to perform signal recovery processing on the original voice set, the enhanced first voice, and the second voice to generate target voice; and controlling the target vehicle based on the target voice.
[0008] According to one aspect of the embodiments of this application, a voice processing method is provided, comprising: a cloud server receiving an original voice set uploaded by a client, wherein the original voice set is acquired by a sound pickup device, the original voice set includes a first voice and a second voice, wherein the first voice is a voice signal emitted from a sound source located within an effective interaction area, and the second voice is a voice signal emitted from a sound source other than the sound source located within the effective interaction area, the sound source located within the effective interaction area being a voice interaction object identified directionally by the sound pickup device; the cloud server performing enhancement processing on the first voice and the second voice respectively to obtain enhanced first voice and second voice; the cloud server using a deep learning model to perform signal recovery processing on the original voice set, the enhanced first voice, and the second voice to generate target voice, wherein the target voice is voice information directionally picked up by the sound pickup device; and the cloud server outputting the target voice to the client.
[0009] According to one aspect of the embodiments of this application, a speech processing system is provided, comprising: a sound pickup device for acquiring an original speech set, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within an effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is a speech interaction object identified directionally by the sound pickup device; and a processing device connected to the sound pickup device for enhancing the first speech and the second speech respectively to obtain enhanced first speech and second speech, and using a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate target speech, wherein the target speech is speech information directionally picked up by the sound pickup device.
[0010] According to one aspect of the embodiments of this application, an audio-visual communication device is provided, comprising: a sound pickup device disposed on the audio-visual communication device, used to acquire an original speech set, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within an effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is a voice interaction object identified directionally by the sound pickup device; a processor, connected to the sound pickup device, used to perform enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech, and to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech using a deep learning model to generate target speech; and an output device, connected to the processor, used to output the target speech.
[0011] According to one aspect of the embodiments of this application, a vehicle is provided, comprising: a sound pickup device disposed on the vehicle, configured to acquire an original speech set, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within an effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is a voice interaction object identified directionally by the sound pickup device; a controller connected to the sound pickup device, configured to perform enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech, perform signal recovery processing on the original speech set, the enhanced first speech and the second speech using a deep learning model to generate target speech, and control the target vehicle based on the target speech.
[0012] According to one aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is running, it controls the device where the storage medium is located to execute the speech processing method of any one of the above.
[0013] In this embodiment, the original speech set collected by the sound pickup device is first obtained. The original speech set includes a first speech and a second speech. The first speech is speech information emitted by a sound source located within the effective interaction area, and the second speech is speech information emitted by a sound source other than the sound source within the effective interaction area. The sound source within the effective interaction area is the speech interaction object identified directionally by the sound pickup device. The first speech and the second speech are enhanced respectively to obtain enhanced first speech and second speech. A deep learning model is used to recover the speech signal from the original speech set, the enhanced first speech and the second speech to generate the target speech. The target speech is the speech information directionally picked up by the sound pickup device, which effectively suppresses sound source interference and environmental noise outside the effective interaction area and improves the extraction effect of speech information within the effective interaction area. It is noteworthy that the first speech emitted by the sound source within the effective interaction area and the second speech emitted by other sound sources outside the effective interaction area can be enhanced separately. Combined with a deep learning model, the second speech emitted by other sound sources can be effectively suppressed, so that the sound pickup device can pick up the speech information within the effective interaction area in a directional manner. This solves the problem of speech information technology in related technologies that is difficult to pick up sound sources within the effective interaction area. Attached Figure Description
[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0015] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a voice processing method according to an embodiment of this application;
[0016] Figure 2 This is a flowchart of the speech processing method according to Embodiment 1 of this application;
[0017] Figure 3 This is a top view of a simulated room environment according to an embodiment of this application;
[0018] Figure 4 This is a schematic diagram of a user interface for training a deep learning model according to an embodiment of this application;
[0019] Figure 5 This is a schematic diagram of a deep learning model according to an embodiment of this application;
[0020] Figure 6 This is a structural block diagram of a voice processing flow according to an embodiment of this application;
[0021] Figure 7This is a flowchart of a speech processing method according to Embodiment 2 of this application;
[0022] Figure 8 This is a flowchart of a speech processing method according to Embodiment 3 of this application;
[0023] Figure 9 This is a flowchart of a speech processing method according to Embodiment 4 of this application;
[0024] Figure 10 This is a schematic diagram of a speech processing system according to Embodiment 5 of this application;
[0025] Figure 11 This is a schematic diagram of an audio / video communication device according to Embodiment 6 of this application;
[0026] Figure 12 This is a schematic diagram of a vehicle according to Embodiment 7 of this application;
[0027] Figure 13 This is a schematic diagram of a voice processing device according to Embodiment 8 of this application;
[0028] Figure 14 This is a schematic diagram of a voice processing device according to Embodiment 9 of this application;
[0029] Figure 15 This is a schematic diagram of a voice processing device according to Embodiment 10 of this application;
[0030] Figure 16 This is a schematic diagram of a voice processing device according to Embodiment 11 of this application;
[0031] Figure 17 This is a structural block diagram of a computer terminal according to Embodiment 12 of this application. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0035] Superdirective beamforming, also known as spatial filtering, is a processing technique that uses a microphone array to receive signals in a directional manner.
[0036] Notch filter (Null beamformer): Primarily used to filter out signals at a specific frequency.
[0037] RIR: Room Impulse Respose, describes the room transfer function from the sound source location to the microphone location.
[0038] Currently, array algorithms such as beamforming or blind source separation are generally used to improve the speech quality of sound sources within the effective interaction area. However, this method based on classical signal processing has limited performance, especially in scenarios with a small number of microphone arrays and many interfering sound sources.
[0039] To address the aforementioned issues, this application incorporates a directional beamforming algorithm based on a deep learning model, which can effectively suppress sound source interference and environmental noise from non-target directions, facilitating the acquisition of speech information from the target direction and thus ensuring a superior voice interaction experience in the target direction.
[0040] Example 1
[0041] According to an embodiment of this application, a voice processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0042] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a voice processing method according to an embodiment of this application. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0043] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). This data processing circuitry serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0044] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the voice processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned voice processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0045] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0046] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0047] Under the aforementioned operating environment, this application provides the following: Figure 2 The speech processing method shown is illustrated. It should be noted that the speech processing method in this embodiment can be derived from... Figure 1 The computer terminal of the illustrated embodiment is used for execution. Figure 2 This is a flowchart of the speech processing method according to Embodiment 1 of this application. Figure 2 As shown, the method may include the following steps:
[0048] Step S202: Obtain the original speech set collected by the sound pickup device.
[0049] The original speech set includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the voice interaction object that the pickup device identifies in a directional manner.
[0050] The aforementioned sound pickup device can be a microphone, which is an accessory used to collect ambient sound. It is an electroacoustic instrument that amplifies sound by receiving sound vibrations. A microphone can consist of a microphone and an audio amplification circuit, and can generally be divided into digital microphones and analog microphones. A digital microphone is a sound sensing device that uses a digital signal processing system to convert analog audio signals into digital signals and perform corresponding digital signal processing; an analog microphone generally amplifies the collected sound using a microphone.
[0051] The aforementioned sound pickup device can also be one or more microphones, and multiple microphones can be a microphone array. The microphone array can be a microphone array composed of several directional microphones arranged linearly. It can be a uniform linear array or a non-uniform linear array. The specific type can be determined according to actual needs.
[0052] The aforementioned voice interaction object can be a sound source within the effective interaction area. This sound source can be the voice of one or more users, a smart speaker, a television, or even a pet. In a driving environment, the sound source can also be the voice of a passenger.
[0053] The aforementioned effective interaction area can be the area within which the microphone can recognize speech. This effective interaction area can be a pre-defined area based on the actual scenario. Optionally, any interaction area can be selected as the effective interaction area, or an interaction area closer to the microphone can be selected. The effective interaction area can be chosen based on actual needs.
[0054] The first speech mentioned above can be speech information emitted by a sound source located within the effective interaction area, and the second speech mentioned above can be speech signals emitted by sound sources other than those within the effective interaction area. These other interaction areas can be considered invalid interaction areas. It should be noted that when the sound pickup device collects the original speech set, interference from the second speech in invalid interaction areas can lead to lower quality of the first speech, thus reducing the effectiveness of speech recognition and worsening the speech interaction experience. Therefore, after collecting the original speech set, it is necessary to process the first and second speech in the original speech set to improve the quality of the first speech and thus improve the speech recognition effect.
[0055] In one optional embodiment, taking a video call scenario as an example, the area near the terminal microphone can be the effective interaction area, and the area away from the terminal microphone can be the ineffective interaction area. The voice signal emitted by the sound source near the terminal microphone can be the first voice, and the voice signal emitted by other sound sources away from the terminal microphone can be the second voice.
[0056] In one optional embodiment, the original speech set collected by the sound pickup device can be obtained, and the speech signals belonging to the first speech and the speech signals belonging to the second speech in the original speech set can be determined according to the pre-set effective interaction area. This is to facilitate the differentiation between the speech signals emitted by the sound source within the effective interaction area and the speech signals emitted by other sound sources outside the effective interaction area, so as to facilitate the subsequent extraction of the speech information picked up directionally by the sound pickup device.
[0057] Step S204: Enhance the first speech and the second speech respectively to obtain the enhanced first speech and the second speech.
[0058] In one optional embodiment, a beamforming algorithm can be used to enhance the speech signal within the effective interaction area to obtain the enhanced first speech. Optionally, a Fourier transform can be performed on the original speech set, and the filter coefficients of the beam can be solved using a convex optimization tool (CVX) in each frequency band based on the microphone array topology and the first speech within the effective interaction area. After filtering the original speech set, the processed frequency domain signal, i.e., the enhanced first speech, can be obtained.
[0059] In another alternative embodiment, a notch filter algorithm can be used to suppress the speech signal within the effective interaction area, which is equivalent to enhancing the speech signal emitted by other sound sources besides the sound source in the effective interaction area. That is, an enhanced second speech can be obtained. Optionally, a Fourier transform can be performed on the original speech set, and the filter coefficients of the notch filter can be solved using a convex optimization tool in each frequency band according to the microphone array topology and the first speech within the effective interaction area. After filtering the original speech set, the processed frequency domain signal is obtained, which is the enhanced second speech.
[0060] In another alternative embodiment, enhancing the first speech and the second speech respectively can make the speech signal emitted by the sound source in the effective interaction area more obvious, and the speech signal emitted by the sound source in the ineffective interaction area also more obvious. The comparison between the two speech signals is also more obvious, which facilitates the subsequent directional picking of speech information in the effective interaction area.
[0061] Step S206: Use a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech, to generate the target speech.
[0062] The target speech is the speech information picked up directionally by the sound pickup device.
[0063] The aforementioned deep learning model can be a pre-trained model capable of recognizing speech emitted by sound sources within an effective interaction region. This deep learning model can be a model that stacks deep feedforward sequential memory networks (DFSMN) and linear layers. It can also be a long short-term memory artificial neural network (LSTM), a convolutional neural network (CNN), or any other neural network model; no limitations are imposed here.
[0064] The model structure of the deep learning model described above can be an input layer, a hidden layer, a linear mapping layer, a sequence memory module, and an output layer. The model structure of the deep learning model can be adjusted according to actual needs.
[0065] In one optional embodiment, features of the original speech set, the enhanced first speech, and the second speech can be extracted firstly. These features are then input into a deep learning model, which performs localization processing on the features to obtain a time-frequency mask for the target speech. This mask is then used to perform signal recovery processing on the enhanced first speech. During the signal recovery process, interference signals in the enhanced first speech can be masked, thereby obtaining a high-quality target speech.
[0066] Time-frequency masking can be divided into time-domain masking and frequency-domain masking. The enhanced first speech can be masked by frequency-domain masking to cover the enhanced second speech occurring simultaneously nearby. Conversely, the enhanced first speech can be masked by time-domain masking to cover the enhanced second speech that is temporally adjacent to it. In other words, the enhanced first speech can mask the enhanced second speech through time-frequency masking to obtain the targeted speech information after enhancement.
[0067] In another alternative embodiment, the information of the effective interaction area can be used as the input of the deep learning model, and the output of the model can be adjusted as the effective interaction area changes, so that the array beam can be dynamically pointed to a non-passing direction.
[0068] Taking a telephone call as an example, in a telephone call scenario, the voice interaction area within a preset range of the terminal microphone can be determined as the effective interaction area, while other areas far from this preset range are considered invalid interaction areas. The first voice can be the voice of the user making the call, and the second voice can be the voice of another speaker or noise from other areas. The original voice set collected by the microphone can be acquired first, and the first voice within the preset range in the original voice set can be enhanced, as can the second voice in other ranges. A deep learning model combined with spatial information can be used to perform signal recovery processing on the original voice set, the enhanced first voice, and the second voice to obtain higher quality voice within the effective interaction area, thereby improving the user's voice recognition results and thus improving the quality of telephone calls.
[0069] Taking a smart speaker as an example, in the interaction scenario of a smart speaker, the voice interaction area within a preset range of the smart speaker's voice acquisition device can be identified as the effective interaction area, while other areas far from this preset range are considered invalid interaction areas. The first voice can be the user's voice used to issue interaction commands to the smart speaker, and the second voice can be noise or the speaker's voice from other areas. The original voice set collected by the voice acquisition device can be obtained first, and the first voice within the preset range in the original voice set can be enhanced, as can the second voice in other ranges. A deep learning model combined with spatial information can be used to perform signal recovery processing on the original voice set, the enhanced first voice, and the second voice, obtaining high-quality voice within the effective interaction area. This can improve the voice recognition results for user-issued interaction commands, thereby improving the voice interaction experience of the smart speaker.
[0070] Taking vehicle control as an example, in the interactive scenario of vehicle control, the voice interaction area within a preset range of the voice acquisition device in the vehicle can be identified as the effective interaction area, while other areas far from this preset range are considered invalid interaction areas. The first voice can be the user's voice issuing an interactive command to the vehicle, and the second voice can be noise or the voice of another speaker from other areas, such as the horn of another vehicle. The original voice set collected by the voice acquisition device can be obtained first, and the first voice within the preset range in the original voice set can be enhanced, as can the second voice in other ranges. A deep learning model combined with spatial information can be used to perform signal recovery processing on the original voice set, the enhanced first voice, and the second voice to obtain higher quality voice within the effective interaction area. This can improve the voice recognition results of the user's interactive commands, thereby improving the accuracy of voice control of the vehicle.
[0071] Through the above steps, the original speech set collected by the sound pickup device is first obtained. The original speech set includes a first speech and a second speech. The first speech is speech information emitted by a sound source located within the effective interaction area, and the second speech is speech information emitted by a sound source other than the sound source within the effective interaction area. The sound source within the effective interaction area is the speech interaction object identified directionally by the sound pickup device. The first speech and the second speech are enhanced respectively to obtain enhanced first speech and second speech. A deep learning model is used to recover the speech signals from the original speech set, the enhanced first speech and the second speech to generate the target speech. The target speech is the speech information directionally picked up by the sound pickup device. This effectively suppresses sound source interference and environmental noise outside the effective interaction area and improves the extraction effect of speech information within the effective interaction area. It is noteworthy that the first speech emitted by the sound source within the effective interaction area and the second speech emitted by other sound sources outside the effective interaction area can be enhanced separately. Combined with a deep learning model, the second speech emitted by other sound sources can be effectively suppressed, so that the sound pickup device can pick up the speech information within the effective interaction area in a directional manner. This solves the problem of speech information technology in related technologies that is difficult to pick up sound sources within the effective interaction area.
[0072] In the above embodiments of this application, enhancing the first speech and the second speech respectively to obtain the enhanced first speech and the second speech includes: superimposing the original speech set using a beamforming algorithm to obtain the enhanced first speech; and filtering the original speech set using a notch filter algorithm to obtain the enhanced second speech.
[0073] The beamforming algorithms mentioned above can be divided into adaptive algorithms based on direction estimation, such as super pointing beamforming algorithms, but are not limited to this.
[0074] In one optional embodiment, the beamforming algorithm and the notch filtering algorithm can be applied in the target direction of the effective interaction area. The effective interaction area can be set to a region approximately 15° to the left and right of 0° directly in front of the pickup device, with the center angle being the target direction. In other words, the directions of beamforming and notch filtering are both 0°. This is only an example and the target direction and effective interaction area can be set according to the actual situation.
[0075] In another alternative embodiment, the original speech set can be subjected to Fourier transform using a beamforming algorithm. In each frequency band, the filter coefficients of the first speech are solved using a convex optimization tool based on the microphone array topology and the direction of the effective interaction area. After filtering the first speech, the array outputs in the target direction can be superimposed in phase to obtain the enhanced first speech.
[0076] In another alternative embodiment, the original speech set can be subjected to Fourier transform using a notch filter algorithm. In each frequency band, the filter coefficients of the first speech are solved using a convex optimization tool based on the microphone array topology and the direction of the effective interaction area. The notch filter is then used to block the first speech in the original speech set, thereby enhancing the second speech in the ineffective interaction area and obtaining the enhanced second speech.
[0077] In the above embodiments of this application, a deep learning model is used to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate target speech. This includes: extracting features from the original speech set, the enhanced first speech, and the second speech respectively to obtain original speech features, first speech features, and second speech features; inputting the original speech features, first speech features, and second speech features into the deep learning model to obtain a time-frequency mask for the target speech; and performing signal recovery processing on the enhanced first speech based on the time-frequency mask to obtain the target speech.
[0078] The original speech features mentioned above include a first speech feature without enhancement and a second speech feature without enhancement.
[0079] The aforementioned time-frequency masking can be phase-sensive masking (PSM). In this case, the time-frequency masking can be applied in the direction of the first speech to perform signal recovery processing on the enhanced first speech, thereby obtaining the target speech.
[0080] In one optional embodiment, the Fbank dimension features (audio dimension features) can be extracted by performing short-time Fourier transform on the original speech set, the enhanced first speech, and the second speech respectively, and the mean and variance of the obtained Fbank dimension features can be normalized to obtain the original speech features, the first speech features, and the second speech features. The Fbank dimension features can be Fbank80 dimension features.
[0081] In another optional embodiment, the original speech features, the first speech features, and the second speech features can be input into a deep learning model to obtain a time-frequency masking that can mask other interfering speech besides the target speech. The time-frequency masking can be applied to the enhanced first speech signal to mask other interfering signals in the enhanced first speech signal, resulting in a high-quality target speech. This allows the target speech to be applied in voice interaction scenarios to improve the accuracy of speech recognition.
[0082] In another alternative embodiment, the deep learning model can also utilize various other features that can represent spatial information of the sound source to obtain time-frequency masking of the target speech, such as interchannel level difference or interchannel phase difference. That is, features of interchannel level difference or interchannel phase difference can be extracted from the original speech set, the enhanced first speech, and the second speech to obtain the original speech features, the first speech features, and the second speech features.
[0083] In the above embodiments of this application, the method further includes: constructing multiple simulated scenarios corresponding to the sound pickup device, wherein the number and type of sound sources contained in different simulated scenarios are different, and the types of sound sources include at least one of the following: target sound source, interference sound source and noise sound source, wherein the target sound source is located in the effective interaction area, and the interference sound source and noise sound source are located in other interaction areas other than the effective interaction area; generating a set of simulated speech corresponding to multiple simulated scenarios; and using the set of simulated speech to train a deep learning model.
[0084] The aforementioned simulation scenarios can be rooms of random sizes.
[0085] The simulation scenario described above can contain one or more sound sources. The types of sound sources in the simulation scenario can be target sound sources, interference sound sources, and noise sound sources. Target sound sources can be sound sources within the effective interaction area, interference sound sources can be sound sources within the ineffective interaction area, and noise sound sources can be random, irregular sound sources within the ineffective interaction area.
[0086] In one optional embodiment, the microphone array can be randomly positioned in the simulated scene, keeping the sound source directly in front of the array. Interference or noise sources can be randomly placed in designated invalid interaction areas, maintaining a certain proportion of random occurrences of the target sound source, interference sources, and noise sources. Optionally, various real-world scenarios involving sound sources can be simulated, such as target sound source + interference source, interference source + noise source, target sound source + noise source, and scenarios where each sound source exists independently.
[0087] Figure 3 This is a top view of a simulated room environment according to an embodiment of this application. For example... Figure 3The diagram shows a microphone array placed in a simulated room. The area where the microphone array collects sound can be divided into an effective interaction area and an ineffective interaction area. One sound source can be placed in the effective interaction area, and two sound sources can be placed in the ineffective interaction area. This simulates the collection of simulated speech from these sound sources. The simulated speech set corresponding to the simulated scene can be generated by changing the number and type of sound sources in the simulated room environment. It can also be generated by changing the size of the simulated room environment and the placement of the microphone array. This simulated speech set can be used to train a deep learning model.
[0088] In another alternative embodiment, a simulated speech set and a simulated scene can be used as training samples. The simulated speech set can be input into a deep learning model, which can obtain a time-frequency masking of the effective interaction region. A loss function can be constructed based on the sound sources within the effective interaction region in the simulated scene and the time-frequency masking. This loss function is then used to update the model parameters of the deep learning model, so that the deep learning model can obtain a more accurate time-frequency masking of the target speech, that is, a time-frequency masking of the speech signal within the effective interaction region. It should be noted that time-frequency masking generally masks sound sources outside the effective interaction region. Therefore, a loss function can be constructed based on the time-frequency masking obtained by the deep learning model and the sound sources within the effective interaction region in the simulated scene to determine whether the time-frequency masking obtained by the deep learning model is accurate.
[0089] Figure 4 This is a schematic diagram of a user interface for training a deep learning model according to an embodiment of this application. Figure 4 As shown, users can upload multiple simulated scene files by clicking the upload control on the user interface. Users can also drag multiple simulated scene files to the dotted box on the user interface to upload them. After successful upload, the images or thumbnails of the uploaded simulated scenes can be displayed in the upper right display box. Users can check whether the uploaded content needs to be changed by checking the images displayed in the display box. If no changes are needed, users can click the generate control to generate a set of simulated voices corresponding to multiple simulated scenes. The generated set of simulated voices corresponding to multiple simulated scenes can be displayed in the lower right display box.
[0090] In the above embodiments of this application, generating a set of simulated speech corresponding to multiple simulated scenarios includes: determining the simulated speech emitted by each sound source in each simulated scenario; determining the transfer function corresponding to each sound source in each simulated scenario using the mirror method; and convolving the simulated speech and the transfer function to obtain a set of simulated speech corresponding to each simulated scenario.
[0091] The image method described above can be used to calculate electrostatic or stable magnetic fields. Specifically, the image method can be implemented using the open-source tool RIR-Generator (room transfer function).
[0092] The transfer function mentioned above can be a room transfer function, which represents the distance from the sound source location to the microphone location.
[0093] In one alternative embodiment, the simulated speech emitted by each sound source in each simulated scene can be determined first. Then, the simulated speech and the room transfer function are convolved to determine the position of the simulated speech from the microphone. This allows it to be determined whether the simulated speech is within the effective interaction area or the ineffective interaction area, thereby obtaining the set of simulated speech corresponding to each simulated scene.
[0094] In the above embodiments of this application, feature extraction is performed on the original speech set, the enhanced first speech, and the second speech to obtain the original speech features, the first speech features, and the second speech features, respectively. This includes: performing short-time Fourier transform on the original speech set, the enhanced first speech, and the second speech to obtain the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal; performing feature extraction on the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal to obtain the original frequency domain features, the first frequency domain features, and the second frequency domain features; and normalizing the original frequency domain features, the first frequency domain features, and the second frequency domain features to obtain the original speech features, the first speech features, and the second speech features.
[0095] In one optional embodiment, a short-time Fourier transform can be performed on the original speech set, the enhanced first speech, and the second speech to obtain the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal. The short-time Fourier transform is a mathematical transformation related to the Fourier transform, used to determine the frequency and phase of the sinusoidal wave in a local region of the time-varying signal. Feature extraction is then performed on the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal to obtain the original frequency domain features, the first frequency domain features, and the second frequency domain features. These features can be 80-dimensional features. Mean-variance normalization is then performed on the original frequency domain features, the first frequency domain features, and the second frequency domain features to obtain the original speech features, the first speech features, and the second speech features. Mean-variance normalization can be performed by normalizing the features using their mean and variance, with the aim of ensuring that the mean of all features is near 0.
[0096] In the above embodiments of this application, the deep learning model includes: an input layer, multiple compact feedforward sequential storage networks, multiple first hidden layers, a first linear mapping layer, and an output layer connected in sequence. Each compact feedforward sequential storage network includes: a second hidden layer, a second linear mapping layer, and a memory module. The memory module in the previous compact feedforward sequential storage network is connected to the memory module in the next compact feedforward sequential storage network.
[0097] The aforementioned multiple first hidden layers are used to transform the output of the compact feedforward sequential storage network into an input that can be used by the first linear mapping layer. Similarly, the second hidden layer is used to transform the output of the input layer into an input that can be used by the compact feedforward sequential storage network.
[0098] The aforementioned compact feedforward sequential memory network can be a compact feedforward sequential memory network (cFSMN). It maps the output of the second hidden layer to a low-dimensional vector through a second linear mapping layer. This low-dimensional vector is then input into a memory module. The memory module performs a weighted sum of multiple input low-dimensional vectors, followed by an affine transformation and a nonlinear function to obtain the output of the compact feedforward sequential memory network. The memory module can also input the obtained low-dimensional vector into the memory module of the next compact feedforward sequential network. The memory modules in the preceding and following compact feedforward sequential memory network structures are connected, allowing the memory module to perform a weighted sum of multiple low-dimensional vectors obtained from memory modules in different compact feedforward sequential networks.
[0099] The first linear mapping layer described above is used to map the outputs of multiple first hidden layers to a low-dimensional vector. The low-dimensional vector can be transformed by an affine transformation and a nonlinear function and then output through the output layer.
[0100] Figure 5 This is a schematic diagram of a deep learning model according to an embodiment of this application. Figure 5 As shown, the original speech features, the first speech features, and the second speech features can be input into a sequentially connected input layer. The low-dimensional vector of the input features is weighted and summed through a compact feedforward sequential storage network. The weighted sum of the low-dimensional vector is then subjected to an affine transformation and a nonlinear function to obtain the output of the compact feedforward sequential storage network. This output can be transformed into an input that can be used by a second linear mapping layer through multiple first hidden layers. The input of the first linear mapping layer can be mapped to a low-dimensional vector, and the low-dimensional vector is subjected to an affine transformation and a nonlinear function to output the time-frequency masking of the target speech.
[0101] Figure 6This is a structural block diagram of a speech processing flow according to an embodiment of this application. The sound pickup device can be a microphone array, and the original speech set can be the microphone signals collected by the microphone array. The microphone signals can include speech signals within the effective interaction area and speech signals from other sound sources besides those within the effective interaction area. Target-direction notch filtering and target-direction beamforming can be applied to the microphone signals. Target-direction notch filtering mainly uses a notch filtering algorithm to suppress the speech signals within the effective interaction area to enhance the signals from other sound sources besides those within the effective interaction area, thus obtaining the enhanced second speech. Target-direction beamforming mainly uses a directional beamforming algorithm to enhance the speech signals within the effective interaction area, thus obtaining the enhanced first speech. Feature extraction and normalization can be performed on the enhanced first and second speech, and the processing results can be input into a deep neural network model (i.e., the deep learning model mentioned above) to obtain the time-frequency masking of the signal within the effective interaction area. Based on this time-frequency masking, signal recovery processing can be performed on the enhanced first speech to obtain a high-quality target signal within the effective interaction area (i.e., the target speech mentioned above), thereby improving the recognition accuracy of the target speech and thus enhancing the user experience in the voice interaction scenario.
[0102] Existing solutions for improving voice quality in voice interaction scenarios are geared towards extracting speech information from speakers in any direction. They only consider relatively simple multi-speaker scenarios and have limited ability to handle situations where speaker interference and noise coexist. Furthermore, these methods only consider large array scenarios with multiple microphones and lack analysis for scenarios with fewer microphones (e.g., two microphones) and smaller array spacing (e.g., less than 4cm).
[0103] The deep learning model in this application comprehensively utilizes the beam pointing to the effective interaction area and the notch signal of the effective interaction area to improve the model's spatial filtering capability. It can be applied in scenarios with a small number of microphones and microphone array spacing. At the same time, the model training data simulation fully considers various real-world scenarios: target sound source + interference sound source, interference sound source + noise sound source, target sound source + noise sound source, and scenarios where each sound source exists independently. This effectively improves the model's practical performance and enables it to handle situations where speaker interference and noise occur simultaneously. In real-world scenarios, the model can achieve a controllable effective interaction area of 20° on a two-microphone array, significantly surpassing the array pointing effect of classic beamforming and blind source separation algorithms.
[0104] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0105] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0106] Example 2
[0107] According to an embodiment of this application, a voice processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0108] Figure 7 This is a flowchart of a speech processing method according to Embodiment 2 of this application, as follows: Figure 7 As shown, the method may include the following steps:
[0109] Step S702: Capture the raw voice data collected by the audio pickup device set on the audio / video communication device.
[0110] The original speech set includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the voice interaction object that the pickup device identifies in a directional manner.
[0111] The audio and video communication devices mentioned above can be audio and video conferencing equipment, smart speakers, smart home appliances (such as TVs and refrigerators with voice control functions), but are not limited to these.
[0112] Step S704: Enhance the first speech and the second speech respectively to obtain the enhanced first speech and the second speech.
[0113] Step S706: Use a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech, to generate the target speech.
[0114] Step S708: Control the audio and video communication equipment to output the target voice.
[0115] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0116] Example 3
[0117] According to an embodiment of this application, a voice processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0118] Figure 8 This is a flowchart of a speech processing method according to Embodiment 3 of this application, as follows: Figure 8 As shown, the method may include the following steps:
[0119] Step S802: Capture the raw voice data collected by the sound pickup device set on the target vehicle.
[0120] The original speech set includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the voice interaction object that the pickup device identifies in a directional manner.
[0121] The target vehicles mentioned above can be fuel vehicles, new energy vehicles, autonomous vehicles, driverless vehicles, etc., and there is no limitation here.
[0122] Step S804: Enhance the first speech and the second speech respectively to obtain the enhanced first speech and the second speech.
[0123] Step S806: Use a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech, to generate the target speech.
[0124] Step S808: Control the target vehicle based on the target voice.
[0125] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0126] Example 4
[0127] According to an embodiment of this application, a voice processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0128] Figure 9 This is a flowchart of a speech processing method according to Embodiment 4 of this application, as follows: Figure 9 As shown, the method may include the following steps:
[0129] Step S902: The cloud server receives the original audio data uploaded by the client.
[0130] The original speech set is acquired by a sound pickup device. The original speech set includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the speech interaction object that the sound pickup device identifies in a specific direction.
[0131] In step S904, the cloud server performs enhancement processing on the first speech and the second speech respectively to obtain the enhanced first speech and the second speech.
[0132] In step S906, the cloud server uses a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate the target speech.
[0133] The target speech is the speech information picked up directionally by the sound pickup device.
[0134] Step S908: The cloud server outputs the target voice to the client.
[0135] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0136] Example 5
[0137] According to an embodiment of this application, a voice processing system is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0138] Figure 10 This is a schematic diagram of a speech processing system according to Embodiment 5 of this application, as shown below. Figure 10 As shown, the speech processing system 1000 includes:
[0139] The sound pickup device 1002 is used to collect a raw speech set, wherein the raw speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is the speech interaction object identified by the sound pickup device.
[0140] The processing device 1004 is connected to the sound pickup device and is used to enhance the first speech and the second speech respectively to obtain the enhanced first speech and the second speech. It also uses a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech to generate the target speech, wherein the target speech is the speech information picked up directionally by the sound pickup device.
[0141] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0142] Example 6
[0143] According to an embodiment of this application, an audio and video communication device is also provided. Figure 11 This is a schematic diagram of an audio / video communication device according to Embodiment 6 of this application, as shown below. Figure 11 As shown, the audio / video communication device 1100 includes:
[0144] The sound pickup device 1102 installed on the audio and video communication device 1100 is used to collect the original speech set, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is the speech interaction object identified by the sound pickup device.
[0145] The processor 1104 is connected to the sound pickup device 1102 and is used to enhance the first speech and the second speech respectively to obtain the enhanced first speech and the second speech. It also uses a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech to generate the target speech.
[0146] The output device 1106 is connected to the processor 1104 and is used to output the target speech.
[0147] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0148] Example 7
[0149] According to an embodiment of this application, a vehicle is also provided. Figure 12 This is a schematic diagram of a vehicle according to Embodiment 7 of this application, as shown below. Figure 12 As shown, the vehicle 1200 includes:
[0150] The sound pickup device 1202 installed on the vehicle 1200 is used to collect an original speech set, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is the speech interaction object identified by the sound pickup device.
[0151] The controller 1204 is connected to the sound pickup device 1202 and is used to enhance the first speech and the second speech respectively to obtain the enhanced first speech and the second speech. The controller uses a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech to generate the target speech, and controls the target vehicle based on the target speech.
[0152] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0153] Example 8
[0154] According to an embodiment of this application, a speech processing apparatus for implementing the above-described speech processing method is also provided. Figure 13 This is a schematic diagram of a voice processing device according to Embodiment 8 of this application, as shown below. Figure 13 As shown, the device 1300 includes: an acquisition module 1302, an enhancement processing module 1304, and a recovery processing module 1306.
[0155] The acquisition module is used to acquire the original speech set collected by the sound pickup device. The original speech set includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the speech interaction object that the sound pickup device identifies in a directional manner. The enhancement processing module is used to enhance the first speech and the second speech respectively to obtain the enhanced first speech and the second speech. The recovery processing module is used to use a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech, to generate the target speech. The target speech is the speech information picked up directionally by the sound pickup device.
[0156] It should be noted that the acquisition module 1302, enhancement processing module 1304, and recovery processing module 1306 mentioned above correspond to steps S202 to S206 of Embodiment 1. The three modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computing terminal 10 provided in Embodiment 1.
[0157] In this embodiment of the application, the enhancement processing module includes: a superposition unit and a filtering unit.
[0158] The superposition unit is used to superimpose the original speech set using a beamforming algorithm to obtain the enhanced first speech; the filtering unit is used to filter the original speech set using a notch filtering algorithm to obtain the enhanced second speech.
[0159] In this embodiment of the application, the recovery processing module includes: an extraction unit, an input unit, and a recovery unit.
[0160] The extraction unit is used to extract features from the original speech set, the enhanced first speech and the second speech respectively, to obtain the original speech features, the first speech features and the second speech features; the input unit is used to input the original speech features, the first speech features and the second speech features into the deep learning model to obtain the time-frequency mask of the target speech; the recovery unit is used to perform signal recovery processing on the enhanced first speech based on the time-frequency mask to obtain the target speech.
[0161] In this embodiment of the application, the device further includes: a construction module, a generation module, and a training module.
[0162] The construction module is used to construct multiple simulated scenarios corresponding to the sound pickup device. The number and type of sound sources contained in different simulated scenarios are different. The types of sound sources include at least one of the following: target sound source, interference sound source, and noise sound source. The target sound source is located in the effective interaction area, and the interference sound source and noise sound source are located in other interaction areas other than the effective interaction area. The generation module is used to generate a set of simulated speech corresponding to multiple simulated scenarios. The training module is used to train a deep learning model using the set of simulated speech.
[0163] In this embodiment of the application, the generation module includes: a determination unit and a convolution unit.
[0164] The determination unit is used to determine the simulated speech emitted by each sound source in each simulated scenario; the determination unit is also used to determine the transfer function corresponding to each sound source in each simulated scenario by means of mirroring; the convolution unit is used to convolve the simulated speech and the transfer function to obtain the set of simulated speech corresponding to each simulated scenario.
[0165] In this embodiment of the application, the extraction unit includes: a processing subunit, an extraction subunit, and a regularization subunit.
[0166] The processing subunit performs short-time Fourier transforms on the original speech set, the enhanced first speech, and the second speech to obtain the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal, respectively. The extraction subunit extracts features from the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal to obtain the original frequency domain features, the first frequency domain features, and the second frequency domain features, respectively. The normalization subunit normalizes the original frequency domain features, the first frequency domain features, and the second frequency domain features to obtain the original speech features, the first speech features, and the second speech features, respectively.
[0167] In this embodiment of the application, the deep learning model includes: an input layer, multiple compact feedforward sequential storage networks, multiple first hidden layers, a first linear mapping layer, and an output layer connected in sequence. Each compact feedforward sequential storage network includes: a second hidden layer, a second linear mapping layer, and a memory module. The memory module in the previous compact feedforward sequential storage network is connected to the memory module in the next compact feedforward sequential storage network.
[0168] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0169] Example 9
[0170] According to an embodiment of this application, a speech processing apparatus for implementing the above-described speech processing method is also provided. Figure 14This is a schematic diagram of a voice processing device according to Embodiment 9 of this application, as shown below. Figure 14 As shown, the device 1400 includes: a capture module 1402, an enhancement processing module 1404, a recovery processing module 1406, and a control module 1408.
[0171] The system includes a capture module for capturing the original speech set collected by a pickup device on the audio / video communication device. The original speech set includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the speech interaction object identified by the pickup device. The enhancement processing module enhances the first speech and the second speech respectively to obtain enhanced first speech and second speech. The recovery processing module uses a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate the target speech. The control module controls the audio / video communication device to output the target speech.
[0172] It should be noted that the capture module 1402, enhancement processing module 1404, recovery processing module 1406, and control module 1408 mentioned above correspond to steps S902 to S908 in Embodiment 2. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computing terminal 10 provided in Embodiment 1.
[0173] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0174] Example 10
[0175] According to an embodiment of this application, a speech processing apparatus for implementing the above-described speech processing method is also provided. Figure 15 This is a schematic diagram of a voice processing device according to Embodiment 10 of this application, as shown below. Figure 15 As shown, the device 1500 includes: a capture module 1502, an enhancement processing module 1504, a recovery processing module 1506, and a control module 1508.
[0176] The system includes a capture module for capturing the original speech set collected by a sound pickup device installed on the target vehicle. The original speech set includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the speech interaction object identified by the sound pickup device. The enhancement processing module enhances the first speech and the second speech respectively to obtain enhanced first speech and second speech. The recovery processing module uses a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech to generate the target speech. The control module controls the target vehicle based on the target speech.
[0177] It should be noted that the capture module 1502, enhancement processing module 1504, recovery processing module 1506, and control module 1508 mentioned above correspond to steps S1002 to S1008 in Embodiment 3. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computing terminal 10 provided in Embodiment 1.
[0178] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0179] Example 11
[0180] According to an embodiment of this application, a speech processing apparatus for implementing the above-described speech processing method is also provided. Figure 16 This is a schematic diagram of a voice processing device according to Embodiment 11 of this application, as shown below. Figure 16 As shown, the device 1600 includes: a receiving module 1602, an enhancement processing module 1604, a recovery processing module 1606, and an output module 1608.
[0181] The receiving module receives the original speech set uploaded by the client through a cloud server. The original speech set is acquired by a sound pickup device and includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the speech interaction object identified directionally by the sound pickup device. The enhancement processing module enhances the first speech and the second speech separately through the cloud server to obtain enhanced first speech and second speech. The recovery processing module uses a deep learning model through the cloud server to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate the target speech. The target speech is the speech information directionally picked up by the sound pickup device. The output module outputs the target speech to the client through the cloud server.
[0182] It should be noted that the receiving module 1602, enhancement processing module 1604, recovery processing module 1606, and output module 1608 mentioned above correspond to steps S1102 to S1108 in Embodiment 4. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computing terminal 10 provided in Embodiment 1.
[0183] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0184] Example 12
[0185] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.
[0186] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0187] In this embodiment, the computer terminal described above can execute the program code for the following steps in the speech processing method: acquiring the original speech set collected by the sound pickup device, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is the speech interaction object identified directionally by the sound pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; using a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate target speech, wherein the target speech is the speech information directionally picked up by the sound pickup device.
[0188] Optionally, Figure 17 This is a structural block diagram of a computer terminal according to Embodiment 12 of this application. Figure 17 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors and memory.
[0189] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the voice processing method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the aforementioned voice processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0190] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: acquiring a set of raw speech collected by the sound pickup device, wherein the raw speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is the speech interaction object identified directionally by the sound pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; using a deep learning model to perform signal recovery processing on the raw speech set, the enhanced first speech, and the second speech to generate target speech, wherein the target speech is the speech information directionally picked up by the sound pickup device.
[0191] Optionally, the processor may also execute program code for the following steps: superimposing the original speech set using a beamforming algorithm to obtain the enhanced first speech; and filtering the original speech set using a notch filtering algorithm to obtain the enhanced second speech.
[0192] Optionally, the processor may also execute program code for the following steps: extracting features from the original speech set, the enhanced first speech, and the second speech respectively to obtain original speech features, first speech features, and second speech features; inputting the original speech features, first speech features, and second speech features into a deep learning model to obtain the time-frequency masking of the target speech; and performing signal recovery processing on the enhanced first speech based on the time-frequency masking to obtain the target speech.
[0193] Optionally, the processor may also execute program code for the following steps: constructing multiple simulated scenarios corresponding to the sound pickup device, wherein the number and type of sound sources contained in different simulated scenarios are different, and the types of sound sources include at least one of the following: target sound source, interference sound source, and noise sound source, wherein the target sound source is located in the effective interaction area, and the interference sound source and noise sound source are located in other interaction areas other than the effective interaction area; generating a set of simulated speech corresponding to multiple simulated scenarios; and training a deep learning model using the set of simulated speech.
[0194] Optionally, the processor may also execute program code that performs the following steps: determining the simulated speech emitted by each sound source in each simulated scenario; determining the transfer function corresponding to each sound source in each simulated scenario using the mirror method; and convolving the simulated speech and the transfer function to obtain the set of simulated speech corresponding to each simulated scenario.
[0195] Optionally, the processor may also execute program code for the following steps: performing short-time Fourier transforms on the original speech set, the enhanced first speech, and the second speech to obtain the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal; extracting features from the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal to obtain the original frequency domain features, the first frequency domain features, and the second frequency domain features; and normalizing the original frequency domain features, the first frequency domain features, and the second frequency domain features to obtain the original speech features, the first speech features, and the second speech features.
[0196] Optionally, the processor may also execute program code that performs the following steps: The deep learning model includes: an input layer, multiple compact feedforward sequential storage networks, multiple first hidden layers, a first linear mapping layer, and an output layer connected in sequence, wherein each compact feedforward sequential storage network includes: a second hidden layer, a second linear mapping layer, and a memory module, and the memory module in the previous compact feedforward sequential storage network is connected to the memory module in the next compact feedforward sequential storage network.
[0197] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: capturing a set of raw speech collected by a pickup device on an audio / video communication device, wherein the raw speech set includes a first speech and a second speech, the first speech being a speech signal emitted from a sound source located within the effective interaction area, and the second speech being a speech signal emitted from a sound source other than the sound source located within the effective interaction area, the sound source located within the effective interaction area being the speech interaction object identified by the pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; using a deep learning model to perform signal recovery processing on the raw speech set, the enhanced first speech, and the second speech to generate target speech; and controlling the audio / video communication device to output the target speech.
[0198] The processor can access information and applications stored in memory via a transmission device to perform the following steps: capturing a set of raw speech collected by a sound pickup device installed on the target vehicle, wherein the raw speech set includes a first speech and a second speech, the first speech being a speech signal emitted from a sound source located within the effective interaction area, and the second speech being a speech signal emitted from a sound source other than the sound source located within the effective interaction area, the sound source located within the effective interaction area being the voice interaction object identified by the sound pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; using a deep learning model to perform signal recovery processing on the raw speech set, the enhanced first speech, and the second speech to generate the target speech; and controlling the target vehicle based on the target speech.
[0199] The processor can access information and applications stored in the memory via a transmission device to execute the following steps: The cloud server receives a raw voice set uploaded by the client, wherein the raw voice set is acquired by a sound pickup device and includes a first voice and a second voice. The first voice is a voice signal emitted from a sound source located within the effective interaction area, and the second voice is a voice signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the voice interaction object identified directionally by the sound pickup device. The cloud server performs enhancement processing on the first voice and the second voice respectively to obtain enhanced first voice and second voice. The cloud server uses a deep learning model to perform signal recovery processing on the raw voice set, the enhanced first voice, and the second voice to generate target voice, wherein the target voice is the voice information directionally picked up by the sound pickup device. The cloud server outputs the target voice to the client.
[0200] In this embodiment, the original speech set collected by the sound pickup device is first obtained. The original speech set includes a first speech and a second speech. The first speech is speech information emitted by a sound source located within the effective interaction area, and the second speech is speech information emitted by a sound source other than the sound source within the effective interaction area. The sound source within the effective interaction area is the speech interaction object identified directionally by the sound pickup device. The first speech and the second speech are enhanced respectively to obtain enhanced first speech and second speech. A deep learning model is used to recover the speech signals from the original speech set, the enhanced first speech, and the second speech to generate the target speech. The target speech is the speech information directionally picked up by the sound pickup device, which effectively suppresses sound source interference and environmental noise outside the effective interaction area and improves the extraction effect of speech information within the effective interaction area. It is noteworthy that the first speech emitted by the sound source within the effective interaction area and the second speech emitted by other sound sources outside the effective interaction area can be enhanced separately. Combined with a deep learning model, the second speech emitted by other sound sources can be effectively suppressed, so that the sound pickup device can pick up the speech information within the effective interaction area in a directional manner. This solves the problem of speech information technology in related technologies that is difficult to pick up sound sources within the effective interaction area.
[0201] Those skilled in the art will understand that Figure 17 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a mobile internet device (MID), a PAD, and other terminal devices. Figure 17 This does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include components that are more... Figure 17 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 17 The different configurations shown.
[0202] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0203] Example 13
[0204] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the speech processing method provided in Embodiment 1.
[0205] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0206] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: acquiring an original speech set collected by the sound pickup device, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, and the sound source located within the effective interaction area is the speech interaction object identified directionally by the sound pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; using a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate target speech, wherein the target speech is the speech information directionally picked up by the sound pickup device.
[0207] Optionally, the storage medium is further configured to store program code for performing the following steps: superimposing the original speech set using a beamforming algorithm to obtain enhanced first speech; and filtering the original speech set using a notch filtering algorithm to obtain enhanced second speech.
[0208] Optionally, the storage medium is further configured to store program code for performing the following steps: extracting features from the original speech set, the enhanced first speech, and the second speech respectively to obtain original speech features, first speech features, and second speech features; inputting the original speech features, first speech features, and second speech features into a deep learning model to obtain a time-frequency masking of the target speech; and performing signal recovery processing on the enhanced first speech based on the time-frequency masking to obtain the target speech.
[0209] Optionally, the aforementioned storage medium is further configured to store program code for performing the following steps: constructing multiple simulated scenarios corresponding to the sound pickup device, wherein the number and type of sound sources contained in different simulated scenarios are different, and the types of sound sources include at least one of the following: target sound source, interference sound source, and noise sound source, wherein the target sound source is located in the effective interaction area, and the interference sound source and noise sound source are located in other interaction areas other than the effective interaction area; generating a set of simulated speech corresponding to multiple simulated scenarios; and training a deep learning model using the set of simulated speech.
[0210] Optionally, the aforementioned storage medium is further configured to store program code for performing the following steps: determining the simulated speech emitted by each sound source in each simulated scenario; determining the transfer function corresponding to each sound source in each simulated scenario using the mirror method; and convolving the simulated speech and the transfer function to obtain a set of simulated speech corresponding to each simulated scenario.
[0211] Optionally, the storage medium is further configured to store program code for performing the following steps: performing short-time Fourier transforms on the original speech set, the enhanced first speech, and the second speech to obtain the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal; extracting features from the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal to obtain the original frequency domain features, the first frequency domain features, and the second frequency domain features; and normalizing the original frequency domain features, the first frequency domain features, and the second frequency domain features to obtain the original speech features, the first speech features, and the second speech features.
[0212] Optionally, the storage medium is further configured to store program code for performing the following steps: the deep learning model includes: an input layer, multiple compact feedforward sequential storage networks, multiple first hidden layers, a first linear mapping layer, and an output layer connected in sequence, wherein each compact feedforward sequential storage network includes: a second hidden layer, a second linear mapping layer, and a memory module, and the memory module in the preceding compact feedforward sequential storage network is connected to the memory module in the following compact feedforward sequential storage network.
[0213] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: capturing a set of original speech collected by a pickup device on an audio-visual communication device, wherein the original speech set includes a first speech and a second speech, the first speech being a speech signal emitted from a sound source located within the effective interaction area, and the second speech being a speech signal emitted from a sound source other than the sound source located within the effective interaction area, the sound source located within the effective interaction area being the voice interaction object identified by the pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; using a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate target speech; and controlling the audio-visual communication device to output the target speech.
[0214] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: capturing an original speech set collected by a sound pickup device installed on the target vehicle, wherein the original speech set includes a first speech and a second speech, the first speech being a speech signal emitted from a sound source located within the effective interaction area, and the second speech being a speech signal emitted from a sound source other than the sound source located within the effective interaction area, the sound source located within the effective interaction area being the voice interaction object identified directionally by the sound pickup device; performing enhancement processing on the first speech and the second speech respectively to obtain enhanced first speech and second speech; using a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech to generate target speech; and controlling the target vehicle based on the target speech.
[0215] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: the cloud server receives an original voice set uploaded by the client, wherein the original voice set is acquired by a sound pickup device, and the original voice set includes a first voice and a second voice. The first voice is a voice signal emitted from a sound source located within the effective interaction area, and the second voice is a voice signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the voice interaction object identified directionally by the sound pickup device. The cloud server performs enhancement processing on the first voice and the second voice respectively to obtain enhanced first voice and second voice. The cloud server uses a deep learning model to perform signal recovery processing on the original voice set, the enhanced first voice, and the second voice to generate target voice, wherein the target voice is the voice information directionally picked up by the sound pickup device. The cloud server outputs the target voice to the client.
[0216] In this embodiment, the original speech set collected by the sound pickup device is first obtained. The original speech set includes a first speech and a second speech. The first speech is speech information emitted by a sound source located within the effective interaction area, and the second speech is speech information emitted by a sound source other than the sound source within the effective interaction area. The sound source within the effective interaction area is the speech interaction object identified directionally by the sound pickup device. The first speech and the second speech are enhanced respectively to obtain enhanced first speech and second speech. A deep learning model is used to recover the speech signals from the original speech set, the enhanced first speech, and the second speech to generate the target speech. The target speech is the speech information directionally picked up by the sound pickup device, which effectively suppresses sound source interference and environmental noise outside the effective interaction area and improves the extraction effect of speech information within the effective interaction area. It is noteworthy that the first speech emitted by the sound source within the effective interaction area and the second speech emitted by other sound sources outside the effective interaction area can be enhanced separately. Combined with a deep learning model, the second speech emitted by other sound sources can be effectively suppressed, so that the sound pickup device can pick up the speech information within the effective interaction area in a directional manner. This solves the problem of speech information technology in related technologies that is difficult to pick up sound sources within the effective interaction area.
[0217] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0218] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0219] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0220] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0221] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0222] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0223] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A speech processing method, characterized in that, include: Acquire the original speech set collected by the sound pickup device, wherein the original speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within the effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, wherein the sound source located within the effective interaction area is the speech interaction object identified by the sound pickup device, and the effective interaction area is an interaction area located in the target direction that is pre-set according to the actual scene; The original speech set is superimposed using a beamforming algorithm to obtain the enhanced first speech. The original speech set is filtered using a notch filter algorithm to obtain the enhanced second speech, wherein the filtering process is used to represent the obstruction of the first speech in the original speech set in order to enhance the second speech; A deep learning model is used to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech using spatial information to generate target speech. The target speech is the speech information picked up directionally by the sound pickup device, and the spatial information is used to represent the relevant speech features between channels.
2. The method according to claim 1, characterized in that, Using a deep learning model, signal recovery processing is performed on the original speech set, the enhanced first speech, and the second speech using spatial information to generate target speech, including: The spatial information is used to extract features from the original speech set, the enhanced first speech and the second speech, respectively, to obtain the original speech features, the first speech features and the second speech features. The original speech features, the first speech features, and the second speech features are input into the deep learning model to obtain the time-frequency masking of the target speech; Based on the time-frequency masking, signal recovery processing is performed on the enhanced first speech to obtain the target speech.
3. The method according to claim 2, characterized in that, The method further includes: Multiple simulated scenarios corresponding to the sound pickup device are constructed, wherein the number and type of sound sources contained in different simulated scenarios are different. The types of sound sources include at least one of the following: target sound source, interference sound source, and noise sound source. The target sound source is located in the effective interaction area, and the interference sound source and the noise sound source are located in other interaction areas other than the effective interaction area. Generate a set of simulated speech corresponding to the multiple simulated scenarios; The deep learning model is trained using the simulated speech set.
4. The method according to claim 3, characterized in that, Generating the simulated speech set corresponding to the multiple simulated scenarios includes: Determine the simulated speech emitted by each sound source in each simulated scenario; The transfer function corresponding to each sound source in each simulated scenario is determined by the mirror method; The simulated speech and the transfer function are convolved to obtain the set of simulated speech corresponding to each simulated scenario.
5. The method according to claim 2, characterized in that, Feature extraction is performed on the original speech set, the enhanced first speech, and the second speech, respectively, to obtain the original speech features, the first speech features, and the second speech features, including: Short-time Fourier transforms are performed on the original speech set, the enhanced first speech, and the second speech to obtain the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal, respectively. Feature extraction is performed on the original frequency domain signal, the first frequency domain signal, and the second frequency domain signal respectively to obtain the original frequency domain features, the first frequency domain features, and the second frequency domain features; The original frequency domain features, the first frequency domain features, and the second frequency domain features are normalized respectively to obtain the original speech features, the first speech features, and the second speech features.
6. The method according to claim 1, characterized in that, The deep learning model includes: an input layer, multiple compact feedforward sequential storage networks, multiple first hidden layers, a first linear mapping layer, and an output layer connected in sequence. Each compact feedforward sequential storage network includes: a second hidden layer, a second linear mapping layer, and a memory module. The memory module in the previous compact feedforward sequential storage network is connected to the memory module in the next compact feedforward sequential storage network.
7. A speech processing method, characterized in that, include: The system captures a set of raw speech collected by a sound pickup device on an audio-visual communication device. The raw speech set includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within an effective interaction area. The second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the voice interaction object identified by the sound pickup device. The effective interaction area is an interaction area located in the target direction that is pre-set according to the actual scene. The original speech set is superimposed using a beamforming algorithm to obtain the enhanced first speech. The original speech set is filtered using a notch filter algorithm to obtain the enhanced second speech, wherein the filtering process is used to represent the obstruction of the first speech in the original speech set in order to enhance the second speech; A deep learning model is used to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech using spatial information to generate the target speech, wherein the spatial information is used to represent the relevant speech features between channels; Control the audio and video communication device to output the target voice.
8. A speech processing method, characterized in that, include: The system captures a set of raw speech data collected by a sound pickup device installed on the target vehicle. The raw speech data includes a first speech and a second speech. The first speech is a speech signal emitted from a sound source located within the effective interaction area. The second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the speech interaction object identified by the sound pickup device. The effective interaction area is an interaction area located in the target direction that is pre-set according to the actual scene. The original speech set is superimposed using a beamforming algorithm to obtain the enhanced first speech. The original speech set is filtered using a notch filter algorithm to obtain the enhanced second speech, wherein the filtering process is used to represent the obstruction of the first speech in the original speech set in order to enhance the second speech; A deep learning model is used to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech using spatial information to generate the target speech, wherein the spatial information is used to represent the relevant speech features between channels; The target vehicle is controlled based on the target voice.
9. A speech processing method, characterized in that, include: The cloud server receives a raw voice collection uploaded by the client. The raw voice collection is acquired by a sound pickup device and includes a first voice and a second voice. The first voice is a voice signal emitted from a sound source located within the effective interaction area, and the second voice is a voice signal emitted from a sound source other than the sound source located within the effective interaction area. The sound source located within the effective interaction area is the voice interaction object identified by the sound pickup device. The effective interaction area is an interaction area located in the target direction that is pre-set according to the actual scene. The cloud server uses a beamforming algorithm to superimpose the original speech set to obtain an enhanced first speech, and uses a notch filtering algorithm to filter the original speech set to obtain an enhanced second speech. The filtering is used to represent the obstruction of the first speech in the original speech set in order to enhance the second speech. The cloud server uses a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech and the second speech through spatial information to generate target speech. The target speech is the speech information picked up directionally by the sound pickup device. The spatial information is used to represent the relevant speech features between channels. The cloud server outputs the target voice to the client.
10. A speech processing system, characterized in that, include: A sound pickup device is used to collect a raw speech set, wherein the raw speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within an effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, wherein the sound source located within the effective interaction area is a voice interaction object identified by the sound pickup device, and the effective interaction area is an interaction area located in the target direction that is pre-set according to the actual scene; A processing device, connected to the sound pickup device, is used to superimpose the original speech set using a beamforming algorithm to obtain an enhanced first speech, and to filter the original speech set using a notch filtering algorithm to obtain the enhanced second speech. The filtering is used to impede the first speech in the original speech set to enhance the second speech. A deep learning model is then used to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech using spatial information to generate a target speech. The target speech is the speech information directionally picked up by the sound pickup device, and the spatial information represents the relevant speech features between channels.
11. An audio and video communication device, characterized in that, include: A sound pickup device installed on an audio-visual communication device is used to collect a raw voice set, wherein the raw voice set includes a first voice and a second voice, wherein the first voice is a voice signal emitted from a sound source located within an effective interaction area, and the second voice is a voice signal emitted from a sound source other than the sound source located within the effective interaction area, wherein the sound source located within the effective interaction area is the voice interaction object identified by the sound pickup device, and the effective interaction area is an interaction area located in the target direction that is pre-set according to the actual scene; The processor, connected to the sound pickup device, is used to superimpose the original speech set using a beamforming algorithm to obtain an enhanced first speech, and to filter the original speech set using a notch filtering algorithm to obtain the enhanced second speech, wherein the filtering is used to represent the obstruction of the first speech in the original speech set to enhance the second speech, and to use a deep learning model to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech using spatial information to generate target speech, wherein the spatial information is used to represent the relevant speech features between channels; An output device, connected to the processor, is used to output the target speech.
12. A vehicle, characterized in that, include: A sound pickup device installed on a vehicle is used to collect a raw speech set, wherein the raw speech set includes a first speech and a second speech, wherein the first speech is a speech signal emitted from a sound source located within an effective interaction area, and the second speech is a speech signal emitted from a sound source other than the sound source located within the effective interaction area, wherein the sound source located within the effective interaction area is the speech interaction object identified by the sound pickup device, and the effective interaction area is an interaction area located in the target direction that is pre-set according to the actual scene; A controller, connected to the sound pickup device, is used to superimpose the original speech set using a beamforming algorithm to obtain an enhanced first speech, and to filter the original speech set using a notch filtering algorithm to obtain the enhanced second speech. The filtering is used to impede the first speech in the original speech set to enhance the second speech. A deep learning model is used to perform signal recovery processing on the original speech set, the enhanced first speech, and the second speech using spatial information to generate a target speech. The controller then controls the vehicle based on the target speech. The spatial information represents the relevant speech features between channels.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the storage medium is located to perform the speech processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Microphone array speech enhancement method and device, electronic equipment and storage medium
CN113889137A
Method, medium, and apparatus for extracting target sound from mixed sound
US20090097670A1