Speech processing method, training method of speech processing model, device and medium
By integrating the characteristics of multiple audio signals and performing classification processing, the poor positioning effect and memory usage of existing sound source positioning methods are solved, and higher positioning accuracy and computing speed are achieved.
Patent Information
- Application Number
- CN202111675104.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The existing sound source positioning methods have problems such as poor positioning effect, susceptibility to noise interference and high memory usage.
By obtaining the first acoustic feature information based on the multiple audio signal, and classifying it to obtain spatial position information, fusing the features of the multiple audio signal to obtain the first acoustic feature information, using the feature extraction module and the feature fusion module for feature extraction and fusion, and identifying the wake-up word in combination with the acoustic wake-up module.
It reduces cascade errors, improves the accuracy and computing speed of sound source positioning, and reduces memory usage.
Smart Images

Figure CN114242066B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a speech processing method, a training method for a speech processing model, a device, and a medium. Background Art
[0002] In voice interaction scenarios, the interactive device locates the user's location based on the received voice, allowing it to better capture the user's voice at that location. This prevents noise from other locations from interfering with the voice interaction process and improves the device's intelligence. However, existing sound source localization methods suffer from poor localization performance and high memory usage. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide a speech processing method, a training method, a device and a medium for a speech processing model, which can reduce cascade errors, improve the accuracy of sound source localization and increase the computing speed.
[0004] In a first aspect, an embodiment of the present application provides a speech processing method, comprising: obtaining first acoustic feature information based on multi-channel audio signals, wherein the first acoustic feature information integrates the features of the multi-channel audio signals, and the multi-channel audio signals are obtained based on the original audio signals collected by at least two audio collection devices; classifying and processing the first acoustic feature information to obtain spatial position information, and the spatial position information is used to characterize the direction of the sound source.
[0005] In certain embodiments of the present application, the speech processing method is applied to a speech processing model, which includes: a feature extraction module, a feature fusion module and an acoustic wake-up module.
[0006] In certain embodiments of the present application, obtaining first acoustic feature information based on multiple audio signals includes: using a feature extraction module to extract feature information corresponding to the multiple audio signals and feature information corresponding to the original audio signals collected by one or more audio collection devices among at least two audio collection devices; using a feature fusion module to fuse the feature information corresponding to the multiple audio signals and the feature information corresponding to the original audio signals to obtain the first acoustic feature information.
[0007] In certain embodiments of the present application, obtaining first acoustic feature information based on multiple audio signals includes: using a feature extraction module to extract feature information corresponding to the multiple audio signals; using a feature fusion module to fuse the feature information corresponding to the multiple audio signals and the phase difference between the original audio signals collected by any two audio collection devices among at least two audio collection devices to obtain the first acoustic feature information.
[0008] In certain embodiments of the present application, obtaining first acoustic feature information based on multiple audio signals includes: using a feature extraction module to extract feature information corresponding to the multiple audio signals and feature information corresponding to the amplitude difference between original audio signals collected by any two audio collection devices among at least two audio collection devices; using a feature fusion module to fuse the feature information corresponding to the multiple audio signals and the feature information corresponding to the amplitude difference to obtain the first acoustic feature information.
[0009] In certain embodiments of the present application, obtaining first acoustic feature information based on multiple audio signals includes: using a feature extraction module to extract feature information corresponding to the multiple audio signals, feature information corresponding to the original audio signals collected by one or more audio collection devices among at least two audio collection devices, and feature information corresponding to the amplitude difference between the original audio signals collected by any two audio collection devices among at least two audio collection devices; using a feature fusion module to fuse the feature information corresponding to the multiple audio signals, the feature information corresponding to the original audio signals, the feature information corresponding to the amplitude difference, and the phase difference between the original audio signals collected by any two audio collection devices among at least two audio collection devices to obtain the first acoustic feature information.
[0010] In some embodiments of the present application, the speech processing method further includes: using an acoustic wake-up module to identify a wake-up word based on the first acoustic feature information.
[0011] In certain embodiments of the present application, using the acoustic wake-up module to identify the wake-up word based on the first acoustic feature information includes: using the acoustic wake-up module to combine spatial position information and the first acoustic feature information to identify the wake-up word.
[0012] In certain embodiments of the present application, an acoustic wake-up module is used in combination with spatial position information and first acoustic feature information to identify a wake-up word, including: using the gating of the acoustic wake-up module to perform gating selection on the first acoustic feature information based on the spatial position information to obtain second acoustic feature information; using the acoustic wake-up module to process the second acoustic feature information to obtain an acoustic score, wherein the acoustic wake-up module is obtained by adversarial training using a gradient flip layer, and the gradient flip layer points to an energy classifier; and identifying the wake-up word based on the acoustic score.
[0013] In some embodiments of the present application, identifying a wake-up word based on an acoustic score includes: inputting the acoustic score into a decoding network to identify the wake-up word.
[0014] In some embodiments of the present application, the multi-channel audio signals are obtained by performing beamforming processing on original audio signals collected by at least two audio collection devices according to multiple preset sound source directions.
[0015] In certain embodiments of the present application, the speech processing model also includes a sound source classifier and a training module, and the speech processing model is obtained through the following training method: using a feature extraction module and a feature fusion module to obtain first sample acoustic feature information based on sample data, wherein the sample data includes multiple sample audio signals, and the multiple sample audio signals are obtained based on the original sample audio signals collected by at least two sample audio collection devices, and the first sample acoustic feature information fuses the features of the multiple sample audio signals; using a sound source classifier to obtain spatial position information corresponding to the sample data based on the first sample acoustic feature information; using a training module to train the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
[0016] In some embodiments of the present application, the sample data also includes at least one of the following items: original sample audio signals collected by one or more sample audio collection devices among at least two sample audio collection devices, sample phase difference between original sample audio signals collected by any two sample audio collection devices among at least two sample audio collection devices, and sample amplitude difference between original sample audio signals collected by any two sample audio collection devices among at least two sample audio collection devices.
[0017] In certain embodiments of the present application, the training method further includes: using the acoustic wake-up module to obtain acoustic information based on the acoustic feature information of the first sample; and using the training module to train the acoustic wake-up module according to the acoustic information and the acoustic annotation information corresponding to the sample data.
[0018] In certain embodiments of the present application, the acoustic wake-up module is used to obtain acoustic information based on the first sample acoustic feature information, including: using the acoustic wake-up module to obtain acoustic information based on the first sample acoustic feature information and spatial position information corresponding to the sample data.
[0019] In certain embodiments of the present application, a gradient flip layer, an energy classifier, an acoustic classifier and a gating are provided in the acoustic wake-up module, and the acoustic wake-up module is used to obtain acoustic information based on the first sample acoustic feature information and the spatial position information corresponding to the sample data, including: using the gating to obtain the second sample acoustic feature information based on the first sample acoustic feature information and the spatial position information corresponding to the sample data; using the acoustic classifier to obtain acoustic information based on the second sample acoustic feature information, wherein the training method also includes: using the gradient flip layer to perform adversarial training on the acoustic wake-up module based on the second acoustic feature information and the energy labeling information, and the gradient flip layer points to the energy classifier.
[0020] In some embodiments of the present application, the multi-channel sample audio signals are obtained by performing beamforming processing on original sample audio signals collected by at least two sample audio collection devices according to multiple preset sound source directions.
[0021] In the second aspect, an embodiment of the present application provides a training method for a speech processing model, wherein the speech processing model includes a feature extraction module, a feature fusion module, a sound source classifier and a training module. The training method includes: using the feature extraction module and the feature fusion module to obtain first sample acoustic feature information based on sample data, wherein the sample data includes multiple sample audio signals, and the multiple sample audio signals are obtained based on the original sample audio signals collected by at least two sample audio collection devices, and the first sample acoustic feature information fuses the features of the multiple sample audio signals; using the sound source classifier to obtain spatial position information corresponding to the sample data based on the first sample acoustic feature information; and using the training module to train the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
[0022] In the third aspect, an embodiment of the present application provides a speech processing device, including: an acquisition module, used to obtain first acoustic feature information based on multi-channel audio signals, wherein the first acoustic feature information integrates the characteristics of the multi-channel audio signals, and the multi-channel audio signals are obtained based on the original audio signals collected by at least two audio collection devices; a classification module, used to classify the first acoustic feature information to obtain spatial position information, and the spatial position information is used to characterize the direction of the sound source.
[0023] In a fourth aspect, an embodiment of the present application provides a training device for a speech processing model, comprising: a first acquisition module for acquiring first sample acoustic feature information based on sample data, wherein the sample data includes multiple channels of sample audio signals, and the multiple channels of sample audio signals are acquired based on original sample audio signals acquired by at least two sample audio acquisition devices, and the first sample acoustic feature information integrates the features of the multiple channels of sample audio signals; a second acquisition module for acquiring spatial position information corresponding to the sample data based on the first sample acoustic feature information; and a training module for training the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
[0024] In a fifth aspect, an embodiment of the present application provides an electronic device comprising: a processor; and a memory for storing processor-executable instructions, wherein the processor is used to execute the speech processing method described in the first aspect above, or to execute the speech processing model training method described in the second aspect.
[0025] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is used to execute the speech processing method described in the first aspect above, or to execute the training method of the speech processing model described in the second aspect.
[0026] Embodiments of the present application provide a speech processing method, a training method for a speech processing model, a device, and a medium. The method obtains multiple audio signals based on original audio signals collected by at least two audio acquisition devices, fuses the features of the multiple audio signals to obtain first acoustic feature information corresponding to the multiple audio signals, and then classifies and processes the first acoustic feature information to obtain spatial position information that can characterize the direction of the sound source. Because the features of the multiple audio signals are fused to obtain the first acoustic feature information, rather than processing each audio signal separately, the speech processing method can reduce memory usage, increase computing speed, reduce cascade errors, and improve the accuracy of sound source localization. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Shown is a schematic diagram of the system architecture of a speech processing system provided by an exemplary embodiment of the present application.
[0028] Figure 2 Shown is a flow chart of a speech processing method provided by an exemplary embodiment of the present application.
[0029] Figure 3 Shown is a structural diagram of a speech processing model provided by an exemplary embodiment of the present application.
[0030] Figure 4 Shown is a flow chart of a speech processing method provided by another exemplary embodiment of the present application.
[0031] Figure 5 Shown is a structural diagram of a speech processing model provided by another exemplary embodiment of the present application.
[0032] Figure 6 Shown is a flowchart of a method for training a speech processing model provided by an exemplary embodiment of the present application.
[0033] Figure 7 Shown is a flowchart of a method for training a speech processing model provided by another exemplary embodiment of the present application.
[0034] Figure 8 FIG2 is a schematic diagram showing the structure of a speech processing device provided by an exemplary embodiment of the present application.
[0035] Figure 9 Shown is a structural diagram of a training device for a speech processing model provided by an exemplary embodiment of the present application.
[0036] Figure 10 Shown is a block diagram of an electronic device for executing a speech processing method or a speech processing model training method provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] Application Overview
[0039] With the development of intelligent level, more and more devices can be controlled by voice. In some scenarios, the device can locate the user's position by collecting the user's voice (audio signal), so that it can facilitate the subsequent effective collection of the user's voice at that position, thereby identifying the user's needs through the user's voice to achieve the interaction process with the user and provide better services for the user. For example, some chat robots, after locating the user's position by collecting the user's voice, can continue to collect the voice of the user at that position without being disturbed by noise from other positions, and identify the user's needs through voice to maintain continuous chat with the user.
[0040] In other scenarios, the device uses the user's voice (audio signal) to locate the user's position, so that it can determine whether the user is in a specific position and provide services to users in specific positions. For example, some in-vehicle interactive devices locate the user's position by collecting sound and determine whether the position is a specific position (such as the main driving direction). If so, the device recognizes the content of the sound and performs the corresponding operation. If not, the device does not recognize the content of the sound to avoid identifying the voice of users in other positions (such as the front passenger seat or the rear seat) and performing the corresponding operation, which may interfere with the driver's driving process or cause confusion in vehicle control.
[0041] A sound source localization method is to perform beamforming processing on the collected original audio signals to obtain multi-channel audio signals, process each audio signal in the multi-channel audio signals separately to obtain the confidence score and energy score corresponding to each audio signal, and make a decision based on the confidence scores and energy scores of the multi-channel audio signals to determine the direction of the sound source.
[0042] This sound source localization method uses a multi-step cascade, which can easily lead to error accumulation. Moreover, this sound source localization method based on energy scores cannot effectively solve the problem of sound source localization in noisy scenes, because the direction with high energy may be noise interference. Even if combined with confidence scores, it cannot effectively solve this type of problem. In addition, this sound source localization method performs the same processing (calculating confidence scores and energy scores) on each audio signal, resulting in an exponential increase in memory and CPU usage.
[0043] That is, the above-mentioned sound source localization method has the problems of easy error accumulation, easy interference by noise, and high memory usage.
[0044] Exemplary Systems
[0045] Figure 1 Figure 1 shows a schematic diagram of the system architecture of a speech processing system 100 provided in accordance with an exemplary embodiment of the present application, illustrating an application scenario for sound source localization of collected raw audio signals. The speech processing system 100 includes at least two audio acquisition devices 110 and a computing device 120. The computing device 120 can be communicatively coupled to the audio acquisition devices 110, or the audio acquisition devices 110 can be integrated into the computing device 120.
[0046] In one embodiment, the audio capture device 110 may be a microphone or a pickup for capturing sounds in the environment (e.g., a user's voice). At least two audio capture devices 110 may be located in different directions to capture the user's voice from different directions. The computing device 120 processes the sounds (audio signals) captured by the audio capture devices 110 to determine the direction of the sound source. The specific processing process is described in the exemplary method section below.
[0047] The computing device 120 may be a mobile phone, tablet, notebook, desktop computer, vehicle-mounted controller, or other device.
[0048] Furthermore, after locating the direction of the sound source, the computing device 120 can perform wake-up word recognition processing on the audio signal in combination with information about the direction of the sound source (such as spatial position information). If a wake-up word is recognized, the operation corresponding to the wake-up word is executed.
[0049] It should be noted that the above application scenarios are only provided to facilitate understanding of the spirit and principles of the present application, and the embodiments of the present application are not limited thereto. On the contrary, the embodiments of the present application can be applied to any applicable scenario.
[0050] Exemplary Methods
[0051] Figure 2 Shown is a flow chart of a speech processing method provided by an exemplary embodiment of the present application. Figure 2 The method may be executed by a computing device. Figure 2 As shown, the voice processing method includes the following contents.
[0052] 210: Obtain first acoustic feature information based on multiple audio signals, wherein the first acoustic feature information integrates features of the multiple audio signals, and the multiple audio signals are obtained based on original audio signals collected by at least two audio collection devices.
[0053] Multi-channel audio signals are obtained according to original audio signals collected by at least two audio collection devices.
[0054] Specifically, the audio acquisition device can be a microphone or a pickup for acquiring original audio signals. The at least two audio acquisition devices can be distributed in different directions to acquire original audio signals from different directions. For example, the at least two audio acquisition devices include four microphones, and the four microphones are respectively arranged at the main driving position, the co-pilot position, the left position of the rear seat, and the right position of the rear seat in the vehicle. There may be phase differences and / or amplitude differences between the original audio signals acquired by the microphones in different directions.
[0055] The original audio signals collected by each microphone are integrated and processed to obtain a multi-channel audio signal. For example, the original audio signals collected by each microphone can be simply combined to obtain the integrated original audio signal, or the integrated original audio signal can be obtained by weighted processing or other methods.
[0056] The number of multi-channel audio signals can be the same as the number of sound source directions that can be actually located by the speech processing method. For example, the number of sound source directions that can be actually located by the speech processing method is three, namely the driver's seat direction, the front passenger seat direction, and the rear seat direction. Processing the original audio signals collected by the four microphones can generate three audio signals corresponding to the three sound source directions.
[0057] It should be noted that there is a certain relationship between the number of audio capture devices and the number of multi-channel audio signals, but it is not an absolute relationship. A one-to-one correspondence between the two is not required, but the greater the number of audio capture devices, the greater the number of multi-channel audio signals that can be acquired. For example, three-channel audio signals can be acquired through two or more audio capture devices, while four or more audio signals can be acquired through three or more audio capture devices.
[0058] The number of audio acquisition devices and the number of multi-channel audio signals can be set according to actual needs, and the embodiment of the present application does not limit this.
[0059] Specifically, first acoustic feature information about the multi-channel audio signals can be obtained by combining the multi-channel audio signals. For example, the multi-channel audio signals can be input into a neural network model to obtain the first acoustic feature information. Alternatively, features from the multi-channel audio signals can be extracted using some algorithm, and these features can be fused to obtain the first acoustic feature information.
[0060] 220: Classify and process the first acoustic feature information to obtain spatial position information, where the spatial position information is used to represent the direction of the sound source.
[0061] Specifically, the first acoustic feature information can be classified and processed by a classifier to obtain spatial position information. The spatial position information can be a matrix or a vector, which is used to characterize the sound source direction of the original audio signal, that is, to characterize the direction of the user corresponding to the original audio signal. For example, three audio signals are obtained based on the original audio signals collected by at least two audio collection devices, and the first acoustic feature information is obtained based on the three audio signals. The spatial position information obtained by classifying the first acoustic feature information is a three-dimensional vector. Each dimension of the vector represents a preset sound source direction, and the value on each dimension represents the probability that the corresponding preset sound source direction is the sound source direction of the original audio signal. The preset sound source direction corresponding to the maximum value in the vector can be determined as the sound source direction finally located.
[0062] An embodiment of the present application provides a speech processing method that obtains multiple audio signals based on original audio signals collected by at least two audio acquisition devices, fuses the features of the multiple audio signals to obtain first acoustic feature information corresponding to the multiple audio signals, and then classifies and processes the first acoustic feature information to obtain spatial position information that can represent the direction of the sound source. Because the features of the multiple audio signals are fused to obtain the first acoustic feature information, rather than processing each audio signal separately, the speech processing method can reduce memory usage, increase computing speed, reduce cascade errors, and improve the accuracy of sound source localization.
[0063] According to an embodiment of the present application, the multi-channel audio signals are obtained by performing beamforming processing on original audio signals collected by at least two audio collection devices according to multiple preset sound source directions.
[0064] Specifically, the multiple preset sound source directions can be multiple sound source directions that can actually be located by the speech processing method. For example, original audio signals collected by at least two audio collection devices can be directionally enhanced according to the multiple preset sound source directions to obtain multi-path beamformed audio signals. Beamforming is a combination of antenna technology and digital signal processing technology for directional signal transmission or reception.
[0065] Of course, other methods may also be used to obtain multi-channel audio signals enhanced in multiple preset sound source directions based on original audio signals collected by at least two audio collection devices.
[0066] Multiple preset sound source directions can be set according to actual needs. The embodiment of the present application does not limit the specific number and specific directions of the multiple preset sound source directions.
[0067] The speech processing method provided in this embodiment performs beamforming processing on raw audio signals collected by at least two audio acquisition devices according to multiple preset sound source directions to generate multi-channel audio signals. This enhances audio signals from the preset sound source directions and suppresses audio signals from other directions, thereby improving the accuracy of sound source localization. Furthermore, the speech processing method provided in this embodiment can reasonably set the number of preset sound source directions as needed, thereby improving the adaptability of the method and, to a certain extent, increasing the computational speed.
[0068] According to an embodiment of the present application, the speech processing method is applied to a speech processing model, which includes: a feature extraction module, a feature fusion module and an acoustic wake-up module.
[0069] For example, Figure 3 As shown, the speech processing model 300 may include a feature extraction module (such as a first feature extraction module 310 and a second feature extraction module 320), a feature fusion module 330, and an acoustic wakeup module. The acoustic wakeup module may further include a gate 360, a fourth feature extraction module 370, and an acoustic classifier 380.
[0070] According to one embodiment of the present application, obtaining first acoustic feature information based on multiple audio signals includes: using a feature extraction module to extract feature information corresponding to the multiple audio signals and feature information corresponding to the original audio signals collected by one or more audio collection devices among at least two audio collection devices; using a feature fusion module 330 to fuse the feature information corresponding to the multiple audio signals and the feature information corresponding to the original audio signals to obtain the first acoustic feature information.
[0071] Specifically, the first acoustic feature information can be obtained based on the beamformed multi-channel audio signal and the original audio signal collected by one or more audio collection devices of the at least two audio collection devices. Here, the beamformed multi-channel audio signal can be regarded as the main feature, and the original audio signal collected by one or more audio collection devices can be regarded as the auxiliary feature. For example, when there are N audio collection devices, the original audio signal collected by one or M audio collection devices can be used as the auxiliary feature, where 1<M and ≤N.
[0072] The speech processing method provided in this embodiment uses the original audio signals collected by one or more audio acquisition devices as auxiliary features, which can avoid the impact of abnormalities in the beamforming algorithm on the sound source localization results. That is, the original audio signal can avoid large errors in the first acoustic feature information when an abnormality occurs in the beamforming algorithm, thereby avoiding excessive deviations between the sound source localization results and the actual situation. In addition, when the original audio signal collected by one audio acquisition device is used as an auxiliary feature, the accuracy of sound source localization can be improved while ensuring a high computing speed. When the original audio signals collected by multiple audio acquisition devices are used as auxiliary features, the accuracy of sound source localization can be further improved.
[0073] According to one embodiment of the present application, obtaining first acoustic feature information based on multiple audio signals includes: using a feature extraction module to extract feature information corresponding to the multiple audio signals; using a feature fusion module 330 to fuse the feature information corresponding to the multiple audio signals and the phase difference between the original audio signals collected by any two audio collection devices among at least two audio collection devices to obtain the first acoustic feature information.
[0074] Specifically, the first acoustic feature information can be obtained based on the beamformed multi-channel audio signal and the phase difference between the original audio signals collected by any two audio collection devices of at least two audio collection devices. Here, the beamformed multi-channel audio signal can be regarded as the main feature, and the phase difference between the original audio signals collected by any two audio collection devices can be regarded as the auxiliary feature. The phase difference between the original audio signals collected by the two audio collection devices can reflect the position distribution relationship between the two audio collection devices, and the direction of the sound source can be estimated based on the phase difference. The multi-channel audio signal contains implicit features about the direction of the sound source, while the phase difference contains explicit features about the direction of the sound source relative to the multi-channel audio signal. Therefore, the phase difference can further provide supplementary information for determining the direction of the sound source.
[0075] For example, when there are three audio acquisition devices, the phase difference (one phase difference) between the original audio signals collected by two of the audio acquisition devices can be used as an auxiliary feature, or the phase difference (three phase differences) between the original audio signals collected by every two audio acquisition devices among the three can be used as an auxiliary feature. Similarly, two phase differences can also be obtained as auxiliary features. When one phase difference is used as an auxiliary feature, the accuracy of sound source localization can be improved while ensuring a high computing speed. When multiple phase differences are used as auxiliary features, the accuracy of sound source localization can be further improved.
[0076] The speech processing method provided in this embodiment uses the phase difference between the original audio signals collected by any two audio collection devices as an auxiliary feature, which can further improve the accuracy of sound source localization.
[0077] According to one embodiment of the present application, obtaining first acoustic feature information based on multiple audio signals includes: using a feature extraction module to extract feature information corresponding to the multiple audio signals and feature information corresponding to the amplitude difference between the original audio signals collected by any two audio collection devices among at least two audio collection devices; using a feature fusion module 330 to fuse the feature information corresponding to the multiple audio signals and the feature information corresponding to the amplitude difference to obtain the first acoustic feature information.
[0078] Specifically, the first acoustic feature information can be obtained based on the beamformed multi-channel audio signal and the amplitude difference between the original audio signals collected by any two audio collection devices of at least two audio collection devices. Here, the beamformed multi-channel audio signal can be regarded as the main feature, and the amplitude difference between the original audio signals collected by any two audio collection devices can be regarded as the auxiliary feature. When the position distributions of the two audio collection devices are different, there will be an amplitude difference between the original audio signals collected by the two audio collection devices, that is, the amplitude difference can reflect the position distribution relationship between the two audio collection devices. The direction of the sound source can be estimated based on the amplitude difference. The multi-channel audio signal contains implicit features about the direction of the sound source, while the amplitude difference contains explicit features about the direction of the sound source relative to the multi-channel audio signal. Therefore, the amplitude difference can further provide supplementary information for determining the direction of the sound source.
[0079] For example, when there are three audio acquisition devices, the amplitude difference (one amplitude difference) between the original audio signals collected by two of the audio acquisition devices can be used as an auxiliary feature, or the amplitude difference (three amplitude differences) between the original audio signals collected by every two audio acquisition devices among the three can be used as an auxiliary feature. Similarly, two amplitude differences can also be obtained as auxiliary features. When one amplitude difference is used as an auxiliary feature, the accuracy of sound source localization can be improved while ensuring a high computing speed. When multiple amplitude differences are used as auxiliary features, the accuracy of sound source localization can be further improved.
[0080] The speech processing method provided in this embodiment uses the amplitude difference between the original audio signals collected by any two audio collection devices as an auxiliary feature, which can further improve the accuracy of sound source localization.
[0081] According to an embodiment of the present application, obtaining first acoustic feature information based on multiple audio signals includes: using a feature extraction module to extract feature information corresponding to the multiple audio signals, feature information corresponding to the original audio signals collected by one or more audio collection devices among at least two audio collection devices, and feature information corresponding to the amplitude difference between the original audio signals collected by any two audio collection devices among the at least two audio collection devices; using a feature fusion module 330 to fuse the feature information corresponding to the multiple audio signals, the feature information corresponding to the original audio signals, the feature information corresponding to the amplitude difference, and the phase difference between the original audio signals collected by any two audio collection devices among the at least two audio collection devices to obtain the first acoustic feature information.
[0082] Specifically, the first acoustic feature information can be obtained based on multi-channel audio signals, original audio signals collected by one or more audio acquisition devices, the phase difference between any two original audio signals collected by any two audio acquisition devices, and the amplitude difference between any two original audio signals collected by any two audio acquisition devices. Here, the beamformed multi-channel audio signals can be considered the primary feature, while the original audio signals, phase difference, and amplitude difference can be considered auxiliary features. The original audio signals can prevent large errors in the first acoustic feature information when anomalies occur in the beamforming algorithm, thereby preventing the sound source localization results from deviating significantly from the actual situation. The phase difference and amplitude difference can reflect the positional distribution relationship between the two audio acquisition devices and can be used to estimate the sound source direction. Multi-channel audio signals contain implicit features related to the sound source direction, while the phase difference and amplitude difference contain explicit features related to the sound source direction relative to the multi-channel audio signals. Therefore, the phase difference and amplitude difference can further provide supplementary information for determining the sound source direction.
[0083] For example, when there are three audio acquisition devices, the original audio signals collected by one or more of the audio acquisition devices, the phase difference (one phase difference) between the original audio signals collected by any two audio acquisition devices, and the amplitude difference (one amplitude difference) between the original audio signals collected by any two audio acquisition devices can be used as auxiliary features, or the original audio signals collected by one or more of the three audio acquisition devices, the phase difference (three phase differences) between the original audio signals collected by every two audio acquisition devices, and the amplitude difference (three amplitude differences) between the original audio signals collected by every two audio acquisition devices can be used as auxiliary features. Of course, the number of original audio signals, phase differences, and amplitude differences in the auxiliary features can be set as needed. When the number of original audio signals, phase differences, and amplitude differences in the auxiliary features is small, the accuracy of sound source localization can be improved to a certain extent while ensuring a high computing speed. When the number of original audio signals, phase differences, and amplitude differences in the auxiliary features is large, the accuracy of sound source localization can be further improved.
[0084] The speech processing method provided in this embodiment obtains first acoustic feature information by combining multiple audio signals, original audio signals, phase differences between original audio signals collected by audio acquisition devices, and amplitude differences between original audio signals collected by audio acquisition devices, which can further improve the accuracy of sound source positioning.
[0085] Figure 3 FIG. 3 is a structural diagram of a speech processing model 300 provided by an exemplary embodiment of the present application. The speech processing model 300 can be used to perform the above-mentioned speech processing method, that is, to locate the direction of the sound source. Figure 3 As shown, the speech processing model 300 includes a first feature extraction module 310 , a second feature extraction module 320 , a feature fusion module 330 , a third feature extraction module 340 and a sound source classifier 350 .
[0086] The speech processing model 300 is provided with a plurality of first feature extraction modules 310, and each audio signal, each original audio signal and each amplitude difference corresponds to a first feature extraction module 310. The first feature extraction module 310 can be used to extract shallow features of the input data. The speech processing model 300 is provided with a plurality of second feature extraction modules 320, and each audio signal and each original audio signal corresponds to a second feature extraction module 320, and all amplitude differences can correspond to one or more second feature extraction modules 320. The output of the first feature extraction module 310 can be used as the input of the second feature extraction module 320, and the second feature extraction module 320 can be used to extract deep features of the input data. The second feature extraction module 320 can be a deep neural network (DNN), a long short-term memory (LSTM) network, a convolutional neural network (CNN) or a Transformer deep learning architecture, etc., and the first feature extraction module 310 is similar to it.
[0087] The output and phase difference of the second feature extraction module 320 can be used as the input of the feature fusion module 330. The feature fusion module 330 can be a concat module that can splice the input feature information. The third feature extraction module 340 can further process the spliced feature information to facilitate the classification process of the subsequent classifier. The third feature extraction module 340 can be a feedforward neural network or other suitable network. The classification result of the classifier (sound source classifier 350) can be a three-dimensional vector (spatial position information). Here, the number of classification nodes of the classifier is 3, and the classifier can adopt a discriminative loss function, such as cross entropy. The parameters of multiple second feature extraction modules 320 can be the same or different. When the parameters of multiple second feature extraction modules 320 are the same, the purpose of reducing the number of parameters can be achieved.
[0088] Taking two audio capture devices and three preset sound source directions as an example, the specific process of the speech processing model 300 in locating the sound source direction is described. Specifically, the speech processing model 300 can use multiple audio signals (one-channel, two-channel, and three-channel audio signals), the original audio signals collected by any audio capture device, the phase difference between the original audio signals collected by two audio capture devices, and the amplitude difference between the original audio signals collected by two audio capture devices as input. The first feature extraction module 310 can perform a Fourier transform on the input data, such as extracting frequency domain features. The amplitude difference can be the difference between the frequency domain features of the original audio signals collected by the audio capture devices. In one implementation, each first feature extraction module 310 can output a 40-dimensional filter bank feature vector. For two audio capture devices, the phase difference can be represented by a vector with a dimension of 1. If the second feature extraction module 320 does not change the dimensionality of the input feature information, the feature information spliced by the feature fusion module 330 is a vector with a dimension of 201.
[0089] Optionally, a first feature fusion module can be provided between the first feature extraction module 310 and the second feature extraction module 320, and a second feature fusion module can be provided between the second feature extraction module 320 and the third feature extraction module 340. That is, a second feature extraction module 320 can be provided in the speech processing model. The output of the first feature extraction module 310 can serve as the input of the first feature fusion module. The first feature fusion module splices the input feature information, and the second feature extraction module 320 can further process the spliced feature information. The output and phase difference of the second feature extraction module 320 can serve as the input of the second feature fusion module. The third feature extraction module 340 and the sound source classifier 350 can be connected after the second feature fusion module.
[0090] Taking two audio acquisition devices and three preset sound source directions as an example, each first feature extraction module 310 can output a 40-dimensional filter bank feature vector, and the first feature fusion module can splice the four 40-dimensional vectors (corresponding to the three-way audio signal, the original audio signal, and the amplitude difference) to obtain a 200-dimensional vector. If the second feature extraction module does not change the dimension of the input feature information, the second feature fusion module can splice the output of the second feature extraction module and the phase difference to obtain a vector with a dimension of 201. Of course, if the phase difference is omitted in the input data, the second feature fusion module can be omitted in the speech processing model.
[0091] According to an embodiment of the present application, the speech processing method further includes: using an acoustic wake-up module to identify a wake-up word based on the first acoustic feature information.
[0092] Specifically, after locating the direction of the sound source, the wake-up word can be further identified based on the first acoustic feature information. In one implementation, when the located sound source direction meets the preset conditions, the step of identifying the wake-up word based on the first acoustic feature information can be performed, otherwise it will not be performed. Here, the preset condition may be that the sound source direction is the main driving direction. After identifying the wake-up word, relevant operations can be performed according to the instructions corresponding to the wake-up word. The voice processing process can perform corresponding operations according to the voice instructions of the user in the main driving direction, and will not respond to the voice instructions of users in other directions.
[0093] The speech processing method provided in this embodiment first locates the direction of the sound source and then identifies the wake-up word. The sound source localization result can be used as a prerequisite for wake-up word recognition. The wake-up word recognition process is only started when the sound source localization result meets the preset conditions. This can avoid confusion in the control process of the interactive device.
[0094] According to an embodiment of the present application, using an acoustic wake-up module to identify a wake-up word based on the first acoustic feature information includes: using the acoustic wake-up module to combine spatial position information and the first acoustic feature information to identify the wake-up word.
[0095] Specifically, after locating the direction of the sound source, the first acoustic feature information can be used to identify the wake-up word according to the direction of the sound source. That is, the wake-up word can be identified by combining the spatial position information and the first acoustic feature information. For example, the acoustic features in the direction of the sound source can be extracted from the first acoustic feature information to perform wake-up word recognition.
[0096] In one implementation, the spatial position information is a three-dimensional vector (a1, a2, a3), each dimension of which corresponds to a beamformed audio signal corresponding to a preset sound source direction. The first acoustic feature information can be a concatenated feature vector Contains the acoustic characteristics of three audio signals Acoustic characteristics of the original audio signal Phase difference Sum Amplitude Difference in, to Can be vectors respectively. The direction of the sound source is the maximum value a in the spatial position information i (i=1, 2 or 3) corresponds to the preset sound source direction, which can be based on the first acoustic feature information and a i Corresponding In other implementations, the spatial position information (a1, a2, a3) can be combined with the acoustic features. The second acoustic feature information is obtained by multiplication, and the wake-up word recognition is performed based on the second acoustic feature information.
[0097] The speech processing method provided in this embodiment can extract information in the direction of the sound source from the first acoustic feature information by combining spatial position information for wake-up word recognition. This can effectively filter out noise in other directions and achieve the effect of acoustic purification, thereby improving the efficiency and accuracy of wake-up word recognition.
[0098] According to one embodiment of the present application, an acoustic wake-up module is used to identify a wake-up word in combination with spatial position information and first acoustic feature information, including: using the gate 360 of the acoustic wake-up module to perform gate selection on the first acoustic feature information based on the spatial position information to obtain second acoustic feature information; using the acoustic wake-up module to process the second acoustic feature information to obtain an acoustic score, wherein the acoustic wake-up module is obtained by adversarial training using a gradient flip layer, and the gradient flip layer points to an energy classifier; and identifying the wake-up word based on the acoustic score.
[0099] Specifically, the spatial position information can characterize the probability that each preset sound source direction is the target sound source direction. The probability that each preset sound source direction is the target sound source direction can be used as the weight of the information corresponding to each audio signal in the first acoustic feature information, and the information corresponding to each audio signal in the first acoustic feature information is fused according to the weight to obtain the second acoustic feature information. It is equivalent to using the spatial position information as the gate of the acoustic wake-up module to perform gated selection on the first acoustic feature information. The Gradient Reversal Layer (GRL) can point to the energy classifier and pass the gradient back to the gate of the acoustic wake-up module, so that the acoustic wake-up module after training is insensitive to energy information, and the output information corresponding to the input information of various energies is distributed consistently, which can solve the problem of low acoustic scores corresponding to audio signals with smaller sounds. Here, the energy in the energy classifier can refer to the volume of the audio signal.
[0100] In one implementation, the spatial position information is a three-dimensional vector (a1, a2, a3), each dimension of which corresponds to a beamformed audio signal corresponding to a preset sound source direction. The first acoustic feature information can be a concatenated feature vector Contains the acoustic characteristics of three audio signals Acoustic characteristics of the original audio signal Phase difference Sum Amplitude Difference in, to Using the spatial position information as the gate of the acoustic wake-up module, the first acoustic feature information is gated and selected to obtain the second acoustic feature information, which is a vector.
[0101] In other implementations, the second acoustic feature information is a vector Here, the three-dimensional vector (a1, a2, a3) is the spatial position information, and a4, a5, and a6 can be parameters obtained through model training or preset parameters.
[0102] The second acoustic feature information is processed using the acoustic wake-up module to obtain an acoustic score. The classification target of the acoustic wake-up module may be an acoustic unit, which may be a phoneme, a syllable, or a word. The acoustic score may be represented by a vector or other forms. For example, the acoustic score includes the probability of the acoustic unit, and the acoustic score may be a one-dimensional or multi-dimensional vector, where the value on each dimension represents the probability of the corresponding target acoustic unit. The wake-up word can be identified based on the acoustic score.
[0103] The speech processing method provided in this embodiment uses the first acoustic feature information as the input of the acoustic wake-up module and the spatial position information as the gate of the acoustic wake-up module, so that the acoustic wake-up module can obtain the second acoustic feature information by gating the first acoustic feature information based on the gate, and perform wake-up word recognition based on the second acoustic feature information. In this way, the effect of acoustic purification can be achieved and the accuracy of wake-up word recognition can be improved. Furthermore, by using the gradient flip layer to train the acoustic wake-up module, the trained acoustic wake-up module can be made insensitive to energy information, which can effectively solve the problem of low acoustic scores corresponding to audio signals with smaller sounds.
[0104] According to an embodiment of the present application, identifying a wake-up word based on an acoustic score includes: inputting the acoustic score into a decoding network to identify the wake-up word.
[0105] Specifically, the acoustic unit can be a word, so the wake-up word can be directly identified based on the acoustic score. For example, the target word with the highest probability can be used as the wake-up word to be identified. This recognition method is simple and efficient.
[0106] Optionally, the acoustic unit can be a phoneme or a syllable. In this case, the acoustic score can be input into the decoding network, which can then obtain a candidate word sequence based on the acoustic score. For example, the word corresponding to the target acoustic unit with a probability greater than or equal to a threshold can be used as a candidate word. The threshold can be set as needed. Furthermore, the confidence module can perform a confidence judgment on each word in the candidate word sequence to obtain a confidence score corresponding to each word. Based on the confidence score, it is determined whether the device is awakened. If awakened, the corresponding operation can be further performed based on the recognized wake-up word.
[0107] Of course, other conventional methods can also be used to identify wake-up words based on acoustic scores.
[0108] The speech processing method provided in this embodiment can improve the accuracy of the recognition results by inputting the acoustic score into the decoding network to recognize the wake-up word.
[0109] Figure 4 Shown is a flow chart of a speech processing method provided by another exemplary embodiment of the present application. Figure 4 The embodiment is Figure 2 In order to avoid repetition, the similarities in the examples of the embodiments will not be described again. Figure 4 The method can be Figure 3 The speech processing model 300 is used to perform the above operations. Figure 4 As shown, the speech processing method includes the following contents.
[0110] 410: Obtain first acoustic feature information based on multiple audio signals.
[0111] Beamforming processing is performed on original audio signals collected by at least two audio collection devices according to multiple preset sound source directions to obtain multi-channel audio signals.
[0112] Taking two audio capture devices and three preset sound source directions as an example, beamforming the original audio signals collected by the two audio capture devices according to the three preset sound source directions can generate multiple audio signals: one audio signal, two audio signals, and three audio signals.
[0113] The first acoustic feature information integrates the features of multiple audio signals. Specifically, Figure 3As shown, the multi-channel audio signals, the original audio signals collected by either of the two audio collection devices, the phase difference between the original audio signals collected by the two audio collection devices, and the amplitude difference between the original audio signals collected by the two audio collection devices can serve as inputs to the speech processing model 300. Based on these inputs, the speech processing model 300 can obtain first acoustic feature information. For example, the multi-channel audio signals, the original audio signals, and the amplitude difference can be respectively input into corresponding first feature extraction modules 310, and each first feature extraction module 310 can output a 40-dimensional filter bank feature vector. Each first feature extraction module 310 can be connected to a second feature extraction module 320, and the output of the first feature extraction module 310 can serve as the input to the second feature extraction module 320, which can be used to extract deep features of the input data. The output of the second feature extraction module 320 can also be a 40-dimensional feature vector.
[0114] For two audio acquisition devices, the phase difference can be a vector with a dimension of 1. The output of the second feature extraction module 320 and the phase difference can serve as input to the feature fusion module 330. The feature fusion module 330 can be a concat module that can concatenate the input feature information. The feature information (first acoustic feature information) concatenated by the feature fusion module 330 is a vector with a dimension of 201.
[0115] 420: Classify and process the first acoustic feature information to obtain spatial position information, where the spatial position information is used to represent the direction of the sound source.
[0116] like Figure 3 As shown, the third feature extraction module 340 can further process the spliced feature information, and the output of the third feature extraction module 340 can be used as the input of the sound source classifier 350. The classification result of the sound source classifier 350 can be a three-dimensional vector (spatial position information).
[0117] 430: Identify the wake-up word by combining the spatial position information and the first acoustic feature information.
[0118] Specifically, the gate 360 of the acoustic wake-up module can gate and select the first acoustic feature information based on the spatial position information to obtain the second acoustic feature information; the second acoustic feature information is processed by the acoustic wake-up module to obtain an acoustic score; and the wake-up word is identified based on the acoustic score. Here, the acoustic wake-up module is obtained by adversarial training using a gradient flip layer, and the gradient flip layer points to an energy classifier. The specific process of identifying the wake-up word based on the acoustic score can be found in the description of the above embodiment.
[0119] like Figure 3As shown, the second acoustic feature information can be used as input to the fourth feature extraction module 370, which can further process the second acoustic feature information (such as extracting more advanced acoustic features) to facilitate the classification process of the subsequent acoustic classifier 380. The acoustic classifier 380 can output an acoustic score. The acoustic wake-up module in this embodiment may include Figure 3 The gate 360, the fourth feature extraction module 370, the acoustic classifier 380, the gradient flip layer and the energy classifier in the acoustic wake-up module. In one embodiment, the acoustic wake-up module may further include a decoding network and a confidence module, such as the acoustic classifier 380 may be connected to the decoding network, and the decoding network may be connected to the confidence module. The gradient flip layer and the energy classifier are set in the acoustic wake-up module, which can make the acoustic wake-up module insensitive to energy information, effectively solve the problem of low acoustic scores corresponding to audio signals with smaller sounds, and improve the accuracy of acoustic recognition (acoustic wake-up). In other embodiments, the gradient flip layer and the energy classifier can be omitted in the acoustic wake-up module, which can simplify the model structure.
[0120] Figure 6 Shown is a flowchart of a method for training a speech processing model provided by an exemplary embodiment of the present application. Figure 6 The method can be executed by a computing device. Figure 6 As shown, the training method of the speech processing model includes the following contents.
[0121] 610: Obtain first sample acoustic feature information based on sample data using a feature extraction module and a feature fusion module, wherein the sample data includes multiple channels of sample audio signals, the multiple channels of sample audio signals are obtained based on original sample audio signals collected by at least two sample audio collection devices, and the first sample acoustic feature information fuses features of the multiple channels of sample audio signals.
[0122] The sample audio acquisition device may be an audio acquisition device that acquires original sample audio signals.
[0123] 620: Obtain spatial position information corresponding to the sample data based on the first sample acoustic feature information using a sound source classifier.
[0124] 630: Using a training module to train the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
[0125] like Figure 5As shown, the speech processing model 500 may include a feature extraction module (such as a first feature extraction module 510 and a second feature extraction module 520), a feature fusion module 530, a sound source classifier 550, an acoustic wake-up module, and a training module (not shown in the figure). The speech processing model 500 is similar to the speech processing model 300, and the similarities can be referenced to each other. The acoustic wake-up module may further include a gate 560, a fourth feature extraction module 570, an acoustic classifier 580, a gradient flip layer 590, and an energy classifier 595.
[0126] Specifically, during the training of a speech processing model, sample data can be used as input to the model, and the sound source annotation information corresponding to the sample data can be used as output to train the model. The sound source annotation information can be obtained by pre-labeling the sample data and can indicate the direction of the target sound source.
[0127] After the sample data is input into the speech processing model, the speech processing model can obtain first sample acoustic feature information based on the sample data, and then obtain spatial location information corresponding to the sample data based on the first sample acoustic feature information. The speech processing model can implement a training (learning) process based on the difference between the spatial location information actually analyzed and the pre-labeled sound source labeling information. For example, the speech processing model can implement the training process using a discriminative loss function (such as cross entropy) or other loss function.
[0128] The multi-channel sample audio signal is similar to the multi-channel audio signal in the above-mentioned speech processing method, and the first sample acoustic feature information is similar to the first acoustic feature information in the above-mentioned speech processing method. The speech processing model obtained by the training method of the speech processing model of the embodiment of the present application can be used to execute the speech processing method of any of the above-mentioned embodiments. Some terms and steps in the training method of the speech processing model are similar to those in the above-mentioned speech processing method, and the similarities can be referenced to each other.
[0129] An embodiment of the present application provides a method for training a speech processing model. The method obtains first sample acoustic feature information that incorporates the features of multiple sample audio signals, then obtains spatial position information corresponding to the sample data based on the first sample acoustic feature information, and trains the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data. Because the features of multiple sample audio signals are fused to obtain the first sample acoustic feature information, rather than processing each sample audio signal separately, the method for training the speech processing model can reduce memory usage, increase computing speed, reduce cascade errors, and improve the accuracy of sound source localization.
[0130] According to an embodiment of the present application, the multi-channel sample audio signals are obtained by performing beamforming processing on original sample audio signals collected by at least two sample audio collection devices according to multiple preset sound source directions.
[0131] Specifically, the multiple preset sound source directions may be multiple sound source directions that the speech processing model is expected to locate. For example, the original sample audio signals collected by at least two audio acquisition devices may be directionally enhanced according to the multiple preset sound source directions to obtain multi-channel beamformed sample audio signals. Of course, other methods may also be used to obtain multi-channel sample audio signals enhanced in multiple preset sound source directions based on the original sample audio signals collected by at least two audio acquisition devices.
[0132] Multiple preset sound source directions can be set according to actual needs. The embodiment of the present application does not limit the specific number and specific directions of the multiple preset sound source directions.
[0133] The speech processing model training method provided in this embodiment performs beamforming processing on raw sample audio signals collected by at least two audio acquisition devices according to multiple preset sound source directions to obtain multi-channel sample audio signals. This can enhance audio signals in the relevant preset sound source directions and suppress audio signals in other directions, thereby improving the accuracy of the model's sound source localization. In addition, the speech processing model training method provided in this embodiment can reasonably set the number of preset sound source directions as needed, thereby improving the adaptability of the method and, to a certain extent, increasing the calculation speed.
[0134] According to one embodiment of the present application, the sample data also includes at least one of the following items: original sample audio signals collected by one or more sample audio collection devices among at least two sample audio collection devices, sample phase difference between original sample audio signals collected by any two sample audio collection devices among at least two sample audio collection devices, and sample amplitude difference between original sample audio signals collected by any two sample audio collection devices among at least two sample audio collection devices.
[0135] Specifically, the beamformed multi-channel audio sample signals can be considered the primary features, while the original audio sample signals collected by one or more audio acquisition devices can be considered the auxiliary features. The original audio sample signals can prevent large errors in the first sample acoustic feature information when anomalies occur in the beamforming algorithm, thereby preventing the sound source localization results from deviating significantly from the actual situation.
[0136] The sample phase difference between the original audio samples collected by any two audio acquisition devices can be considered an auxiliary feature. This sample phase difference between the original audio samples collected by two audio acquisition devices can reflect the positional distribution relationship between the two audio acquisition devices, and the sound source direction can be estimated based on this sample phase difference. Multi-channel audio sample signals contain implicit features related to the sound source direction, while the sample phase difference contains explicit features related to the sound source direction relative to the multi-channel audio sample signals. Therefore, the sample phase difference can further provide supplementary information for determining the sound source direction.
[0137] The sample amplitude difference between the original sample audio signals collected by any two audio acquisition devices can be regarded as an auxiliary feature. When the position distribution of the two audio acquisition devices is different, there will be a sample amplitude difference between the original sample audio signals collected by the two audio acquisition devices, that is, the sample amplitude difference can reflect the position distribution relationship between the two audio acquisition devices. The direction of the sound source can be estimated based on the sample amplitude difference. The multi-channel sample audio signals contain implicit features about the direction of the sound source, while the sample amplitude difference contains explicit features about the direction of the sound source relative to the multi-channel sample audio signals. Therefore, the sample amplitude difference can further provide supplementary information for determining the direction of the sound source.
[0138] The training method of the speech processing model provided in this embodiment can further improve the accuracy of the model sound source positioning by combining the main features and auxiliary features to obtain the first sample acoustic feature information.
[0139] According to one embodiment of the present application, the training method also includes: using the acoustic wake-up module to obtain acoustic information based on the acoustic feature information of the first sample; and using the training module to train the acoustic wake-up module according to the acoustic information and the acoustic annotation information corresponding to the sample data.
[0140] Specifically, the acoustic annotation information may be obtained by annotating the sample data in advance, and may include syllables, phonemes and / or words corresponding to the wake-up word in the sample audio signal.
[0141] The speech processing model can be used for sound source localization, and further, can also be used to identify wake-up words. For example, the acoustic wake-up module obtains acoustic information based on the acoustic feature information of the first sample. The acoustic wake-up module can implement the training (learning) process based on the difference between the acoustic information obtained by its actual analysis and the pre-labeled acoustic annotation information. For example, the acoustic wake-up module can use a discriminative loss function (such as cross entropy) or other loss functions to implement the training process. In one embodiment, the acoustic information can be an acoustic score.
[0142] In one implementation, the trainable speech processing model executes the step of training the acoustic wake-up module based on the acoustic feature information of the first sample and the acoustic annotation information corresponding to the sample data when the located sound source direction meets the preset conditions, otherwise it is not executed. Here, the preset condition can be that the sound source direction is the main driving direction. After recognizing the wake-up word, the relevant operation can be performed according to the instruction corresponding to the wake-up word. The speech processing model can perform corresponding operations according to the voice instructions of the user in the main driving direction, and does not respond to the voice instructions of users in other directions.
[0143] The speech processing model trained using this method can first locate the direction of the sound source and then identify the wake-up word. The sound source localization result is used as a prerequisite for wake-up word recognition. The wake-up word recognition process is initiated only when the sound source localization result meets the preset conditions. This can avoid confusion in the control process of the interactive device.
[0144] According to an embodiment of the present application, the acoustic information is obtained based on the first sample acoustic feature information using the acoustic wake-up module, including: obtaining the acoustic information based on the first sample acoustic feature information and the spatial position information corresponding to the sample data using the acoustic wake-up module.
[0145] Specifically, the acoustic wake-up module can obtain acoustic information in combination with the spatial position information corresponding to the sample data and the first sample acoustic feature information to extract the acoustic information in the direction of the sound source from the first sample acoustic feature information. For example, the spatial position information may indicate the direction of the sound source, and information related to the sound source direction is extracted from the first sample acoustic feature information to train the acoustic wake-up module. The acoustic annotation information in this embodiment may include syllables, phonemes and / or words corresponding to the wake-up word in the audio signal in the direction of the target sound source.
[0146] The speech processing model trained by this training method can combine the spatial position information to extract information in the direction of the sound source from the first acoustic feature information for wake-up word recognition. This can effectively filter out noise in other directions and achieve the effect of acoustic purification, thereby improving the efficiency and accuracy of wake-up word recognition.
[0147] According to one embodiment of the present application, a gradient flip layer 590, an energy classifier 595, an acoustic classifier 580 and a gate 560 are provided in the acoustic wake-up module, and the acoustic information is obtained based on the first sample acoustic feature information and the spatial position information corresponding to the sample data using the acoustic wake-up module, including: obtaining the second sample acoustic feature information based on the first sample acoustic feature information and the spatial position information corresponding to the sample data using the gate 560; obtaining the acoustic information based on the second sample acoustic feature information using the acoustic classifier 580, wherein the training method also includes: performing adversarial training on the acoustic wake-up module based on the second sample acoustic feature information and the energy labeling information using the gradient flip layer 590, and the gradient flip layer points to the energy classifier 595.
[0148] Specifically, the spatial position information can represent the probability that each preset sound source direction is the target sound source direction. The probability that each preset sound source direction is the target sound source direction can be used as the weight of the information corresponding to each channel of sample audio signals in the first sample acoustic feature information. The information corresponding to each channel of sample audio signals in the first sample acoustic feature information is fused according to the weight to obtain the second sample acoustic feature information. This is equivalent to using the spatial position information as the gate of the acoustic wake-up module to perform gated selection on the first sample acoustic feature information.
[0149] The gradient reversal layer 590 (GRL) can point to the energy classifier 595 and pass the gradient back to the gate 560 of the acoustic wake-up module, so that the acoustic wake-up module after training is insensitive to energy information and the output information corresponding to the input information of various energies is distributed consistently. This can solve the problem of low acoustic scores corresponding to audio signals with smaller sounds and improve the robustness of the model. Here, the energy in the energy classifier can refer to the volume of the audio signal. The energy labeling information can be obtained by labeling the sample data according to the volume.
[0150] The acoustic wakeup module also includes an acoustic classifier 580, which can obtain acoustic information based on the acoustic feature information of the second sample. The acoustic wakeup module is trained based on the difference between the acoustic information and the acoustic annotation information. In addition, the gradient flip layer 590 pointing to the energy classifier 595 returns the gradient, making the acoustic wakeup module insensitive to energy information.
[0151] The speech processing model trained using this method can achieve acoustic purification and improve the accuracy of wake-up word recognition. Furthermore, by using a gradient flip layer to train the acoustic wake-up module, the trained acoustic wake-up module can be made insensitive to energy information, effectively solving the problem of low acoustic scores corresponding to quiet audio signals.
[0152] Figure 7 Shown is a flowchart of a method for training a speech processing model provided by another exemplary embodiment of the present application. Figure 7 The embodiment is Figure 6 In order to avoid repetition, the similarities in the examples of the embodiments will not be described again. Figure 7 The method can be used to train Figure 5 The speech processing model 500 is shown. Figure 7 As shown, the training method of the speech processing model includes the following contents.
[0153] 710: Obtain first sample acoustic feature information based on sample data, where the sample data includes multiple channels of sample audio signals, and the first sample acoustic feature information integrates features of the multiple channels of sample audio signals.
[0154] Beamforming processing is performed on original sample audio signals collected by at least two audio collection devices according to multiple preset sound source directions to obtain multi-channel sample audio signals.
[0155] Take two audio acquisition devices and three preset sound source directions as an example. Performing beamforming processing on the original sample audio signals collected by the two audio acquisition devices according to the three preset sound source directions can obtain multiple channels of sample audio signals: one channel of sample audio signal, two channels of sample audio signal, and three channels of sample audio signal. The multiple channels of sample audio signals, the original sample audio signals collected by any of the two audio acquisition devices, the sample phase difference between the original sample audio signals collected by the two audio acquisition devices, and the sample amplitude difference between the original sample audio signals collected by the two audio acquisition devices can be used as Figure 5 The input (sample data) of the speech processing model 500 is shown. Based on these inputs, the speech processing model 500 can obtain first sample acoustic feature information through the first feature extraction module 510, the second feature extraction module 520 and the feature fusion module 530.
[0156] 720: Obtain spatial position information corresponding to the sample data based on the first sample acoustic feature information, and train the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
[0157] like Figure 5 As shown, the third feature extraction module 540 can further process the spliced feature information, and the output of the third feature extraction module 540 can be used as the input of the sound source classifier 550. The classification result of the sound source classifier 550 can be a three-dimensional vector (spatial position information).
[0158] 730: Train the acoustic wake-up module based on the first sample acoustic feature information, the spatial position information corresponding to the sample data, and the acoustic annotation information.
[0159] The speech processing model 500 is provided with an acoustic wake-up module, such as Figure 5 As shown, the acoustic wake-up module includes a gate 560 , a gradient flip layer 590 , an energy classifier 595 , a fourth feature extraction module 570 , and an acoustic classifier 580 .
[0160] The gate 560 is used to perform gated selection on the acoustic feature information of the first sample to obtain the acoustic feature information of the second sample. The output of the gate 560 can be used as the input of the gradient flip layer 590 and the fourth feature extraction module 570. The fourth feature extraction module 570 can extract more advanced acoustic features based on the second sample acoustic feature information. Based on the output of the fourth feature extraction module 570, the acoustic classifier 580 can obtain acoustic information. The acoustic wake-up module is trained based on the difference between the acoustic information and the acoustic annotation information. In addition, the gradient flip layer 590 pointing to the energy classifier 595 can pass the gradient back to the gate 560 of the acoustic wake-up module, so that the acoustic wake-up module is insensitive to energy information (such as making the acoustic wake-up module unable to learn energy information).
[0161] The third feature extraction module 540 and the sound source classifier 550 in the speech processing model 500 are used for sound source localization, and the acoustic wake-up module is used to obtain acoustic scores or recognize wake-up words. The speech processing model 500 provided in the embodiment of the present application has the ability to process multi-channel inputs, which can avoid the high computational complexity caused by the need to independently process the input of each channel in traditional methods; and the speech processing model 500 can learn the order information, energy information, and spatial position information between each input, thereby improving the robustness of the model and the accuracy of sound source localization and acoustic recognition.
[0162] The gradient flipping layer 590 in the speech processing model 500 may work during training but may not work during testing, for example, Figure 3 The speech processing model 300 shown in the test phase has the gradient flipping layer and the energy classifier disabled. Figure 3 The gradient flipping layer and energy classifier are not shown.
[0163] Exemplary devices
[0164] Figure 8 FIG. 8 is a schematic diagram of the structure of a speech processing device 800 provided by an exemplary embodiment of the present application. Figure 8 As shown, the speech processing device 800 includes: an acquisition module 810 and a classification module 820 .
[0165] Acquisition module 810 is configured to acquire first acoustic feature information based on multiple audio signals, where the first acoustic feature information incorporates features of the multiple audio signals, which are acquired based on original audio signals collected by at least two audio acquisition devices. Classification module 820 is configured to classify the first acoustic feature information to obtain spatial position information, which is used to characterize the direction of the sound source.
[0166] An embodiment of the present application provides a speech processing device that obtains multiple audio signals based on original audio signals collected by at least two audio acquisition devices, fuses the features of the multiple audio signals to obtain first acoustic feature information corresponding to the multiple audio signals, and then classifies and processes the first acoustic feature information to obtain spatial position information that can represent the direction of the sound source. Because the features of the multiple audio signals are fused to obtain the first acoustic feature information, rather than processing each audio signal separately, the speech processing device can reduce memory usage, increase computing speed, reduce cascade errors, and improve the accuracy of sound source localization.
[0167] According to an embodiment of the present application, the multi-channel audio signals are obtained by performing beamforming processing on original audio signals collected by at least two audio collection devices according to multiple preset sound source directions.
[0168] According to an embodiment of the present application, the acquisition module 810 is configured to acquire first acoustic feature information based on multiple audio signals and original audio signals collected by one or more of the at least two audio collection devices.
[0169] According to an embodiment of the present application, the acquisition module 810 is configured to acquire first acoustic feature information based on a phase difference between a multi-channel audio signal and original audio signals collected by any two audio collection devices among at least two audio collection devices.
[0170] According to an embodiment of the present application, the acquisition module 810 is configured to acquire first acoustic feature information based on an amplitude difference between the multi-channel audio signal and original audio signals collected by any two audio collection devices among at least two audio collection devices.
[0171] According to an embodiment of the present application, the acquisition module 810 is used to obtain first acoustic feature information based on multiple audio signals, original audio signals collected by one or more audio collection devices among at least two audio collection devices, a phase difference between the original audio signals collected by any two audio collection devices among the at least two audio collection devices, and an amplitude difference between the original audio signals collected by any two audio collection devices among the at least two audio collection devices.
[0172] According to an embodiment of the present application, the speech processing device 800 further includes a recognition module 830 for recognizing a wake-up word based on the first acoustic feature information.
[0173] According to an embodiment of the present application, the recognition module 830 is used to recognize the wake-up word by combining the spatial position information and the first acoustic feature information.
[0174] According to one embodiment of the present application, the recognition module 830 is used to: use spatial position information as the gate of the acoustic wake-up module to perform gate selection on the first acoustic feature information to obtain the second acoustic feature information; use the acoustic wake-up module to process the second acoustic feature information to obtain an acoustic score, wherein the acoustic wake-up module is obtained by adversarial training using a gradient flip layer, and the gradient flip layer points to an energy classifier; and identify the wake-up word according to the acoustic score.
[0175] According to an embodiment of the present application, the recognition module 830 is configured to input the acoustic score into the decoding network to recognize the wake-up word.
[0176] It should be understood that the operations and functions of the acquisition module 810, the classification module 820, and the identification module 830 in the above embodiment can refer to the above Figure 2 or Figure 4 To avoid repetition, the description of the speech processing method provided in the embodiment will not be repeated here.
[0177] Figure 9 FIG. 9 is a schematic diagram of a training device 900 for a speech processing model according to an exemplary embodiment of the present invention. Figure 9 As shown, the training device 900 includes: a first acquisition module 910 , a second acquisition module 920 and a training module 930 .
[0178] The first acquisition module 910 is configured to acquire first sample acoustic feature information based on sample data, wherein the sample data includes multiple sample audio signals, the multiple sample audio signals being acquired based on original sample audio signals acquired by at least two sample audio acquisition devices, and the first sample acoustic feature information incorporating the features of the multiple sample audio signals. The second acquisition module 920 is configured to acquire spatial position information corresponding to the sample data based on the first sample acoustic feature information. The training module 930 is configured to train a speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
[0179] An embodiment of the present application provides a training device for a speech processing model. The device obtains first sample acoustic feature information that incorporates the features of multiple sample audio signals, then obtains spatial position information corresponding to sample data based on the first sample acoustic feature information, and trains the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data. Because the features of multiple sample audio signals are fused to obtain the first sample acoustic feature information, rather than processing each sample audio signal separately, the training device for the speech processing model can reduce memory usage, increase computing speed, reduce cascade errors, and improve the accuracy of sound source localization.
[0180] According to an embodiment of the present application, the multi-channel sample audio signals are obtained by performing beamforming processing on original sample audio signals collected by at least two sample audio collection devices according to multiple preset sound source directions.
[0181] According to one embodiment of the present application, the sample data also includes at least one of the following items: original sample audio signals collected by one or more sample audio collection devices among at least two sample audio collection devices, sample phase difference between original sample audio signals collected by any two sample audio collection devices among at least two sample audio collection devices, and sample amplitude difference between original sample audio signals collected by any two sample audio collection devices among at least two sample audio collection devices.
[0182] According to one embodiment of the present application, an acoustic wake-up module is provided in the speech processing model, and the training module 930 is further used to train the acoustic wake-up module based on the acoustic feature information of the first sample and the acoustic annotation information corresponding to the sample data.
[0183] According to an embodiment of the present application, the training module 930 is used to train the acoustic wake-up module based on the first sample acoustic feature information, the spatial position information corresponding to the sample data, and the acoustic annotation information.
[0184] According to one embodiment of the present application, a gradient flip layer is provided in the acoustic wake-up module, and the training module 930 is used to: obtain the second sample acoustic feature information based on the first sample acoustic feature information and the spatial position information corresponding to the sample data; use the gradient flip layer to perform adversarial training on the acoustic wake-up module based on the second acoustic feature information and the acoustic labeling information, and the gradient flip layer points to the energy classifier.
[0185] It should be understood that the operations and functions of the first acquisition module 910, the second acquisition module 920 and the training module 930 in the above embodiment can refer to the above Figure 6 or Figure 7 To avoid repetition, the description of the training method of the speech processing model provided in the embodiment will not be repeated here.
[0186] Figure 10 FIG2 is a block diagram of an electronic device 1000 for executing a speech processing method or a speech processing model training method provided by an exemplary embodiment of the present application.
[0187] Reference Figure 10 The electronic device 1000 includes a processing component 1010, which further includes one or more processors, and a memory resource represented by a memory 1020 for storing instructions executable by the processing component 1010, such as an application. The application stored in the memory 1020 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1010 is configured to execute the instructions to perform the above-mentioned speech processing method or speech processing model training method.
[0188] The electronic device 1000 may further include a power supply component configured to perform power management of the electronic device 1000, a wired or wireless network interface configured to connect the electronic device 1000 to a network, and an input / output (I / O) interface. The electronic device 1000 may be operated based on an operating system stored in the memory 1020, such as Windows Server 2003. TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or similar.
[0189] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device 1000, enables the electronic device 1000 to execute a speech processing method or a method for training a speech processing model. The speech processing method includes: obtaining first acoustic feature information based on multiple audio signals, wherein the first acoustic feature information integrates the features of the multiple audio signals, and the multiple audio signals are obtained based on original audio signals collected by at least two audio acquisition devices; classifying and processing the first acoustic feature information to obtain spatial position information, and the spatial position information is used to characterize the direction of the sound source. The training method of the speech processing model includes: obtaining first sample acoustic feature information based on sample data, wherein the sample data includes multiple sample audio signals, and the multiple sample audio signals are obtained based on original sample audio signals collected by at least two sample audio acquisition devices, and the first sample acoustic feature information integrates the features of the multiple sample audio signals; obtaining spatial position information corresponding to the sample data based on the first sample acoustic feature information; and training the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
[0190] All of the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, and will not be described in detail here.
[0191] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0192] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0193] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0194] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0195] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0196] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program check codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0197] It should be noted that, in the description of this application, the terms "first," "second," "third," etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0198] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A speech processing method, characterized in that: include: Acquiring first acoustic feature information based on multiple audio signals, wherein the first acoustic feature information incorporates features of the multiple audio signals, and the multiple audio signals are acquired based on original audio signals acquired by at least two audio acquisition devices; The first acoustic feature information is classified and processed to obtain spatial position information, where the spatial position information is used to characterize the direction of the sound source. The first acoustic feature is obtained by fusing feature information corresponding to multiple audio signals and feature information corresponding to original audio signals collected by one or more of at least two audio collection devices.
2. The speech processing method according to claim 1, wherein: The speech processing method is applied to a speech processing model, which includes: a feature extraction module, a feature fusion module and an acoustic wake-up module.
3. The speech processing method according to claim 2, wherein: The acquiring first acoustic feature information based on the multi-channel audio signals includes: Extracting feature information corresponding to the multiple audio signals and feature information corresponding to original audio signals collected by one or more of the at least two audio collection devices using the feature extraction module; The feature fusion module is used to fuse the feature information corresponding to the multi-channel audio signals and the feature information corresponding to the original audio signal to obtain the first acoustic feature information.
4. The speech processing method according to claim 2, wherein: The acquiring first acoustic feature information based on the multi-channel audio signals includes: Extracting feature information corresponding to the multi-channel audio signals using the feature extraction module; The feature fusion module is used to fuse the feature information corresponding to the multi-channel audio signals and the phase difference between the original audio signals collected by any two audio collection devices among the at least two audio collection devices to obtain the first acoustic feature information.
5. The speech processing method according to claim 2, wherein: The acquiring first acoustic feature information based on the multi-channel audio signals includes: Extracting, by the feature extraction module, feature information corresponding to the multi-channel audio signals and feature information corresponding to an amplitude difference between original audio signals collected by any two of the at least two audio collection devices; The feature fusion module is used to fuse the feature information corresponding to the multi-channel audio signals and the feature information corresponding to the amplitude difference to obtain the first acoustic feature information.
6. The speech processing method according to claim 2, wherein: The acquiring first acoustic feature information based on the multi-channel audio signals includes: extracting, using the feature extraction module, feature information corresponding to the multiple audio signals, feature information corresponding to original audio signals collected by one or more of the at least two audio collection devices, and feature information corresponding to an amplitude difference between original audio signals collected by any two of the at least two audio collection devices; The feature fusion module is used to fuse feature information corresponding to the multiple audio signals, feature information corresponding to the original audio signals, feature information corresponding to the amplitude difference, and a phase difference between the original audio signals collected by any two audio collection devices among the at least two audio collection devices to obtain the first acoustic feature information.
7. The speech processing method according to claim 2, wherein: Also includes: The acoustic wake-up module is used to identify a wake-up word based on the first acoustic feature information.
8. The speech processing method according to claim 7, characterized in that: The identifying the wake-up word based on the first acoustic feature information by using the acoustic wake-up module includes: The acoustic wake-up module is used to identify the wake-up word in combination with the spatial position information and the first acoustic feature information.
9. The speech processing method according to claim 8, characterized in that: The identifying the wake-up word by using the acoustic wake-up module in combination with the spatial position information and the first acoustic feature information includes: Performing gated selection on the first acoustic feature information based on the spatial position information using the gating of the acoustic wake-up module to obtain second acoustic feature information; Processing the second acoustic feature information using the acoustic wake-up module to obtain an acoustic score, wherein the acoustic wake-up module is obtained by adversarial training using a gradient flip layer, and the gradient flip layer points to an energy classifier; The wake-up word is identified according to the acoustic score.
10. The speech processing method according to claim 9, characterized in that: The identifying the wake-up word according to the acoustic score includes: The acoustic score is input into a decoding network to identify the wake word.
11. The speech processing method according to any one of claims 1 to 10, characterized in that: The multi-channel audio signals are obtained by performing beamforming processing on original audio signals collected by the at least two audio collection devices according to multiple preset sound source directions.
12. The speech processing method according to claim 2, characterized in that: The speech processing model further includes a sound source classifier and a training module, and the speech processing model is obtained by the following training method: Acquiring first sample acoustic feature information based on sample data using the feature extraction module and the feature fusion module, wherein the sample data includes multiple channels of sample audio signals, the multiple channels of sample audio signals being acquired based on original sample audio signals acquired by at least two sample audio acquisition devices, and the first sample acoustic feature information being an integration of features of the multiple channels of sample audio signals; Obtaining spatial position information corresponding to the sample data based on the first sample acoustic feature information using the sound source classifier; The training module is used to train the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
13. The speech processing method according to claim 12, characterized in that: The sample data also includes at least one of the following items: original sample audio signals collected by one or more sample audio collection devices among the at least two sample audio collection devices, a sample phase difference between the original sample audio signals collected by any two sample audio collection devices among the at least two sample audio collection devices, and a sample amplitude difference between the original sample audio signals collected by any two sample audio collection devices among the at least two sample audio collection devices.
14. The speech processing method according to claim 12, characterized in that: The training method further comprises: Obtaining acoustic information based on the first sample acoustic feature information using the acoustic wake-up module; The training module is used to train the acoustic wake-up module according to the acoustic information and the acoustic annotation information corresponding to the sample data.
15. The speech processing method according to claim 14, characterized in that: The obtaining of acoustic information based on the first sample acoustic feature information by using the acoustic wake-up module includes: The acoustic information is obtained by using the acoustic wake-up module based on the first sample acoustic feature information and the spatial position information corresponding to the sample data.
16. The speech processing method according to claim 15, characterized in that: The acoustic wake-up module is provided with a gradient flip layer, an energy classifier, an acoustic classifier, and a gate. The acoustic wake-up module is used to obtain the acoustic information based on the first sample acoustic feature information and the spatial position information corresponding to the sample data, including: Obtaining second sample acoustic feature information based on the first sample acoustic feature information and spatial position information corresponding to the sample data using the gating; Obtaining the acoustic information based on the second sample acoustic feature information using the acoustic classifier, The training method further comprises: The acoustic wake-up module is subjected to adversarial training using the gradient flip layer based on the acoustic feature information and energy labeling information of the second sample, and the gradient flip layer points to the energy classifier.
17. The speech processing method according to any one of claims 12 to 16, characterized in that: The multi-channel sample audio signals are obtained by performing beamforming processing on original sample audio signals collected by at least two sample audio collection devices according to multiple preset sound source directions.
18. A method for training a speech processing model, characterized in that: The speech processing model includes a feature extraction module, a feature fusion module, a sound source classifier and a training module, and the training method includes: Acquiring first sample acoustic feature information based on sample data using the feature extraction module and the feature fusion module, wherein the sample data includes multiple channels of sample audio signals, the multiple channels of sample audio signals being acquired based on original sample audio signals acquired by at least two sample audio acquisition devices, and the first sample acoustic feature information being an integration of features of the multiple channels of sample audio signals; Obtaining, using the sound source classifier, spatial position information corresponding to the sample data based on the first sample acoustic feature information, where the first sample acoustic feature information is obtained by fusing feature information corresponding to multiple channels of sample audio signals and feature information corresponding to original sample audio signals collected by one or more of at least two sample audio collection devices; The training module is used to train the speech processing model based on the spatial position information corresponding to the sample data and the sound source annotation information corresponding to the sample data.
19. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor, The processor is used to execute the speech processing method described in any one of claims 1 to 17 or the training method of the speech processing model described in claim 18.
20. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which is used to execute the speech processing method according to any one of claims 1 to 17 or the training method of the speech processing model according to claim 18.
Citation Information
Patent Citations
Pick-up method and system based on microphone array
CN106782585A
End-to-end far-field speech recognition method and system
CN111179920A
Voice awakening method and device, electronic equipment and storage medium
CN111933111A