Data simulation method and electronic equipment
Through automatic simulation, the processed audio samples are generated, and the problems of low audio sample acquisition efficiency and instability of speech recognition models in the prior art are solved, and the effect of improving the recognition ability and stability of speech recognition models is achieved.
Patent Information
- Application Number
- CN202311738846.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-12-15
AI Technical Summary
In the prior art, recording audio samples is low efficiency, and the trained speech recognition model has unstable speech recognition effect in practical applications.
Through automatic simulation, a large number of processed audio samples are generated, which reduces the difficulty of obtaining training samples and improves the ability of the trained speech recognition model to recognize processed audio. The specific method includes obtaining the first audio data and the second audio data, generating mixed audio, eliminating the second audio data in the mixed audio, obtaining a target audio sample, and using the target audio sample to train the speech recognition model.
Generating audio samples through automatic simulation improves the recognition ability and stability of the speech recognition model and reduces the difficulty of obtaining training samples.
Smart Images

Figure CN120199238A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of human-computer interaction, and in particular, to a data simulation method and an electronic device. Background Art
[0002] With the development of speech recognition technology, speech interaction functions are widely used in various terminal devices and have become a common interaction function for most users. The key to realizing the speech interaction function lies in the model for recognizing speech, such as a speech recognition model.
[0003] In related technologies, recorded audio samples are used to train a speech recognition model to improve the model's ability to recognize speech. However, during this process, the efficiency of recording audio samples is low. In addition, for the trained model, the speech recognition effect is also unstable during actual application. Summary of the Invention
[0004] In view of this, the present application provides a data simulation method and an electronic device, which generate a large number of processed audio samples through an automatic simulation method, reduce the difficulty of obtaining training samples, and improve the ability of the trained speech recognition model to recognize the processed audio.
[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0006] In a first aspect, a data simulation method provided by an embodiment of the present application, the method includes:
[0007] Obtain first audio data and second audio data. The first audio data corresponds to a first label, and the first label includes the content text of the first audio data. Exemplarily, the first audio data and the second audio data may come from different audio databases or the same audio database. The above first audio data may be an audio with actual semantics, for example, a recorded sentence. The above second audio data may be an audio with actual semantics or other audio without semantics.
[0008] Then, generate a mixed audio corresponding to the first audio data and the second audio data. The mixed audio is an audio simulated to be collected by a microphone. Among them, the mixed audio may be an audio simulated to be collected by the microphone during the propagation of the first audio data and the second audio data. In the mixed audio, the second audio data is interference audio for the first audio data.
[0009] The generated mixed audio (audio with interference) through simulation has a higher generation efficiency compared to actually recording audio data with interference.
[0010] Then, eliminate the second audio data in the mixed audio to obtain the third audio data. Among them, the third audio data is the processed audio, which is different from the original first audio data. For example, there is audio damage, or it also includes residual data points of the second audio data, etc. After that, associate the third audio data with the first label to obtain the target audio sample. Using the target audio sample to train the speech recognition model helps to improve the recognition ability of the speech recognition model for the processed audio.
[0011] In the above embodiment, by improving the acquisition method of the target audio sample, it helps to train a speech recognition model with more stable speech recognition effect. For example, the speech recognition model can stably recognize various processed audio data.
[0012] In some embodiments, generating the mixed audio corresponding to the first audio data and the second audio data includes: randomly generating first feature information, where the first feature information includes first size information for indicating the space size and / or reverberation time RT60 within the space; respectively determining the first room impulse response RIR corresponding to the first audio data, and the second RIR corresponding to the second audio data; where the first RIR is to simulate the acoustic response of the first audio data in the target space indicated by the first feature information; the second RIR is to simulate the acoustic response of the second audio data in the target space indicated by the first feature information; after obtaining the first RIR and the second RIR, generate the mixed audio, and the mixed audio is the audio obtained by superimposing the first reverberant audio and the second reverberant audio, the first reverberant audio is the audio determined by the first audio data and the first RIR, and the second reverberant audio is the audio determined by the second audio data and the second RIR.
[0013] In the above embodiment, when simulating the mixed audio, the influence of the audio propagation space on the mixed audio is fully considered. Through different first feature information, environments with different influences on audio propagation can be simulated, such as what is called an audio scene.
[0014] In addition, simulating the difficulty of the mixed audio collected under different audio scenes is lower than the difficulty of restoring the actual audio scene for audio collection, which is beneficial to obtaining mixed audio covering more audio scenes.
[0015] In some embodiments, before separately determining the first room impulse response (RIR) corresponding to the first audio data and the second RIR corresponding to the second audio data, the method further includes: simulating the first position information of the microphone for collecting audio in the target space; simulating the second position information of the sound source corresponding to the first audio data in the target space; simulating the third position information of the sound source corresponding to the second audio data in the target space; determining the first room impulse response (RIR) corresponding to the first audio data includes: determining the first RIR corresponding to the first audio data according to the first characteristic information, the first position information, and the second position information; determining the second RIR corresponding to the second audio data includes: determining the second RIR corresponding to the second audio data according to the first characteristic information, the first position information, and the third position information.
[0016] In the above embodiments, when simulating the mixed audio, the influence of the sound source positions of the first audio data and the second audio data on the mixed audio is fully considered, so that the simulated mixed audio is more in line with the actual situation.
[0017] In some embodiments, simulating the first position information of the microphone for collecting audio in the target space includes: obtaining the second dimension information of the first device configured with the microphone, where the second dimension information includes the configured position of the microphone on the body of the first device; generating the fourth position information of the first device in the target space; and determining the first position information corresponding to the microphone of the first device in a random posture according to the fourth position information and the configured position of the microphone in the second dimension information.
[0018] In other embodiments, after the first position information is determined, it can be checked whether the position indicated by the first position information belongs to the target space. If not, a new fourth position information of the first device is regenerated, or a new random posture is determined, and the first position information of the microphone is determined again.
[0019] In the above embodiments, the mixed audio collected when the microphone is at different positions in the target space can be simulated, so that the simulated mixed audio can cover various audio scenarios.
[0020] In some embodiments, the first device further includes a speaker, and the sound source of the second audio data is the speaker of the first device.
[0021] In some embodiments, the second position information for simulating the sound source corresponding to the first audio data in the target space includes: constructing a first spherical coordinate system with the origin of the first spherical coordinate system being the position point indicated by the fourth position information; randomly determining a first coordinate in the first spherical coordinate system, where the first coordinate includes a distance, an azimuth angle, and a polar angle; wherein the first coordinate indicates the position of the sound source of the first audio data relative to the first device, the azimuth angle is greater than -180° and less than 180°; the polar angle is greater than -30° and less than 180°; converting the first coordinate information into a second coordinate, where the second coordinate is a rectangular coordinate; and determining the second position information of the sound source of the first audio data in the target space based on the second coordinate and the fourth position information.
[0022] In the above embodiments, when the sound source of the first audio data is simulated at different and reasonable positions in the target space, the mixed audio collected can be obtained, so that the simulated mixed audio can cover various audio scenarios and be closer to reality.
[0023] In other embodiments, after the second position information is determined, it can be checked whether the position indicated by the second position information belongs to the target space. If not, the second position information is re-determined to ensure that the simulated mixed audio is more realistic.
[0024] In some embodiments, the method for determining the first reverberant audio from the first audio data and the first RIR includes: performing convolution processing based on the first audio data and the first RIR to obtain the first reverberant audio; the method for determining the second reverberant audio from the second audio data and the second RIR includes: performing convolution processing based on the second audio data and the second RIR to obtain the second reverberant audio.
[0025] In some embodiments, after performing convolution processing based on the second audio data and the second RIR and before obtaining the second reverberant audio, the method further includes: obtaining the convolution audio corresponding to the second audio data and the second RIR; performing audio alignment on the convolution audio and the second audio data; determining a first duration, where the first duration belongs to a preset delay interval; and obtaining the second reverberant audio includes: generating the second reverberant audio according to the first duration and the aligned convolution audio; wherein there is a delay of the first duration between the second reverberant audio and the aligned convolution audio.
[0026] In the above embodiments, after audio alignment, delay processing is performed. While solving the unreasonable audio delay caused by convolution processing, a reasonable audio delay that appears during actual acquisition is simulated.
[0027] In some embodiments, eliminating the second audio data in the mixed audio to obtain third audio data includes: under preset conditions, inputting the mixed audio and the second audio into an echo cancellation module to obtain the third audio data; the preset conditions include: the mixed audio simulates the audio collected by the microphone of a first device, and the second audio data simulates the audio played by the speaker of the first device.
[0028] In the above embodiments, a mixed audio with echo interference can be simulated.
[0029] In a second aspect, a method for training a speech recognition model provided by an embodiment of the present application includes: obtaining a target audio sample; where the target audio sample is an audio simulated according to the method provided in the first aspect and its possible embodiments; using the target audio sample to train a pre-configured speech recognition model until the speech recognition model meets the pre-configured model convergence condition.
[0030] In the above embodiments, by improving the acquisition method of the target audio sample, it helps to train a speech recognition model with more stable speech recognition effect. For example, the speech recognition model can stably recognize various processed audio data.
[0031] In some embodiments, the method further includes: obtaining fourth audio data with a second label, where the fourth audio data is an un-denoised original audio, and the second label includes the content text of the fourth audio data; using the fourth audio data to train a pre-configured speech recognition model until the speech recognition model meets the pre-configured model convergence condition.
[0032] In some embodiments, the method further includes: obtaining fifth audio data with a third label, where the fifth audio data is a recorded and denoised audio, and the third label includes the content text of the fifth audio data; using the fifth audio data to train a pre-configured speech recognition model until the speech recognition model meets the pre-configured model convergence condition.
[0033] In a third aspect, a speech recognition method provided by an embodiment of the present application is applied to a second device. The second device is configured with a speech recognition model and an echo cancellation module. The speech recognition model is a model obtained by using the speech recognition model training method provided in the second aspect and other possible embodiments. The speech recognition method includes: when the second device plays sixth audio data, it collects seventh audio data; after inputting the seventh audio data and the sixth audio data into the echo cancellation module, it obtains the eighth audio data; using the speech recognition model to recognize the content text corresponding to the eighth audio data.
[0034] In the above embodiments, the second device can stably identify the audio data to be identified in various actual audio scenarios, improving the audio recognition ability of the second device.
[0035] Fourthly, an electronic device provided by an embodiment of the present application includes one or more processors and a memory; the memory is coupled to the processor, and the memory is used to store computer program code, and the computer program code includes computer instructions. When one or more processors execute the computer instructions, the one or more processors are used for the methods in the first aspect, the second aspect, the third aspect and their possible embodiments above.
[0036] Fifthly, a computer storage medium provided by an embodiment of the present application includes computer instructions. When the computer instructions run on an electronic device, the electronic device is caused to execute the methods in the first aspect, the second aspect, the third aspect and their possible embodiments above.
[0037] Sixthly, a computer program product provided by the present application, when the computer program product runs on the above-mentioned electronic device, causes the electronic device to execute the methods in the first aspect, the second aspect, the third aspect and their possible embodiments above.
[0038] In addition, the electronic devices, computer storage media, and computer program products provided in the above aspects are all applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here. Description of the Drawings
[0039] Figure 1 One of the schematic diagrams of the voice recognition scenario provided by an embodiment of the present application;
[0040] Figure 2 Another schematic diagram of the voice recognition scenario provided by an embodiment of the present application;
[0041] Figure 3 Another schematic diagram of the voice recognition scenario provided by an embodiment of the present application;
[0042] Figure 4 Schematic diagram of the implementation principle of voice interaction provided by an embodiment of the present application;
[0043] Figure 5 Schematic diagram of the principle of training a voice recognition model in some embodiments;
[0044] Figure 6 Schematic diagram of the audio data with echo collected by the microphone of the terminal device;
[0045] Figure 7 Schematic diagram of the terminal device eliminating echo;
[0046] Figure 8 It is a schematic diagram of the principle for training a speech recognition model provided by an embodiment of the present application;
[0047] Figure 9 It is a schematic diagram for generating a simulated audio sample provided by an embodiment of the present application;
[0048] Figure 10 It is one of the flowcharts of the steps for generating simulated audio provided by an embodiment of the present application;
[0049] Figure 11 It is another flowchart of the steps for generating simulated audio provided by an embodiment of the present application;
[0050] Figure 12 It is the third flowchart of the steps for generating simulated audio provided by an embodiment of the present application;
[0051] Figure 13 It is the fourth flowchart of the steps for generating simulated audio provided by an embodiment of the present application;
[0052] Figure 14 It is a schematic diagram for establishing an extended coordinate system provided by an embodiment of the present application;
[0053] Figure 15 It is the fifth flowchart of the steps for generating simulated audio provided by an embodiment of the present application;
[0054] Figure 16 It is a schematic diagram of the spherical coordinate system provided by an embodiment of the present application;
[0055] Figure 17 It is the sixth flowchart of the steps for generating simulated audio provided by an embodiment of the present application;
[0056] Figure 18 It is a schematic diagram of RIR provided by an embodiment of the present application;
[0057] Figure 19 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0058] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.
[0059] The implementation manners of this embodiment will be described in detail below with reference to the accompanying drawings.
[0060] The above voice interaction means that when the user utters a statement containing an instruction, after the terminal device recognizes the statement, it can make a corresponding response.
[0061] Exemplarily, the above terminal device can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a smart TV, a smart watch, a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc. The embodiments of the present application do not impose special restrictions on the specific form of the terminal device.
[0062] Taking the terminal device as a mobile phone as an example. An exemplary scenario is as Figure 1 shown, the user holds the mobile phone 101. The mobile phone 101 is displaying a playback interface 102 provided by a short video application. Exemplarily, the playback interface 102 displays the picture content 103 of video data a. In addition, the mobile phone 101 synchronously plays the audio corresponding to the video data a, for example, music 1. Exemplarily, in the playback interface 102, an identifier of the currently playing audio is also displayed, such as identifier 104, indicating that music 1 is being played.
[0063] During the display of the playback interface 102, the user utters the statement 105, for example, the statement "Play the next video".
[0064] In this scenario, the mobile phone 101 collects the audio data of the statement 105 uttered by the user, can recognize the semantics of the statement 105, and determine a control instruction that matches the semantics. For example, in the scenario where the short video application is running in the foreground on the mobile phone 101, it is recognized that the semantics of the statement 105 is to switch the played video. According to the business logic of the short video application, the determined control instruction that matches it is an instruction to indicate playing video data b. Among them, both video data b and video data a are video data determined by the short video application to be recommended to the user, and the recommendation order of video data b is after that of video data a.
[0065] In this way, the mobile phone 101 can respond to the statement 105 uttered by the user and switch to display the playback interface 106. Exemplarily, the above playback interface 106 displays the picture 107 of video data b. In addition, the mobile phone 101 synchronously plays the audio corresponding to the video data b, for example, music 2. Exemplarily, in the playback interface 106, an identifier of the currently playing audio is also displayed, such as identifier 108, indicating that music 2 is being played.
[0066] Taking the terminal device as a smart watch as an example, an exemplary scenario is as follows Figure 2 As shown, the user wears the smart watch 201, and the smart watch 201 executes an alarm reminder service. For example, an alarm reminder capsule 203 is displayed on the dial interface 202, an alarm reminder sound effect is played through the speaker, and / or the body of the watch vibrates, etc.
[0067] During the execution of the alarm reminder service by the smart watch 201, the user speaks a statement 204, for example, the statement "Turn off the alarm".
[0068] In this scenario, the smart watch 201 collects the audio data of the statement 204 spoken by the user, can recognize the semantics of the statement 204, and determine a control instruction that matches the semantics. For example, in the scenario where the smart watch 201 is executing the alarm reminder service, it is recognized that the semantics of the statement 204 is to stop executing the alarm reminder service. According to the logic of the alarm reminder service, the determined control instruction that matches it is an instruction to cancel the reminder items related to the alarm.
[0069] In this way, the smart watch 201 can respond to the statement 204 spoken by the user, cancel the display of the alarm reminder capsule 203, cancel the playback of the alarm reminder sound effect, and / or cancel the vibration of the body.
[0070] Continuing to take the terminal device as a smart TV as an example, an exemplary scenario is as follows Figure 3 As shown, the smart TV 301 is playing video data from TV channel a. Exemplarily, a picture 302 of the video data from TV channel a is displayed on the smart TV 301.
[0071] During the display of playing video data c, the user speaks a statement 303, for example, the statement "Switch to the next channel".
[0072] In this scenario, the smart TV 301 collects the audio data of the statement 303 spoken by the user, can recognize the semantics of the statement 303, and determine a control instruction that matches the semantics. For example, in the scenario where the smart TV 301 is playing video data from TV channel a, it is recognized that the semantics of the statement 303 is to switch channels. According to the TV service logic, the determined control instruction that matches it is: an instruction to play the video data from TV channel b. Among them, TV channel b is the next adjacent channel to TV channel a.
[0073] In this way, the smart TV 301 can respond to the statement 303 spoken by the user and switch to the picture 304 of the video data from TV channel b.
[0074] As can be seen from the above examples, based on the voice interaction function, it is more convenient for users to operate the terminal device, improving the human-computer interaction efficiency of the terminal device. The implementation of the above voice interaction function depends on speech recognition technology. In some embodiments, a speech recognition model can be pre-configured in the terminal device. The semantic recognition of the collected statements is completed using the speech recognition model.
[0075] Taking Figure 1 the scenario shown as an example, as Figure 4 shown, a speech recognition model 401 is configured in the mobile phone 101. Among them, the above speech recognition model 401 can be a neural network model. The model structure of the neural network model can refer to related technologies and will not be elaborated here. After the above neural network model is trained, it has the ability of speech recognition. As Figure 4 shown, after the mobile phone 101 collects the audio data of the statement 105, the target audio corresponding to the statement 105 is obtained.
[0076] Exemplarily, the above target audio can be the original audio data collected by the mobile phone 101 (for example, the audio data of the statement 105). Also exemplarily, the above target audio can also be the audio obtained after the audio data collected by the mobile phone 101 is processed. The embodiments of the present application do not make specific limitations on this.
[0077] The mobile phone 101 can input the target audio corresponding to the statement 105 into the speech recognition model 401, and the speech recognition model 401 determines the recognition result corresponding to the target audio. When the recognition result is text 402 (for example, "Play the next video"), the mobile phone 101 can instruct the short video application running in the foreground to switch to play the next video.
[0078] Obviously, the recognition accuracy of the above speech recognition model 401 will affect the voice interaction effect. In some embodiments, the speech recognition model can be trained using a recorded audio sample set. Among them, the above audio sample set includes a large number of original audio samples. Each original audio sample includes an audio data and a corresponding audio label, and the audio label contains the semantic content of the audio data. In addition, the audio data in the original audio sample can be recorded interference-free audio corpus.
[0079] As Figure 5 shown, the training process of the above speech recognition model can be executed by a model training device. Exemplarily, the model training device can be a device with more computing resources, such as a server, a desktop computer, a mobile phone, etc.
[0080] An initial model of the speech recognition model can be configured in the model training device. As Figure 5As shown, the model training device can iteratively train the initial model based on the original audio samples in the audio sample set. Among them, the above-mentioned audio sample set can be configured in the model training device or in other devices communicating with the model training device.
[0081] Among them, the process of iterative training can refer to related technologies and will not be elaborated here one by one. As Figure 5 shown, the model training device can input the audio data of the original audio sample into the initial model, and the initial model predicts the predicted semantics corresponding to the audio data. Then, the model training device determines whether the initial model converges. For example, the model training device calculates the loss value between the predicted semantics pre-output by the initial model and the audio label, and then the model training device compares the loss value with the preset value to determine whether the initial model converges. When the model convergence condition is not met, the model training device updates the model parameters according to the loss function during the iteration process. When the model convergence condition is met, the model training device stops the iterative training and obtains the speech recognition model.
[0082] In the above example, whether the loss value is less than the preset value (a model convergence condition) is used to evaluate whether the model converges. In other examples, other model convergence conditions can also be used to evaluate whether the model converges. For example, other model convergence conditions can be that the number of model iterations reaches a specified number. In actual applications, the model training device can pre-configure the corresponding model convergence conditions.
[0083] The speech recognition model has the ability to recognize the semantics of audio data. After implanting the speech recognition model into the terminal device, the terminal device also has the ability to recognize the semantics of audio data.
[0084] In some possible embodiments, the model training device can also be the terminal device that actually uses the speech recognition model. That is, after training the speech recognition model, the model training device can also have the ability to recognize the semantics of audio data based on the speech recognition model. Continuing with the example of the mobile phone 101 configured with the speech recognition model, as Figure 6 and Figure 7 shown, in the actual usage scenario, among the audio data (original audio data) actually collected by the mobile phone 101 through the microphone, it not only contains the audio of the sentence to be recognized but also mixed with interference audio. Among them, the above-mentioned interference audio is the audio data that does not need to participate in speech recognition. For example, the echo audio generated by the music 1 that has been played on the mobile phone 101 (which can be simply referred to as the played music 1). The above-mentioned sentence to be recognized can be the sentence spoken by the user.
[0085] As Figure 6As shown, when the mobile phone 101 is playing Music 1 through the speaker 601, the original audio data collected by the microphone 602 can be a mixed audio between the audio containing "play the next video" and the echo audio of Music 1.
[0086] In Figure 6 the scenario shown, the interfering audio can be Music 1 played by the speaker 601. In other scenarios, the interfering audio can also include audio information emitted by other sound sources.
[0087] To avoid the influence of the interfering audio on the speech recognition result. In some embodiments, before the mobile phone 101 actually enables the speech recognition model to perform speech recognition, it is necessary to preprocess the original audio data to remove the interfering audio.
[0088] As Figure 7 shown, the mobile phone 101 can play Music 1 through the speaker according to the audio data 702. For example, the advanced digital signal processor (ADSP) in the mobile phone 101 can determine the drive signal for the speaker 601 according to the audio data 702. Among them, the drive signal can be a digital signal. When the mobile phone 101 is configured with multiple speakers 601, the drive signals for different speakers 601 can be different to achieve better sound effects. After obtaining the drive signal for the speaker 601, the drive signal can also be converted into an analog signal and transmitted to the speaker 601 to make the speaker 601 play the sound of Audio 1.
[0089] During the mobile phone 101 playing Music 1, the microphone 602 of the mobile phone 101 can collect the original audio data 701. Among them, the original audio data 701 can be a mixed audio between the audio "play the next video" and the echo audio of Music 1.
[0090] After collecting the original audio data 701, the original audio data 701 (audio data collected by the microphone) and the audio data 702 of the interfering audio (also called the reference audio) can be input into the acoustic echo cancellation (AEC) module to eliminate the echo generated by Music 1 in the original audio data 701.
[0091] In Figure 7 the scenario shown, the mobile phone 101 can read the audio data 702 corresponding to the Music 1 being played by the speaker 601 from the ADSP.
[0092] After that, the mobile phone 101 inputs the original audio data 701 and the audio data 702 into the AEC module. Based on the original audio data 701 and the audio data 702, the AEC module can output the target audio 703. The above target audio 703 can be audio data with interference audio eliminated. Among them, the principle of the AEC module for eliminating echo can refer to the related technology and will not be elaborated here for the time being.
[0093] In other embodiments, when the interference audio also includes the audio influence from other interference sources (other sound sources outside the mobile phone 101), other denoising modules can also be used to eliminate other interference audio, and the embodiments of the present application do not make specific limitations on this.
[0094] Whether it is the AEC module or other denoising modules, during use, it will inevitably cause audio damage to the audio data. The audio data with data damage is different from the original audio samples for training the speech recognition model. And the current speech recognition model has limited ability to recognize damaged audio.
[0095] In addition, the mixing situations of the audio to be recognized and the interference audio are different, and after being processed by the AEC module, the degrees of audio damage are different. After different audio data are processed by the AEC module, the positions (time positions) where audio damage occurs are also different.
[0096] In this way, even if the mobile phone 101 is configured with a high-precision speech recognition model, in different actual application scenarios, the speech recognition effect is still unstable.
[0097] To improve the above problems, the embodiments of the present application provide a method for training a speech recognition model. While retaining the training of the speech recognition model using the recorded audio sample set, it also adds the training of the speech recognition model using simulated audio samples and re-recorded audio samples.
[0098] As Figure 8 shown, the model training device can extract the original audio samples (the fourth audio data) from the recorded audio sample set to iteratively train the initial model until the model converges. For details, refer to Figure 5 , which will not be elaborated here for the time being. Among them, the fourth audio data from the recorded audio sample set corresponds to an audio label, that is, the second label.
[0099] Continuing as Figure 8 shown, the model training device can also receive the re-recorded audio samples (the fifth audio data) from the recording device and use the re-recorded audio samples to iteratively train the initial model until the model converges.
[0100] Among them, the above-mentioned recorded audio samples can come from one or more recording devices. Among them, the recording device can be various types of terminal devices mentioned in the foregoing embodiments, or can be a device only used for recording audio. An AEC module (and / or other noise reduction modules) is configured in each recording device. Exemplarily, the recording device actually records the original audio data in various actual sound scenarios (including the audio to be recognized and interfering audio). The original audio data is processed by the AEC module (and / or other noise reduction modules) and associated with the corresponding audio label (the third label) to obtain the corresponding recorded audio sample.
[0101] In the above example, the audio quality of the above-mentioned recorded audio sample is high and conforms to reality. Using the recorded audio sample to train the speech recognition model can improve the ability of the speech recognition model to recognize damaged audio in various sound scenarios from the dimension of quality.
[0102] Continue as Figure 8 shown, the model training device can also use a large number of simulated audio samples to iteratively train the initial model until the model converges.
[0103] Among them, the simulated audio sample includes the audio that has been processed by the AEC module (and / or other noise reduction modules) simulated based on the real audio data. Exemplarily, the model training device can extract audio data from the audio set, and then, based on the extracted audio data and the simulation module, simulate the simulated audio processed by the AEC module (and / or other noise reduction modules). After that, the simulated audio is associated with the corresponding audio label to obtain the corresponding simulated audio sample.
[0104] In the above example, the generation efficiency of the above-mentioned simulated audio sample is high, and a large number of audio that meets the requirements can be quickly generated, reducing the difficulty of obtaining samples. Using the simulated audio sample to train the speech recognition model can improve the ability of the speech recognition model to recognize damaged audio in various sound scenarios from the dimension of the number of samples.
[0105] In other embodiments, one or a combination of the recorded audio sample, the simulated audio sample, and the original audio sample can also be used to train the initial module of the speech recognition model to obtain a speech recognition model with stable recognition accuracy.
[0106] The following describes the implementation details of a data simulation method provided by the application embodiments in conjunction with the accompanying drawings. This data simulation method can be applied to a device configured with a simulation module, such as a simulation device, which can generate the simulation audio samples mentioned in the foregoing embodiments. Exemplarily, the above simulation device can be the model training device mentioned in the foregoing embodiments. In this way, data simulation and model training can be executed by the same device. Additionally, the above simulation device can also be other devices other than the model training device provided in the foregoing embodiments. In this way, data simulation and model training can be executed by different devices, and the application embodiments do not make specific limitations in this regard.
[0107] In some embodiments, as Figure 9 shown, the above simulation module may include an audio simulation sub-module and an AEC module. The steps for the above simulation device to execute the above data simulation method are as follows:
[0108] S101, obtain at least two pieces of audio data.
[0109] Among them, the above at least two pieces of audio data include interfering audio (such as the second audio data) and the audio to be recognized (such as the first audio data).
[0110] In some embodiments, as Figure 9 shown, both the above interfering audio and the audio to be recognized can come from the same audio set. For example, the simulation device can randomly select two different original audio samples from the recorded audio sample set. Among them, the audio data of one original audio sample is used as the audio to be recognized, and the other original audio sample is used as the interfering audio. The audio data from the audio sample set all correspond to audio tags. For example, the first audio data corresponds to the first tag. Among them, the first tag is an audio tag, and the first tag includes the content text corresponding to the first audio data. For example, if the first audio data is the statement "Switch to the next video", the corresponding first tag is the text "Switch to the next video".
[0111] In other embodiments, both the above interfering audio and the audio to be recognized can come from different audio sets. For example, the audio to be recognized comes from the audio sample set, and the interfering audio comes from other audio sets, such as a music database, a video database, an audio set of various reminder sound effects, etc.
[0112] In other embodiments, the above interfering audio may also include multiple pieces of audio data. The application embodiments do not make specific limitations in this regard. In the subsequent embodiments, one audio to be recognized and one interfering audio are used as examples for description.
[0113] S102, based on the above at least two pieces of audio data, simulate the mixed audio collected by the microphone of the terminal device.
[0114] Among them, the above mixed audio can be the audio data simulated by the microphone, and the audio data is related to the interference audio and the audio to be recognized.
[0115] In some embodiments, the audio simulation sub-module can be used to directly mix the audio to be recognized and the interference audio to obtain the corresponding mixed audio.
[0116] In other embodiments, the audio simulation sub-module can be used to simulate the target space for collecting audio data, and then, when simulating the propagation of the audio to be recognized and the interference audio in the target space, the corresponding mixed audio. For example, the room impulse response (RIR) corresponding to the audio to be recognized and the interference audio is randomly determined, and the determined RIR is related to the target space. Then, based on the RIR corresponding to the audio to be recognized and the interference audio, the mixed audio is generated, and the specific implementation details can be referred to the subsequent embodiments.
[0117] In other embodiments, the audio simulation sub-module can also be used to simulate the target space for collecting audio data, and the positions of the sound sources corresponding to the audio to be recognized and the interference audio. Among them, the positions of the sound sources corresponding to the audio to be recognized and the interference audio can be different positions. In addition, the positions of the sound sources corresponding to the simulated audio to be recognized and the interference audio can be located within the target space. Then, when simulating the sound sources of the audio to be recognized and the interference audio at different positions in the target control, the mixed audio collected by the microphone is simulated.
[0118] As an implementation manner, as Figure 10 shown, the implementation steps of the above S102 can be the following sub-steps:
[0119] S102-1, generate the first feature information for indicating the target space.
[0120] Among them, the above first feature information can be the feature information that affects the propagation of the audio signal in the target space.
[0121] Exemplarily, the size of the target space can affect the propagation range of the audio signal, and the above first feature information can include the size information of the target space. For example, the first feature information includes the length, width and height of the target space, also known as the first size information.
[0122] Exemplarily, the spatial material of the target space can affect the attenuation rate of audio energy during the propagation of an audio signal. For example, the above first feature information may include the reverberation time (RT) 60 corresponding to the spatial material of the target space. Among them, the above RT60 refers to the time taken for the sound field in the target space to attenuate by 60 dB. The larger the RT60, the longer the audio propagation time in the target space. The smaller the RT60 value, the shorter the audio propagation time in the target space. In addition, the RT60 value of the real space is affected by the size of the space and also by the building materials used to construct the space.
[0123] In some embodiments, as Figure 11 shown, the simulation device can randomly generate a set of the above length, width, height, and RT60 values, where the length, width, height, and RT60 values all belong to the first feature information. In this example, the generated length, width, height, and RT60 values represent the simulated target space 1101.
[0124] In other embodiments, the simulation device can randomly generate length, width, height, and RT60 values within the value ranges corresponding to various first feature information, so that the simulated target space 1101 is closer to the actual scenario.
[0125] Among them, the random value ranges corresponding to the above first feature information such as length, width, height, and RT60 can be empirical values, and this application does not make specific limitations on this. Exemplarily, the random value range of the RT60 value is from 1 / n seconds to m seconds, where n and m are positive integers. After determining the length, width, and height of the target space, the size of the target space is determined. In this case, the simulation device can simulate the target space built with different building materials through the random RT60 value.
[0126] S102-2, generate the position information 1 of the terminal device in the target space.
[0127] In some embodiments, before S102-2, the simulation device constructs a three-dimensional coordinate system (such as called the target coordinate system) relative to the target space, and each position point in the target space can be mapped to this target coordinate system, so that each position point in the target space corresponds to a target coordinate system. In addition, the target coordinate system includes the x-axis 1, the y-axis 1, and the z-axis 1, and the three-dimensional coordinates of any point in the target space 1101 in the target coordinate system satisfy the following conditions: the coordinate value on the x-axis 1 does not exceed the length of the target space 1101, the coordinate value on the y-axis 1 does not exceed the width of the target space 1101, and the coordinate value on the z-axis 1 does not exceed the height of the target space 1101.
[0128] In some embodiments, the above S102-2 may be: The simulation device randomly determines a position point a in the target space, and determines the position coordinates of the position point a, for example, determines (a, b, c) as the position information 1 (the fourth position information) of the terminal device in the target space, where a is not greater than the length of the target space 1101, b is not greater than the width of the target space 1101, and c is not greater than the height of the target space 1101.
[0129] For example, as Figure 12 shown, a position point 1201 is randomly determined in the target space, and the position coordinates of the position point 1201 are used as the position information 1 of the terminal device 1202 in the target space 1101.
[0130] In Figure 12 , the position point 1201 overlaps with the center point of the simulated terminal device 1202, that is, the position coordinates of the center point of the terminal device 1202 are used to represent the position of the terminal device 1202 in the target space 1101. In other examples, the position point 1201 may also overlap with other positions of the simulated terminal device 1202, for example, overlap with the end point of any edge on the body of the terminal device 1202. In the subsequent embodiments, the case where the position point 1201 overlaps with the center point of the terminal device 1202 is taken as an example for description.
[0131] S102-3, based on the position information 1 of the terminal device, determine the position information 2 of the microphone of the terminal device in the target space.
[0132] In some embodiments, the simulation device may use the position point 1201 as the origin to construct a two-dimensional coordinate system relative to the terminal device 1202.
[0133] Among them, constructing a two-dimensional coordinate system relative to the terminal device 1202 may be to establish a coordinate system on the specified surface of the terminal device 1202, and the origin of the coordinate system is the position point 1201. Among them, the above-mentioned specified surface may be the plane with the largest area on the terminal device 1202. For example, the specified surface may be the plane where the display screen is located, or the surface opposite to the display screen.
[0134] Taking the position point 1201 as the center point of the terminal device 1202 and the specified surface as the plane where the display screen of the terminal device 1202 is located as an example.
[0135] For example, as Figure 13As shown, the simulation device can construct a coordinate system 1301 relative to the terminal device 1202, where the coordinate system 1301 includes an x-axis 2 and a y-axis 2. The terminal device 1202 can be projected onto the coordinate system 1301, that is, each point in the terminal device 1202 corresponds to a two-dimensional coordinate in the coordinate system 1301. Among them, the two-dimensional coordinate of the center point of the terminal device 1202 in the coordinate system 1301 is (0, 0).
[0136] Exemplarily, the microphone 1302 of the terminal device 1202 can also be projected onto the coordinate system 1301, and the simulation device can determine the two-dimensional coordinate of the microphone 1302 in the coordinate system 1301 according to the size information of the terminal device 1202.
[0137] As an implementation, after determining the product type of the terminal device (the first device) to be simulated, the simulation device can obtain the size information of the terminal device from the product server according to the product model of the terminal device, such as the second size information. Among them, the above size information includes information such as the length, width, thickness, the configuration position of the speaker on the body, and the configuration position of the microphone on the body of the terminal device. In this way, according to the size information of the terminal device, the relative position relationship between the microphone of the terminal device and the center point of the terminal device can be determined, and further the projection position of the microphone in the two-dimensional coordinate system corresponding to the terminal device can be determined.
[0138] For example, the size information of the terminal device 1202 includes: the width of the terminal device is 40 cm, the length is 160 cm, and the distance from the microphone at the bottom to the central axis of the terminal device 1202 is 10 cm. As Figure 14 shown, after projecting the terminal device 1202 onto the coordinate system 1301, the area 1401 is the projection range of the terminal device 1202. The width of this area 1401 is 40 cm and the length is 160 cm. Correspondingly, the distance from the microphone 1302 to the x-axis 2 of the coordinate system 1301 is 80 cm, and the distance from the microphone 1302 to the y-axis 2 of the coordinate system 1301 is 10 cm. In the Figure 14 scenario shown, the two-dimensional coordinate of the microphone 1302 in the coordinate system 1301 is (10, -80).
[0139] The microphone 1302 of the above terminal device 1202 not only corresponds to the two-dimensional coordinate in the coordinate system 1301, but also corresponds to the three-dimensional coordinate in the three-dimensional coordinate system (target coordinate system) of the target space 1101. Among them, after the coordinate system 1301 is established with the size information of the terminal device 1202 unchanged, the two-dimensional coordinate of the microphone 1302 remains fixed. When the terminal device 1202 is in different postures in the target space 1101, the corresponding three-dimensional coordinates can be different.
[0140] In some embodiments, after the two-dimensional coordinates of the microphone 1302 are determined, the simulation device may, based on the two-dimensional coordinates, simulate the terminal device 1202 in a random posture to obtain the three-dimensional coordinates of the corresponding microphone 1302 in the target coordinate system, which is used as the position information 2 of the microphone 1302, also known as the first position information.
[0141] As an implementation manner, the process of simulating the terminal device 1202 in a random posture to obtain the three-dimensional coordinates of the corresponding microphone 1302 in the target coordinate system may be as follows:
[0142] (1) Expand the coordinate system 1301 into a three-dimensional coordinate system (such as called the expanded coordinate system). Among them, as Figure 14 shown, compared with the coordinate system 1301, the above-mentioned expanded coordinate system 1402 adds a z-axis 2. The origin of the expanded coordinate system 1402 is still the center point of the terminal device 1202, and the three-dimensional coordinates of this center point in the expanded coordinate system 1402 are (0, 0, 0). In the expanded coordinate system 1402, the three-dimensional coordinates of the microphone 1302 are (10, -80, 0).
[0143] (2) Randomly generate the angle 1 of the terminal device 1202 rotating relative to the x-axis 2 of the expanded coordinate system 1402, the angle 2 of rotating relative to the y-axis 2 of the expanded coordinate system 1402, and the angle 2 of rotating relative to the z-axis 2 of the expanded coordinate system 1402. The above-mentioned randomly generated angles 1, 2, and 3 can simulate the terminal device 1202 in a random posture.
[0144] In three-dimensional coordinate systems such as the expanded coordinate system 1402 and the target coordinate system, the position of the microphone 1302 changes following the posture of the terminal device 1202. For the angle 1 of the terminal device 1202 rotating relative to the x-axis 2 of the expanded coordinate system 1402, the microphone 1302 also rotates by the angle 1 relative to the x-axis 2 of the expanded coordinate system 1402. For the angle 2 of the terminal device 1202 rotating relative to the y-axis 2 of the expanded coordinate system 1402, the microphone 1302 also rotates by the angle 2 relative to the y-axis 2 of the expanded coordinate system 1402. For the angle 3 of the terminal device 1202 rotating relative to the z-axis 2 of the expanded coordinate system 1402, the microphone 1302 also rotates by the angle 3 relative to the z-axis 2 of the expanded coordinate system 1402.
[0145] (3) According to the angles 1, 2, 3 and the initial three-dimensional coordinates (10, -80, 0) of the microphone 1302, combined with the three-dimensional space coordinate rotation calculation formula, after the terminal device 1202 switches to a random posture, the rotated three-dimensional coordinates of the microphone 1302 in the expanded coordinate system 1402 can be determined.
[0146] Among them, the principle of three-dimensional space coordinate rotation is as follows:
[0147] 1) Obtain the rotation matrices about the x-axis 2, y-axis 2, and z-axis 2 respectively.
[0148] Exemplarily, the rotation matrix for rotating an angle 1 about the x-axis 2, e.g., rotating clockwise by γ, is as follows:
[0149]
[0150] where, R x (γ) is the rotation matrix for rotating clockwise by γ, and γ is the angle of clockwise rotation with respect to the x-axis 2, for example, equal to angle 1.
[0151] Exemplarily, the rotation matrix for rotating an angle 2 about the y-axis 2, e.g., rotating clockwise by β, is as follows:
[0152]
[0153] where, R y (β) is the rotation matrix for rotating clockwise by β, and β is the angle of clockwise rotation with respect to the y-axis 2, for example, equal to angle 2.
[0154] Exemplarily, the rotation matrix for rotating an angle 3 about the z-axis 2, e.g., rotating clockwise by α, is as follows:
[0155]
[0156] where, R z (α) is the rotation matrix for rotating clockwise by α, and α is the angle of clockwise rotation with respect to the z-axis 2, for example, equal to angle 3.
[0157] 2) Generate the corresponding target rotation matrix in the order of rotation about the x-axis 2, y-axis 2, and z-axis 2.
[0158] For example, in the scenario of first rotating about the x-axis 2, then about the y-axis 2, and finally about the z-axis 2, the corresponding target rotation matrix is:
[0159] R = R z (α)R y (β)R x (γ);
[0160] where, R is the target rotation matrix, R z (α) is the rotation matrix corresponding to the z-axis 2, R y (β) is the rotation matrix corresponding to the y-axis 2, R x (γ) is the rotation matrix corresponding to the x-axis 2.
[0161] Correspondingly,
[0162]
[0163] For another example, in the scenario of first rotating by 2 around the z-axis, then rotating by 2 around the y-axis, and finally rotating by 2 around the x-axis, the corresponding target rotation matrix is:
[0164] R = R x (γ)R y (β)R z (α).
[0165] (3) Based on the target rotation matrix and the initial three-dimensional coordinates (10, -80, 0) of the microphone 1302 in the extended coordinate system 1402, combined with the formula:
[0166]
[0167] Determine the three-dimensional coordinates of the microphone 1302 in the extended coordinate system 1402 after rotation. Among them, R is the corresponding target rotation matrix, (x0, y0, z0) are the initial three-dimensional coordinates of the microphone 1302 in the extended coordinate system 1402, that is, x0 = 10, y0 = -80, z0 = 0, and the above (x′, y′, z′) are the three-dimensional coordinates of the microphone 1302 in the extended coordinate system 1402 after rotation.
[0168] (4) According to the three-dimensional coordinates of the microphone 1302 after rotation in the extended coordinate system 1402, the three-dimensional coordinates of the center point of the terminal device 1202 in the extended coordinate system 1402, and the three-dimensional coordinates of the center point of the terminal device 1202 in the target coordinate system, determine the three-dimensional coordinates of the microphone 1302 in the target coordinate system.
[0169] Among them, during the rotation of the terminal device 1202, the three-dimensional coordinates of the center point of the terminal device 1202 in the extended coordinate system 1402 remain unchanged, still being (0, 0, 0).
[0170] Whether in the extended coordinate system 1402 or in the target coordinate system, the relative position relationship between the microphone 1302 and the center point of the terminal device 1202 remains unchanged. That is, A - B = A1 - B1. Among them, A is the three-dimensional coordinates of the center point of the terminal device 1202 in the target coordinate system, that is, in the foregoing embodiment, the determined position information 1 is (a, b, c). B is the three-dimensional coordinates of the microphone 1302 in the target coordinate system, that is, the position information 2 of the microphone 1302 to be determined. A1 is the three-dimensional coordinates of the center point of the terminal device 1202 in the extended coordinate system 1402, that is, (0, 0, 0) mentioned in the foregoing embodiment. B1 is the three-dimensional coordinates of the microphone 1302 in the extended coordinate system 1402, that is, the three-dimensional coordinates obtained after rotation of the microphone 1302 in the extended coordinate system 1402 mentioned in the foregoing embodiment.
[0171] As an implementation manner, the position information 2 corresponding to the microphone 1302 can be determined as B = (A - A1) + B1.
[0172] For example, the three-dimensional coordinates of the microphone 1302 after rotation in the extended coordinate system 1402 are (73, 6, 31), the three-dimensional coordinates of the center point of the terminal device 1202 in the target coordinate system are (1000, 1000, 1000), and the three-dimensional coordinates in the extended coordinate system 1402 are (0, 0, 0). It can be determined that the three-dimensional coordinates of the microphone 1302 in the target coordinate system are (1073, 1006, 1031).
[0173] In other possible embodiments, after the position information 2 of the microphone 1302 is determined, it can also be determined whether the position information 2 is within the target space 1101. If the position point indicated by the position information 2 is not within the target space 1101, for example, the x coordinate value in the position information 2 is greater than the length corresponding to the target space 1101, the y coordinate value in the position information 2 is greater than the width of the target space 1101, and / or the z coordinate value in the position information 2 is greater than the height of the target space 1101, then S102-2 and S102-3 are executed again. If the position point indicated by the position information 2 is within the target space 1101, the process proceeds to S102-4.
[0174] S102-4: Randomly generate the position information 3 of the target sound source within the target space.
[0175] Wherein, the above target sound source can be a sound source that emits the audio to be recognized. For example, the target sound source can be a simulated user.
[0176] In some embodiments, the simulation device can randomly determine another position point b in the target space 1101, which is different from the position point a, that is, the three-dimensional coordinates corresponding to the position point b are different from the position information 1. The simulation device can use the three-dimensional coordinates of the position point b as the position information 3 of the target sound source, such as called the second position information.
[0177] For example, as Figure 15 shown, the position point 1501 is randomly determined in the target space 1101, and the position point 1501 is different from the position point 1201. To simulate the user (target sound source) speaking the audio to be recognized at the position point 1501. In this way, the three-dimensional coordinates of the above position point 1501 are the position information 3 of the target sound source.
[0178] As an implementation manner, the process of randomly determining the position point 1501 and obtaining the corresponding position information 3 is as follows:
[0179] (1) As Figure 16As shown, based on the center point of the terminal device 1202, a corresponding spherical coordinate system 1601 (the first spherical coordinate system) is established. The above-mentioned spherical coordinate system 1601 includes an x-axis 3, a y-axis 3, and a z-axis 3.
[0180] (2) Randomly generate a set of ρ, θ, and φ as the first coordinates randomly generated in the first spherical coordinate system. Among them, as Figure 16 shown, ρ is the distance value of the randomly generated position point 1501 from the origin of the spherical coordinate system 1601 (the center point of the terminal device 1202), that is, the modulus length of the position vector corresponding to the position point 1501. θ is the angle between the projection of the position vector corresponding to the randomly generated position point 1501 on the x-y plane and the x-axis 3, also known as the azimuth angle. In addition, the x-y plane is the plane determined by the x-axis 3 and the y-axis 3. φ is the angle between the position vector corresponding to the randomly generated position point 1501 and the z-axis 3, also known as the polar angle. The random value range of θ is (-180°, 180°), the random value range of φ is (-30°, 180°), and the random value range of ρ is (0, infinity). The numerical values of the above random value ranges are only examples. In other embodiments, other random value ranges can also be pre-configured.
[0181] (3) Based on the spherical coordinates (ρ, θ, φ) and the following formulas:
[0182] x = ρ * sinθ * cosφ,
[0183] y = ρ * sinθ * sinφ,
[0184] z = ρ * cosθ;
[0185] Project the position point 1501 onto the conversion coordinate system to obtain the corresponding three-dimensional coordinates (the second coordinates). Among them, the conversion coordinate system is a rectangular coordinate system established with the center point of the terminal device 1202 as the origin. The above x is the coordinate value on the x-axis of the conversion coordinate system after the spherical coordinates (ρ, θ, φ) are converted, y is the coordinate value on the y-axis of the conversion coordinate system after the spherical coordinates (ρ, θ, φ) are converted, and z is the coordinate value on the z-axis of the conversion coordinate system after the spherical coordinates (ρ, θ, φ) are converted.
[0186] (4) Project the three-dimensional coordinates of the position point 1501 in the conversion coordinate system onto the target coordinate system to obtain the position information 3 corresponding to the position point 1501. The process of calculating the position information 3 can refer to the process of converting the three-dimensional coordinates of the microphone 1302 in the extended coordinate system 1402 into the three-dimensional coordinate system in the target coordinate system in the foregoing embodiments, which will not be elaborated here.
[0187] In other embodiments, after obtaining the position information 3 corresponding to the randomly generated position point 1501, it is possible to detect whether the position point 1501 is within the target space 1101. Exemplarily, the ways to check whether the position point 1501 is within the target space 1101 include: checking whether the coordinate value of the position point 1501 on the x-axis 1 is less than or equal to the length of the target space 1101, checking whether the coordinate value of the position point 1501 on the y-axis 1 is less than or equal to the width of the target space 1101, and checking whether the coordinate value of the position point 1501 on the z-axis 1 is less than or equal to the height of the target space 1101. If all of the above check results are yes, it is determined that the position point 1501 is located within the target space 1101. If any one of the above check results is no, it is determined that the position point 1501 is not within the target space 1101. In the case where it is determined that the position point 1501 is not within the target space 1101, S102-4 is re-executed.
[0188] As another implementation, a three-dimensional coordinate value can be randomly generated, including the coordinate value on the x-axis 1, the coordinate value on the y-axis 1, and the coordinate value on the z-axis 1, and this three-dimensional coordinate value belongs to the target space 1101, and the randomly generated three-dimensional coordinate value is used as the position information 3 of the target sound source within the target space.
[0189] S102-5, determine the position information 4 of the interfering sound source within the target space.
[0190] Among them, the above interfering sound source can be a sound source that emits interfering audio. Exemplarily, as Figure 17 shown, the above interfering sound source can be the speaker 1701 of the terminal device 1202. Correspondingly, the way to determine the position information 4 (the third position information) of the speaker 1701 can refer to the implementation details of determining the position information 2 of the microphone 1302 in S102-3, which will not be elaborated here.
[0191] Exemplarily again, the above interfering sound source can also be a sound source independent of the terminal device 1202. Correspondingly, the way to determine the position information 4 of this type of interfering sound source can refer to the implementation details of determining the position information 3 of the target sound source in S102-4, which will not be elaborated here.
[0192] In some embodiments, there is no necessary sequence among the above S102-2 to S102-5.
[0193] S102-6, determine the RIR1 corresponding to the audio to be recognized according to the first characteristic information, the position information 3 of the target sound source, and the position information 2 of the microphone.
[0194] Among them, the Room Impulse Response (RIR) can characterize the acoustic response of a space and can be used for room equalization and calculating acoustic parameters in the target space, etc. For example, it is used to calculate the reverberant audio formed during the propagation of sound waves in the target space. Figure 18 shows the time-domain graph corresponding to the RIR when audio data propagates in the target space. As Figure 18 shown, as time increases, the amplitude value of the room impulse response generated by the audio data in the target space becomes smaller.
[0195] Exemplarily, each propagating audio in the target space corresponds to an RIR, where the RIR corresponding to the audio to be recognized is RIR1, also known as the first RIR.
[0196] In some embodiments, the method for calculating the RIR1 corresponding to the audio to be recognized can be: input the first feature information, the position information of the target sound source 3, and the position information of the microphone 2 into a pre-configured RIR generation tool to obtain the corresponding RIR1.
[0197] Among them, available RIR generation tools can refer to related technologies. For example, they include pyroomacoustics or gpuRIR tools, which will not be elaborated here for the time being.
[0198] S102-7, determine the RIR2 corresponding to the interfering audio according to the first feature information, the position information of the interfering source 4, and the position information of the microphone 2.
[0199] In some embodiments, the above S102-7 can refer to S102-6, which will not be elaborated here for the time being. In the embodiments of the present application, when there are multiple interfering audios, the RIR2 of each interfering audio, also known as the second RIR, can be determined in the same way.
[0200] S102-8, simulate the mixed audio collected by the microphone according to the audio to be recognized, RIR1, the interfering audio, and RIR2.
[0201] In some embodiments, the simulation device can simulate the reverberant audio formed during the propagation of the audio to be recognized in the target space 1101 based on the audio to be recognized and RIR1, such as the first reverberant audio.
[0202] In some embodiments, the simulation device can also simulate the reverberant audio formed during the propagation of the interfering audio (such as the audio played by the speaker) in the target space 1101 by combining the interfering audio and RIR2. Then, the reverberant audios of the above audio to be recognized and the interfering audio in the target space are superimposed to obtain the simulated mixed audio, such as the second reverberant audio.
[0203] Exemplarily, the simulation device can obtain the corresponding reverberant audio by convolving the audio with the RIR. Correspondingly, the way to collect the mixed audio by the analog microphone can be: according to the audio to be recognized, RIR1, the interfering audio, RIR2, combined with the formula:
[0204]
[0205] Generate the mixed audio. Wherein, the above x r is the generated mixed audio, t is time, and x r (t) includes the audio amplitude corresponding to each time point in the mixed audio. The above x1 is the audio to be recognized, x1(t) includes the audio amplitude corresponding to each time point in the audio to be recognized, RIR1 is the room impulse response of the audio to be recognized in the target space, and RIR1(t) includes the amplitude of the impulse response corresponding to the audio to be recognized at each time point in the target space. n i is the i-th interfering audio, and n i (t) includes the audio amplitude corresponding to each time point in the i-th interfering audio, and i is a positive integer. RIR i+1 is the room impulse response of the i-th interfering audio in the target space. RIR i+1 (t) includes the amplitude of the impulse response corresponding to the i-th interfering audio at each time point in the target space.
[0206] In summary, when there are multiple interfering audios, the reverberant audios corresponding to the multiple interfering audios can be obtained, for example, the convolution result between the interfering audio and the corresponding RIR. Then, the reverberant audios corresponding to all the interfering audios and the reverberant audio of the audio to be recognized are superimposed as the mixed audio collected by the simulated microphone.
[0207] In other embodiments, before generating the mixed audio, the convolution result between the interfering audio and the corresponding RIR (such as called the convolution audio) can also be aligned with the interfering audio to obtain the aligned audio corresponding to the convolution audio. Then, random delay processing is performed on the above aligned audio to obtain the reverberant audio corresponding to the interfering audio. To eliminate the problem that the audio signal has too large a delay after convolution processing, resulting in the finally simulated mixed audio not conforming to the actual situation.
[0208] As an implementation manner, the way to align the convolution result between the interfering audio and the corresponding RIR (such as called the convolution audio) with the interfering audio can be:
[0209] (1) Based on the interfering audio and the convolution audio, call the np.correlate function to obtain the cross-correlation sequence.
[0210] Wherein, the above np.correlate function is used to calculate the cross-correlation value between different digital signals.
[0211] In some embodiments, the interfering audio and the convolutional audio may be numerical signals. The digital signal includes a plurality of data points, where the value of the above data points may indicate the audio amplitude, and when the value of the data point is 0, it indicates no amplitude. Additionally, each data point corresponds to a time point. The data points in the above digital signal are sorted from left to right in the order of the corresponding time points. For example, the interfering audio is (0, 0, 2, 3, 5, 4, 8, 6, 9, 4, 6), and the convolutional audio is (0, 0, 0, 0, 0, 3, 4, 4, 5, 9, 7). The above is only an example, and the interfering audio and the convolutional audio may also correspond to more or fewer data points.
[0212] In some embodiments, the simulation device may execute the command np.correlate(interfering audio, convolutional audio, mode = "full") to obtain the cross-correlation sequence returned by the np.correlate function. The cross-correlation sequence includes a plurality of data points. The number of data points in the cross-correlation series is a + b - 1, where a is the number of data points of the convolutional audio and b is the number of data points of the interfering audio.
[0213] (2) Find the maximum cross-correlation value in the above cross-correlation sequence and determine the position corresponding to the maximum cross-correlation value. The positions of the data points in the cross-correlation sequence start from 0, that is, in the cross-correlation sequence, the position of the data point arranged on the leftmost side is 0, and the positions of other data points increase sequentially from left to right.
[0214] (3) Determine the target offset 1 according to the position of the maximum cross-correlation value and the number of data points of the convolutional audio. For example, according to the position of the maximum cross-correlation value and the number of data points of the convolutional audio, use the formula:
[0215] P = g - a + 1;
[0216] Calculate the target offset 1. Where P is the target offset 1, g is the position of the maximum cross-correlation value, and a is the number of data points of the convolutional audio.
[0217] (4) Perform offset processing on the convolutional audio according to the target offset 1 to align the convolutional audio with the interfering audio.
[0218] Exemplarily, when P is greater than 0, shift the convolutional audio to the left by P data points. When P is less than 0, shift the convolutional audio to the right by |P| data points. |P| is the absolute value of P.
[0219] As an implementation, shifting the convolutional audio to the left by P data points can be as follows: A new signal A=(0, 0, 0, 0, 0, 0, 0) with all data points having a value of 0 is created, and the data length of this new signal A is the same as that of the convolutional audio. The (P + 1)-th data point of the convolutional audio is assigned to the first data point of signal A; the (P + m)-th data point of the convolutional audio is assigned to the m-th data point of signal A, where m is a positive integer greater than 1 and P + m is not greater than a. In this way, the signal A obtained after assignment is the audio obtained by shifting the convolutional audio to the left by P data points, that is, the aligned audio mentioned in the foregoing embodiment. For example, if the convolutional audio is (0, 0, 0, 1, 2, 3, 4), after shifting to the left by 3 data points (i.e., P equals 3), (1, 2, 3, 4, 0, 0, 0) is obtained.
[0220] After obtaining the aligned audio, the simulation device can randomly delay the above-mentioned aligned audio.
[0221] As an implementation, the method of randomly delaying the audio can include:
[0222] Within a preset time delay interval, a time length is randomly determined as the delay duration. Among them, the above-mentioned time delay interval can be a time interval determined by statistically analyzing the duration values 1 corresponding to various types of terminal devices. Among them, the above-mentioned duration value 1 is the time consumption between the actual playback of an audio by the speaker of the terminal device and the actual acquisition of the audio by the microphone of the device.
[0223] Then, according to the delay duration, the number of data points to be shifted to the right is determined, which can also be referred to as the target offset 2. Among them, there is a fixed time interval between two adjacent data points in the audio data. According to the delay duration and the time interval between two adjacent data points, the number of data points to be shifted to the right is determined. For example, calculate the quotient between the two, and subtract 1 from the integer part of the quotient to obtain the number of data points to be shifted to the right.
[0224] Finally, the above-mentioned aligned video is shifted to the right according to the target offset 2.
[0225] Exemplarily, a new signal B=(0, 0, 0, 0, 0, 0, 0) with all data points having a value of 0 is established. The data length of this new signal B is the same as that of the aligned audio. Taking the target offset 2 as h, for example, the value of the first data point of the aligned audio can be assigned to the (h + 1)-th data point of signal B. Here, h is a positive integer. The value of the second data point of the interfering audio is assigned to the (h + 2)-th data point of signal B, and so on. The value of the (j - h)-th data point of the interfering audio is assigned to the j-th data point of signal B, where j is the total number of data points of the aligned audio and is a positive integer. In this way, the signal B obtained after assignment is the audio of the aligned audio shifted to the right by h data points. For example, when the aligned audio is (1, 2, 3, 4, 0, 0, 0) and the target offset 2 is equal to 1 data point, after shifting to the right, the audio (0, 1, 2, 3, 4, 0, 0) is obtained, that is, the reverberant audio corresponding to the interfering audio mentioned in the foregoing embodiments.
[0226] S103. Eliminate the interfering audio in the mixed audio to obtain a simulated audio sample.
[0227] In some embodiments, under preset conditions, the simulation device can input the interfering audio and the mixed audio into the AEC module. Among them, the interfering audio is used as the reference audio in the echo cancellation process, and the mixed audio is used as the audio collected by the microphone. After being processed by the AEC module, the AEC module outputs the corresponding simulated audio (the third audio data). Among them, there are differences between the above-mentioned simulated audio and the audio to be recognized. Compared with the audio to be recognized, the simulated audio has partial audio damage. Then, the audio label corresponding to the audio to be recognized is associated with the obtained simulated audio to obtain a simulated audio sample (the target audio sample). The above-mentioned preset conditions are the scenario conditions indicating that echo collection will occur. For example, the preset conditions include: the mixed audio is simulated as the audio collected by the microphone of the first device, and the second audio data is simulated as the audio played by the speaker of the first device.
[0228] In this way, the simulation device can randomly generate a large number of different simulated audio samples by randomly selecting the audio to be recognized, the interfering audio, the first feature information of the target space, the position of the target sound source, the position of the interfering sound source, and / or the position of the microphone of the terminal device.
[0229] In some embodiments, the simulation device uses a large number of simulated audio samples (target audio samples) obtained by simulation to train a speech recognition model, so as to improve the recognition ability of the speech recognition model for audio data that has been processed (processed by the AEC module and / or other denoising modules). After that, the trained speech recognition model can be configured into various terminal devices. Then, the terminal device can perform speech recognition based on the speech recognition model.
[0230] Exemplarily, an embodiment of the present application provides a speech recognition method, and the above method may further include: after the microphone of the terminal device (the second device) collects audio data, preprocess the collected audio data to obtain target audio. For example, based on the reference audio played by the speaker (for example, Figure 7 in, the audio data 702 corresponding to the interfering audio that causes echo, such as called the sixth audio data), the audio data collected by the microphone (the seventh audio data), and the AEC module, perform echo cancellation on the audio data to obtain target audio (the eighth audio data). The terminal device inputs the target audio into the speech recognition model to obtain the corresponding speech recognition result.
[0231] In the above embodiment, after the terminal device configures the speech recognition model trained by the above method, in different audio scenarios, the audio recognition effect is better than that of the terminal device in the related art. The corresponding data is shown in Table 1:
[0232] Table 1
[0233]
[0234] Among them, A% is greater than B%, C% is greater than A%, and in addition, B% is greater than 12%. The above recognition effect is: the recognition accuracy rate of the audio in the medium noise scenario; the above relative improvement is: the improvement of the speech recognition accuracy rate compared with the terminal device in the related art.
[0235] An embodiment of the present application also provides an electronic device. The above electronic device can execute one or more of the data simulation method, speech recognition model training method, and speech recognition method mentioned in the foregoing embodiments. Exemplarily, the above electronic device may be the model training device for training the speech recognition model in the foregoing embodiments. Again exemplarily, the above electronic device may also be the terminal device actually using the speech recognition model in the foregoing embodiments. Again exemplarily, the above electronic device may also be the simulation device for generating simulation audio samples in the foregoing embodiments.
[0236] The electronic device may include: a memory and one or more processors. The memory and the processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can execute the respective steps in the above embodiments. Of course, the electronic device includes but is not limited to the above memory and one or more processors.
[0237] Please refer to Figure 19 , Figure 19 which shows a possible schematic diagram of the hardware structure of the electronic device 100:
[0238] As Figure 19As shown in the figure, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0239] Among them, the above-mentioned sensor module 180 may include sensors such as a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, and a bone conduction sensor.
[0240] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0241] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0242] The controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0243] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the said memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0244] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0245] It can be understood that the interface connection relationships between the modules illustrated in this embodiment are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In some other embodiments, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0246] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change the display information.
[0247] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc.
[0248] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0249] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element (image sensor). The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0250] The camera 193 is used to capture static images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In some embodiments, the electronic device 100 can include N cameras 193, where N is a positive integer greater than 1.
[0251] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0252] A video codec is used to compress or decompress digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0253] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0254] The embodiment of the present application also provides a chip system, which can be applied to the electronic device in the foregoing embodiment. The chip system includes at least one processor and at least one interface circuit. The processor can be the processor in the above-mentioned electronic device. The processor and the interface circuit can be interconnected by a line. The processor can receive and execute computer instructions from the memory of the above-mentioned electronic device through the interface circuit. When the computer instructions are executed by the processor, the electronic device can execute each step in the foregoing embodiment. Of course, the chip system can also include other discrete devices, and the embodiment of the present application does not make specific limitations thereto.
[0255] In some embodiments, through the description of the above implementation manners, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0256] In each embodiment of the embodiments of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0257] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.
[0258] As described above, the above is only the specific implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present application should be covered by the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.
Claims
1. A data simulation method, characterized in that, The method includes: Obtaining first audio data and second audio data, where the first audio data corresponds to a first tag, and the first tag includes the content text of the first audio data; Generating mixed audio corresponding to the first audio data and the second audio data, where the mixed audio is audio simulated to be collected by a microphone; Eliminating the second audio data from the mixed audio to obtain third audio data, where the third audio data is different from the first audio data; Associating the third audio data with the first tag to obtain a target audio sample, where the target audio sample is used to train a speech recognition model.
2. The method according to claim 1, characterized in that, The generating of the mixed audio corresponding to the first audio data and the second audio data includes: Randomly generating first feature information, where the first feature information includes first dimension information for indicating a spatial size and / or reverberation time RT60 within the space; Respectively determining a first room impulse response RIR corresponding to the first audio data, and a second RIR corresponding to the second audio data; where the first RIR is an acoustic response simulating the first audio data in the target space indicated by the first feature information; the second RIR is an acoustic response simulating the second audio data in the target space indicated by the first feature information; After obtaining the first RIR and the second RIR, generating the mixed audio, where the mixed audio is audio obtained by superimposing a first reverberant audio and a second reverberant audio, the first reverberant audio is audio determined by the first audio data and the first RIR, and the second reverberant audio is audio determined by the second audio data and the second RIR.
3. The method according to claim 2, wherein Before the respectively determining the first room impulse response RIR corresponding to the first audio data, and the second RIR corresponding to the second audio data, the method further includes: Simulating first position information of the microphone for collecting audio in the target space; Simulating second position information of the sound source corresponding to the first audio data in the target space; Simulating third position information of the sound source corresponding to the second audio data in the target space; The determining of the first room impulse response RIR corresponding to the first audio data includes: determining the first RIR corresponding to the first audio data according to the first feature information, the first position information, and the second position information; The determining of the second RIR corresponding to the second audio data includes: determining the second RIR corresponding to the second audio data according to the first feature information, the first position information, and the third position information.
4. The method according to claim 3, characterized in that, The simulating of the first position information of the microphone for collecting audio in the target space includes: Obtaining second dimension information of a first device configured with the microphone, where the second dimension information includes the configuration position of the microphone on the body of the first device; Generating fourth position information of the first device in the target space; Determine the first position information corresponding to the microphone of the first device in a random posture according to the configured position of the microphone in the fourth position information and the second dimension information.
5. The method according to claim 4, characterized in that The first device further includes a speaker, and the sound source of the second audio data is the speaker of the first device.
6. The method according to claim 4, characterized in that, The simulating the second position information of the sound source corresponding to the first audio data in the target space includes: Construct a first spherical coordinate system, and the origin of the first spherical coordinate system is the position point indicated by the fourth position information; In the first spherical coordinate system, randomly determine a first coordinate, and the first coordinate includes a distance, an azimuth angle, and a polar angle; wherein, the first coordinate indicates the position of the sound source of the first audio data relative to the first device, the azimuth angle is greater than -180° and less than 180°; the polar angle is greater than -30° and less than 180°; Convert the first coordinate information into a second coordinate, and the second coordinate is a rectangular coordinate; Based on the second coordinate and the fourth position information, determine the second position information of the sound source of the first audio data in the target space.
7. The method according to any one of claims 2-6, characterized in that, The manner of determining the first reverberant audio from the first audio data and the first RIR includes: performing convolution processing based on the first audio data and the first RIR to obtain the first reverberant audio; The manner of determining the second reverberant audio from the second audio data and the second RIR includes: performing convolution processing based on the second audio data and the second RIR to obtain the second reverberant audio.
8. The method according to claim 7, wherein After performing the convolution processing based on the second audio data and the second RIR and before obtaining the second reverberant audio, the method further includes: Obtain the convolution audio corresponding to the second audio data and the second RIR; Perform audio alignment on the convolution audio and the second audio data; Determine a first duration, and the first duration belongs to a preset delay interval; The obtaining the second reverberant audio includes: generating the second reverberant audio according to the first duration and the aligned convolution audio; wherein, there is a delay of the first duration between the second reverberant audio and the aligned convolution audio.
9. The method according to any one of claims 1 to 8, characterized in that The eliminating the second audio data in the mixed audio to obtain the third audio data includes: Under preset conditions, input the mixed audio and the second audio into an echo cancellation module to obtain the third audio data; The preset conditions include: the mixed audio is an audio simulated to be collected by the microphone of the first device, and the second audio data is an audio simulated to be played by the speaker of the first device.
10. A method for training a speech recognition model, characterized in that, The method includes: Obtain a target audio sample; wherein, the target audio sample is an audio simulated according to the method of any one of claims 1-9; Use the target audio sample to train a pre-configured speech recognition model until the speech recognition model meets the pre-configured model convergence condition.
11. The method according to claim 10, wherein The method further includes: Obtain fourth audio data with a second tag, where the fourth audio data is the original audio without denoising, and the second tag contains the content text of the fourth audio data; Utilize the fourth audio data to train a pre-configured speech recognition model until the speech recognition model meets the pre-configured model convergence condition.
12. The method according to claim 10, characterized in that, The method further includes: Obtain fifth audio data with a third tag, where the fifth audio data is the recorded and denoised audio, and the third tag contains the content text of the fifth audio data; Utilize the fifth audio data to train a pre-configured speech recognition model until the speech recognition model meets the pre-configured model convergence condition.
13. A speech recognition method, characterized in that, Applied to a second device, where the second device is configured with a speech recognition model and an echo cancellation module, and the speech recognition model is a model obtained by using the speech recognition model training method according to any one of claims 10-12, and the speech recognition method includes: When the second device plays sixth audio data, seventh audio data is collected; After inputting the seventh audio data and the sixth audio data into the echo cancellation module, the eighth audio data is obtained; Utilize the speech recognition model to recognize the content text corresponding to the eighth audio data.
14. An electronic device, characterized in that, The electronic device includes: a processor and a memory, where the memory is used to store computer instructions, and when the processor executes the computer instructions, the electronic device is caused to execute the method according to any one of claims 1-13.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instruction, and when the computer program or instruction runs on a computer, the computer is caused to execute the method according to any one of claims 1-13.