Data simulation method and electronic device

CN120199238BActive Publication Date: 2026-09-08HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311738846.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2026-09-08
Estimated Expiration
2043-12-15

AI Technical Summary

Technical Problem

然而,在此过程中,录制音频样本的效率较低

Benefits of technology

[0020] In the above embodiments, the mixed audio collected when the sound source of the first audio data is in different and reasonable positions in the target space can be simulated, so that the simulated mixed audio can cover various audio scenarios and is closer to reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199238B_ABST
    Figure CN120199238B_ABST
Patent Text Reader

Abstract

The application provides a data simulation method and an electronic device, and relates to the field of human-computer interaction. The problem of poor speech recognition effect of a speech recognition model is solved. The specific scheme is as follows: first audio data and second audio data are obtained, the first audio data corresponds to a first label, and the first label contains content text of the first audio data; a mixed audio corresponding to the first audio data and the second audio data is generated, the mixed audio is audio collected by a microphone; the second audio data in the mixed audio is eliminated to obtain third audio data, the third audio data is different from the first audio data; the third audio data is associated with the first label to obtain a target audio sample, and the target audio sample is used for training a speech recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction, and more particularly to a data simulation method and an electronic device. Background Technology

[0002] With the development of speech recognition technology, voice interaction functions are widely used in various terminal devices and have become a commonly used interaction function for most users. The key to realizing voice interaction functions lies in the model used to recognize speech, such as a speech recognition model.

[0003] In related technologies, recorded audio samples are used to train speech recognition models, thereby improving the models' speech recognition capabilities. However, recording audio samples is inefficient in this process. Furthermore, the trained models do not perform consistently well in practical applications. Summary of the Invention

[0004] In view of this, this application provides a data simulation method and electronic device that generates a large number of processed audio samples through automatic simulation, thereby reducing the difficulty of obtaining training samples and improving the ability of the trained speech recognition model to recognize processed audio.

[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: In a first aspect, an embodiment of this application provides a data simulation method, the method comprising: The system acquires first audio data and second audio data. The first audio data corresponds to a first tag, which contains the text content of the first audio data. For example, the first and second audio data can come from different audio databases or the same audio database. The first audio data can be audio with actual semantic meaning, such as recorded statements. The second audio data can be audio with actual semantic meaning or other audio without semantic meaning.

[0006] Then, a mixed audio is generated corresponding to the first and second audio data. This mixed audio simulates the audio captured by the microphone. Specifically, the mixed audio can simulate the audio captured by the microphone during the propagation of the first and second audio data. In the mixed audio, the second audio data is considered interference audio relative to the first audio data.

[0007] The simulated mixed audio (audio with interference) is more efficient to generate than the actual recorded audio data with interference.

[0008] Then, the second audio data in the mixed audio is removed to obtain the third audio data. This third audio data is the processed audio and differs from the original first audio data; for example, it may contain audio artifacts or residual data points from the second audio data. Next, the third audio data is associated with the first label to obtain the target audio sample. Training the speech recognition model using the target audio sample helps improve the model's ability to recognize processed audio.

[0009] In the above embodiments, by improving the method of acquiring target audio samples, it is helpful to train a speech recognition model with more stable speech recognition performance. For example, the speech recognition model can stably recognize various types of processed audio data.

[0010] In some embodiments, generating a mixed audio corresponding to the first audio data and the second audio data includes: randomly generating first feature information, wherein the first feature information includes first size information for indicating the size of a space and / or indicating the reverberation time RT60 within the space; determining a first room impulse response RIR corresponding to the first audio data, and a second RIR corresponding to the second audio data; wherein the first RIR simulates the acoustic response of the first audio data in the target space indicated by the first feature information; the second RIR simulates the acoustic response of the second audio data in the target space indicated by the first feature information; after obtaining the first RIR and the second RIR, generating the mixed audio, wherein the mixed audio is an audio obtained by superimposing the first reverberation audio and the second reverberation audio, wherein the first reverberation audio is the audio determined by the first audio data and the first RIR, and the second reverberation audio is the audio determined by the second audio data and the second RIR.

[0011] In the above embodiments, the impact of audio propagation space on the mixed audio is fully considered when simulating mixed audio. By using different first feature information, environments with different effects on audio propagation can be simulated, such as audio scenes.

[0012] In addition, simulating mixed audio collected under different audio scenarios is less difficult than reproducing actual audio scenarios for audio collection, which is beneficial for obtaining mixed audio that covers more audio scenarios.

[0013] In some embodiments, before determining the first room impulse response (RIR) corresponding to the first audio data and the second RIR corresponding to the second audio data, the method further includes: simulating first position information of the microphone used for acquiring audio in the target space; simulating second position information of the sound source corresponding to the first audio data in the target space; simulating third position information of the sound source corresponding to the second audio data in the target space; determining the first room impulse response (RIR) corresponding to the first audio data includes: determining the first RIR corresponding to the first audio data based on the first feature information, the first position information, and the second position information; determining the second RIR corresponding to the second audio data includes: determining the second RIR corresponding to the second audio data based on the first feature information, the first position information, and the third position information.

[0014] In the above embodiments, when simulating mixed audio, the influence of the sound source position of the first audio data and the sound source position of the second audio data on the mixed audio is fully considered, so that the simulated mixed audio is more realistic.

[0015] In some embodiments, simulating the first position information of a microphone used for audio acquisition in the target space includes: obtaining second size information of a first device configured with the microphone, the second size information including the configuration position of the microphone on the body of the first device; generating fourth position information of the first device in the target space; and determining the first position information corresponding to the microphone in a random posture of the first device based on the fourth position information and the configuration position of the microphone in the second size information.

[0016] In other embodiments, after the first location information is determined, it can be checked whether the location indicated by the first location information belongs to the target space. If it does not belong, a fourth location information of the first device is regenerated, or a random posture is re-determined and the first location information of the microphone is re-determined.

[0017] In the above embodiments, the mixed audio collected by the microphone at different positions in the target space can be simulated, so that the simulated mixed audio can cover various audio scenarios.

[0018] In some embodiments, the first device further includes a speaker, and the source of the second audio data is the speaker of the first device.

[0019] In some embodiments, simulating the second location information of the sound source corresponding to the first audio data in the target space includes: constructing a first spherical coordinate system, the origin of which is the location point indicated by the fourth location information; randomly determining a first coordinate in the first spherical coordinate system, the first coordinate including distance, azimuth, and polar angle; wherein the first coordinate indicates the position of the sound source of the first audio data relative to the first device, the azimuth is greater than -180° and less than 180°, and the polar angle is greater than -30° and less than 180°; converting the first coordinate information into a second coordinate, the second coordinate being a rectangular coordinate; and determining the second location information of the sound source of the first audio data in the target space based on the second coordinate and the fourth location information.

[0020] In the above embodiments, the mixed audio collected when the sound source of the first audio data is in different and reasonable positions in the target space can be simulated, so that the simulated mixed audio can cover various audio scenarios and is closer to reality.

[0021] In other embodiments, after the second location information is determined, it can be checked whether the location indicated by the second location information belongs to the target space. If it does not belong, the second location information is re-determined to ensure that the simulated mixed audio is more realistic.

[0022] In some embodiments, determining a first reverberant audio from the first audio data and the first RIR includes: performing convolution processing based on the first audio data and the first RIR to obtain the first reverberant audio; determining a second reverberant audio from second audio data and a second RIR includes: performing convolution processing based on the second audio data and the second RIR to obtain the second reverberant audio.

[0023] In some embodiments, after performing convolution processing based on the second audio data and the second RIR, before obtaining the second reverberant audio, the method further includes: acquiring the convolutional audio corresponding to the second audio data and the second RIR; aligning the convolutional audio with the second audio data; determining a first duration, the first duration belonging to a preset delay interval; obtaining the second reverberant audio includes: generating the second reverberant audio based on the first duration and the aligned convolutional audio; wherein, there is a delay of the first duration between the second reverberant audio and the aligned convolutional audio.

[0024] In the above embodiments, audio alignment is performed before delay processing. This addresses the unreasonable audio delay caused by convolution processing while simulating a reasonable audio delay that occurs during actual data acquisition.

[0025] In some embodiments, eliminating the second audio data in the mixed audio to obtain the third audio data includes: under preset conditions, inputting the mixed audio and the second audio into an echo cancellation module to obtain the third audio data; the preset conditions include: the mixed audio is audio that is simulated by a microphone of the first device, and the second audio data is audio that is simulated by a speaker of the first device.

[0026] In the above embodiments, mixed audio with echo interference can be simulated.

[0027] Secondly, an embodiment of this application provides a speech recognition model training method, the method comprising: acquiring a target audio sample; wherein the target audio sample is audio simulated according to the method provided in the first aspect and its possible embodiments; and using the target audio sample to train a pre-configured speech recognition model until the speech recognition model satisfies the pre-configured model convergence condition.

[0028] In the above embodiments, by improving the method of acquiring target audio samples, it is helpful to train a speech recognition model with more stable speech recognition performance. For example, the speech recognition model can stably recognize various types of processed audio data.

[0029] In some embodiments, the method further includes: acquiring fourth audio data with a second label, the fourth audio data being undenoised raw audio, the second label containing the content text of the fourth audio data; and using the fourth audio data to train a pre-configured speech recognition model until the speech recognition model meets a pre-configured model convergence condition.

[0030] In some embodiments, the method further includes: acquiring fifth audio data with a third tag, the fifth audio data being recorded and denoised audio, the third tag containing the content text of the fifth audio data; and using the fifth audio data to train a pre-configured speech recognition model until the speech recognition model meets a pre-configured model convergence condition.

[0031] Thirdly, an embodiment of this application provides a speech recognition method applied to a second device. The second device is configured with a speech recognition model and an echo cancellation module. The speech recognition model is a model obtained by using the speech recognition model training method provided in the second aspect and other possible embodiments. The speech recognition method includes: the second device collecting seventh audio data while playing sixth audio data; inputting the seventh audio data and the sixth audio data into the echo cancellation module to obtain the eighth audio data; and using the speech recognition model to recognize the content text corresponding to the eighth audio data.

[0032] In the above embodiments, the second device can stably identify the required audio data in various real-world audio scenarios, thereby improving the audio recognition capability of the second device.

[0033] Fourthly, an electronic device is provided in the embodiments of this application. The electronic device includes one or more processors and a memory. The memory is coupled to the processor and is used to store computer program code, which includes computer instructions. When one or more processors execute the computer instructions, the one or more processors are used for the methods described in the first aspect, the second aspect, the third aspect, and their possible embodiments.

[0034] Fifthly, embodiments of this application provide a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the methods described in the first, second, and third aspects and their possible embodiments.

[0035] In a sixth aspect, this application provides a computer program product that, when run on the aforementioned electronic device, causes the electronic device to perform the methods described in the first aspect, second aspect, third aspect, and possible embodiments thereof.

[0036] Furthermore, the electronic devices, computer storage media, and computer program products provided in the above-mentioned aspects are all applied to the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here. Attached Figure Description

[0037] Figure 1 One of the schematic diagrams of a speech recognition scenario provided in the embodiments of this application; Figure 2 The second schematic diagram of a speech recognition scenario provided in the embodiments of this application; Figure 3 The third schematic diagram of a speech recognition scenario provided in the embodiments of this application; Figure 4 A schematic diagram illustrating the implementation principle of voice interaction provided in the embodiments of this application; Figure 5 This is a schematic diagram illustrating the principle of training a speech recognition model in some embodiments; Figure 6 An example diagram showing the acquisition of echo-containing audio data by the microphone of a terminal device; Figure 7 Example diagram for eliminating echoes in terminal devices; Figure 8 A schematic diagram illustrating the principle of training a speech recognition model provided in an embodiment of this application; Figure 9Example diagrams for generating simulated audio samples provided in embodiments of this application; Figure 10 This is one of the flowcharts illustrating the steps for generating simulated audio provided in this application embodiment; Figure 11 This is the second flowchart illustrating the steps for generating simulated audio in an embodiment of this application. Figure 12 The third flowchart of the steps for generating simulated audio provided in the embodiments of this application; Figure 13 The fourth flowchart of the steps for generating simulated audio provided in the embodiments of this application; Figure 14 Example diagram for establishing an extended coordinate system provided in the embodiments of this application; Figure 15 The fifth flowchart illustrating the steps for generating simulated audio provided in this application embodiment; Figure 16 An example diagram of the spherical coordinate system provided in the embodiments of this application; Figure 17 The sixth flowchart of the steps for generating simulated audio provided in the embodiments of this application; Figure 18 Example diagram of RIR provided for embodiments of this application; Figure 19 Example diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0038] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.

[0039] The implementation of this embodiment will now be described in detail with reference to the accompanying drawings.

[0040] The aforementioned voice interaction refers to the user speaking a sentence containing instructions, and the terminal device recognizing the sentence and responding accordingly.

[0041] For example, the aforementioned terminal device may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), smart TV, smartwatch, netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) / virtual reality (VR) device, etc. The embodiments of this application do not impose any special restrictions on the specific form of the terminal device.

[0042] Taking a mobile phone as an example as the terminal device. An exemplary scenario is... Figure 1 As shown, a user holds a mobile phone 101. The phone 101 is displaying a playback interface 102 provided by a short video application. For example, the playback interface 102 displays the screen content 103 of video data a. Additionally, the phone 101 simultaneously plays the audio corresponding to video data a, such as music 1. For example, the playback interface 102 also displays an indicator of the currently playing audio, such as indicator 104, indicating that music 1 is being played.

[0043] During the display of the playback interface 102, the user utters a statement 105, such as the statement "Play the next video".

[0044] In this scenario, mobile phone 101 collects the audio data of the user's spoken statement 105, can recognize the semantics of statement 105, and determine the control command that matches that semantics. For example, in a scenario where a short video application is running in the foreground on mobile phone 101, the semantics of statement 105 are recognized as switching the playing video. According to the business logic of the short video application, the control command that matches it is determined to be an instruction to play video data b. Here, video data b and video data a are both video data that the short video application determines to recommend to the user, and the recommendation order of video data b is after video data a.

[0045] In this way, the mobile phone 101 can respond to the user's spoken statement 105 and switch to display the playback interface 106. For example, the playback interface 106 displays a screen 107 showing video data b. Additionally, the mobile phone 101 simultaneously plays the audio corresponding to video data b, such as music 2. For example, the playback interface 106 also displays an indicator of the currently playing audio, such as indicator 108, indicating that music 2 is currently playing.

[0046] Taking a smartwatch as an example, an exemplary scenario is as follows: Figure 2As shown, the user wears a smartwatch 201, which performs alarm clock reminders, such as displaying an alarm reminder capsule 203 on the watch face 202, playing alarm clock sound effects through the speaker, and / or vibrating the watch body.

[0047] During the alarm reminder service of smartwatch 201, the user can say a statement 204, such as "Turn off the alarm".

[0048] In this scenario, the smartwatch 201 collects audio data of the user's spoken statement 204, recognizes the semantics of statement 204, and determines the control command that matches that semantics. For example, if the smartwatch 201 is executing an alarm reminder service, the semantics of statement 204 are recognized as stopping the alarm reminder service. According to the logic of the alarm reminder service, the corresponding control command is determined to be an instruction to cancel the reminder related to the alarm.

[0049] In this way, the smartwatch 201 can respond to the user's spoken words 204, such as canceling the display of the alarm reminder capsule 203, canceling the playback of the alarm reminder sound effect, and / or canceling the vibration of the watch body.

[0050] Continuing with the example of a smart TV as the terminal device, an exemplary scenario is as follows: Figure 3 As shown, the smart TV 301 is playing video data from TV channel a. For example, the smart TV 301 displays a screen 302 showing video data from TV channel a.

[0051] During the display of video playback data c, the user utters statement 303, for example, the statement "Switch to the next channel".

[0052] In this scenario, the smart TV 301 collects the audio data of the user's spoken statement 303, can recognize the semantics of statement 303, and determine the control command that matches that semantics. For example, if the smart TV 301 is playing video data from TV channel a, the semantics of statement 303 are recognized as "channel switching." According to the TV service logic, the corresponding control command is determined to be: an instruction to play video data from TV channel b. Here, TV channel b is the next adjacent channel after TV channel a.

[0053] In this way, the smart TV 301 can respond to the user's spoken words 303 and switch the video data 304 from TV channel b.

[0054] As can be seen from the above examples, voice interaction makes it more convenient for users to operate terminal devices, improving the efficiency of human-computer interaction. The implementation of this voice interaction function relies on speech recognition technology. In some embodiments, a speech recognition model can be pre-configured in the terminal device. The speech recognition model is then used to perform semantic recognition of the collected statements.

[0055] by Figure 1 Taking the scenario shown as an example, such as Figure 4 As shown, mobile phone 101 is equipped with a speech recognition model 401. This speech recognition model 401 can be a neural network model; the model structure of this neural network model can be found in relevant technologies and will not be elaborated here. After training, this neural network model possesses speech recognition capabilities. Figure 4 As shown, after the mobile phone 101 collects the audio data of statement 105, it obtains the target audio corresponding to statement 105.

[0056] For example, the target audio can be the raw audio data collected by the mobile phone 101 (e.g., the audio data of statement 105). Alternatively, the target audio can be the audio data collected by the mobile phone 101, processed to obtain the audio; this embodiment does not specifically limit this.

[0057] Mobile phone 101 can input the target audio corresponding to statement 105 into speech recognition model 401, and speech recognition model 401 can determine the recognition result corresponding to the target audio. If the recognition result is text 402 (e.g., "Play the next video"), mobile phone 101 can instruct the short video application currently running in the foreground to switch to playing the next video.

[0058] Clearly, the accuracy of the speech recognition model 401 described above will affect the voice interaction effect. In some embodiments, the speech recognition model can be trained using a set of recorded audio samples. This set of audio samples includes a large number of raw audio samples. Each raw audio sample includes an audio data point and a corresponding audio tag, which contains the semantic content of the audio data. Furthermore, the audio data in the raw audio samples can be recorded, interference-free audio corpus.

[0059] like Figure 5 As shown, the training process of the aforementioned speech recognition model can be performed by a model training device. For example, the model training device can be a device with ample computing resources, such as a server, desktop computer, or mobile phone.

[0060] The model training device can be configured with an initial model of a speech recognition model. For example... Figure 5As shown, the model training device can iteratively train the initial model based on the original audio samples from the audio sample set. This audio sample set can be configured within the model training device or within other devices that communicate with it.

[0061] The iterative training process can be found in relevant techniques, which will not be elaborated upon here. For example... Figure 5 As shown, the model training device can input the audio data of the original audio samples into the initial model, which then predicts the semantic meaning corresponding to the audio data. The model training device then determines whether the initial model has converged. For example, the model training device calculates the loss value between the predicted semantic meaning pre-output by the initial model and the audio label. Then, the model training device compares the loss value with a preset value to determine whether the initial model has converged. When the model convergence condition is not met, the model training device updates the model parameters according to the loss function during the iteration process. When the model convergence condition is met, the model training device stops iterative training, obtaining the speech recognition model.

[0062] In the example above, the convergence of the model is evaluated by checking whether the loss value is less than a preset value (one type of model convergence condition). In other examples, other model convergence conditions can be used to evaluate model convergence; for example, another model convergence condition could be that the model has reached a specified number of iterations. In practical applications, the model training equipment can be pre-configured with the corresponding model convergence conditions.

[0063] This speech recognition model has the ability to recognize the semantics of audio data. After the speech recognition model is embedded in a terminal device, the terminal device also has the ability to recognize the semantics of audio data.

[0064] In some possible embodiments, the model training device can also be a terminal device that actually uses the speech recognition model. That is, after training the speech recognition model, the model training device can also possess the ability to recognize the semantics of audio data based on the speech recognition model. Continuing with the example of a mobile phone 101 configured with a speech recognition model, as follows... Figure 6 and Figure 7 As shown, in actual use cases, the audio data (raw audio data) actually collected by the microphone of mobile phone 101 not only contains the audio of the statement to be recognized, but also includes interference audio. This interference audio is audio data that does not need to participate in speech recognition, such as the echo audio generated by music 1 already played on mobile phone 101 (which can be simply referred to as the played music 1). The statement to be recognized can be a statement spoken by the user.

[0065] like Figure 6As shown, when the mobile phone 101 is playing music 1 through the speaker 601, the raw audio data collected by the microphone 602 can be a mixture of audio containing the audio of "play the next video" and the echo audio of music 1.

[0066] exist Figure 6 In the scenario shown, the interfering audio can be music 1 played by speaker 601. In other scenarios, the interfering audio may also include audio information emitted by other sound sources.

[0067] To avoid the influence of interfering audio on the speech recognition results, in some embodiments, the raw audio data needs to be preprocessed to remove interfering audio before the mobile phone 101 actually enables the speech recognition model for speech recognition.

[0068] like Figure 7 As shown, mobile phone 101 can play music 1 through its speaker based on audio data 702. For example, the advanced digital signal processor (ADSP) in mobile phone 101 can determine the driving signal for speaker 601 based on audio data 702. The driving signal can be a digital signal; if mobile phone 101 is equipped with multiple speakers 601, the driving signal for each speaker 601 can be different to achieve better sound effects. After obtaining the driving signal for speaker 601, the driving signal can be converted into an analog signal and transmitted to speaker 601, causing speaker 601 to play the sound of audio 1.

[0069] While the phone 101 is playing music 1, the phone 101's microphone 602 can capture raw audio data 701. The raw audio data 701 can be a mixture of audio between the audio "play next video" and the echo audio of music 1.

[0070] After acquiring the original audio data 701, the original audio data 701 (audio data acquired by the microphone) and the audio data 702 of the interfering audio (also known as the reference audio) can be input into the acoustic echo cancellation (AEC) module to eliminate the echo generated by music 1 in the original audio data 701.

[0071] exist Figure 7 In the scenario shown, mobile phone 101 can read the audio data 702 corresponding to the music 1 being played by speaker 601 from ADSP.

[0072] Next, mobile phone 101 inputs the original audio data 701 and audio data 702 into the AEC module. Based on the original audio data 701 and audio data 702, the AEC module can output the target audio 703. The target audio 703 can be audio data with interference eliminated. The principle of echo cancellation by the AEC module can be found in relevant technologies, and will not be elaborated here.

[0073] In other embodiments, when the interfering audio also includes audio from other interference sources (sound sources other than mobile phone 101), other noise reduction modules can be used to eliminate the other interfering audio. This application embodiment does not specifically limit this.

[0074] Whether it's the AEC module or other noise reduction modules, they all inevitably cause audio data degradation during use. Audio data with data degradation differs from the original audio samples used to train the speech recognition model. And current speech recognition models have limited ability to recognize damaged audio.

[0075] Furthermore, the mixing of the audio to be identified and the interfering audio varies, resulting in different degrees of audio impairment after processing by the AEC module. The location (time position) of audio impairment also differs depending on the audio data processed by the AEC module.

[0076] Thus, even with a high-precision speech recognition model, the speech recognition performance of the 101 phone remains unstable in different real-world application scenarios.

[0077] To address the aforementioned issues, this application provides a speech recognition model training method. While retaining the use of recorded audio sample sets for training the speech recognition model, it also adds the use of simulated audio samples and re-recorded audio samples for training the speech recognition model.

[0078] like Figure 8 As shown, the model training device can extract original audio samples (fourth audio data) from the recorded audio sample set to iteratively train the initial model until the model converges. For details, please refer to [link / reference]. Figure 5 This will not be elaborated upon here. Among them, the fourth audio data from the recorded audio sample set corresponds to an audio tag, that is, the second tag.

[0079] Continue as Figure 8 As shown, the model training device can also receive recorded audio samples (fifth audio data) from the recording device and use the recorded audio samples to iteratively train the initial model until the model converges.

[0080] The aforementioned recorded audio samples can come from one or more recording devices. These recording devices can be various terminal devices mentioned in the preceding embodiments, or devices used solely for recording audio. Each recording device is equipped with an AEC module (and / or other noise reduction modules). For example, the recording devices actually record raw audio data in various real-world sound scenarios (including audio to be identified and interfering audio). The raw audio data is then processed by the AEC module (and / or other noise reduction modules) and associated with corresponding audio tags (third tags) to obtain the corresponding recorded audio samples.

[0081] In the above example, the audio quality of the recorded audio sample is high and consistent with reality. By using the recorded audio sample to train the speech recognition model, the ability of the speech recognition model to recognize damaged audio in various sound scenarios can be improved from the perspective of quality.

[0082] Continue as Figure 8 As shown, the model training device can also use a large number of simulated audio samples to iteratively train the initial model until the model converges.

[0083] The simulated audio samples include audio that has been processed by the AEC module (and / or other denoising modules) based on real audio data. For example, the model training device can extract audio data from an audio set, then, based on the extracted audio data and the simulation module, simulate the audio processed by the AEC module (and / or other denoising modules), and then associate the simulated audio with corresponding audio tags to obtain the corresponding simulated audio samples.

[0084] In the above example, the generation efficiency of the simulated audio samples is high, and a large number of qualified audio samples can be generated quickly, reducing the difficulty of obtaining samples. Using simulated audio samples to train a speech recognition model can improve the model's ability to recognize damaged audio in various sound scenarios, thus increasing the number of samples.

[0085] In other embodiments, the initial module of the speech recognition model can be trained using one or a combination of recorded audio samples, simulated audio samples, and original audio samples to obtain a speech recognition model with stable recognition accuracy.

[0086] The following describes the implementation details of a data simulation method provided in the embodiments of this application, with reference to the accompanying drawings. This data simulation method can be applied to a device configured with a simulation module, such as a simulation device, which can generate the simulated audio samples mentioned in the foregoing embodiments. Exemplarily, the simulation device can be the model training device mentioned in the foregoing embodiments, thus data simulation and model training can be performed by the same device. More exemplaryly, the simulation device can also be other devices besides the model training device provided in the foregoing embodiments, thus data simulation and model training can be performed by different devices. This application does not specifically limit the specific implementation of this method.

[0087] In some embodiments, such as Figure 9 As shown, the simulation module described above may include an audio simulation submodule and an AEC module. The steps by which the simulation device performs the data simulation method described above are as follows: S101, acquire at least two audio data points.

[0088] Among them, the above at least two audio data include interfering audio (such as the second audio data) and audio to be identified (such as the first audio data).

[0089] In some embodiments, such as Figure 9 As shown, both the interference audio and the audio to be identified can come from the same audio set. For example, the simulation device can randomly select two different original audio samples from the recorded audio sample set. One original audio sample's audio data is used as the audio to be identified, and the other original audio sample is used as the interference audio. The audio data from the audio sample set all correspond to audio tags. For example, the first audio data corresponds to the first tag. The first tag is an audio label that includes the text content corresponding to the first audio data. For example, if the first audio data is the statement "Switch to the next video," the corresponding first tag is the text "Switch to the next video."

[0090] In other embodiments, the interfering audio and the audio to be identified can both come from different audio sets. For example, the audio to be identified comes from an audio sample set, while the interfering audio comes from other audio sets, such as music databases, video databases, and audio sets of various reminder sound effects.

[0091] In other embodiments, the aforementioned interference audio may also include multiple audio data. This application does not specifically limit this, and in subsequent embodiments, one audio to be identified and one interference audio will be used as an example for description.

[0092] S102, based on the above at least two audio data, simulate the mixed audio collected by the microphone of the terminal device.

[0093] The aforementioned mixed audio can be analogous to audio data collected by a microphone, and this audio data is related to both the interfering audio and the audio to be identified.

[0094] In some embodiments, the audio simulation submodule can be used to directly mix the audio to be identified and the interference audio to obtain the corresponding mixed audio.

[0095] In other embodiments, an audio simulation submodule can be used to simulate the target space for collecting audio data. Then, the mixed audio generated when the audio to be identified and the interfering audio propagate within the target space can be simulated. For example, the room impulse response (RIR) corresponding to the audio to be identified and the interfering audio can be randomly determined, and the determined RIR is related to the target space. Then, based on the RIRs corresponding to the audio to be identified and the interfering audio, the mixed audio is generated. Specific implementation details can be found in subsequent embodiments.

[0096] In other embodiments, an audio simulation submodule can be used to simulate the target space for acquiring audio data, as well as the locations of the sound sources corresponding to the audio to be identified and the interfering audio. The locations of the sound sources corresponding to the audio to be identified and the interfering audio can be different locations. Furthermore, the simulated locations of the sound sources corresponding to the audio to be identified and the interfering audio can be located within the target space. Then, the mixed audio collected by the microphone is simulated when the sound sources of the audio to be identified and the interfering audio are located at different positions within the target control.

[0097] As one implementation method, such as Figure 10 As shown, the implementation steps of S102 above can be as follows: S102-1, Generate first feature information to indicate the target space.

[0098] The first feature information mentioned above can be feature information that affects the propagation of audio signals within the target space.

[0099] For example, the size of the target space can affect the propagation range of the audio signal, and the aforementioned first feature information may include the size information of the target space. For instance, the first feature information includes the length, width, and height of the target space, also known as the first size information.

[0100] For another example, the material of the target space can affect the rate attenuation of audio energy during audio signal propagation. For instance, the first feature information mentioned above may include the reverberation time (RT) 60 corresponding to the material of the target space. Here, RT60 refers to the time required for the sound field to attenuate by 60 dB within the target space. A larger RT60 indicates a longer audio propagation time within the target space. A smaller RT60 value indicates a shorter audio propagation time within the target space. Furthermore, the RT60 value of a real space is affected by the size of the space and also by the building materials used to construct that space.

[0101] In some embodiments, such as Figure 11 As shown, the simulation device can randomly generate a set of the aforementioned length, width, height, and RT60 values, where the length, width, height, and RT60 values ​​all belong to the first feature information. In this example, the generated length, width, height, and RT60 values ​​refer to the simulated target space 1101.

[0102] In other embodiments, the simulation device can randomly generate length, width, height and RT60 values ​​within the range of values ​​corresponding to various first feature information, so that the simulated target space 1101 is closer to the actual scene.

[0103] The random value range corresponding to the aforementioned length, width, height, and RT60, etc., can be empirical values, and this application does not impose specific limitations on this. For example, the random value range of RT60 is from 1 / n seconds to m seconds, where n and m are positive integers. After determining the length, width, and height of the target space, the size of the target space is determined. In this case, the simulation device can simulate the target space constructed using different building materials using random RT60 values.

[0104] S102-2, Generate the location information of the terminal device in the target space 1.

[0105] In some embodiments, prior to S102-2, the simulation device constructs a three-dimensional coordinate system (such as a target coordinate system) relative to the target space. Each position point in the target space can be mapped to this target coordinate system, so that each position point in the target space corresponds to a target coordinate system. Furthermore, the target coordinate system includes an x-axis (1), a y-axis (1), and a z-axis (1). The three-dimensional coordinates of any point in the target space 1101 in the target coordinate system satisfy the following conditions: the coordinate value on the x-axis (1) does not exceed the length of the target space 1101; the coordinate value on the y-axis (1) does not exceed the width of the target space 1101; and the coordinate value on the z-axis (1) does not exceed the height of the target space 1101.

[0106] In some embodiments, S102-2 may be: the simulation device randomly determines a location point a in the target space, and determines the location coordinates of the location point a, for example, (a, b, c), as the location information 1 (fourth location information) of the terminal device in the target space, wherein a is not greater than the length of the target space 1101, b is not greater than the width of the target space 1101, and c is not greater than the height of the target space 1101.

[0107] For example, such as Figure 12 As shown, a location point 1201 is randomly determined in the target space, and the location coordinates of the location point 1201 are used as the location information 1 of the terminal device 1202 in the target space 1101.

[0108] exist Figure 12 In this example, location point 1201 overlaps with the center point of the simulated terminal device 1202. That is, the position coordinates of the center point of terminal device 1202 indicate the position of terminal device 1202 in target space 1101. In other examples, location point 1201 may also overlap with other locations of the simulated terminal device 1202, such as the endpoint of any edge on the body of terminal device 1202. In subsequent embodiments, the overlap between location point 1201 and the center point of terminal device 1202 will be used as an example for description.

[0109] S102-3, Based on the location information 1 of the terminal device, determine the location information 2 of the microphone of the terminal device in the target space.

[0110] In some embodiments, the simulation device may use position point 1201 as the origin and construct a two-dimensional coordinate system relative to the terminal device 1202.

[0111] The construction of a two-dimensional coordinate system relative to the terminal device 1202 can be achieved by establishing a coordinate system on a designated surface of the terminal device 1202, with the origin of the coordinate system being the position point 1201. This designated surface can be the plane with the largest area on the terminal device 1202; for example, the designated surface can be the plane where the display screen is located, or the surface opposite to the display screen.

[0112] Taking location point 1201 as the center point of terminal device 1202 and the specified surface as the plane where the display screen of terminal device 1202 is located, this example illustrates the concept.

[0113] like Figure 13As shown, the simulation device can construct a coordinate system 1301 relative to the terminal device 1202, where the coordinate system 1301 includes an x-axis and a y-axis. The terminal device 1202 can be projected onto the coordinate system 1301, that is, each point in the terminal device 1202 corresponds to a two-dimensional coordinate in the coordinate system 1301. The center point of the terminal device 1202 has two-dimensional coordinates (0,0) in the coordinate system 1301.

[0114] For example, the microphone 1302 of the terminal device 1202 can also be projected onto the coordinate system 1301, and the simulation device can determine the two-dimensional coordinates of the microphone 1302 in the coordinate system 1301 based on the size information of the terminal device 1202.

[0115] As one implementation method, after determining the product type of the terminal device to be simulated (the first device), the simulation device can obtain the size information of the terminal device, referred to as the second size information, from the product server based on the product model of the terminal device. This size information includes the length, width, thickness of the terminal device, the placement of the speaker on the device body, and the placement of the microphone on the device body. Thus, based on the size information of the terminal device, the relative positional relationship between the microphone and the center point of the terminal device can be determined, and further, the projection position of the microphone in the corresponding two-dimensional coordinate system of the terminal device can be determined.

[0116] For example, the dimensions of terminal device 1202 include: the width of the terminal device is 40cm, the length is 160cm, and the distance from the bottom microphone to the central axis of terminal device 1202 is 10cm. Figure 14 As shown, after projecting the terminal device 1202 onto coordinate system 1301, region 1401 is the projection range of the terminal device 1202. This region 1401 has a width of 40cm and a length of 160cm. Correspondingly, the microphone 1302 is 80cm from the x-axis 2 of coordinate system 1301 and 10cm from the y-axis 2 of coordinate system 1301. Figure 14 In the scenario shown, the two-dimensional coordinates of microphone 1302 in coordinate system 1301 are (10, -80).

[0117] The microphone 1302 of the aforementioned terminal device 1202 corresponds not only to two-dimensional coordinates in coordinate system 1301, but also to three-dimensional coordinates in the target space 1101 (target coordinate system). Specifically, with the size information of the terminal device 1202 unchanged, the two-dimensional coordinates of the microphone 1302 remain fixed after establishing coordinate system 1301. The three-dimensional coordinates of the terminal device 1202 may differ depending on its orientation within the target space 1101.

[0118] In some embodiments, after the two-dimensional coordinates of the microphone 1302 are determined, the simulation device can simulate the terminal device 1202 in a random posture based on the two-dimensional coordinates to obtain the corresponding three-dimensional coordinates of the microphone 1302 in the target coordinate system, which serves as the position information 2 of the microphone 1302, also known as the first position information.

[0119] As one implementation method, the process of simulating the terminal device 1202 being in a random posture to obtain the three-dimensional coordinates of the corresponding microphone 1302 in the target coordinate system can be as follows: (1) Extend coordinate system 1301 into a three-dimensional coordinate system (e.g., called the extended coordinate system). Where, for example... Figure 14 As shown, compared to coordinate system 1301, the extended coordinate system 1402 adds a z-axis 2. The origin of the extended coordinate system 1402 remains the center point of the terminal device 1202, and the three-dimensional coordinates of this center point in the extended coordinate system 1402 are (0, 0, 0). Under the extended coordinate system 1402, the three-dimensional coordinates of the microphone 1302 are (10, -80, 0).

[0120] (2) Randomly generate the rotation angle 1 of the terminal device 1202 relative to the x-axis 2 of the extended coordinate system 1402, the rotation angle 2 of the terminal device 1202 relative to the y-axis 2 of the extended coordinate system 1402, and the rotation angle 3 of the terminal device 1202 relative to the z-axis 2 of the extended coordinate system 1402. The randomly generated angles 1, 2 and 3 can simulate the terminal device 1202 being in a random posture.

[0121] In the extended coordinate system 1402 and the target coordinate system, the position of the microphone 1302 follows the attitude change of the terminal device 1202. The terminal device 1202 rotates by an angle 1 relative to the x-axis 2 of the extended coordinate system 1402, and the microphone 1302 also rotates by an angle 1 relative to the x-axis 2 of the extended coordinate system 1402. The terminal device 1202 rotates by an angle 2 relative to the y-axis 2 of the extended coordinate system 1402, and the microphone 1302 also rotates by an angle 2 relative to the y-axis 2 of the extended coordinate system 1402. The terminal device 1202 rotates by an angle 3 relative to the z-axis 2 of the extended coordinate system 1402, and the microphone 1302 also rotates by an angle 3 relative to the z-axis 2 of the extended coordinate system 1402.

[0122] (3) Based on angle 1, angle 2, angle 3 and the initial three-dimensional coordinates (10, -80, 0) of microphone 1302, combined with the three-dimensional space coordinate rotation calculation formula, the three-dimensional coordinates of microphone 1302 after the terminal device 1202 enters the random posture can be determined.

[0123] The principle of three-dimensional spatial coordinate rotation is as follows: 1) Obtain the rotation matrices around x-axis 2, y-axis 2 and z-axis 2 respectively.

[0124] For example, the rotation matrix for rotating about x-axis 2 by an angle of 1, such as clockwise rotation by γ, is as follows: ; in, Let y be the rotation matrix for clockwise rotation γ, where γ is the angle of clockwise rotation relative to x-axis 2, for example, equal to angle 1.

[0125] For example, rotating by an angle of 2 about the y-axis, such as clockwise rotation. The rotation matrix is ​​as follows: ; in, Rotate clockwise The rotation matrix, This is the angle of clockwise rotation relative to the y-axis, for example, equal to angle 2.

[0126] For example, rotating an angle 3 about the z-axis 2, such as clockwise rotation. The rotation matrix is ​​as follows: ; in, Rotate clockwise The rotation matrix, This is the angle relative to the z-axis 2, rotated clockwise, for example, equal to angle 3.

[0127] 2) Generate the corresponding target rotation matrix according to the order of rotation around x-axis 2, y-axis 2 and z-axis 2.

[0128] For example, in a scenario where the rotation is first around the x-axis (2), then around the y-axis (2), and finally around the z-axis (2), the corresponding target rotation matrix is: ; in, For the target rotation matrix, This is the rotation matrix corresponding to the z-axis 2. This is the rotation matrix corresponding to the y-axis 2. This is the rotation matrix corresponding to x-axis 2.

[0129] Correspondingly, ; For example, in a scenario where the rotation occurs first around the z-axis (2), then around the y-axis (2), and finally around the x-axis (2), the corresponding target rotation matrix is: .

[0130] 3) Based on the target rotation matrix and the initial three-dimensional coordinates (10, -80, 0) of microphone 1302 in the extended coordinate system 1402, combined with the formula: ; After the microphone 1302 is rotated, its three-dimensional coordinates in the extended coordinate system 1402 are determined. Here, R is the corresponding target rotation matrix. , , The initial three-dimensional coordinates of microphone 1302 in extended coordinate system 1402 are x0=10, y0=-80, z0=0. , , () represents the three-dimensional coordinates of the microphone 1302 in the extended coordinate system 1402 after it has been rotated.

[0131] (4) Determine the three-dimensional coordinates of the microphone 1302 in the target coordinate system based on the three-dimensional coordinates of the microphone 1302 after rotation in the extended coordinate system 1402, the three-dimensional coordinates of the center point of the terminal device 1202 in the extended coordinate system 1402, and the three-dimensional coordinates of the center point of the terminal device 1202 in the target coordinate system.

[0132] During the rotation of the terminal device 1202, the three-dimensional coordinates of the center point of the terminal device 1202 in the extended coordinate system 1402 remain unchanged, still being (0, 0, 0).

[0133] Regardless of whether it's in the extended coordinate system 1402 or the target coordinate system, the relative positional relationship between the microphone 1302 and the center point of the terminal device 1202 remains unchanged. That is, AB = A1 - B1. Here, A is the three-dimensional coordinate of the center point of the terminal device 1202 in the target coordinate system, which is (a, b, c) as determined in the aforementioned embodiment. B is the three-dimensional coordinate of the microphone 1302 in the target coordinate system, which is the positional information 2 of the microphone 1302 that needs to be determined. A1 is the three-dimensional coordinate of the center point of the terminal device 1202 in the extended coordinate system 1402, which is (0, 0, 0) as mentioned in the aforementioned embodiment. B1 is the three-dimensional coordinate of the microphone 1302 in the extended coordinate system 1402, which is the three-dimensional coordinate of the microphone 1302 obtained after rotation in the extended coordinate system 1402 as mentioned in the aforementioned embodiment.

[0134] As one implementation method, the position information 2 corresponding to microphone 1302 can be determined as B = (A-A1) + B1.

[0135] For example, the three-dimensional coordinates of microphone 1302 after rotation in extended coordinate system 1402 are (73, 6, 31), the three-dimensional coordinates of the center point of terminal device 1202 in target coordinate system are (1000, 1000, 1000), and the three-dimensional coordinates in extended coordinate system 1402 are (0, 0, 0). It can be determined that the three-dimensional coordinates of microphone 1302 in target coordinate system are (1073, 1006, 1031).

[0136] In other possible embodiments, after determining the position information 2 of the microphone 1302, it can also be determined whether the position information 2 is within the target space 1101. If the position point indicated by the position information 2 is not within the target space 1101, for example, if the x-coordinate value in the position information 2 is greater than the length of the target space 1101, the y-coordinate value in the position information 2 is greater than the width of the target space 1101, and / or the z-coordinate value in the position information 2 is greater than the height of the target space 1101, S102-2 and S102-3 are re-executed. If the position point indicated by the position information 2 is within the target space 1101, the process proceeds to S102-4.

[0137] S102-4, Randomly generate the location information of the target sound source in the target space 3.

[0138] The target sound source can be a sound source that emits the audio to be identified, such as a simulated user.

[0139] In some embodiments, the simulation device can randomly determine another location point b in the target space 1101. This location point b is different from location point a; that is, the three-dimensional coordinates corresponding to location point b are different from the location information 1. The simulation device can use the three-dimensional coordinates of this location point b as the location information 3 of the target sound source, such as referred to as the second location information.

[0140] For example, such as Figure 15 As shown, a location point 1501 is randomly determined in the target space 1101. Location point 1501 is different from location point 1201. A simulated user (target sound source) speaks the audio to be recognized at location point 1501. Thus, the three-dimensional coordinates of location point 1501 represent the location information 3 of the target sound source.

[0141] As one implementation method, the process of randomly determining location point 1501 and obtaining the corresponding location information 3 is as follows: (1) such as Figure 16 As shown, a corresponding spherical coordinate system 1601 (first spherical coordinate system) is established based on the center point of the terminal device 1202. The spherical coordinate system 1601 includes the x-axis 3, the y-axis 3 and the z-axis 3.

[0142] (2) Randomly generate a set , and , as the first coordinate randomly generated in the first spherical coordinate system. Where, for example... Figure 16 As shown, This is the distance value between the randomly generated position point 1501 and the origin of the spherical coordinate system 1601 (the center point of the terminal device 1202), that is, the magnitude of the position vector corresponding to position point 1501. The angle between the projection of the position vector corresponding to the randomly generated position point 1501 onto the xy-plane and the x-axis 3 is also called the azimuth angle. Furthermore, the xy-plane is a surface defined by the x-axis 3 and the y-axis 3. The angle between the position vector corresponding to the randomly generated position point 1501 and the z-axis 3 is also called the polar angle. The random value range is (-180°, 180°). The random value range is (-30°, 180°). The random value range is (0, infinity). The above random value range is only an example. In other embodiments, other random value ranges can be pre-configured.

[0143] (3) Based on spherical coordinates ( , , And the following formula: x= sin cos , y= sin sin , z= cos ; Projecting position point 1501 onto the transformed coordinate system yields the corresponding three-dimensional coordinates (second coordinates). The transformed coordinate system is a Cartesian coordinate system established with the center point of terminal device 1202 as its origin. The x-coordinate mentioned above is a spherical coordinate (…). , , After transformation, the coordinates on the x-axis of the transformed coordinate system are given, and the y-coordinates are given in spherical coordinates. , , After transformation, the coordinates on the y-axis of the transformed coordinate system are given, and the z-coordinates are given in spherical coordinates. , , After transformation, the coordinate values ​​on the z-axis of the transformed coordinate system.

[0144] (4) Project the three-dimensional coordinates of position point 1501 in the transformed coordinate system to the target coordinate system to obtain the position information 3 corresponding to position point 1501. The process of calculating the position information 3 described above can be referred to in the previous embodiment, which is to convert the three-dimensional coordinates of microphone 1302 in the extended coordinate system 1402 into the three-dimensional coordinate system in the target coordinate system. It will not be described again here.

[0145] In other embodiments, after obtaining the position information 3 corresponding to the randomly generated position point 1501, it is possible to detect whether the position point 1501 is within the target space 1101. For example, checking whether the position point 1501 is within the target space 1101 includes: checking whether the coordinate value of the position point 1501 on the x-axis 1 is less than or equal to the length of the target space 1101, checking whether the coordinate value of the position point 1501 on the y-axis 1 is less than or equal to the width of the target space 1101, and checking whether the coordinate value of the position point 1501 on the z-axis 1 is less than or equal to the height of the target space 1101. If all of the above checks are true, it is determined that the position point 1501 is within the target space 1101. If any of the above checks are false, it is determined that the position point 1501 is not within the target space 1101. If it is determined that the position point 1501 is not within the target space 1101, S102-4 is re-executed.

[0146] As another implementation method, a three-dimensional coordinate value can be randomly generated, which includes coordinate values ​​on the x-axis 1, y-axis 1, and z-axis 1, and the three-dimensional coordinate value belongs to the target space 1101. The randomly generated three-dimensional coordinate value is used as the position information 3 of the target sound source in the target space.

[0147] S102-5, Determine the location information of the interference source in the target space 4.

[0148] The aforementioned source of interference can be a source that emits interfering audio, for example, such as... Figure 17 As shown, the aforementioned source of interference can be the speaker 1701 of the terminal device 1202. Correspondingly, the method for determining the position information 4 (third position information) of the speaker 1701 can be referred to the implementation details of determining the position information 2 of the microphone 1302 in S102-3, which will not be elaborated here.

[0149] As another example, the aforementioned interference source can also be a sound source independent of the terminal device 1202. Correspondingly, the method for determining the location information 4 of such interference source can be found in the implementation details of determining the location information 3 of the target sound source in S102-4, which will not be elaborated here.

[0150] In some embodiments, there is no necessary order between S102-2 to S102-5.

[0151] S102-6, Based on the first feature information, the location information 3 of the target sound source, and the location information 2 of the microphone, determine the RIR1 corresponding to the audio to be identified.

[0152] Room Impulse Response (RIR) can characterize the acoustic response of a space and can be used for room equalization and calculating acoustic parameters within a target space. For example, it can be used to calculate the reverberant audio generated as sound waves propagate within a target space. Figure 18 The diagram shows the time-domain plot of the RIR (Radio Interference Ratio) as audio data propagates within the target space. For example... Figure 18 As shown, the amplitude of the room impulse response generated by the audio data in the target space decreases with time.

[0153] For example, each audio being propagated in the target space corresponds to an RIR, wherein the audio to be identified corresponds to RIR1, also known as the first RIR.

[0154] In some embodiments, the method for calculating the RIR1 corresponding to the audio to be identified may be: inputting the first feature information, the location information 3 of the target sound source and the location information 2 of the microphone into a pre-configured RIR generation tool to obtain the corresponding RIR1.

[0155] The available RIR generation tools can be found in relevant technologies, such as pyroomacoustics or gpuRIR tools, which will not be elaborated here.

[0156] S102-7, Based on the first feature information, the location information 4 of the interference source and the location information 2 of the microphone, determine the RIR2 corresponding to the interference audio.

[0157] In some embodiments, S102-7 can be referred to as S102-6, and will not be repeated here. In the embodiments of this application, when there are multiple interfering audios, the RIR2 of each interfering audio, also known as the second RIR, can be determined in the same way.

[0158] S102-8, based on the audio to be identified, RIR1, interference audio, RIR2, and mixed audio collected by the analog microphone.

[0159] In some embodiments, the simulation device may simulate the reverberant audio formed during the propagation of the audio to be identified in the target space 1101, such as referred to as the first reverberant audio, based on the audio to be identified and RIR1.

[0160] In some embodiments, the simulation device can also combine the interfering audio with RIR2 to simulate the reverberation audio formed during the propagation of the interfering audio (e.g., audio played by a speaker) within the target space 1101. Then, the reverberation audio of the audio to be identified and the interfering audio within the target space is superimposed to obtain a simulated mixed audio, referred to as the second reverberation audio.

[0161] For example, the simulation device can obtain the corresponding reverberant audio by convolving the audio with the RIR. Correspondingly, the method for acquiring mixed audio using an analog microphone can be: based on the audio to be identified, RIR1, interference audio, and RIR2, combined with the formula: ; Generate mixed audio. Among them, the above... The generated mixed audio, where t is time. This includes the audio amplitude at each time point in the mixed audio. (The above...) For audio recognition, This includes identifying the audio amplitude corresponding to each time point in the audio. To identify the room impulse response of audio in the target space, This includes identifying the amplitude of the impulse response corresponding to the audio at each time point within the target space. For the i-th interfering audio, This includes the audio amplitude corresponding to each time point in the i-th interfering audio, where i is a positive integer. Let be the room impulse response of the i-th interfering audio in the target space. This includes the amplitude of the impulse response corresponding to the i-th interfering audio at each time point within the target space.

[0162] In summary, when there are multiple interfering audio signals, the reverberation audio signals corresponding to these signals can be obtained, for example, the convolution result between the interfering audio signal and its corresponding RIR signal. Then, the reverberation audio signals corresponding to all the interfering audio signals are superimposed with the reverberation audio signals of the audio signal to be identified, resulting in a simulated mixed audio signal captured by the microphone.

[0163] In other embodiments, before generating the mixed audio, the convolution result between the interfering audio and its corresponding RIR (e.g., called the convolutional audio) can be aligned with the interfering audio to obtain the aligned audio corresponding to the convolutional audio. Then, the aligned audio is subjected to random delay processing to obtain the reverberant audio corresponding to the interfering audio. This eliminates the problem that excessive audio signal delay after convolution processing leads to unrealistic simulated mixed audio.

[0164] One implementation method is to align the convolution result between the interfering audio and the corresponding RIR (such as the convolutional audio) with the interfering audio: (1) Based on the interference audio and convolutional audio, call the np.correlate function to obtain the cross-correlation sequence.

[0165] The np.correlate function mentioned above is used to calculate the cross-correlation value between different digital signals.

[0166] In some embodiments, the interference audio and convolutional audio can be numerical signals. The digital signal includes multiple data points, where the values ​​of these data points indicate audio amplitude; a value of 0 indicates no amplitude. Furthermore, each data point corresponds to a time point. The data points in the digital signal are ordered from left to right according to the chronological order of their corresponding time points. For example, the interference audio might be (0,0,2,3,5,4,8,6,9,4,6), and the convolutional audio might be (0,0,0,0,0,3,4,4,5,9,7). This is just an example; the interference audio and convolutional audio may correspond to more or fewer data points.

[0167] In some embodiments, the simulation device can execute the command `np.correlate(interference audio, convolution audio, mode="full")` to obtain the cross-correlation sequence returned by the `np.correlate` function. The cross-correlation sequence includes multiple data points. The number of data points in the cross-correlation sequence is `a + b - 1`, where `a` is the number of data points in the convolution audio and `b` is the number of data points in the interference audio.

[0168] (2) Find the largest cross-correlation value in the above cross-correlation sequence and determine the position corresponding to the largest cross-correlation value. The position of each data point in the cross-correlation sequence starts from 0. That is, in the cross-correlation sequence, the position of the data point on the far left is 0, and the positions of other data points increase sequentially from left to right.

[0169] (3) Determine the target offset 1 based on the position of the maximum cross-correlation value and the number of data points in the convolutional audio. For example, based on the position of the maximum cross-correlation value and the number of data points in the convolutional audio, use the formula: P = g - a + 1; Calculate the target offset 1. Where P is the target offset 1, g is the position of the maximum cross-correlation value, and a is the number of data points in the convolutional audio.

[0170] (4) The convolutional audio is offset according to the target offset of 1 to align the convolutional audio with the interference audio.

[0171] For example, when P is greater than zero, the convolutional audio is shifted to the left by P data points. When P is less than zero, the convolutional audio is shifted to the right by |P| data points. |P| is the absolute value of P.

[0172] One implementation method is to shift the convolutional audio to the left by P data points: A new signal A = (0,0,0,0,0,0,0,0) is created, with all data points set to 0. The data length of this new signal A is the same as the data length of the convolutional audio. The (P+1)th data point of the convolutional audio is assigned to the first data point of signal A; the (P+m)th data point of the convolutional audio is assigned to the mth data point of signal A, where m is a positive integer greater than 1, and P+m is not greater than a. Thus, the signal A obtained after this assignment is the audio obtained by shifting the convolutional audio to the left by P data points, which is the aligned audio mentioned in the previous embodiment. For example, if the convolutional audio is (0,0,0, 1,2,3,4,), shifting it to the left by 3 data points (i.e., P equals 3) results in (1,2,3,4,0,0,0).

[0173] After obtaining the aligned audio, the simulation device can randomly delay the aligned audio.

[0174] As one implementation method, randomized audio delay methods can include: Within a preset delay interval, a random time length is determined as the delay duration. This delay interval can be calculated by statistically analyzing the duration value 1 corresponding to various terminal devices. The duration value 1 represents the time elapsed between the terminal device's speaker actually playing an audio signal and the device's microphone actually picking up that audio signal.

[0175] Then, based on the delay duration, determine the number of data points that need to be shifted to the right, also known as the target offset. In the audio data, there is a fixed time interval between any two adjacent data points. The number of data points to be shifted to the right is determined based on the delay duration and the time interval between adjacent data points. For example, calculate the quotient between the two, subtract one from the integer part of the quotient, and obtain the number of data points to be shifted to the right.

[0176] Finally, the aligned video is shifted to the right according to the target offset of 2.

[0177] For example, a new signal B = (0,0,0,0,0,0,0,0) with all data points set to 0 is created. The data length of this new signal B is the same as the data length of the aligned audio. Taking a target offset of 2 as h as an example, the value of the first data point of the aligned audio can be assigned to the (h+1)th data point of signal B, where h is a positive integer. The value of the second data point of the interfering audio is assigned to the (h+2)th data point of signal B, and so on, until the value of the jhth data point of the interfering audio is assigned to the jth data point of signal B, where j is the total number of data points in the aligned audio and is a positive integer. Thus, the signal B obtained after the assignment is the audio after the aligned audio is offset to the right by h data points. For example, if the aligned audio is (1,2,3,4,0,0,0), and the target offset of 2 equals 1 data point, after offsetting to the right, the audio obtained is (0,1,2,3,4,0,0), which is the reverberant audio corresponding to the aforementioned interfering audio.

[0178] S103, eliminates interfering audio in the mixed audio to obtain a simulated audio sample.

[0179] In some embodiments, under preset conditions, the simulation device can input interference audio and mixed audio into the AEC module. The interference audio serves as reference audio during the echo cancellation process, and the mixed audio is audio captured by the microphone. After processing by the AEC module, the corresponding simulated audio (third audio data) is output. The simulated audio differs from the audio to be identified; the simulated audio exhibits some audio degradation compared to the audio to be identified. The audio tag corresponding to the audio to be identified is then associated with the obtained simulated audio to obtain a simulated audio sample (target audio sample). The preset conditions indicate the scenario conditions under which echoes may occur. For example, the preset conditions may include: the mixed audio is simulated audio captured by the microphone of the first device, and the second audio data is simulated audio played by the speaker of the first device.

[0180] In this way, the simulation device can randomly generate a large number of different simulated audio samples by randomly selecting the audio to be identified, the interfering audio, the first feature information of the target space, the location of the target sound source, the location of the interfering sound source, and / or the location of the microphone of the terminal device.

[0181] In some embodiments, the simulation device uses a large number of simulated audio samples (target audio samples) to train a speech recognition model, improving the model's ability to recognize processed audio data (processed by an AEC module and / or other denoising modules). The trained speech recognition model can then be configured on various terminal devices. The terminal devices can then perform speech recognition based on this model.

[0182] For example, this application provides a speech recognition method, which may further include: after the microphone of a terminal device (second device) collects audio data, preprocessing the collected audio data to obtain target audio. For example, based on reference audio played through a speaker (e.g., Figure 7 In the process, the audio data 702 corresponding to the interfering audio that causes the echo (referred to as the sixth audio data), the audio data collected by the microphone (seventh audio data), and the AEC module perform echo cancellation on the audio data to obtain the target audio (eighth audio data). The terminal device inputs the target audio into the speech recognition model to obtain the corresponding speech recognition result.

[0183] In the above embodiments, after the terminal device is configured with the speech recognition model trained by the above method, the audio recognition performance is better than that of terminal devices in related technologies under different audio scenarios. The corresponding data is shown in Table 1: Table 1

[0184] Among them, A% is greater than B%, C% is greater than A%, and B% is greater than 12%. The above recognition effect is: the accuracy of audio recognition in a moderate noise scene; the above relative improvement is: the improvement in speech recognition accuracy compared to terminal devices in related technologies.

[0185] This application also provides an electronic device. The electronic device can execute one or more of the data simulation method, speech recognition model training method, and speech recognition method mentioned in the foregoing embodiments. For example, the electronic device can be a model training device used to train a speech recognition model in the foregoing embodiments. Also for example, the electronic device can be a terminal device that actually uses the speech recognition model in the foregoing embodiments. Also for example, the electronic device can be a simulation device used to generate simulated audio samples in the foregoing embodiments.

[0186] The electronic device may include a memory and one or more processors. The memory and processors are coupled. The memory stores computer program code, which includes computer instructions. When the processor executes the computer instructions, it causes the electronic device to perform the steps described in the above embodiments. Of course, the electronic device includes, but is not limited to, the memory and one or more processors described above.

[0187] Please refer to Figure 19 , Figure 19 A schematic diagram of a possible hardware structure for electronic device 100 is shown below: like Figure 19As shown, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0188] The aforementioned sensor module 180 may include sensors such as pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, and bone conduction sensors.

[0189] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0190] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0191] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0192] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0193] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0194] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0195] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0196] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.

[0197] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0198] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element (image sensor). The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0199] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include N cameras 193, where N is a positive integer greater than 1.

[0200] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0201] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0202] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0203] This application also provides a chip system that can be applied to the electronic devices described in the foregoing embodiments. The chip system includes at least one processor and at least one interface circuit. The processor may be the processor in the aforementioned electronic device. The processor and the interface circuit are interconnected via wiring. The processor can receive and execute computer instructions from the memory of the aforementioned electronic device through the interface circuit. When the computer instructions are executed by the processor, the electronic device can perform the various steps in the foregoing embodiments. Of course, the chip system may also include other discrete devices, and this application does not specifically limit this.

[0204] In some embodiments, as described above, those skilled in the art will clearly understand that, for the sake of convenience and brevity, the division of the functional modules described above is merely an example. In practical applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0205] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0206] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.

[0207] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.

Claims

1. A data simulation method, characterized in that, The method includes: Acquire first audio data and second audio data, wherein the first audio data corresponds to a first tag, and the first tag contains the content text of the first audio data; Generate a mixed audio corresponding to the first audio data and the second audio data, wherein the mixed audio is an analog audio sampled by a microphone; The second audio data in the mixed audio is removed to obtain third audio data. Compared with the first audio data, the third audio data has audio impairment and / or includes residual data points from the second audio data. The third audio data is associated with the first label to obtain a target audio sample. The target audio sample is used to train a speech recognition model. The speech recognition model is used to recognize the predicted content text of the third audio data. The loss value between the predicted content text of the third audio data and the first label of the third audio data is used to iterate the speech recognition model.

2. The method according to claim 1, characterized in that, The generation of the mixed audio corresponding to the first audio data and the second audio data includes: Randomly generate first feature information, wherein the first feature information includes first size information for indicating the size of the space and / or indicating the reverberation time RT60 within the space; The first room impulse response (RIR) corresponding to the first audio data and the second RIR corresponding to the second audio data are determined respectively; wherein, the first RIR is the acoustic response of the first audio data in the target space indicated by the first feature information; and the second RIR is the acoustic response of the second audio data in the target space indicated by the first feature information. After obtaining the first RIR and the second RIR, the mixed audio is generated. The mixed audio is the audio obtained by superimposing the first reverb audio and the second reverb audio. The first reverb audio is the audio determined by the first audio data and the first RIR, and the second reverb audio is the audio determined by the second audio data and the second RIR.

3. The method according to claim 2, characterized in that, Before determining the first room impulse response (RIR) corresponding to the first audio data and the second RIR corresponding to the second audio data, the method further includes: Simulate the first position information of the microphone used for audio acquisition in the target space; Simulate the second location information of the sound source corresponding to the first audio data in the target space; The third location information of the sound source corresponding to the second audio data in the target space is simulated; Determining the first room impulse response (RIR) corresponding to the first audio data includes: determining the first RIR corresponding to the first audio data based on the first feature information, the first location information, and the second location information; Determining the second RIR corresponding to the second audio data includes: determining the second RIR corresponding to the second audio data based on the first feature information, the first location information, and the third location information.

4. The method according to claim 3, characterized in that, The first position information of the microphone used for audio acquisition in the simulated target space includes: Obtain second size information of a first device configured with the microphone, the second size information including the configuration position of the microphone on the body of the first device; Generate the fourth location information of the first device in the target space; Based on the configuration position of the microphone in the fourth position information and the second size information, the first position information corresponding to the microphone of the first device in a random posture is determined.

5. The method according to claim 4, characterized in that, The first device also includes a speaker, and the source of the second audio data is the speaker of the first device.

6. The method according to claim 4, characterized in that, The second location information of the sound source corresponding to the simulated first audio data in the target space includes: Construct a first spherical coordinate system, the origin of which is the position point indicated by the fourth position information; In the first spherical coordinate system, a first coordinate is randomly determined, which includes distance, azimuth angle, and polar angle; wherein, the first coordinate indicates the position of the sound source of the first audio data relative to the first device, the azimuth angle is greater than -180° and less than 180°, and the polar angle is greater than -30° and less than 180°; The first coordinate information is converted into a second coordinate, which is a rectangular coordinate. Based on the second coordinates and the fourth position information, the second position information of the sound source of the first audio data in the target space is determined.

7. The method according to any one of claims 2-6, characterized in that, The method of determining the first reverberant audio from the first audio data and the first RIR includes: performing convolution processing based on the first audio data and the first RIR to obtain the first reverberant audio; The method of determining the second reverberant audio from the second audio data and the second RIR includes: performing convolution processing based on the second audio data and the second RIR to obtain the second reverberant audio.

8. The method according to claim 7, characterized in that, After performing convolution processing based on the second audio data and the second RIR, and before obtaining the second reverberant audio, the method further includes: Obtain the convolutional audio corresponding to the second audio data and the second RIR; Align the convolutional audio with the second audio data; Determine a first duration, which belongs to a preset delay interval; Obtaining the second reverberant audio includes: generating the second reverberant audio based on the first duration and the aligned convolutional audio; wherein there is a delay of the first duration between the second reverberant audio and the aligned convolutional audio.

9. The method according to any one of claims 1-8, characterized in that, The process of eliminating the second audio data from the mixed audio to obtain the third audio data includes: Under preset conditions, the mixed audio and the second audio are input into the echo cancellation module to obtain the third audio data; The preset conditions include: the mixed audio is simulated audio collected by the microphone of the first device, and the second audio data is simulated audio played by the speaker of the first device.

10. A method for training a speech recognition model, characterized in that, The method includes: Obtain a target audio sample; wherein the target audio sample is audio simulated according to the method described in any one of claims 1-9; Using the target audio samples, a pre-configured speech recognition model is trained until the speech recognition model meets the pre-configured model convergence conditions.

11. The method according to claim 10, characterized in that, The method further includes: Obtain fourth audio data with a second tag, wherein the fourth audio data is the original audio without noise reduction, and the second tag contains the content text of the fourth audio data; Using the fourth audio data, a pre-configured speech recognition model is trained until the speech recognition model meets the pre-configured model convergence conditions.

12. The method according to claim 10, characterized in that, The method further includes: Obtain fifth audio data with a third tag, wherein the fifth audio data is recorded and denoised audio, and the third tag contains the content text of the fifth audio data; Using the fifth audio data, a pre-configured speech recognition model is trained until the speech recognition model meets the pre-configured model convergence conditions.

13. A speech recognition method, characterized in that, The application is to a second device, which is equipped with a speech recognition model and an echo cancellation module. The speech recognition model is a model obtained using the speech recognition model training method according to any one of claims 10-12, and the speech recognition method includes: The second device acquires the seventh audio data while playing the sixth audio data; After inputting the seventh and sixth audio data into the echo cancellation module, the eighth audio data is obtained; Using the speech recognition model, the content text corresponding to the eighth audio data is identified.

14. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory being used to store computer instructions, which, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1-13.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Voice processing model training method and device, electronic equipment and storage medium

    CN111933164A

  • Feature information mining method and device and electronic equipment

    CN112614484A

  • Audio data processing method and device, medium and equipment

    CN114242097A

  • Audio signal processing method and device, training method and device, equipment and storage medium

    CN114242100A