Surround sound generation method and device, equipment and medium
By extracting dry and wet sound signals from stereo audio signals using a neural network model, and avoiding the introduction of additional reverberation effects, the problem of surround sound distortion in existing technologies is solved, and higher quality surround sound generation is achieved.
Patent Information
- Application Number
- CN202410533525.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-29
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, surround sound generated using methods such as Upmix includes additional reverberation effects in addition to the original audio, resulting in surround sound distortion and poor performance.
The system uses a neural network model to identify dry and wet audio signals in stereo audio signals and then fuses them with the stereo audio signals to generate surround sound, avoiding the introduction of additional reverberation effects.
It reduces surround sound distortion, improves sound quality, and enhances the immersion and realism of the sound.
Smart Images

Figure CN120881501A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of audio processing technology, and in particular to a method, apparatus, device and medium for generating surround sound. Background Technology
[0002] Related technologies can separate different types of sound signals from stereo audio and generate surround sound based on the separated sound signals. For example, upmix is a commonly used method to convert stereo into surround sound. However, in practical applications, the inventors of this application have found that surround sound generated based on methods such as Upmix includes additional reverberation effects in addition to the original audio, resulting in surround sound distortion and poor quality. Therefore, a new surround sound generation method is needed to reduce distortion and improve the generation effect of surround sound. Summary of the Invention
[0003] To address the aforementioned technical problems, this disclosure provides a method, apparatus, device, and medium for generating surround sound.
[0004] A first aspect of this disclosure provides a surround sound generation method, the method comprising:
[0005] Acquire the first stereo audio signal;
[0006] Based on a preset neural network model, the first dry sound signal and the first wet sound signal contained in the first stereo audio signal are identified, and the first dry sound signal and the first wet sound signal are extracted from the first stereo audio signal.
[0007] The first dry sound signal and the first wet sound signal are respectively fused with the first stereo audio signal to obtain a second dry sound signal and a second wet sound signal. The second dry sound signal is the signal obtained by fusing the first dry sound signal and the first stereo audio signal, and the second wet sound signal is the signal obtained by fusing the first wet sound signal and the first stereo audio signal.
[0008] Surround sound is generated based on the second dry sound signal and the second wet sound signal.
[0009] In one embodiment, the step of identifying a first dry sound signal and a first wet sound signal contained in the first stereo audio signal based on a preset neural network model, and extracting the first dry sound signal and the first wet sound signal from the first stereo audio signal, includes:
[0010] Based on the neural network model, the first dry sound signal and the first wet sound signal contained in the first stereo audio signal are separated and masked to obtain a first mask of the first dry sound signal and a second mask of the first wet sound signal. The first mask is used to mark the location of the first dry sound signal, and the second mask is used to mark the location of the first wet sound signal.
[0011] The first dry sound signal is extracted from the location of the first mask, and the first wet sound signal is extracted from the location of the second mask.
[0012] In one embodiment, the step of separating and masking the first dry sound signal and the first wet sound signal contained in the first stereo audio signal based on the neural network model to obtain a first mask for the first dry sound signal and a second mask for the first wet sound signal includes:
[0013] The first stereo audio signal is input into a pre-trained U-Net model, and the U-Net model outputs a first mask for the first dry sound signal and a second mask for the first wet sound signal.
[0014] In one embodiment, the step of fusing the first dry sound signal and the first wet sound signal with the first stereo audio signal to obtain a second dry sound signal and a second wet sound signal includes:
[0015] The second dry sound signal is obtained by multiplying the frequency domain signal of the first dry sound signal at the first masking location with the frequency domain signal of the first stereo audio signal, and the second wet sound signal is obtained by multiplying the frequency domain signal of the first wet sound signal at the second masking location with the frequency domain signal of the first stereo audio signal.
[0016] In one embodiment, generating surround sound based on the second dry sound signal and the second wet sound signal includes:
[0017] The second dry sound signal is subjected to gain processing based on the pre-obtained dry sound signal gain to obtain the target dry sound signal;
[0018] The second wet sound signal is processed by gain processing based on the pre-obtained wet sound signal gain to obtain the target wet sound signal;
[0019] The target dry sound signal and the target wet sound signal are superimposed to generate surround sound.
[0020] In one embodiment, the training method of the preset neural network model includes: performing data augmentation processing on the dry sound sample signal based on the preset data augmentation network to obtain the wet sound sample signal;
[0021] The neural network model is trained based on dry and wet sound sample signals to obtain the preset neural network model.
[0022] In one embodiment, the process of performing data enhancement processing on the dry sound sample signal based on a preset data enhancement network to obtain the wet sound sample signal includes:
[0023] Based on a preset feedback delay network (FDN) and / or convolutional reverberation network, the dry sound sample signal is subjected to data enhancement processing to obtain the wet sound sample signal.
[0024] In one implementation, based on a preset feedback delay network (FDN) and a convolutional reverberation network, the dry sound sample signal is subjected to data enhancement processing to obtain a wet sound sample signal, including:
[0025] The dry sound sample signal is enhanced by a convolutional reverberation network to obtain an enhanced result. The enhanced result is an audio signal obtained by superimposing a reverberation effect on the dry sound sample signal.
[0026] Based on the FDN, the enhancement result is subjected to data enhancement processing to obtain the wet sound sample signal.
[0027] A second aspect of this disclosure provides a surround sound generation apparatus, the apparatus comprising:
[0028] The acquisition module is used to acquire the first stereo audio signal;
[0029] The extraction module is used to identify the first dry sound signal and the first wet sound signal contained in the first stereo audio signal based on a preset neural network model, and to extract the first dry sound signal and the first wet sound signal from the first stereo audio signal.
[0030] The fusion module is used to fuse the first dry sound signal and the first wet sound signal with the first stereo audio signal respectively to obtain a second dry sound signal and a second wet sound signal, wherein the second dry sound signal is the signal obtained by fusing the first dry sound signal and the first stereo audio signal, and the second wet sound signal is the signal obtained by fusing the first wet sound signal and the first stereo audio signal.
[0031] The generation module is used to generate surround sound based on the second dry sound signal and the second wet sound signal.
[0032] In one implementation, the extraction module is used for
[0033] Based on the neural network model, the first dry sound signal and the first wet sound signal contained in the first stereo audio signal are separated and masked to obtain a first mask of the first dry sound signal and a second mask of the first wet sound signal. The first mask is used to mark the location of the first dry sound signal, and the second mask is used to mark the location of the first wet sound signal.
[0034] The first dry sound signal is extracted from the location of the first mask, and the first wet sound signal is extracted from the location of the second mask.
[0035] In one implementation, the extraction module is used for
[0036] The first stereo audio signal is input into a pre-trained U-Net model, and the U-Net model outputs a first mask for the first dry sound signal and a second mask for the first wet sound signal.
[0037] In one implementation, the fusion module is used for
[0038] The second dry sound signal is obtained by multiplying the frequency domain signal of the first dry sound signal at the first masking location with the frequency domain signal of the first stereo audio signal, and the second wet sound signal is obtained by multiplying the frequency domain signal of the first wet sound signal at the second masking location with the frequency domain signal of the first stereo audio signal.
[0039] In one implementation, a generation module is used for
[0040] The second dry sound signal is subjected to gain processing based on the pre-obtained dry sound signal gain to obtain the target dry sound signal;
[0041] The second wet sound signal is processed by gain processing based on the pre-obtained wet sound signal gain to obtain the target wet sound signal;
[0042] The target dry sound signal and the target wet sound signal are superimposed to generate surround sound.
[0043] In one embodiment, the surround sound generating device further includes:
[0044] The data enhancement module is used to perform data enhancement processing on the dry sound sample signal based on a preset data enhancement network to obtain the wet sound sample signal;
[0045] The training module is used to train the neural network model based on dry sound sample signals and wet sound sample signals to obtain the preset neural network model.
[0046] In one implementation, the data enhancement module is used for
[0047] Based on a preset feedback delay network (FDN) and / or convolutional reverberation network, the dry sound sample signal is subjected to data enhancement processing to obtain the wet sound sample signal.
[0048] In one implementation, the data enhancement module is used for
[0049] The dry sound sample signal is enhanced by a convolutional reverberation network to obtain an enhanced result. The enhanced result is an audio signal obtained by superimposing a reverberation effect on the dry sound sample signal.
[0050] Based on the FDN, the enhancement result is subjected to data enhancement processing to obtain the wet sound sample signal.
[0051] A third aspect of this disclosure provides a computer device, the device comprising:
[0052] Memory;
[0053] Processor; and
[0054] A computer program, wherein the computer program is stored in memory and configured to be executed by a processor to implement the method described in the first aspect above.
[0055] A fourth aspect of this disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect above.
[0056] A fifth aspect of this disclosure provides a vehicle including the aforementioned computer equipment.
[0057] The technical solution provided in this disclosure has the following advantages compared with the prior art:
[0058] The surround sound generation method, apparatus, device, and medium provided in this disclosure improve the accuracy of dry and wet sound signal extraction by, after acquiring a first stereo audio signal, identifying a first dry sound signal and a first wet sound signal contained in the first stereo audio signal based on a preset neural network model, extracting the first dry sound signal and the first wet sound signal from the first stereo audio signal, and fusing the first dry sound signal and the first wet sound signal with the first stereo audio signal respectively. Furthermore, since the second dry sound signal and the second wet sound signal are signals originally carried in the first stereo audio signal without adding additional reverberation effects, generating surround sound based on the second dry sound signal and the second wet sound signal can reduce surround sound distortion and improve sound quality. Attached Figure Description
[0059] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0060] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0061] Figure 1 This is a flowchart of a model training method provided in an embodiment of this disclosure;
[0062] Figure 2 This is a flowchart of a surround sound generation method provided in an embodiment of this disclosure;
[0063] Figure 3 This is a schematic diagram of a surround sound generation method provided in an embodiment of this disclosure;
[0064] Figure 4 This is a schematic diagram of the structure of a surround sound generation device provided in an embodiment of this disclosure;
[0065] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0066] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0067] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0068] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0069] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0070] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0071] As described in the background section, Upmix is a commonly used method for converting stereo to surround sound. In related technologies, Upmix generally comes in two forms: one is an Upmix model based on mid / side recording techniques, which can only separate mid / side audio signals and cannot separate other types of audio signals. The other form of Upmix uses correlation and coherence as indicators; however, this model cannot cover all musical styles, and its correlation and coherence parameters depend on two-channel signals, failing to process mono or other multi-channel signals. Furthermore, regardless of the form of Upmix, the generated surround sound adds extra reverberation to the original audio base, resulting in surround sound distortion.
[0072] To address the problems existing in related technologies, this disclosure provides a scheme for extracting dry and wet signals from stereo audio signals using a neural network model, and a scheme for generating surround sound based on the dry and wet signals carried in the stereo audio signal. In this scheme, since the surround sound is generated based on the dry and wet signals inherent in the stereo audio signal itself, no additional audio data other than the original audio signal is introduced. Therefore, surround sound distortion can be reduced, and the immersiveness and realism of the sound can be improved. Furthermore, by training the neural network model described in this application with rich training samples (audio of various musical styles, audio with different numbers of channels, audio signals generated or recorded using different techniques, etc.), the applicability of the model can be improved, the limitations of audio separation can be reduced, and the accuracy of audio separation can be improved, providing a data foundation for generating realistic and immersive surround sound. The scheme provided by this disclosure can be widely applied to various media such as movies, television, and games, providing a more realistic and natural surround sound effect.
[0073] The solutions of the embodiments of this disclosure will be described below with reference to exemplary examples.
[0074] Example, Figure 1 This is a flowchart illustrating a model training method provided in this disclosure, through which a neural network model, as described in this disclosure, can be trained. In this disclosure, the method can be executed by a computer device, which can be rationally understood as a computer, in-vehicle system, or other device with computing and processing capabilities. Figure 1 As shown, the model training method provided in this embodiment of the disclosure may include the following steps 101-102.
[0075] Step 101: Perform data enhancement processing on the dry sound sample signal based on the preset data enhancement network to obtain the wet sound sample signal.
[0076] The data augmentation network in this disclosure can be any type of artificial intelligence network, such as a Feedback Delay Network (FDN) or a Convolutional Reverb. Alternatively, in some embodiments, the data augmentation network referred to in this disclosure can be composed of multiple networks. For example, in a feasible example, a Convolutional Reverb Network and an FDN Network can be used in series.
[0077] The data augmentation network described in this embodiment is trained to perform data augmentation processing on the input dry audio signal to obtain the corresponding wet audio signal. The data augmentation processing can be understood as superimposing a reverberation effect (such as reverberation effects of different styles or effects) on the dry audio signal to obtain the corresponding wet audio signal.
[0078] For example, in some implementations, dry acoustic sample signals generated or acquired using different technologies can be input into the FDN network, which then generates wet acoustic signals of different styles or effects. The FDN network can be trained based on dry acoustic signals generated or acquired using different technologies and wet acoustic signals of different styles or effects. Through training, the parameters of the FDN (e.g., number of delay lines, delay time, pre-delay time, decay, absorption, and mix ratio) are optimized so that the FDN can generate wet acoustic signals of different styles or effects based on the dry acoustic signals. Specific training methods or procedures for the FDN network can be found in related technologies and will not be elaborated here.
[0079] For example, in some other implementations, dry sound sample signals generated or acquired using different technologies can be input into a convolutional reverberation network. The convolutional reverberation network then generates wet sound signals of different styles or effects; that is, it superimposes reverberation effects of different styles or effects onto the dry sound sample signals to obtain the corresponding wet sound signals. The convolutional reverberation network can also be trained based on dry sound signals generated or acquired using different technologies and wet sound signals of different styles or effects. Training optimizes the parameters of the convolutional reverberation network so that it can generate wet sound signals of different styles or effects based on the dry sound signals. Specific training methods or processes for the convolutional reverberation network can be found in related technologies and will not be elaborated here.
[0080] For example, in another feasible implementation, the dry sound sample signal can be enhanced first using a convolutional reverberation network to obtain the enhanced result, and then the enhanced result of the convolutional reverberation network can be further enhanced using an FDN network to obtain the wet sound sample signal. By using the convolutional reverberation network and the FDN network in series, the richness of the data can be enhanced, providing a new way to achieve a more natural surround sound effect. This provides rich data for training the neural network model, improves the generalization ability of the neural network model, reduces overfitting in model training, and makes the final generated surround sound more natural.
[0081] Step 102: Train the neural network model based on the dry sound sample signal and the wet sound sample signal to obtain the trained neural network model.
[0082] The dry sound sample signal and wet sound sample signal provided in this embodiment can be either time-domain signals or frequency-domain signals. When the dry sound sample signal and wet sound sample signal are understood as time-domain signals, before inputting the dry sound sample signal and wet sound sample signal into the neural network model for training, the dry sound sample signal and wet sound sample signal can be converted into frequency-domain signals, and then the frequency-domain dry sound signal sample and wet sound signal sample can be input into the neural network model for training.
[0083] The neural network model referred to in this disclosure can be any structure or type of neural network model in related technologies. As a preferred embodiment, this disclosure specifically uses a U-Net model. In this disclosure, the U-Net model is trained to generate masks for dry and wet sound signals. The training method or process for the U-Net model can refer to related technologies. For example, dry sound sample signals and their labels, as well as wet sound sample signals and their labels, can be input into the U-Net model to be trained. The labels of the dry sound sample signals are used to indicate the masking positions of the dry and wet sound sample signals. The masks for the dry and wet sound sample signals output by the U-Net model are compared with their corresponding labels. If they are inconsistent, the parameters of the U-Net model are adjusted until the accuracy of the U-Net model reaches a preset accuracy. In fact, the training method or process of the U-Net model in this disclosure is similar to related technologies and will not be repeated here.
[0084] In this embodiment of the disclosure, the data augmentation network can enhance the richness of the data, providing a new approach to achieve a more natural surround sound effect. It provides rich data for the training of neural network models, improves the generalization ability and signal separation accuracy of neural network models, reduces overfitting in model training, and makes the final generated surround sound more natural.
[0085] Figure 2 This is a flowchart of a surround sound generation method provided in an embodiment of this disclosure. This method can be executed by a computer device, such as a computer, in-vehicle infotainment system, or other device with audio processing capabilities. Figure 2 As shown, the surround sound generation method provided in this embodiment may include steps 201-204.
[0086] Step 201: Obtain the first stereo audio signal.
[0087] The first stereo audio signal referred to in this embodiment can be understood as a frequency domain signal. If the first stereo audio signal is a time domain signal, this embodiment may further include a step of converting the acquired time-domain first stereo audio signal into a frequency domain signal after acquiring the first stereo audio signal. For ease of understanding, the first stereo audio signal in this embodiment is specifically referred to as a frequency domain signal.
[0088] The first stereo audio signal mentioned in the embodiments of this disclosure can be acquired by an audio acquisition device or transmitted from other devices such as storage devices, processing devices or playback devices through data transmission or exchange technology.
[0089] Step 202: Identify the first dry sound signal and the first wet sound signal contained in the first stereo audio signal based on the preset neural network model, and extract the first dry sound signal and the first wet sound signal from the first stereo audio signal.
[0090] In one implementation, the neural network model referred to in the embodiments of this disclosure can be understood as... Figure 1 The model is trained in the manner described in this embodiment. In this embodiment, the first stereo audio signal obtained in step 201 can be used as input to a neural network model. The neural network model separates and masks the dry sound signal (hereinafter referred to as the first dry sound signal for easy distinction) and the wet sound signal (hereinafter referred to as the first wet sound signal for easy distinction) contained in the first stereo audio signal, obtaining a first mask for the first dry sound signal and a second mask for the first wet sound signal. Thus, the first dry sound signal is extracted from the location of the first mask, and the first wet sound signal is extracted from the location of the second mask. The first mask and the second mask are used to mark the locations of the first dry sound signal and the first wet sound signal. The neural network model referred to in this embodiment can be any kind of neural network model. For example, it can be a U-Net model. The first stereo audio signal is input into a pre-trained U-Net model, and the U-Net model outputs the first mask for the first dry sound signal and the second mask for the first wet sound signal.
[0091] In another exemplary embodiment, the neural network model described in this disclosure can also be trained to directly identify and extract dry and wet signals from stereo audio signals. Specifically, this refers to the neural network model's ability to directly extract dry and wet signals from stereo audio signals. After acquiring the first stereo audio signal, this disclosure embodiment can directly extract the first dry and first wet signals through the neural network model.
[0092] Step 203: The first dry sound signal and the first wet sound signal are fused with the first stereo audio signal to obtain the second dry sound signal and the second wet sound signal. The second dry sound signal is the signal obtained by fusing the first dry sound signal and the first stereo audio signal, and the second wet sound signal is the signal obtained by fusing the first wet sound signal and the first stereo audio signal.
[0093] There are various methods for fusing the first dry sound signal, the first wet sound signal, and the first stereo audio signal, such as multiplication, weighted multiplication, or addition. For ease of understanding, in the embodiments of this disclosure, the fusion processing of the first dry sound signal, the first wet sound signal, and the first stereo audio signal can be understood as multiplying the frequency domain signal of the first dry sound signal at the first mask with the frequency domain signal of the first stereo audio signal to obtain the second dry sound signal, and multiplying the frequency domain signal of the first wet sound signal at the second mask with the frequency domain signal of the first stereo audio signal to obtain the second wet sound signal.
[0094] It should be noted that in practical scenarios, the method for fusing the first dry sound signal and the first stereo audio signal can be the same as or different from the method for fusing the first wet sound signal and the first stereo audio signal.
[0095] Step 204: Generate surround sound based on the second dry sound signal and the second wet sound signal.
[0096] For example, in embodiments of this disclosure where the second dry sound signal and the second wet sound signal obtained by the aforementioned method are frequency domain signals, when generating surround sound based on the second dry sound signal and the second wet sound signal, the second dry sound signal and the second wet sound signal can first be converted into time domain signals, and then the time domain dry sound signal and wet sound signal can be superimposed to obtain surround sound. Alternatively, the second dry sound signal and the second wet sound signal can be directly superimposed based on the frequency domain to obtain frequency domain surround sound, and then the frequency domain surround sound can be converted into time domain surround sound.
[0097] Alternatively, in other embodiments, the second dry signal can be amplified using a pre-obtained dry signal gain (e.g., by multiplying or adding the dry signal gain to the second dry signal) to obtain a target dry signal, and the second wet signal can be amplified using a pre-obtained wet signal gain (e.g., by multiplying or adding the wet signal gain to the second wet signal) to obtain a target wet signal. The target dry signal and the target wet signal are then superimposed to obtain surround sound. Similarly, the gain processing can be performed in the frequency domain or the time domain. If performed in the frequency domain, after obtaining the frequency domain surround sound, a step of converting the frequency domain surround sound into time domain surround sound may be included. In the embodiments of this disclosure, the dry signal gain and wet signal gain can be default settings or user-defined settings. By amplifying the second dry signal and the second wet signal using the dry signal gain and wet signal gain, the sound quality of the surround sound can be improved, satisfying personalized user settings.
[0098] In this embodiment, after acquiring the first stereo audio signal, a first dry sound signal and a first wet sound signal are identified based on a preset neural network model. The first dry sound signal and the first wet sound signal are then extracted from the first stereo audio signal. These signals are then fused with the first stereo audio signal to obtain a second dry sound signal and a second wet sound signal, thereby improving the accuracy of the dry and wet sound signal extraction. Furthermore, since the second dry sound signal and the second wet sound signal are signals originally carried in the first stereo audio signal and do not add any additional reverberation effect, generating surround sound based on the second dry sound signal and the second wet sound signal can reduce surround sound distortion and improve sound quality.
[0099] Figure 3 This is a schematic diagram of a surround sound generation method provided in an embodiment of this disclosure. See also... Figure 3In one implementation scenario, the time-domain stereo audio signal is processed by a Short-Time Fourier Transform (STFT) to obtain a frequency-domain stereo audio signal (which can be understood as the first stereo audio signal in the above embodiment). Further, the frequency-domain stereo audio signal is processed using a U-Net neural network to obtain a mask for the first dry sound signal and a mask for the first wet sound signal, thereby obtaining the first dry sound signal and the first wet sound signal from the corresponding masks. A second dry sound signal is obtained by multiplying the first dry sound signal with the first stereo audio signal, and a second wet sound signal is obtained by multiplying the first wet sound signal with the first stereo audio signal. Further, the second dry sound signal and the second wet sound signal can be processed by an Inverse Short-Time Fourier Transform (ISTFT) to obtain the time-domain dry sound signal dry(n) and the time-domain wet sound signal dry(n). The target dry signal is obtained by multiplying the dry signal dry(n) with the pre-obtained dry signal gain, and the target wet signal is obtained by multiplying the wet signal wet(n) with the wet signal gain. The target dry signal and the wet signal are then superimposed to obtain surround sound, which enhances the immersion and realism of the surround sound.
[0100] In this embodiment, after acquiring the first stereo audio signal, a first dry sound signal and a first wet sound signal are extracted from the first stereo audio signal based on a preset U-Net neural network model. The first dry sound signal and the first wet sound signal are then multiplied with the first stereo audio signal to obtain a second dry sound signal and a second wet sound signal, thereby improving the accuracy of the dry and wet sound signal extraction. Furthermore, since the second dry sound signal and the second wet sound signal are signals originally carried in the first stereo audio signal and do not add any additional reverberation effect, generating surround sound based on the second dry sound signal and the second wet sound signal can reduce surround sound distortion and improve sound quality.
[0101] Figure 4 This is a schematic diagram of a surround sound generation device provided in an embodiment of this disclosure. The surround sound generation device provided in this embodiment can be understood as the aforementioned computer device or a functional module within the aforementioned computer device. Figure 4 As shown, in some embodiments, the surround sound generating device 40 includes:
[0102] Acquisition module 41 is used to acquire the first stereo audio signal;
[0103] Extraction module 42 is used to identify the first dry sound signal and the first wet sound signal contained in the first stereo audio signal based on a preset neural network model, and to extract the first dry sound signal and the first wet sound signal from the first stereo audio signal.
[0104] The fusion module 43 is used to fuse the first dry sound signal and the first wet sound signal with the first stereo audio signal respectively to obtain a second dry sound signal and a second wet sound signal, wherein the second dry sound signal is the signal obtained by fusing the first dry sound signal and the first stereo audio signal, and the second wet sound signal is the signal obtained by fusing the first wet sound signal and the first stereo audio signal.
[0105] The generation module 44 is used to generate surround sound based on the second dry sound signal and the second wet sound signal.
[0106] In one embodiment, the extraction module 42 is used for
[0107] Based on the neural network model, the first dry sound signal and the first wet sound signal contained in the first stereo audio signal are separated and masked to obtain a first mask of the first dry sound signal and a second mask of the first wet sound signal. The first mask is used to mark the location of the first dry sound signal, and the second mask is used to mark the location of the first wet sound signal.
[0108] The first dry sound signal is extracted from the location of the first mask, and the first wet sound signal is extracted from the location of the second mask.
[0109] In one implementation, the extraction module is used for
[0110] The first stereo audio signal is input into a pre-trained U-Net model, and the U-Net model outputs a first mask for the first dry sound signal and a second mask for the first wet sound signal.
[0111] In one embodiment, the fusion module 43 is used for
[0112] The second dry sound signal is obtained by multiplying the frequency domain signal of the first dry sound signal at the first masking location with the frequency domain signal of the first stereo audio signal, and the second wet sound signal is obtained by multiplying the frequency domain signal of the first wet sound signal at the second masking location with the frequency domain signal of the first stereo audio signal.
[0113] In one implementation, the generation module 44 is used for
[0114] The second dry sound signal is subjected to gain processing based on the pre-obtained dry sound signal gain to obtain the target dry sound signal;
[0115] The second wet sound signal is processed by gain processing based on the pre-obtained wet sound signal gain to obtain the target wet sound signal;
[0116] The target dry sound signal and the target wet sound signal are superimposed to generate surround sound.
[0117] In one implementation, the neural network model includes the U-Net model.
[0118] In one embodiment, the surround sound generating device 40 may further include:
[0119] Data enhancement module 51 is used to perform data enhancement processing on dry sound sample signals based on a preset data enhancement network to obtain wet sound sample signals;
[0120] Training module 52 is used to train the neural network model based on dry sound sample signals and wet sound sample signals to obtain the preset neural network model.
[0121] In one implementation, the data enhancement module 51 is used for
[0122] Based on a preset feedback delay network (FDN) and / or convolutional reverberation network, the dry sound sample signal is subjected to data enhancement processing to obtain the wet sound sample signal.
[0123] In one implementation, the data enhancement module 51 is used for
[0124] The dry sound sample signal is enhanced by a convolutional reverberation network to obtain an enhanced result. The enhanced result is an audio signal obtained by superimposing a reverberation effect on the dry sound sample signal.
[0125] Based on the FDN, the enhancement result is subjected to data enhancement processing to obtain the wet sound sample signal.
[0126] The model training apparatus provided in this disclosure is capable of performing... Figures 1-3 The methods in any of the embodiments are similar in execution and beneficial effects, and will not be described again here.
[0127] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown.
[0128] In this embodiment of the disclosure, Figure 5 The computer equipment shown can be a server, in-vehicle system, terminal, or server cluster. Specifically, the terminal can include a mobile phone, computer, tablet computer, in-vehicle terminal, or any device capable of being used for... Figures 1-3 Any device, etc., of any embodiment of the method is not limited herein.
[0129] like Figure 5 As shown, the computer device 170 may include a memory 171 storing computer program instructions and a processor 172.
[0130] Specifically, the processor 172 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0131] Memory 171 may include a large-capacity storage for information or instructions. For example, and not limitingly, memory 171 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 171 may include removable or non-removable (or fixed) media. Where appropriate, memory 171 may be internal or external to the integrated gateway device. In a particular embodiment, memory 171 is a non-volatile solid-state memory. In a particular embodiment, memory 171 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (Electrically Programmable ROM, EPROM), an electrically erasable programmable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0132] The processor 172 performs the steps of the method embodiments provided in this disclosure by reading and executing computer program instructions stored in the memory 171.
[0133] In one example, the electronic device may also include a communication interface 173. Wherein, such as Figure 5 As shown, the processor 172, memory 171 and communication interface 173 are connected via a bus and communicate with each other.
[0134] A bus can be hardware, software, or both. For example, and not limited to, a bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, a bus can include one or more buses.
[0135] This disclosure also provides a computer-readable storage medium that can store a computer program, which, when executed by a processor, causes the processor to implement the embodiments of this disclosure. Figures 1-3 The method of any of the embodiments.
[0136] The aforementioned storage medium may, for example, include a memory 171 containing computer program instructions, which can be executed by a processor 172 of an electronic device to perform the methods described in the embodiments of this disclosure. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), compact disc-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device.
[0137] This disclosure also provides a vehicle that includes the computer equipment described in the above embodiments, which can implement the various processes and effects described in the above embodiments of this disclosure, and will not be elaborated here.
[0138] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0139] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating surround sound, characterized in that, The method includes: Acquire the first stereo audio signal; Based on a preset neural network model, the first dry sound signal and the first wet sound signal contained in the first stereo audio signal are identified, and the first dry sound signal and the first wet sound signal are extracted from the first stereo audio signal. The first dry sound signal and the first wet sound signal are respectively fused with the first stereo audio signal to obtain a second dry sound signal and a second wet sound signal, wherein the second dry sound signal is the signal obtained by fusing the first dry sound signal and the first stereo audio signal, and the second wet sound signal is the signal obtained by fusing the first wet sound signal and the first stereo audio signal. Surround sound is generated based on the second dry sound signal and the second wet sound signal.
2. The method according to claim 1, characterized in that, The step of identifying the first dry sound signal and the first wet sound signal contained in the first stereo audio signal based on a preset neural network model, and extracting the first dry sound signal and the first wet sound signal from the first stereo audio signal, includes: Based on the neural network model, the first dry sound signal and the first wet sound signal contained in the first stereo audio signal are separated and masked to obtain a first mask for the first dry sound signal and a second mask for the first wet sound signal. The first mask is used to mark the location of the first dry sound signal, and the second mask is used to mark the location of the first wet sound signal. The first dry sound signal is extracted from the location of the first mask, and the first wet sound signal is extracted from the location of the second mask.
3. The method according to claim 2, characterized in that, The step of separating and masking the first dry sound signal and the first wet sound signal contained in the first stereo audio signal based on the neural network model to obtain a first mask for the first dry sound signal and a second mask for the first wet sound signal includes: The first stereo audio signal is input into a pre-trained U-Net model, and the U-Net model outputs a first mask for the first dry sound signal and a second mask for the first wet sound signal.
4. The method according to claim 2 or 3, characterized in that, The step of fusing the first dry sound signal and the first wet sound signal with the first stereo audio signal to obtain the second dry sound signal and the second wet sound signal includes: The second dry sound signal is obtained by multiplying the frequency domain signal of the first dry sound signal at the first masking location with the frequency domain signal of the first stereo audio signal, and the second wet sound signal is obtained by multiplying the frequency domain signal of the first wet sound signal at the second masking location with the frequency domain signal of the first stereo audio signal.
5. The method according to claim 1, characterized in that, The generation of surround sound based on the second dry sound signal and the second wet sound signal includes: The second dry sound signal is subjected to gain processing based on the pre-obtained dry sound signal gain to obtain the target dry sound signal; The second wet sound signal is processed by gain processing based on the pre-obtained wet sound signal gain to obtain the target wet sound signal; The target dry sound signal and the target wet sound signal are superimposed to generate surround sound.
6. The method according to claim 1, characterized in that, The training method for the preset neural network model includes: The dry sound sample signal is enhanced by a preset data enhancement network to obtain the wet sound sample signal. The neural network model is trained based on dry and wet sound sample signals to obtain the preset neural network model.
7. The method according to claim 6, characterized in that, The method of performing data enhancement processing on the dry sound sample signal based on a preset data enhancement network to obtain a wet sound sample signal includes: Based on a preset feedback delay network (FDN) and / or convolutional reverberation network, the dry sound sample signal is subjected to data enhancement processing to obtain the wet sound sample signal.
8. The method according to claim 7, characterized in that, Based on a pre-defined feedback delay network (FDN) and a convolutional reverberation network, the dry sound sample signal is subjected to data enhancement processing to obtain a wet sound sample signal, including: The dry sound sample signal is enhanced by a convolutional reverberation network to obtain an enhanced result. The enhanced result is an audio signal obtained by superimposing a reverberation effect on the dry sound sample signal. Based on the FDN, the enhancement result is subjected to data enhancement processing to obtain the wet sound sample signal.
9. A surround sound generating device, characterized in that, include: The acquisition module is used to acquire the first stereo audio signal; The extraction module is used to identify the first dry sound signal and the first wet sound signal contained in the first stereo audio signal based on a preset neural network model, and to extract the first dry sound signal and the first wet sound signal from the first stereo audio signal. The fusion module is used to fuse the first dry sound signal and the first wet sound signal with the first stereo audio signal respectively to obtain a second dry sound signal and a second wet sound signal, wherein the second dry sound signal is the signal obtained by fusing the first dry sound signal and the first stereo audio signal, and the second wet sound signal is the signal obtained by fusing the first wet sound signal and the first stereo audio signal. The generation module is used to generate surround sound based on the second dry sound signal and the second wet sound signal.
10. A computer device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.