Audio data processing method, electronic equipment, vehicle and medium
By using a stereo imaging deep learning model to render 3D audio-visual data from 2D audio-visual data in an in-vehicle scene, the problem of poor audio-visual rendering effect in existing technologies is solved, and a richer and more three-dimensional auditory experience is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BYD CO LTD
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing 3D audio-visual rendering methods are not suitable for in-vehicle scenes, resulting in poor rendering effects and impacting the user's auditory experience.
By acquiring two-dimensional audio-visual data and inputting it into a stereo audio-visual deep learning model for three-dimensional audio-visual rendering, and by using the deep learning model for training and optimization, three-dimensional audio-visual data is output to improve the audio-visual rendering effect.
It improves the sound and image rendering effect, provides a richer and more three-dimensional auditory experience, increases the spatial sense of sound, and makes the user's listening experience more realistic and immersive.
Smart Images

Figure CN121968006A_ABST
Abstract
Description
Audio data processing methods, electronic devices, vehicles and media Technical Field
[0001] This application belongs to the field of vehicle technology, and specifically relates to a method for processing audio data, electronic devices, vehicles, and media. Background Technology
[0002] With the continuous advancement of intelligent car cockpits, users' demands for in-vehicle audio experience are increasing, especially in scenarios such as music, games, and movies, where people are increasingly pursuing an auditory experience with a strong sense of spatial immersion.
[0003] Currently, vehicle manufacturers have provided the hardware foundation for enhancing the user's auditory experience by deploying audio equipment in multiple locations within the vehicle. However, existing technologies often result in poor spatial sound quality in the output audio. Summary of the Invention
[0004] This application provides an audio data processing method, electronic device, vehicle, and medium to solve the problem that existing three-dimensional audio-visual rendering methods are not suitable for in-vehicle scenes and have poor rendering effects.
[0005] In a first aspect, embodiments of this application provide a method for processing audio data, including:
[0006] Acquire two-dimensional audio-visual data, wherein the two-dimensional audio-visual data is at least one two-dimensional audio-visual data corresponding to the two-dimensional original audio data;
[0007] The two-dimensional audio-visual data is input into the stereo audio-visual deep learning model, and three-dimensional audio-visual data is output.
[0008] Optionally, the stereo image deep learning model is obtained by pre-training based on stereo version data samples.
[0009] Optionally, it also includes: the training method of the stereo image deep learning model includes:
[0010] The stereo version data sample is input into the stereo image deep learning model for three-dimensional sound image rendering to obtain predicted three-dimensional sound image data.
[0011] Based on the predicted 3D acoustic image data and the target 3D acoustic image data, the stereo acoustic image deep learning model is updated;
[0012] Until the stereoscopic deep learning model converges, the target three-dimensional stereoscopic data is the preset benchmark three-dimensional stereoscopic data.
[0013] Optionally, after updating the stereo audio-visual deep learning model based on the predicted 3D audio-visual data and the target 3D audio-visual data, the method further includes: if the stereo audio-visual deep learning model does not converge, inputting the stereo version data sample into the updated stereo audio-visual deep learning model to obtain the updated predicted 3D audio-visual data; and updating the stereo audio-visual deep learning model based on the updated predicted 3D audio-visual data and the target 3D audio-visual data.
[0014] Optionally, updating the stereoscopic deep learning model based on the predicted 3D acoustic image data and the target 3D acoustic image data includes:
[0015] Based on the predicted 3D audio-visual data, the target 3D audio-visual data, and the loss function, the stereo audio-visual deep learning model is updated. The loss function is a function that measures the difference between the predicted 3D audio-visual data and the target 3D audio-visual data.
[0016] Optionally, the stereo version data sample includes: two-dimensional sound source object data and two-dimensional sound image coordinate data, wherein the two-dimensional sound source object data is the two-dimensional channel information of the training sound source object, and the two-dimensional sound image coordinate data is the two-dimensional coordinate data of the training sound image object.
[0017] Optionally, the predicted three-dimensional acoustic image data includes: predicted three-dimensional sound source object data and predicted three-dimensional acoustic image coordinate data, wherein the predicted three-dimensional sound source object data is the three-dimensional channel information of the training sound source object, and the three-dimensional acoustic image coordinate data is the three-dimensional coordinate data of the training acoustic image object.
[0018] Optional, also includes:
[0019] The target's three-dimensional acoustic image data is obtained based on multi-channel version data samples.
[0020] Optionally, the target three-dimensional acoustic image data includes: target three-dimensional sound source object data and target three-dimensional acoustic image coordinate data, wherein the target three-dimensional sound source object data is multi-channel sound source signal data of the training sound source object, and the target three-dimensional acoustic image coordinate data is three-dimensional coordinate data of the training sound source object.
[0021] Optionally, the target three-dimensional acoustic coordinates include the pitch angle.
[0022] Optionally, the method further includes:
[0023] During the training process of the stereo imaging deep learning model, a first performance parameter of the stereo imaging deep learning model is evaluated based on a validation dataset, and the hyperparameters of the stereo imaging deep learning model are adjusted based on the first performance parameter until the stereo imaging deep learning model converges. The first performance parameter is used to evaluate the predictive ability of the stereo imaging deep learning model on the validation dataset, and the hyperparameters are used to control the behavior and performance of the model.
[0024] Optionally, the method further includes:
[0025] After the stereo imaging deep learning model converges, the stereo imaging deep learning model is evaluated based on the test dataset to obtain a second performance parameter of the stereo imaging deep learning model, wherein the second performance parameter is used to evaluate the prediction ability of the stereo imaging deep learning model.
[0026] Optionally, acquiring two-dimensional acoustic image data includes:
[0027] Based on two-dimensional raw audio data and an environmental sound field impulse response database, the two-dimensional sound image data is obtained, wherein the environmental sound field impulse response database includes the correspondence between the two-dimensional raw audio data and the two-dimensional sound image data.
[0028] Optionally, obtaining the two-dimensional sound image data based on the two-dimensional raw audio data and the environmental sound field impulse response database includes:
[0029] The left and right channel signals of a two-dimensional audio-visual object are obtained based on the two-dimensional raw audio data.
[0030] Based on the left channel signal, the right channel signal, and the ambient sound field impulse response database, the two-dimensional sound image data is obtained. The ambient sound field impulse response database includes: the left channel signal and the right channel signal of the two-dimensional original audio data, and the correspondence between them and the two-dimensional sound image data.
[0031] Optionally, acquiring the two-dimensional acoustic image data based on the left channel signal, the right channel signal, and the ambient sound field impulse response database includes:
[0032] Based on the left channel signal and the right channel signal, the binaural time difference and binaural sound level difference are obtained;
[0033] The two-dimensional acoustic image data is obtained by matching the binaural time difference and binaural sound level difference with the environmental sound field impulse response database.
[0034] Optionally, before obtaining the left and right channel signals of the two-dimensional sound image object based on the two-dimensional raw audio data, the method further includes:
[0035] Sound and image separation is performed based on the original two-dimensional audio data to obtain data of independent two-dimensional sound and image objects.
[0036] Optionally, before acquiring the two-dimensional sound image data based on the left channel signal, the right channel signal, and the ambient sound field impulse response database, the method further includes:
[0037] Establish the environmental sound field impulse response database.
[0038] Optionally, establishing the environmental sound field impulse response database includes:
[0039] Acquire the impulse response signal of the two-dimensional raw audio data sample within the coverage area of the ambient sound field;
[0040] The binaural time difference and binaural sound level difference of the two-dimensional raw audio data sample are obtained based on the impulse response signal;
[0041] Based on the binaural time difference and binaural sound level difference of the two-dimensional original audio data sample, and the two-dimensional sound image data corresponding to the two-dimensional original audio data sample, the environmental sound field impulse response database is established.
[0042] Secondly, embodiments of this application provide an audio data processing apparatus, comprising:
[0043] The acquisition module is used to acquire two-dimensional audio-visual data, wherein the two-dimensional audio-visual data is at least one two-dimensional audio-visual data corresponding to the two-dimensional original audio data;
[0044] The processing module is used to input the two-dimensional audio-visual data into the stereo audio-visual deep learning model and output three-dimensional audio-visual data.
[0045] Optionally, the stereo image deep learning model is obtained by pre-training based on stereo version data samples.
[0046] Optionally, the processing module is further configured to input the stereo version data sample into the stereo image deep learning model for three-dimensional audio-visual rendering to obtain predicted three-dimensional audio-visual data; update the stereo image deep learning model based on the predicted three-dimensional audio-visual data and the target three-dimensional audio-visual data; until the stereo image deep learning model converges, and the target three-dimensional audio-visual data is the preset benchmark three-dimensional audio-visual data.
[0047] Optionally, the processing module is further configured to, if the stereo image deep learning model does not converge, input the stereo version data sample into the updated stereo image deep learning model to obtain the updated predicted three-dimensional sound image data; and update the stereo image deep learning model based on the updated predicted three-dimensional sound image data and the target three-dimensional sound image data.
[0048] Optionally, the processing module is specifically used to update the stereoscopic deep learning model based on the predicted 3D audio-visual data, the target 3D audio-visual data, and the loss function, wherein the loss function is a function that measures the difference between the predicted 3D audio-visual data and the target 3D audio-visual data.
[0049] Optionally, the stereo version data sample includes: two-dimensional sound source object data and two-dimensional sound image coordinate data, wherein the two-dimensional sound source object data is the two-dimensional channel information of the training sound source object, and the two-dimensional sound image coordinate data is the two-dimensional coordinate data of the training sound image object.
[0050] Optionally, the predicted three-dimensional acoustic image data includes: predicted three-dimensional sound source object data and predicted three-dimensional acoustic image coordinate data, wherein the predicted three-dimensional sound source object data is the three-dimensional channel information of the training sound source object, and the three-dimensional acoustic image coordinate data is the three-dimensional coordinate data of the training acoustic image object.
[0051] Optionally, the processing module is also used to acquire the target three-dimensional acoustic image data based on multi-channel version data samples.
[0052] Optionally, the target three-dimensional acoustic image data includes: target three-dimensional sound source object data and target three-dimensional acoustic image coordinate data, wherein the target three-dimensional sound source object data is multi-channel sound source signal data of the training sound source object, and the target three-dimensional acoustic image coordinate data is three-dimensional coordinate data of the training sound source object.
[0053] Optionally, the target three-dimensional acoustic coordinates include the pitch angle.
[0054] Optionally, the processing module is further configured to evaluate a first performance parameter of the stereo imaging deep learning model based on a validation dataset during the training process of the stereo imaging deep learning model, and adjust the hyperparameters of the stereo imaging deep learning model based on the first performance parameter until the stereo imaging deep learning model converges, wherein the first performance parameter is used to evaluate the predictive ability of the stereo imaging deep learning model on the validation dataset, and the hyperparameters are used to control the behavior and performance of the model.
[0055] Optionally, the processing module is further configured to evaluate the stereo imaging deep learning model based on a test dataset after the stereo imaging deep learning model converges, and obtain a second performance parameter of the stereo imaging deep learning model, wherein the second performance parameter is used to evaluate the predictive ability of the stereo imaging deep learning model.
[0056] Optionally, the acquisition module is specifically used to acquire the two-dimensional sound image data based on the two-dimensional raw audio data and the environmental sound field impulse response database, wherein the environmental sound field impulse response database includes the correspondence between the two-dimensional raw audio data and the two-dimensional sound image data.
[0057] Optionally, the acquisition module is specifically used to acquire the left channel signal and right channel signal of the two-dimensional sound image object based on the two-dimensional original audio data; and to acquire the two-dimensional sound image data based on the left channel signal, the right channel signal, and an environmental sound field impulse response database, wherein the environmental sound field impulse response database includes: the left channel signal and right channel signal of the two-dimensional original audio data, and the correspondence between them and the two-dimensional sound image data.
[0058] Optionally, the acquisition module is specifically used to acquire the binaural time difference and binaural sound level difference based on the left channel signal and the right channel signal; and to match the binaural time difference and binaural sound level difference with the environmental sound field impulse response database to obtain the two-dimensional sound image data.
[0059] Optionally, the acquisition module is further configured to perform sound-image separation based on the two-dimensional raw audio data to obtain data of independent two-dimensional sound-image objects.
[0060] Optionally, the processing module is also used to establish the environmental sound field impulse response database.
[0061] Optionally, the processing module is specifically used to acquire the impulse response signal of the two-dimensional original audio data sample within the coverage area of the ambient sound field; acquire the binaural time difference and binaural sound level difference of the two-dimensional original audio data sample based on the impulse response signal; and establish the ambient sound field impulse response database based on the binaural time difference and binaural sound level difference of the two-dimensional original audio data sample, as well as the two-dimensional sound image data corresponding to the two-dimensional original audio data sample.
[0062] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the audio data processing method as described in any of the first aspects.
[0063] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the audio data processing method as described in any of the first aspects.
[0064] Fifthly, embodiments of this application provide a computer program product that, when executed by a processor of a vehicle or a cloud server, implements the steps of the audio data processing method as described in any of the first aspects.
[0065] In a sixth aspect, embodiments of this application provide a vehicle, including: a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the audio data processing method as described in any of the first aspects.
[0066] The audio data processing method, electronic device, vehicle, and medium provided in this application embodiment acquire two-dimensional audio-visual data, wherein the two-dimensional audio-visual data is at least one two-dimensional audio-visual data corresponding to the two-dimensional original audio data; input the two-dimensional audio-visual data into a stereo audio-visual deep learning model, and output three-dimensional audio-visual data. That is, based on the stereo audio-visual deep learning model, three-dimensional audio-visual rendering is performed on the two-dimensional audio-visual data to obtain three-dimensional audio-visual data, thereby improving the audio-visual rendering effect, providing users with a richer and more three-dimensional auditory experience, increasing the spatial sense of sound, and making the user's listening experience more realistic and immersive. Attached Figure Description
[0067] Figure 1 is a schematic diagram of a two-dimensional sound source provided in an embodiment of this application;
[0068] Figure 2 is a schematic diagram of a three-dimensional sound source provided in an embodiment of this application;
[0069] Figure 3 is a flowchart illustrating an audio data processing method provided in an embodiment of this application;
[0070] Figure 4 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0071] Figure 5 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0072] Figure 6 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0073] Figure 7 is a schematic diagram of an environmental sound field impulse response database provided in an embodiment of this application;
[0074] Figure 8 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0075] Figure 9 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0076] Figure 10 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0077] Figure 11 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0078] Figure 12 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0079] Figure 13 is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0080] Figure 14 is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application. Detailed Implementation
[0081] The technical solutions in the embodiments of this application will be clearly described below with reference to the accompanying drawings. All other embodiments obtained by those skilled in the art based on the embodiments in this application are within the scope of protection of this application.
[0082] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and that the objects distinguished by "first" and "second" are generally of the same class and are not limited in number.
[0083] In related technologies, rendering is performed based on the original three-dimensional sound source through decoding. For example, the sound field in three-dimensional space can be represented and processed through spherical harmonic transformation. When playing back audio, the encoded sound field needs to be decoded through a speaker array or headphones. However, at certain frequencies, the spherical harmonic function may be zero, resulting in the inability to obtain sound field information at certain frequencies. Consequently, the rendering effect is poor, affecting the user experience.
[0084] This application provides an audio data processing method to improve the audio-visual rendering effect and provide users with a richer and more three-dimensional auditory experience. The audio data processing method can use a stereo audio-visual deep learning model to perform three-dimensional audio-visual rendering based on at least one two-dimensional audio-visual data corresponding to the two-dimensional original audio data, so as to obtain three-dimensional audio-visual data and improve the audio-visual rendering effect.
[0085] This application also acquires two-dimensional sound image data from the two-dimensional raw audio data based on the left channel signal, the right channel signal, and an ambient sound field impulse response database. This improves the efficiency and accuracy of acquiring two-dimensional sound image data.
[0086] This application also establishes the environmental sound field impulse response database based on the impulse response signal received from the two-dimensional raw audio data sample transmitted within the stereo sound field coverage area.
[0087] Audio data can be divided into two-dimensional audio and three-dimensional audio. Two-dimensional audio typically uses stereo (two-channel) to provide a left and right channel audio experience, such as when using a mobile phone to listen to music, broadcast, or watch TV. Three-dimensional audio can simulate a realistic auditory experience, allowing users to perceive sounds from different directions, such as when watching a movie in a theater. This application uses an in-vehicle application scenario as an example to describe the audio data processing method.
[0088] The embodiments described below can be implemented in a vehicle. Existing vehicles typically contain multiple audio systems. Taking a five-channel system as an example, the five-channel system includes: front left, front right, rear left, rear right, and center speakers, with each speaker corresponding to one channel. Compared to the dual-channel layout of stereo, using multi-channel audio playback can more accurately simulate the propagation of sound in three-dimensional space, providing a more realistic and immersive listening experience. Therefore, rendering a two-dimensional sound source as a three-dimensional sound source can improve the user's listening experience.
[0089] The "two-dimensional" and "three-dimensional" descriptions of a sound source typically refer to its spatial distribution and the way sound propagates. Figure 1 is a schematic diagram of a two-dimensional sound source provided in an embodiment of this application. As shown in Figure 1, a two-dimensional sound source generally refers to a sound source distributed on a plane, without including height sound information. For a sound source S in a two-dimensional plane, the main components are: the sound source distance R and the azimuth angle θ. The sound source distance R refers to the distance between the sound source S and the origin O, which can be considered as a loudspeaker. The azimuth angle θ refers to the angle between the sound source S and the vertical direction in the two-dimensional plane, i.e., the angle between the sound source S and the Y-axis. The audio signal of a two-dimensional sound source typically only contains sound localization in the horizontal direction, such as the left and right channels.
[0090] Figure 2 is a schematic diagram of a three-dimensional sound source provided in an embodiment of this application. As shown in Figure 2, a three-dimensional sound source refers to a sound source distributed in three-dimensional space. A three-dimensional sound source can simulate the propagation of sound in the horizontal plane (left, right) and the vertical plane (front, back, up, down). For the sound source S in the two-dimensional plane, the main components include: the sound source distance R, the azimuth angle θ, and the pitch angle. Wherein, the sound source distance R refers to the distance between the sound source S and the origin O, and the origin O can be regarded as a loudspeaker; the azimuth angle θ refers to the angle between the sound source S and the Y-axis, that is, the angle between the sound source S' mapped onto the two-dimensional plane and the Y-axis; and the pitch angle... It refers to the angle between the line connecting the sound source S and the origin O and the horizontal plane.
[0091] The following describes several specific embodiments of the audio data processing method. Figure 3 is a flowchart illustrating an audio data processing method provided in this application.
[0092] S32: Acquire two-dimensional audio-visual data, which is at least one two-dimensional audio-visual data corresponding to the two-dimensional original audio data.
[0093] Specifically, two-dimensional sound image data can be obtained based on two-dimensional raw audio data and an environmental sound field impulse response database. The environmental sound field impulse response database includes the correspondence between the two-dimensional raw audio data and the two-dimensional sound image data.
[0094] Among them, the two-dimensional raw audio data is generally a stereo sound source, and the two-dimensional raw audio data is dual-channel. The two-dimensional sound image coordinates can be represented as (R, θ), where R is the distance to the sound source and θ is the azimuth angle.
[0095] S34: Input two-dimensional audio-visual data into the stereo audio-visual deep learning model and output three-dimensional audio-visual data.
[0096] The three-dimensional acoustic image data can be represented as (R, θ, ... R is the distance to the sound source, and θ is the azimuth angle. The pitch angle is represented by the audio signal from multiple channels in the 3D audio-visual data. Each channel corresponds to an in-vehicle speaker, and each channel can contain different audio information.
[0097] In this embodiment, two-dimensional audio-visual data is acquired, which is at least one two-dimensional audio-visual data corresponding to the original two-dimensional audio data. The two-dimensional audio-visual data is input into a stereo audio-visual deep learning model, and three-dimensional audio-visual data is output. Three-dimensional audio-visual data can provide a richer auditory experience than mono or stereo, reproduce sound through multiple channels, thereby simulating a more realistic sound field environment and improving the audio-visual rendering effect.
[0098] Figure 4 is a flowchart illustrating another audio data processing method provided in an embodiment of this application. Figure 4, based on the embodiment shown in Figure 3, further includes, optionally, the following method before executing S32:
[0099] S30: Perform sound-image separation based on two-dimensional raw audio data to obtain data of independent two-dimensional sound-image objects.
[0100] Preprocessing of the raw 2D audio data, including denoising and normalization, improves the model's separation performance. Features are automatically extracted using a deep learning model. The deep learning model is trained using a labeled dataset. The dataset should contain mixed sounds and their corresponding clean sound sources.
[0101] Deep learning models are used to separate two-dimensional raw audio data. A stereo sound source is an audio signal that can provide a richer auditory experience than a mono sound source, reproducing sound through at least two channels (usually the left and right channels).
[0102] Assume we use a deep multichannel dealiasing model (Demucs-Music Source Separation). Demucs is an encoder-decoder architecture consisting of a convolutional encoder, a bidirectional long short-term memory network, and a convolutional decoder, connected via skip connections. The Demucs model predicts an ideal binary mask (IBM) and an ideal ratio mask (IRM) for each sound source to separate the target sound source from the mixed audio. IBM is a binary mask, typically using 0 and 1 to represent the presence of the target sound source. If the signal-to-noise ratio (SNR) within a time-frequency unit is higher than a certain threshold, the unit is considered speech-dominated, and IBM is marked as 1; if the SNR is low, the unit is considered noise-dominated, and IBM is marked as 0. IRM is a continuous-value mask representing the intensity of the target sound source relative to noise, ranging from 0 to 1. A higher value indicates a higher proportion of speech within the time-frequency unit.
[0103] Suppose we have an original stereo source containing vocals, guitar sounds, and ambient sounds. We use Demucs to separate the vocal and guitar tracks. Since Demucs primarily processes mono audio, for stereo audio, it is converted to mono before separation, preserving the original left and right channel information in the separated source signals. The input to Demucs is the features of the mixed audio, and the output is an IBM or IRM for the vocals and guitar sounds. For the new mixed audio, the Demucs model predicts masks for the vocals and guitar sounds. For example, IBM predicts a 1 for each vocal-dominant time-frequency unit and a 0 for each guitar-dominant unit, and vice versa. Using the predicted IBM or IRM, the masks are applied to the features of the mixed audio. For the vocal IBM, the corresponding time-frequency units are extracted from the mixed audio, reconstructed to obtain the vocal signal, and its left and right channel signals are obtained. For the guitar sound, the corresponding time-frequency units are extracted from the mixed audio, reconstructed to obtain the guitar signal, and its left and right channel signals are obtained.
[0104] In this embodiment, sound-image separation is performed on the two-dimensional raw audio data to obtain data of independent two-dimensional sound-image objects, which facilitates the acquisition of the left and right channel signals of independent two-dimensional sound-image objects, thus laying the foundation for extracting information of two-dimensional sound-image objects.
[0105] Figure 5 is a flowchart illustrating another audio data processing method provided in an embodiment of this application. Based on the embodiment shown in Figure 3 or Figure 4, Figure 5 further includes, optionally, the following step before executing S32:
[0106] S31: Establish an environmental sound field impulse response database.
[0107] Optionally, an environmental sound field impulse response database can be established based on two-dimensional raw audio data samples.
[0108] In this embodiment, by establishing an environmental sound field impulse response database, a foundation is laid for obtaining two-dimensional sound image data of audio data.
[0109] Figure 6 is a flowchart illustrating another audio data processing method provided in an embodiment of this application. As shown in Figure 6, based on the embodiment shown in Figure 3, a possible implementation of step S31 is as follows:
[0110] S311: Acquire the impulse response signal of the two-dimensional original audio signal sample within the coverage area of the ambient sound field.
[0111] It receives impulse response signals from two-dimensional raw audio data samples transmitted within the stereo sound field coverage area. When sound waves propagate in enclosed spaces such as vehicles, they interact with the geometry of the space, forming unique impulse responses. These impulse responses can be used to analyze the location and characteristics of sound sources. An environmental sound field impulse response database can be obtained through measured or simulated impulse responses.
[0112] S312: Obtain the binaural time difference and binaural sound level difference of two-dimensional raw audio signal samples based on impulse response signals.
[0113] Interaural time difference (ITD) is the time difference between when a sound source enters the two ears, providing information about the location of the sound image. Interaural intensity difference (ILD) is the difference in sound level when a sound source enters the two ears, i.e., the difference in signal intensity between the left and right acoustic channels, and also provides information about the location of the sound image. ITD mainly affects the localization of low-frequency sounds, while ILD is more effective at localizing high-frequency sounds.
[0114] As shown in Figure 7, which is a schematic diagram of an environmental sound field impulse response database provided in an embodiment of this application, region 701 represents the stereo sound field coverage area, specifically the two-dimensional fan-shaped planar area in front of the passenger in the vehicle. In the figure, W represents the width of the stereo sound field, and D represents the depth of the stereo sound field. W and D correspond to the actual sound field in the vehicle. Sound source S is arranged within region 601, and the azimuth coordinates (R, θ) of the sound source traverse the entire region with O as the origin, sending impulse signals. MicL is the left receiving microphone, and MicR is the right receiving microphone, installed on both sides of the headrest of the passenger seat. The distance L between the left and right receiving microphones is approximately equal to the distance between a person's ears. The impulse responses received by MicL and MicR are recorded as h, respectively. L (R,θ) and h R (R,θ), by comparing h L (R,θ) and h R The time axis (R,θ) can be used to obtain the time difference between the two ears, by comparing h L (R,θ) and h R The impulse response amplitude of (R,θ) yields the binaural sound level difference.
[0115] S313: Based on the binaural time difference and binaural sound level difference of the two-dimensional original audio signal samples, and the two-dimensional sound image data corresponding to the two-dimensional original audio signal samples, establish an environmental sound field impulse response database.
[0116] The received impulse response can be represented as a function h(ILD,ITD,R,θ), thereby constructing an environmental sound field impulse response database.
[0117] In this embodiment, the impulse response signal of a two-dimensional original audio signal sample within the coverage area of an ambient sound field is acquired; based on the impulse response signal, the binaural time difference and binaural sound level difference of the two-dimensional original audio signal sample are acquired; based on the binaural time difference and binaural sound level difference of the two-dimensional original audio signal sample, and the corresponding two-dimensional sound image data, an ambient sound field impulse response database is established. This improves the accuracy of acquiring two-dimensional sound image data.
[0118] Figure 8 is a flowchart illustrating another audio data processing method provided in this application embodiment. Figure 8 further illustrates a possible implementation of step S32 based on the embodiment shown in Figure 3:
[0119] S321: Obtain the left and right channel signals of a two-dimensional audio-visual object based on the two-dimensional raw audio data.
[0120] Using two independent microphones: Place the two microphones at a certain distance, one corresponding to the left channel and the other to the right channel, to capture the stereo effect and thus obtain the left and right channel signals of the two-dimensional sound image object.
[0121] S322: Based on the left channel signal, right channel signal, and ambient sound field impulse response database, acquire two-dimensional sound image data. The ambient sound field impulse response database includes: the left channel signal and right channel signal of the two-dimensional original audio data, and the correspondence between them and the two-dimensional sound image data.
[0122] The ambient sound field impulse response database describes the correspondence between the left and right channel signals of the two-dimensional raw audio data and the two-dimensional sound image data. By obtaining the left and right channel signals of a specific two-dimensional sound image object, the corresponding two-dimensional sound image data can be matched.
[0123] In this embodiment, the left and right channel signals of a two-dimensional audio-visual object are acquired. Based on the left and right channel signals and an environmental sound field impulse response database, two-dimensional audio-visual data is obtained. The environmental sound field impulse response database includes the left and right channel signals of the original two-dimensional audio data and their correspondence with the two-dimensional audio-visual data. That is, by obtaining the left and right channel signals of the two-dimensional audio-visual object, the left and right channel signals of the two-dimensional audio-visual object are matched with the left and right channel signals in the environmental sound field impulse response database. If a match is found, the two-dimensional audio-visual data of the two-dimensional audio-visual object is obtained.
[0124] Figure 9 is a flowchart illustrating another audio data processing method provided in this application embodiment. Figure 9 further illustrates a possible implementation of step S322 based on the embodiment shown in Figure 8.
[0125] S3221: Based on the left and right channel signals, obtain the time difference between the two ears and the sound level difference between the two ears.
[0126] ITD can use the cross-correlation function to obtain the time delay between the left and right channel signals. The cross-correlation function calculates the correlation between two signals at different time offsets, and the formula for the cross-correlation function is:
[0127]
[0128] Where, x * Let y(t) represent the complex conjugate of the left channel signal, y(t) represent the right channel signal, τ be the time delay, and R(τ) be the cross-correlation function, representing the similarity between the two signals under the time delay. The maximum value of the cross-correlation function is found; the corresponding τ is the binaural time difference. A positive result indicates that the left channel signal arrives before the right channel signal; a negative result indicates that the right channel signal arrives before the left channel signal.
[0129] For example, suppose the left channel signal x(t) = sin(2πft), and the right channel signal y(t) = sin(2πf(t-0.5)), and suppose f = 1Hz. Since x(t) is a real function, x... * If x(t) = x(t), then its cross-correlation function is:
[0130]
[0131] Since the integral of the second term cos(4πt+2πτ-π) over the entire domain is 0, we only need to calculate:
[0132]
[0133] R(τ) is at its maximum when τ = 0.5, therefore the left channel signal arrives 0.5 seconds earlier than the right channel signal.
[0134] ILD (Input Mean Square) is calculated by comparing the root mean square (RMS) values of two signals. The difference between the RMS values of the left and right channel signals is the ILD. Specifically, the RMS values are calculated separately for the left and right channel signals. The RMS value is an effective indicator of signal power, and its formula is:
[0135]
[0136] Where s(t) is the signal, and the left and right channel signals are substituted into the calculation, and T is the signal duration. The difference between the RMS value of the left channel signal and the RMS value of the right channel signal is calculated; this difference is the ILD. If the RMS value of the left channel signal is greater than the RMS value of the right channel signal, the ILD is positive; conversely, if the RMS value of the left channel signal is less than the RMS value of the right channel signal, the ILD is negative. For example, suppose there is a left channel signal S... L (t) = Asin(2πft), right channel signal S R (t) = Bsin(2πft + φ), where A and B are amplitudes, f is the common frequency, and φ is the phase difference. According to the definition of RMS, the root mean square of the left channel signal is:
[0137]
[0138] The root mean square of the right channel signal is:
[0139]
[0140] but Assuming A = 1.0 indicates that the amplitude of the left channel is 1, and B = 0.5 indicates that the amplitude of the right channel is 0.5, then... A positive result indicates that the signal strength of the left channel is greater than that of the right channel.
[0141] S3222: Two-dimensional acoustic image data is obtained by matching the binaural time difference and binaural sound level difference with the environmental sound field impulse response database.
[0142] Each record in the environmental sound field impulse response database includes: two-dimensional sound image data, the correspondence between binaural time difference and binaural sound level difference.
[0143] For a given two-dimensional raw audio data, after obtaining ITD and ILD through the left and right channel signals, they are matched in the environmental sound field impulse response database. When a matching ITD and ILD are found, the corresponding R and θ are obtained, and R and θ are recorded as two-dimensional sound image coordinates (R, θ).
[0144] For example, given raw two-dimensional audio data, with an ITD of 0.0005 seconds and an ILD of 1.5 dB calculated, a pre-established environmental sound field impulse response database is used. This database contains ITD and ILD values corresponding to sound sources at different locations, along with their corresponding R and θ values. The database is searched for the entry closest to the measured ITD = 0.0005 seconds and ILD = 1.5 dB. Assuming the database contains the following match: ITD = 0.0005 seconds, ILD = 1.5 dB, corresponding R = 20 cm, θ = 45°, then the recorded two-dimensional sound image coordinates are (20, 45).
[0145] In this embodiment, the left and right channel signals of the two-dimensional raw audio data are acquired; based on the left and right channel signals of each two-dimensional raw audio data, the binaural time difference and binaural sound level difference are obtained; based on the calculated binaural time difference and binaural sound level difference, the two-dimensional sound image data corresponding to the two-dimensional raw audio data are matched with the environmental sound field impulse response database; that is, by matching the obtained binaural time difference and binaural sound level difference with the data in the environmental sound field impulse response database, the two-dimensional sound image data of the two-dimensional raw audio data can be quickly calculated.
[0146] The stereo image deep learning model shown in Figure 3 is obtained by pre-training based on stereo version data samples. Figure 10 is a flowchart illustrating another audio data processing method provided in this application embodiment. Figure 10 further describes a possible implementation of the training method for the stereo image deep learning model based on the embodiment shown in Figure 3, as shown in Figure 10, including:
[0147] S1001: Input the stereo version data sample into the stereo image deep learning model to perform three-dimensional sound image rendering and obtain the predicted three-dimensional sound image data.
[0148] The stereo version data sample includes: two-dimensional sound source object data and two-dimensional sound image coordinate data. The two-dimensional sound source object data is the two-dimensional channel information of the training sound source object, and the two-dimensional sound image coordinate data is the two-dimensional coordinate data of the training sound image object.
[0149] The predicted 3D audio-visual data includes: predicted 3D sound source object data and predicted 3D audio-visual coordinate data. The predicted 3D sound source object data is the 3D channel information of the training sound source object, and the 3D audio-visual coordinate data is the 3D coordinate data of the training audio-visual object.
[0150] The target 3D acoustic image data includes: target 3D sound source object data and target 3D acoustic image coordinate data. The target 3D sound source object data is the multi-channel sound source signal data of the training sound source object, and the target 3D acoustic image coordinate data is the 3D coordinate data of the training sound source object. The target 3D acoustic image coordinates include the pitch angle.
[0151] Optionally, target three-dimensional acoustic image data can be obtained based on multi-channel version data samples.
[0152] Optionally, Pro Tools can be used to obtain the 2D sound source object data and 2D sound image coordinate data corresponding to the stereo version data. In the Pro Tools editing interface, create a new stereo track, import an existing audio file, and use the "Import" function to add the audio file to the stereo track. In the mix view, locate the channel balance control slider (Pan Slider) of the stereo track. Move the Pan Slider to adjust the left and right position of the audio in the stereo field; moving it to the left shifts the sound towards the left channel, and moving it to the right shifts it towards the right channel. In Pro Tools, add a reverb effect to the track, selecting a suitable reverb type, such as room reverb or hall reverb. Adjust the reverb size, pre-delay, time, number of reflections, volume, and hue to simulate the distance of the sound source. After adjusting the sound source localization, record the Pan Slider position, reverb, and delay effect settings. This will give you the sound source distance and azimuth.
[0153] The obtained two-dimensional sound source object data and two-dimensional sound image coordinate data are combined with the channel parameters of the vehicle audio system to train a stereo sound image deep learning model. When using the stereo sound image deep learning model for sound source localization, frequency domain features can be used. For example, steering vectors can be used as input. The input vectors are correlated with the direction of the sound source, and each element has a unit amplitude. The signal subspace is the subspace spanned by the steering vectors.
[0154] Assuming a sound source signal with a distance R = 30cm and an azimuth angle θ = 30°, the candidate two-dimensional sound image coordinate data is (30cm, 30°), and the coordinates mapped to the two-dimensional rectangular coordinate axis after initialization are ( (15cm). The neural network of the stereo imaging deep learning model contains an input layer, multiple hidden layers, and an output layer. The input layer receives preprocessed data, the hidden layers are responsible for extracting features and performing nonlinear transformations, automatically learning the mapping relationship from two-dimensional acoustic imaging information to three-dimensional acoustic imaging data, and the output layer predicts the three-dimensional acoustic imaging data.
[0155] A Convolutional Neural Network (CNN) is used to automatically extract acoustic features, resulting in predicted 3D sound image data as the output. Assume a sound source signal is emitted from the vehicle's smart display screen. The distance between this sound source signal and the first speaker is R = 30cm, and the azimuth angle is θ = 30°. The vehicle's audio system has four channels, with speakers located on either side of the headrests of the driver and passenger seats. The vertical distance difference between the smart display screen and the headrests is 30cm. An input dataset containing these parameters can be constructed. The CNN uses this data as input. First, it determines the position of the sound source relative to the vehicle's coordinate system. We can assume the vehicle's smart display screen is located at the origin (0, 0, 0) of the coordinate system. Its coordinates relative to the first speaker on the two-dimensional Cartesian coordinate axis are (…). 15cm). Given that the smart display screen is 40 cm away from the first headrest in the vertical direction (i.e., its Z-axis coordinate is 40), then the three-dimensional coordinates of the sound source are ( (15cm, 30cm) is converted into predicted three-dimensional acoustic coordinate data as (30cm, 30°, 45°).
[0156] Optionally, target 3D sound image data can be obtained based on multi-channel version data samples. For example, Pro Tools can be used to obtain target 3D sound image data. Pro Tools supports Dolby Atmos and can record multi-channel audio source signals, thus obtaining target 3D sound source object data and target 3D sound image coordinate data, with each channel corresponding to a specific speaker. The Pro Tools 3D Panner plugin can be used to place and move audio objects to simulate the position of the sound source in 3D space. The target 3D sound image coordinate data can be set and recorded using the X, Y, and Z parameters in the 3D Panner.
[0157] S1003: Update the stereoscopic deep learning model based on predicted 3D audio-visual data and target 3D audio-visual data.
[0158] Optionally, one possible implementation is as follows: based on the predicted 3D audio-visual data, the target 3D audio-visual data, and the loss function, update the stereo audio-visual deep learning model, where the loss function is a function that measures the difference between the predicted 3D audio-visual data and the target 3D audio-visual data.
[0159] S1005: Until the stereo imaging deep learning model converges, the target 3D imaging data is the preset benchmark 3D imaging data.
[0160] Once the stereo audio-visual learning model has converged, training is stopped, and the target 3D audio-visual data is set to the preset benchmark 3D audio-visual data.
[0161] Optionally, prior to S1005, the following may also be included:
[0162] S10041: Input the stereo version data sample into the updated stereo image deep learning model to obtain the updated predicted three-dimensional stereo image data.
[0163] In other words, if the stereo image deep learning model does not converge, the stereo version data sample is input into the updated stereo image deep learning model to obtain the updated predicted 3D stereo image data.
[0164] S10042: Determine whether the stereo imaging deep learning model has converged. If it has converged, proceed to S1005; otherwise, proceed to S10043.
[0165] When the loss value meets the preset conditions, it is determined whether the stereo imaging deep learning model has converged.
[0166] S10043: Update the stereoscopic deep learning model based on the updated predicted 3D acoustic image data and the target 3D acoustic image data. Return to execute S10041.
[0167] Specifically, the updated predicted 3D audio-visual data is used as the predicted value, and the target 3D audio-visual data is used as the target value. The loss function between the predicted value and the target value is calculated, and the parameters of the stereo audio-visual deep learning model are updated based on the loss function until the stereo audio-visual learning model converges.
[0168] In this embodiment, stereo version data samples are input into a stereo imaging deep learning model for 3D audio-visual rendering to obtain predicted 3D audio-visual data. Based on the predicted 3D audio-visual data and the target 3D audio-visual data, the stereo imaging deep learning model is updated. The stereo version data samples are then input into the updated stereo imaging deep learning model to obtain updated predicted 3D audio-visual data. It is then determined whether the stereo imaging deep learning model has converged. If it has converged, the target 3D audio-visual data is set to the preset benchmark 3D audio-visual data. If it has not converged, the process returns to inputting the stereo version data samples into the updated stereo imaging deep learning model to obtain updated predicted 3D audio-visual data. This makes the output of the stereo imaging deep learning model closer to the target value, thereby improving the accuracy of the stereo imaging deep learning model in rendering sound source data from two dimensions to three dimensions.
[0169] Figure 11 is a flowchart illustrating another audio data processing method provided in an embodiment of this application. Based on the embodiment shown in Figure 10, Figure 11 further illustrates a possible implementation of step S1003 as follows:
[0170] S1021: Obtain the loss value based on the predicted 3D acoustic image data, the target 3D acoustic image data, and the loss function.
[0171] The loss function is calculated by using the predicted 3D acoustic image data as the predicted value and the target 3D acoustic image data as the target value. Loss functions include mean squared error loss, cross-entropy loss, and contrastive loss. The choice of loss function should match the characteristics of the task and the output type of the model. An appropriate optimizer is selected to minimize the loss function, such as stochastic gradient descent. The hyperparameters of the optimization algorithm, such as learning rate, momentum, and decay rate, are initialized to improve the training speed and stability of the model.
[0172] For example, the loss can be calculated using the Mean Squared Error (MSE) function, and the formula is as follows:
[0173]
[0174] Where N is the number of samples, y i For the target value, These are predicted values.
[0175] S1022: Update the parameters of the stereo imaging deep learning model based on the loss value.
[0176] Optional parameters for the stereo imaging deep learning model include weights and biases.
[0177] The gradient of the model parameters is obtained by calculating the loss function. The gradient is the derivative of the loss function in the parameter space, pointing in the direction in which the loss function increases the most.
[0178] For the MSE loss, the gradient with respect to parameter a can be expressed as:
[0179]
[0180] in, This is the derivative of the predicted value with respect to parameter 'a', where 'a' specifically includes the weights and biases. Since neural networks are multi-layered, the chain rule is needed to calculate the gradient of each layer. For each layer in the network, the gradient can be calculated as:
[0181]
[0182] Among them, h l It is the output of the l-th layer, h l+1 It is the output of the (l+1)th layer.
[0183] After calculating the gradient, the optimizer updates the model parameters, mainly the weights and biases. The model is then iteratively trained. The process returns to step S1001, where the stereo version data samples are input into the stereo deep learning model for 3D audio-visual rendering to obtain the predicted 3D audio-visual data. The parameter update formula can be expressed as:
[0184]
[0185] Where η is the learning rate, a new This is the updated parameter, a old These are the parameters before the update.
[0186] S1023: Determine whether the loss value meets the preset conditions.
[0187] If the loss value meets the preset conditions, the stereo image learning model is determined to have converged, and a stereo image deep learning model is obtained; if the loss value does not meet the preset conditions, the stereo image learning model has not converged, and the process returns to execute S1001.
[0188] The stereo image learning model is considered converged when the loss value meets a preset condition. For example, if the preset condition is that the loss value is less than 0.45, then during the training process, the stereo image learning model is considered converged when the loss value is less than 0.45.
[0189] In this embodiment, a loss value is obtained based on the predicted 3D audio-visual data, the target 3D audio-visual data, and a loss function. The parameters of the stereo audio-visual deep learning model are updated based on the loss value. The model then returns to its previous state, inputting stereo version data samples into the stereo audio-visual deep learning model for 3D audio-visual rendering to obtain the predicted 3D audio-visual data. This process continues until the loss value meets a preset condition, at which point the stereo audio-visual learning model is considered converged. This provides a basic model for rendering audio-visual objects from 2D to 3D, thereby realizing the rendering of audio-visual objects from a 2D audio-visual plane to a 3D audio-visual space, improving the user's listening experience.
[0190] Figure 12 is a flowchart illustrating another audio data processing method provided in an embodiment of this application. As shown in Figure 12, based on the embodiment shown in Figure 10, it may further include:
[0191] S121: Divide the master files of the 3D immersive sound source into training sample sets, validation datasets, and test datasets.
[0192] The proportions of the dataset should be determined. Common proportions include 70% training dataset, 15% validation dataset, and 15% test dataset, or 60% training dataset, 20% validation dataset, and 20% test dataset. These proportions can be adjusted according to actual circumstances and needs; this application does not impose any restrictions. The validation and test datasets are not involved in the model training parameter tuning process. The validation dataset is used to test the model's predictive ability on new data during model training, while the test dataset is used to test the model's predictive ability on new data after model convergence.
[0193] S122: Train a stereo deep learning model based on the training sample set.
[0194] The implementation method of this step is similar to that of the embodiment shown in Figure 10, and will not be described again here.
[0195] S123: During the training process of the stereo imaging deep learning model, the first performance parameter of the stereo imaging deep learning model is evaluated based on the validation dataset, and the hyperparameters of the stereo imaging deep learning model are adjusted based on the first performance parameter until the stereo imaging deep learning model converges.
[0196] The first performance parameter is used to evaluate the predictive ability of the stereo imaging deep learning model on the validation dataset, while the hyperparameters are used to control the model's behavior and performance.
[0197] The primary performance parameters include accuracy, precision, and recall. In other words, after training the model for a period of time, its predictive ability on unseen data is tested using a validation dataset. Based on this, the hyperparameters are adjusted, and the stereo imaging deep learning model is then trained again.
[0198] The hyperparameters adjusted include the learning rate and regularization parameter. The candidate deep learning model for audio-visual rendering is evaluated using a validation dataset, and performance metrics such as accuracy and loss are calculated.
[0199] For example, cross-validation methods, such as k-fold cross-validation, can be used to evaluate the model's generalization ability. The learning rate is a key hyperparameter controlling the magnitude of model weight updates; different learning rate values can be tried to evaluate the predictive ability of the stereo imaging deep learning model. Adjusting the values of regularization parameters, such as the weight decay coefficient, can control model complexity and prevent overfitting. Automated hyperparameter optimization methods, such as Bayesian optimization, can be used to find the optimal regularization parameters.
[0200] For example, in the training process of a stereo imaging deep learning model, after setting the learning rate to 0.001 and iterating 100 times, the stereo imaging deep learning model is evaluated using a validation dataset, and the returned accuracy is 70%. The learning rate is then adjusted to 0.005 and the stereo imaging deep learning model is trained again until iterates 100 times again. After evaluating the stereo imaging deep learning model using a validation dataset, a new returned accuracy of 75% is obtained. By continuously adjusting the hyperparameters set in the stereo imaging deep learning model and evaluating the stereo imaging deep learning model using a validation dataset after a certain number of iterations, the stereo imaging deep learning model is adjusted to suitable hyperparameters.
[0201] In this embodiment, the master file of the 3D immersive sound source is divided into a training sample set, a validation dataset, and a test dataset. The training sample set is used to train the stereo imaging deep learning model, the validation dataset is used to evaluate the first performance parameter of the stereo imaging deep learning model, and the hyperparameters of the stereo imaging deep learning model are adjusted based on the first performance parameter, thereby avoiding overfitting of the stereo imaging deep learning model.
[0202] Figure 13 is a flowchart illustrating another audio data processing method provided in an embodiment of this application. Figure 13, based on the embodiment shown in Figure 10, optionally further includes:
[0203] S124: After the stereo imaging deep learning model converges, evaluate the stereo imaging deep learning model based on the test dataset to obtain the second performance parameter of the stereo imaging deep learning model.
[0204] The second performance parameter is used to evaluate the predictive ability of the stereo imaging deep learning model.
[0205] After model convergence, a stereo imaging deep learning model is obtained. This model is then evaluated using a test dataset to obtain its second performance measure, which reflects its performance on unseen data. The model is then evaluated using the test dataset, and sound quality evaluation metrics are calculated.
[0206] For example, spectral flux, spectral roll-off, and binaural cross-correlation coefficients reflect the spatial distribution and localization perception of sound. Spectral flux is an indicator of the degree of spectral variation, obtained by calculating the difference in the spectrum between consecutive frames. Specifically, for each frame of signal, a Fast Fourier Transform is performed to obtain the spectrum, the sum of squares of the differences between the spectra of adjacent frames is calculated, and the square root of the sum of squares is taken to obtain the spectral flux value. A larger spectral flux indicates more drastic spectral variation, which is usually related to dynamic changes in audio signals, such as rhythmic changes in music or pitch changes in speech. Spectral roll-off refers to the highest frequency corresponding to a certain proportion (e.g., 85% or 90%) of energy accumulated in the spectrum. Spectral roll-off reflects the width of the signal energy distribution and can be used to distinguish the spectral characteristics of different instruments or sounds. In sound imaging, spectral roll-off can help identify the brightness and clarity of sound, thus affecting the spatial localization of sound.
[0207] For example, after the stereo imaging deep learning model converges, the performance of the model can be evaluated using the Jensen-Shannon Divergence (JSD). The test dataset is input, and the JSD value is obtained. JSD is an indicator that measures the similarity between two probability distributions, with a value range of [0, 1]. Assuming the evaluation result is JSD = 0.1, it indicates that the distribution of the generated predicted 3D stereo imaging data is very close to that of the target 3D stereo imaging data, meaning that the stereo imaging deep learning model has good predictive ability for unknown data.
[0208] In this embodiment, the stereo imaging deep learning model is evaluated based on a test dataset, which can provide the model's performance on unseen data.
[0209] Figure 14 is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application, including: an acquisition module 1401 and a processing module 1402, wherein the acquisition module 1401 is used to acquire two-dimensional audio-visual data, and the two-dimensional audio-visual data is at least one two-dimensional audio-visual data corresponding to two-dimensional original audio data;
[0210] The processing module 1402 is used to input the two-dimensional audio-visual data into the stereo audio-visual deep learning model and output three-dimensional audio-visual data.
[0211] Optionally, the stereo image deep learning model is obtained by pre-training based on stereo version data samples.
[0212] Optionally, the processing module 1402 is further configured to input the stereo version data sample into the stereo image deep learning model for three-dimensional audio-visual rendering to obtain predicted three-dimensional audio-visual data; update the stereo image deep learning model based on the predicted three-dimensional audio-visual data and the target three-dimensional audio-visual data; until the stereo image deep learning model converges, and the target three-dimensional audio-visual data is the preset benchmark three-dimensional audio-visual data.
[0213] Optionally, the processing module 1402 is further configured to, when the stereo image deep learning model does not converge, input the stereo version data sample into the updated stereo image deep learning model to obtain the updated predicted three-dimensional sound image data; and update the stereo image deep learning model based on the updated predicted three-dimensional sound image data and the target three-dimensional sound image data.
[0214] Optionally, the processing module 1402 is specifically used to update the stereoscopic deep learning model based on the predicted three-dimensional audio-visual data, the target three-dimensional audio-visual data, and the loss function, wherein the loss function is a function that measures the difference between the predicted three-dimensional audio-visual data and the target three-dimensional audio-visual data.
[0215] Optionally, the stereo version data sample includes: two-dimensional sound source object data and two-dimensional sound image coordinate data, wherein the two-dimensional sound source object data is the two-dimensional channel information of the training sound source object, and the two-dimensional sound image coordinate data is the two-dimensional coordinate data of the training sound image object.
[0216] Optionally, the predicted three-dimensional acoustic image data includes: predicted three-dimensional sound source object data and predicted three-dimensional acoustic image coordinate data, wherein the predicted three-dimensional sound source object data is the three-dimensional channel information of the training sound source object, and the three-dimensional acoustic image coordinate data is the three-dimensional coordinate data of the training acoustic image object.
[0217] Optionally, the processing module 1402 is further configured to acquire the target three-dimensional acoustic image data based on multi-channel version data samples.
[0218] Optionally, the target three-dimensional acoustic image data includes: target three-dimensional sound source object data and target three-dimensional acoustic image coordinate data, wherein the target three-dimensional sound source object data is multi-channel sound source signal data of the training sound source object, and the target three-dimensional acoustic image coordinate data is three-dimensional coordinate data of the training sound source object.
[0219] Optionally, the target three-dimensional acoustic coordinates include the pitch angle.
[0220] Optionally, the processing module 1402 is further configured to, during the training process of the stereo imaging deep learning model, evaluate a first performance parameter of the stereo imaging deep learning model based on a validation dataset, and adjust the hyperparameters of the stereo imaging deep learning model based on the first performance parameter until the stereo imaging deep learning model converges, wherein the first performance parameter is used to evaluate the predictive ability of the stereo imaging deep learning model on the validation dataset, and the hyperparameters are used to control the behavior and performance of the model.
[0221] Optionally, the processing module 1402 is further configured to evaluate the stereo imaging deep learning model based on a test dataset after the stereo imaging deep learning model converges, and obtain a second performance parameter of the stereo imaging deep learning model, wherein the second performance parameter is used to evaluate the prediction ability of the stereo imaging deep learning model.
[0222] Optionally, the acquisition module 1401 is specifically used to acquire the two-dimensional sound image data based on the two-dimensional raw audio data and the environmental sound field impulse response database, wherein the environmental sound field impulse response database includes the correspondence between the two-dimensional raw audio data and the two-dimensional sound image data.
[0223] Optionally, the acquisition module 1401 is specifically used to acquire the left channel signal and right channel signal of the two-dimensional sound image object based on the two-dimensional original audio data; and to acquire the two-dimensional sound image data based on the left channel signal, the right channel signal, and an environmental sound field impulse response database, wherein the environmental sound field impulse response database includes: the left channel signal and right channel signal of the two-dimensional original audio data, and the correspondence between them and the two-dimensional sound image data.
[0224] Optionally, the acquisition module 1401 is specifically used to acquire the binaural time difference and binaural sound level difference based on the left channel signal and the right channel signal; and to match the binaural time difference and binaural sound level difference with the environmental sound field impulse response database to obtain the two-dimensional sound image data.
[0225] Optionally, the acquisition module 1401 is further configured to perform sound-image separation based on the two-dimensional raw audio data to obtain data of independent two-dimensional sound-image objects.
[0226] Optionally, the processing module 1402 is also used to establish the environmental sound field impulse response database.
[0227] Optionally, the processing module 1402 is specifically used to acquire the impulse response signal of the two-dimensional original audio data sample within the coverage area of the ambient sound field; acquire the binaural time difference and binaural sound level difference of the two-dimensional original audio data sample based on the impulse response signal; and establish the ambient sound field impulse response database based on the binaural time difference and binaural sound level difference of the two-dimensional original audio data sample, as well as the two-dimensional sound image data corresponding to the two-dimensional original audio data sample.
[0228] The apparatus in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0229] This application also provides an electronic device, which includes a processor and a memory. The memory stores programs or instructions that can run on the processor, and when executed by the processor, the programs or instructions implement the steps of the audio data processing method shown in Figures 3 to 13.
[0230] This application also provides a vehicle, which includes a processor and a memory. The memory stores programs or instructions that can run on the processor. When the program or instructions are executed by the processor, they implement the steps of the audio data processing method shown in Figures 3 to 13.
[0231] This application also provides a computer-readable storage medium storing a program or instructions, which, when executed by a processor, implement the steps of the audio data processing method shown in Figures 3 to 13.
[0232] This application also provides a computer program product, which, when executed by a processor of a vehicle or a cloud server, implements the steps of the audio data processing method shown in Figures 3 to 13.
[0233] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0234] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. A method for processing audio data, characterized in that, include: Acquire two-dimensional audio-visual data, wherein the two-dimensional audio-visual data is at least one two-dimensional audio-visual data corresponding to the two-dimensional original audio data; The two-dimensional audio-visual data is input into the stereo audio-visual deep learning model, and three-dimensional audio-visual data is output.
2. The method according to claim 1, characterized in that, The stereo deep learning model is obtained by pre-training based on stereo version data samples.
3. The method according to claim 2, characterized in that, Also includes: The training method of the stereo imaging deep learning model includes: inputting the stereo version data sample into the stereo imaging deep learning model to perform three-dimensional sound and image rendering to obtain predicted three-dimensional sound and image data; Based on the predicted 3D audio-visual data and the target 3D audio-visual data, the stereo audio-visual deep learning model is updated until the stereo audio-visual deep learning model converges, and the target 3D audio-visual data is the preset benchmark 3D audio-visual data.
4. The method according to claim 3, characterized in that, After updating the stereo audio-visual deep learning model based on the predicted 3D audio-visual data and the target 3D audio-visual data, the method further includes: if the stereo audio-visual deep learning model does not converge, inputting the stereo version data sample into the updated stereo audio-visual deep learning model to obtain the updated predicted 3D audio-visual data; and updating the stereo audio-visual deep learning model based on the updated predicted 3D audio-visual data and the target 3D audio-visual data.
5. The method according to claim 3, characterized in that, The step of updating the stereo deep learning model based on the predicted 3D audio-visual data and the target 3D audio-visual data includes: updating the stereo deep learning model based on the predicted 3D audio-visual data, the target 3D audio-visual data, and a loss function, wherein the loss function is a function that measures the difference between the predicted 3D audio-visual data and the target 3D audio-visual data.
6. The method according to claim 2, characterized in that, The stereo version data sample includes: two-dimensional sound source object data and two-dimensional sound image coordinate data. The two-dimensional sound source object data is the two-dimensional channel information of the training sound source object, and the two-dimensional sound image coordinate data is the two-dimensional coordinate data of the training sound image object.
7. The method according to claim 3, characterized in that, The predicted three-dimensional acoustic image data includes: predicted three-dimensional sound source object data and predicted three-dimensional acoustic image coordinate data. The predicted three-dimensional sound source object data is the three-dimensional channel information of the training sound source object, and the three-dimensional acoustic image coordinate data is the three-dimensional coordinate data of the training acoustic image object.
8. The method according to claim 3, characterized in that, Also includes: The target's three-dimensional acoustic image data is obtained based on multi-channel version data samples.
9. The method according to claim 3, characterized in that, The target three-dimensional acoustic image data includes: target three-dimensional sound source object data and target three-dimensional acoustic image coordinate data. The target three-dimensional sound source object data is multi-channel sound source signal data of the training sound source object, and the target three-dimensional acoustic image coordinate data is three-dimensional coordinate data of the training sound source object.
10. The method according to claim 9, characterized in that, The target's three-dimensional acoustic coordinates include the pitch angle.
11. The method according to claim 2, characterized in that, The method further includes: during the training process of the stereo imaging deep learning model, evaluating a first performance parameter of the stereo imaging deep learning model based on a validation dataset, and adjusting the hyperparameters of the stereo imaging deep learning model based on the first performance parameter until the stereo imaging deep learning model converges, wherein the first performance parameter is used to evaluate the predictive ability of the stereo imaging deep learning model on the validation dataset, and the hyperparameters are used to control the behavior and performance of the model.
12. The method according to claim 2, characterized in that, The method further includes: after the stereo imaging deep learning model converges, evaluating the stereo imaging deep learning model based on a test dataset to obtain a second performance parameter of the stereo imaging deep learning model, wherein the second performance parameter is used to evaluate the predictive ability of the stereo imaging deep learning model.
13. The method according to any one of claims 1-12, characterized in that, The acquisition of two-dimensional acoustic image data includes: acquiring the two-dimensional acoustic image data based on two-dimensional raw audio data and an environmental sound field impulse response database, wherein the environmental sound field impulse response database includes: the correspondence between the two-dimensional raw audio data and the two-dimensional acoustic image data.
14. The method according to claim 13, characterized in that, The process of obtaining the two-dimensional sound image data based on two-dimensional raw audio data and an environmental sound field impulse response database includes: obtaining the left channel signal and right channel signal of the two-dimensional sound image object based on the two-dimensional raw audio data; and obtaining the two-dimensional sound image data based on the left channel signal, the right channel signal, and the environmental sound field impulse response database, wherein the environmental sound field impulse response database includes: the correspondence between the left channel signal and the right channel signal of the two-dimensional raw audio data and the two-dimensional sound image data.
15. The method according to claim 14, characterized in that, The step of obtaining the two-dimensional sound image data based on the left channel signal, the right channel signal, and the ambient sound field impulse response database includes: obtaining the binaural time difference and binaural sound level difference based on the left channel signal and the right channel signal; and matching the binaural time difference and binaural sound level difference with the ambient sound field impulse response database to obtain the two-dimensional sound image data.
16. The method according to claim 15, characterized in that, Before obtaining the left and right channel signals of the two-dimensional sound image object based on the two-dimensional raw audio data, the method further includes: performing sound image separation based on the two-dimensional raw audio data to obtain data of the independent two-dimensional sound image object.
17. The method according to claim 15, characterized in that, Before acquiring two-dimensional sound image data based on the left channel signal, the right channel signal, and the ambient sound field impulse response database, the method further includes: establishing the ambient sound field impulse response database.
18. The method according to claim 17, characterized in that, The establishment of the environmental sound field impulse response database includes: acquiring the impulse response signal of a two-dimensional original audio data sample within the coverage area of the environmental sound field; acquiring the binaural time difference and binaural sound level difference of the two-dimensional original audio data sample based on the impulse response signal; and establishing the environmental sound field impulse response database based on the binaural time difference and binaural sound level difference of the two-dimensional original audio data sample, as well as the two-dimensional sound image data corresponding to the two-dimensional original audio data sample.
19. An electronic device, characterized in that, include: A processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the audio data processing method as described in any one of claims 1 to 18.
20. A computer-readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio data processing method as described in any one of claims 1 to 18.
21. A computer program product, which, when executed by a processor of a vehicle or a cloud server, implements the steps of the audio data processing method as described in any one of claims 1 to 18.
22. A vehicle, characterized in that, include: A processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the audio data processing method as described in any one of claims 1 to 18.