Training method of sound effect adjustment model, sound effect adjustment method, device and equipment
By training the environmental spatial information prediction model and the sound wave attenuation model, the sound effect parameters are automatically adjusted to adapt to different environments, solving the problem that traditional sound effect adjustment solutions cannot automatically adjust, and improving the accuracy of sound effect adjustment and user experience.
Patent Information
- Application Number
- CN202510974885.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-24
AI Technical Summary
Traditional sound effect adjustment solutions cannot automatically adjust sound effect parameters according to environmental changes, resulting in a reduced user experience.
By training the environmental space information prediction model and the sound wave attenuation model, we obtain sample environmental space information and sound pressure data. Based on these data, we train the initial sound effect adjustment model and automatically adjust the sound effect parameters to adapt to different environments.
Improves the accuracy and automation of sound effect adjustments, enhances the user experience, and avoids the need for users to manually select audio parameter modes.
Smart Images

Figure CN120833792A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of sound effect adjustment, and in particular to a sound effect adjustment model training method, a sound effect adjustment method, a device and equipment. BACKGROUND
[0002] In current audio processing technology, the adjustment of audio quality (AQ) usually depends on the user manually selecting the type of the current environment, and then manually setting the corresponding sound effect parameters.
[0003] However, the traditional sound effect adjustment scheme only provides limited preset modes, such as "standard", "classroom", "meeting", and "audio-visual" modes. The user needs to manually select the audio parameter mode matching the current environment according to his / her own judgment, in order to achieve the best sound effect performance. Since the setting mode of this traditional sound effect adjustment scheme is fixed and single, the user experience is reduced. SUMMARY
[0004] The present application provides a sound effect adjustment model training method, a sound effect adjustment method, a device and equipment, to solve the defect that the traditional sound effect adjustment scheme in the prior art cannot automatically adjust the sound effect parameters according to the change of the environment, reducing the user experience.
[0005] The present application provides a sound effect adjustment model training method, comprising the following steps: inputting sound pressure data into an environment space information prediction model to obtain sample environment space information output by the environment space information prediction model; the sample environment space information includes space size and space type; training an initial sound effect adjustment model based on the sample environment space information to obtain the sound effect adjustment model.
[0006] According to the sound effect adjustment model training method provided by the present application, the training step of the environment space information prediction model comprises: obtaining an initial environment space information prediction model, sample sound pressure data, and labeled environment space information of the sample sound pressure data; inputting the sample sound pressure data into the initial environment space information prediction model to obtain predicted environment space information output by the initial environment space information prediction model; based on the predicted environment space information and the labeled environment space information, performing parameter iteration on the initial environment space information prediction model to obtain the environment space information prediction model.
[0007] According to the sound effect adjustment model training method provided by the application, the label environment space information includes label space size and label space type; and the predicted environment space information includes predicted space size and predicted space type. The parameter iteration of the initial environment space information prediction model based on the predicted environment space information and the label environment space information includes: Based on the label space size and the predicted space size, a first loss is determined. Based on the label space type and the predicted space type, a second loss is determined. Based on the first loss and the second loss, a prediction loss is determined, and the parameter iteration of the initial environment space information prediction model is performed based on the prediction loss.
[0008] According to the sound effect adjustment model training method provided by the application, the sample sound pressure data acquisition step includes: Sample sound power data of a sound source corresponding to sample microphone array data is acquired. The sample sound power data is input into the sound wave attenuation model to obtain the sample sound pressure data output by the sound wave attenuation model.
[0009] According to the sound effect adjustment model training method provided by the application, the sound wave attenuation model training step includes: Label sound pressure data corresponding to the sample sound power data and an initial sound wave attenuation model are acquired. The sample sound power data is input into the initial sound wave attenuation model to obtain predicted sound pressure data output by the initial sound wave attenuation model. Based on the label sound pressure data and the predicted sound pressure data, the parameter iteration of the initial sound wave attenuation model is performed to obtain the sound wave attenuation model.
[0010] According to the sound effect adjustment model training method provided by the application, the parameter iteration of the initial sound wave attenuation model based on the label sound pressure data and the predicted sound pressure data includes: Based on the difference between the label sound pressure data and the predicted sound pressure data, a similarity loss is determined. Based on the similarity loss, a target loss is determined, and the parameter iteration of the initial sound wave attenuation model is performed based on the target loss.
[0011] According to the sound effect adjustment model training method provided by the application, the determination of the target loss based on the similarity loss includes: Obtaining a positive sample pair and a negative sample pair; the positive sample pair includes a first sound power data pair of a sound source corresponding to first microphone array data, the first microphone array data being collected based on the same sound source or a similar sound source; the negative sample pair includes a second sound power data pair of a sound source corresponding to second microphone array data, the second microphone array data being collected based on a different sound source or a dissimilar sound source; Inputting the first sound power data pair into the initial sound wave attenuation model to obtain first predicted sound pressure data and second predicted sound pressure data output by the initial sound wave attenuation model; Inputting the second sound power data pair into the initial sound wave attenuation model to obtain third predicted sound pressure data and fourth predicted sound pressure data output by the initial sound wave attenuation model; determining a contrast loss based on the first predicted sound pressure data and the second predicted sound pressure data, and the third predicted sound pressure data and the fourth predicted sound pressure data; The target loss is determined based on the similarity loss and the contrast loss.
[0012] According to a training method for a sound effect adjustment model provided by the present invention, determining the target loss based on the similarity loss and the contrast loss includes: Determining an impulse response of a current room corresponding to the sample microphone array data, and a label impulse response corresponding to the impulse response; An impulse response loss is determined based on the impulse response and the label impulse response, and a target loss is determined based on the impulse response loss, the similarity loss, and the contrast loss.
[0013] According to a training method for a sound effect adjustment model provided by the present invention, determining the impulse response loss based on the impulse response and the label impulse response includes: determining a phase loss based on a first phase corresponding to the impulse response and a second phase corresponding to the tag impulse response; determining an amplitude loss based on an amplitude corresponding to the impulse response and a tag amplitude corresponding to the tag impulse response; The impulse response loss is determined based on the phase loss and the amplitude loss.
[0014] The present invention also provides a sound effect adjustment method, comprising the following steps: Obtain the ambient space information corresponding to the sound effect to be adjusted; Inputting the environmental space information into a sound effect adjustment model to obtain a sound effect adjustment parameter output by the sound effect adjustment model, and applying the sound effect adjustment parameter to perform sound effect adjustment on the sound effect to be adjusted; The sound effect adjustment model is obtained based on the training method of the sound effect adjustment model.
[0015] The application further provides a sound effect adjustment model training device, comprising the following units: The first obtaining unit is configured to input the sound pressure data into the environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model; the sample environmental space information comprises a space size and a space type. The training unit is configured to train an initial sound effect adjustment model based on the sample environmental space information to obtain the sound effect adjustment model.
[0016] The application further provides a sound effect adjustment device, comprising the following units: The second obtaining unit is configured to obtain environmental space information corresponding to the sound effect to be adjusted. The first input unit is configured to input the environmental space information into the sound effect adjustment model to obtain sound effect adjustment parameters output by the sound effect adjustment model, and apply the sound effect adjustment parameters to the sound effect adjustment of the sound effect to be adjusted. The sound effect adjustment model is obtained based on the training method of the sound effect adjustment model.
[0017] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the sound effect adjustment model training method or the sound effect adjustment method when executing the program.
[0018] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the sound effect adjustment model training method or the sound effect adjustment method.
[0019] The application further provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the sound effect adjustment model training method or the sound effect adjustment method.
[0020] The application provides a sound effect adjustment model training method, a sound effect adjustment method, a device and equipment. On the one hand, sound pressure data is input into an environmental space information prediction model to obtain sample environmental space information, and then an initial sound effect adjustment model is trained based on the sample environmental space information to obtain a sound effect adjustment model. The sound effect adjustment model can learn the optimal sound effect parameters in different environments, thereby improving the accuracy of sound effect adjustment. On the other hand, compared with the prior art, in which a user needs to manually select an audio parameter mode matching the current environment according to the user's judgment of the current environment, so as to achieve the best sound effect performance, the sound effect adjustment model in the method can automatically adjust the sound effect parameters according to the change of the environment, without the need for the user to manually select the audio parameter mode, thereby improving the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0022] Figure 1 Fig. 1 is a flowchart of a sound effect adjustment model training method provided by the application.
[0023] Figure 2 Fig. 2 is another flowchart of a sound effect adjustment model training method provided by the application.
[0024] Figure 3 Fig. 3 is a flowchart of step 230 in the sound effect adjustment method provided by the application.
[0025] Figure 4 Fig. 4 is a third flowchart of a sound effect adjustment model training method provided by the application.
[0026] Figure 5 Fig. 5 is a fourth flowchart of a sound effect adjustment model training method provided by the application.
[0027] Figure 6 Fig. 6 is a flowchart of a sound effect adjustment method provided by the application.
[0028] Figure 7 Fig. 7 is a structural schematic diagram of a sound effect adjustment model training device provided by the application.
[0029] Figure 8 Fig. 8 is a structural schematic diagram of a sound effect adjustment device provided by the application.
[0030] Figure 9 Fig. 9 is a structural schematic diagram of an electronic device provided by the application. DETAILED DESCRIPTION
[0031] The technical solutions of the present application will be described clearly and completely below by combining the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0032] The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second" and the like are generally of a kind.
[0033] Figure 1 is one of the flowcharts of the training method of the sound effect adjustment model provided by the present application, as shown in Figure 1 The method comprises steps 110 and 120.
[0034] Step 110: input the sound pressure data into the environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model; the sample environmental space information comprises a space size and a space type; Step 120: training an initial sound effect adjustment model based on the sample environmental space information to obtain the sound effect adjustment model.
[0035] Specifically, considering that the echo effect of different environmental spaces will vary due to various factors. The space size is one of the key factors affecting the echo, and a larger space will generally produce more obvious echo and longer reverberation time. In addition, the space type (such as a classroom, a conference room, a video room, etc.) will also have a significant impact on the echo effect. Therefore, these factors work together to cause differences in the sound effects heard by the human ear in different environmental spaces. At the same time, sound pressure data is also one of the important factors affecting environmental space information.
[0036] Based on the above considerations, the sound pressure data can be input into the environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model. In this way, the environmental space information prediction model can more accurately predict and describe the acoustic characteristics of different environmental spaces. The sample environmental space information includes space size and space type. The space size refers to the physical dimensions of the current space, which can be measured by three dimensions of length, width and height. In sound effect adjustment, the space size has a significant impact on the propagation, reflection and reverberation of sound. For example, the length of the space determines the distance of sound propagation in the space, the width of the space affects the lateral propagation and reflection of sound, and the height of the space affects the vertical propagation and reflection of sound.
[0037] The space type refers to the category according to the use and function of the space. The space type includes conference room, audio-visual, classroom and general category, etc. The embodiments of the present application do not make specific limitation on this.
[0038] Here, the sound pressure data refers to the sound pressure value in the environmental space. The sound pressure data can be obtained based on a sound wave attenuation model. For example, sample sound power data of a sound source corresponding to sample microphone (MIC) array data is input into a sound wave attenuation model to obtain sound pressure data output by the sound wave attenuation model.
[0039] In an optional embodiment, the sample environmental space information can also be obtained by field measurement, sensor network and user input. Field measurement is the most direct way. By using a tape measure, a laser range finder and other tools, the length, width and height of the space are accurately measured, which is suitable for small indoor spaces. The sensor network estimates the space size by deploying Wi-Fi, Bluetooth and other sensors in the space combined with signal strength and positioning algorithms. In addition, user input and annotation are also an effective way. Users can manually annotate the space size and space type.
[0040] Here, the initial sound effect adjustment model can be a multi-layer convolutional neural network (CNN) with a cascade structure, or a deep neural network (DNN), a long short-term memory network (LSTM), or a combination structure of CNN and DNN, etc. The embodiments of the present application do not make specific limitation on this.
[0041] Here, the parameters of the initial sound effect adjustment model can be pre-set or randomly generated. The embodiments of the present application do not make specific limitation on this.
[0042] After obtaining the sample environment space information, the initial sound effect adjustment model can be trained based on the sample environment space information, and the initial sound effect adjustment model after training is completed is taken as the sound effect adjustment model.
[0043] It should be noted that the sound effect parameter is an important parameter at the large screen end, which determines the sound effect of the large screen loudspeaker. The sound effect adjustment involves multiple aspects of audio processing and optimization, and the sound effect adjustment can include gain adjustment, equalizer adjustment, reverberation adjustment, noise suppression, etc., and the sound effect adjustment in the embodiment of the present application mainly lies in the gain range of each frequency band. For example, the gain ranges of the low frequency parameter, the medium frequency parameter and the high frequency parameter of the sound effect can be adjusted based on the sample environment space information, wherein the low frequency parameter can be 120hz, the medium frequency parameter can be 500hz, 1.5khz, and the high frequency parameter can be 5khz, 10khz. For each frequency band, the gain range can be -12dB to +12dB.
[0044] Here, the low frequency parameter, the medium frequency parameter and the high frequency parameter in the embodiment of the present application, and the gain range of each frequency band can be adjusted according to actual conditions, which is not specifically limited in the embodiment of the present application.
[0045] It can be understood that adjusting the gain ranges of the low frequency parameter, the medium frequency parameter and the high frequency parameter of the sound effect meets the needs of professional users for gain adjustment of each frequency band, and allows professional users to have better product experience.
[0046] The method provided by the embodiment of the present application, on the one hand, inputs the sound pressure data into the environment space information prediction model to obtain the sample environment space information, and then trains the initial sound effect adjustment model based on the sample environment space information to obtain the sound effect adjustment model. The sound effect adjustment model can learn the best sound effect parameters in different environments, thereby improving the accuracy of sound effect adjustment; on the other hand, compared with the prior art in which the user needs to manually select the audio parameter mode matched with the current environment according to his own judgment to achieve the best sound effect performance, the sound effect adjustment model in this method can automatically adjust the sound effect parameters according to the change of the environment, without the need for the user to manually select the audio parameter mode, thereby improving the user experience.
[0047] Based on the above embodiment, Figure 2 is a flowchart of the training method of the sound effect adjustment model provided by the present application, as Figure 2 shown, the method comprises: Step 210, obtaining an initial environment space information prediction model, sample sound pressure data, and labeled environment space information of the sample sound pressure data; Step 220, inputting the sample sound pressure data into the initial environment space information prediction model to obtain the predicted environment space information output by the initial environment space information prediction model; In step 230, the initial environment space information prediction model is iterated in parameters based on the predicted environment space information and the labeled environment space information, to obtain the environment space information prediction model. In step 240, sample environment space information and an initial sound effect adjustment model are obtained; the sample environment space information includes a space size and a space type; the sample environment space information is obtained based on the environment space information prediction model. In step 250, the initial sound effect adjustment model is trained based on the sample environment space information, to obtain the sound effect adjustment model.
[0048] Specifically, the sample environment space information is obtained based on the environment space information prediction model, that is, the sample sound pressure data can be input into the environment space information prediction model to obtain the sample environment space information output by the environment space information prediction model.
[0049] In order to better obtain the environment space information prediction model and apply the environment space information prediction model to determine the environment space information, the environment space information prediction model can be trained based on the following steps: First, an initial environment space information prediction model, sample sound pressure data, and labeled environment space information of the sample sound pressure data are obtained. Here, the parameters of the initial environment space information prediction model can be pre-set or randomly generated, and the present embodiment is not limited in this regard.
[0050] Here, the sample sound pressure data can be obtained based on a sound wave attenuation model, can be extracted from voice data, or can be directly collected by a special acoustic measurement device. The voice data can be obtained by a sound pickup device, which can be a smart phone, a tablet computer, or a smart electric appliance such as a sound box, a television, and an air conditioner. The sound pickup device can also amplify and denoise the voice data after obtaining the voice data through a microphone array, and the present embodiment is not limited in this regard.
[0051] After obtaining the sample sound pressure data, the sample sound pressure data can be input into the initial environment space information prediction model to obtain predicted environment space information output by the initial environment space information prediction model. The predicted environment space information includes a predicted space size and a predicted space type. The initial environment space information prediction model can be a Transformer model, a multi-layer convolutional neural network with a cascade structure, a deep neural network, a long short-term memory network, or the like, and the present embodiment is not limited in this regard.
[0052] After obtaining the predicted environmental space information, a prediction loss can be determined based on the predicted environmental space information and the labeled environmental space information, and a parameter iteration is performed on the initial environmental space information prediction model based on the prediction loss, and the initial environmental space information prediction model after the parameter iteration is completed is taken as the environmental space information prediction model.
[0053] It can be understood that the greater the difference between the predicted environmental space information and the labeled environmental space information, the greater the prediction loss; the smaller the difference between the predicted environmental space information and the labeled environmental space information, the smaller the prediction loss.
[0054] After obtaining the sample environmental space information, an initial sound effect adjustment model can be trained based on the sample environmental space information, and the initial sound effect adjustment model after the training is completed is taken as the sound effect adjustment model.
[0055] Here, the initial sound effect adjustment model can be a multi-layer convolutional neural network with a cascade structure, or can be a deep neural network, a long short-term memory network, or a combination structure of CNN and DNN, and the present embodiment is not limited in this regard.
[0056] Here, the parameters of the initial sound effect adjustment model can be pre-set or randomly generated, and the present embodiment is not limited in this regard.
[0057] The method provided by the present embodiment can learn the mapping relationship between the input sound pressure data and the environmental space information based on the predicted environmental space information and the labeled environmental space information, and the supervised learning method can significantly improve the prediction accuracy of the environmental space information prediction model. On the other hand, the sample environmental space information is obtained based on the environmental space information prediction model, so that the sound effect adjustment model can automatically adjust the sound effect parameters based on better input environmental space information, thereby improving the accuracy and automation of sound effect adjustment and further enhancing the user experience.
[0058] Based on the above embodiment, the labeled environmental space information includes a labeled space size and a labeled space type; and the predicted environmental space information includes a predicted space size and a predicted space type. Figure 3 is a flowchart of step 230 in the sound effect adjustment method provided by the present application, as shown in Figure 3 Step 230 includes: Step 231, determining a first loss based on the labeled space size and the predicted space size; Step 232, determining a second loss based on the labeled space type and the predicted space type; Step 233: Determine a predicted loss based on the first loss and the second loss, and perform parameter iteration on the initial environmental spatial information prediction model based on the predicted loss.
[0059] Specifically, label space information includes the label space size and label space type, while predicted space information includes the predicted space size and predicted space type. Label space information refers to the known, annotated features of the environment space during training. Label space size refers to the known physical dimensions of the environment space, typically expressed in length, width, and height. Label space type refers to the known functional category of the environment space.
[0060] Here, the predicted ambient spatial information refers to the ambient spatial characteristics predicted by the ambient spatial information prediction model based on the input sample sound pressure data. The predicted spatial size refers to the spatial size predicted by the ambient spatial information prediction model based on the input sample sound pressure data. The predicted spatial type refers to the spatial type predicted by the ambient spatial information prediction model based on the input sample sound pressure data.
[0061] After predicting the space size and the space type, the first loss can be determined based on the label space size and the predicted space size, and the second loss can be determined based on the label space type and the predicted space type. Finally, the prediction loss is determined based on the first loss and the second loss, and the parameters of the initial environmental space information prediction model are iterated based on the prediction loss. The initial environmental space information prediction model after the parameter iteration is completed is used as the environmental space information prediction model.
[0062] It can be understood that the larger the difference between the label space size and the prediction space size, the larger the first loss; the smaller the difference between the label space size and the prediction space size, the smaller the first loss.
[0063] It can be understood that the greater the difference between the label space type and the prediction space type, the greater the second loss; and the smaller the difference between the label space type and the prediction space type, the smaller the second loss.
[0064] Here, the predicted loss may be determined based on the sum of the first loss and the second loss, or the weighted sum of the first loss and the second loss, which is not specifically limited in this embodiment of the present invention.
[0065] Here, the first loss can use a mean squared error loss function (MSE), can also use a mean absolute error loss function (MAE), can also use a root mean squared error loss function (RMSE), etc., and the second loss can use a cross entropy loss function (Cross Entropy Loss Function), can also use a KL divergence loss, and the embodiments of the present application do not make specific limitations on this.
[0066] The method provided by the embodiments of the present application trains the environment space information prediction model based on the predicted space size and the label space size, and the predicted space type and the label space type, the model can learn the mapping relationship between the input sound pressure data and the space size and the space type, and the supervised learning method can significantly improve the prediction accuracy of the environment space information prediction model. On the other hand, the sample environment space information is obtained based on the environment space information prediction model, so that the environment space information prediction model can automatically predict the space size and the space type, and the subsequent sound effect adjustment model adjusts the sound effect parameters accordingly, thereby improving the accuracy and automation of sound effect adjustment, and further enhancing the user experience.
[0067] Based on the above embodiments, the sample sound pressure data acquisition step comprises: Step 310, obtaining sample sound power data of a sound source corresponding to sample microphone array data; Step 320, inputting the sample sound power data into the sound wave attenuation model to obtain the sample sound pressure data output by the sound wave attenuation model.
[0068] Specifically, the sound effect adjustment model further comprises a sound wave attenuation model, which is used to predict the attenuation characteristics of sound waves in different environments.
[0069] Correspondingly, the sample sound pressure data acquisition step is as follows: First, sample sound power data of a sound source corresponding to sample microphone array data is obtained, wherein the sample microphone array data is obtained by playing audio of a specific frequency range through a loudspeaker and collecting the audio by a microphone array.
[0070] Here, the number of loudspeakers can be one or multiple, and the number of loudspeakers is different for different space environments. The number of microphone arrays can be 6, 8, 1, etc., and the embodiments of the present application do not make specific limitations on this. The specific frequency range refers to 120-10khz, which can also be 20 Hz-20 kHz, etc., and the embodiments of the present application do not make specific limitations on this.
[0071] Here, the sample sound power data of the sound source corresponding to the sample microphone array data is to locate the sound source position using a signal processing algorithm for the microphone array data and reconstruct the sound pressure distribution of the sound source surface, and estimate the sound intensity distribution through the sound pressure gradient of the sound pressure distribution, and then the reconstructed sound intensity distribution is numerically integrated on a virtual closed curved surface (such as a hemispherical surface or a measurement grid) to obtain the sound power data. The signal processing algorithm can be beamforming, or acoustic holography or inverse method, etc., and the embodiments of the present application do not make specific limitations.
[0072] After obtaining the sample sound power data, the sample sound power data can be input into the sound wave attenuation model to obtain the sample sound pressure data output by the sound wave attenuation model, wherein the formula of the sound wave attenuation model is as follows: wherein, represents the sample sound pressure data, that is, represents the sound pressure at a distance of from the sound source (horn), represents the sample sound power data of the sound source corresponding to the sample microphone array data, represents the center position of the microphone array, represents the air absorption coefficient, represents the reflection correction term.
[0073] It should be noted that in the case of an 8-microphone array, the in the sound wave attenuation model is the center position of the 4th and 5th microphones of the 8-microphone array, in the case of a 6-microphone array, the in the sound wave attenuation model is the center position of the 3rd and 4th microphones, and in the case of a microphone, the in the sound wave attenuation model is the position of the microphone.
[0074] It can be understood that in the embodiments of the present application, the sound wave attenuation model is mapped into a trainable neural network model, and since different environments have different reflection correction terms and air absorption coefficients, by taking the reflection correction term and the air absorption coefficient as trainable parameters, by training the neural network model to learn these parameters, the sound wave attenuation model obtained by training can adapt to different environments, so as to output more accurate sample sound pressure data based on the sound wave attenuation model.
[0075] Here, the sound wave attenuation model can be mapped into a fully connected neural network (FCNN), or a convolutional neural network, or an LSTM model, etc., and the present embodiment is not limited in this regard.
[0076] Here, the sample microphone array data acquisition step is as follows: Figure 4 is a flowchart of the third of the training methods of the sound effect adjustment model provided by the present application, as shown in Figure 4 The SOC (System on Chip) initializes the related resources, controls the loudspeaker, and plays the audio linear sound in the specified range (120-10khz). Specifically, the related resources include the audio DSP (Digital Signal Processor), the I2S (Inter-IC Sound) bus, the acoustic configuration file, and the PCIe DMA (Peripheral Component Interconnect Express Direct Memory Access) transmission channel established with the NPU (Neural Processing Unit), etc., and the present embodiment is not limited in this regard.
[0077] Here, the audio DSP coprocessor is usually integrated inside the SOC and is specifically used for processing audio signals. The audio DSP coprocessor is responsible for executing audio processing tasks such as audio codec, filtering, echo cancellation, etc.
[0078] The SOC starts the audio DSP coprocessor, i.e., initializes its internal registers and configurations, so that it enters the working state. The I2S bus controller is also integrated inside the SOC and is used for processing the digital transmission of audio data. The I2S bus is a standard audio digital interface used for transmitting audio data between audio devices. By initializing the I2S bus, it can be ensured that the audio data can be correctly transmitted to the external audio device (such as the loudspeaker). When the system starts, the acoustic configuration file is loaded into the audio DSP coprocessor for initializing the audio system. These configuration files define the sampling rate, quantization bits, audio processing algorithm, etc. of the audio system, ensuring that the audio system can work correctly when starting, for example, the sampling rate can be set to 48khz, 24bit quantization, the sampling rate can also be set to 96khz, 24bit quantization, etc.
[0079] Here, the SOC establishes a PCIe DMA transmission channel with the NPU, which refers to establishing a high-speed data transmission channel through a PCIe bus, for transmitting audio data from an audio processing unit (such as an audio DSP coprocessor) to the NPU for AI processing. Here, the transmission bandwidth of the PCIe DMA transmission channel between the SOC and the NPU is ≥ 5 Gbps.
[0080] Then, the SOC controls the microphone array to cancel the echo cancellation, and then the microphone array captures the stereo sound of the loudspeaker to obtain sample microphone array data.
[0081] In a preferred embodiment, the microphone array is configured as a symmetrical 8-microphone array, and the radius of the 8-microphone array can be set to 35 mm, and the 8-microphone array is uniformly distributed.
[0082] It should be noted that in the audio processing system, the AEC (Acoustic Echo Cancellation) module is usually configured as a standard. However, this module will interfere with the microphone's collection of the sound played by the large screen's own loudspeaker. Therefore, it is necessary to close the AEC module, that is, to stop performing the echo cancellation function. This is because when the computer plays audio, if the computer's sound recorder is used for recording at the same time, the played audio content usually cannot be effectively recorded.
[0083] In the embodiment of the application, the microphone array is enabled, and the synchronization sampling accuracy is ensured to reach PPS (Pulse Per Second) accuracy ± 10 nanoseconds. In addition, the system is configured with an 80 decibel dynamic range automatic gain control (Automatic Gain Control, AGC), while ensuring that the signal-to-noise ratio (Signal-to-Noise Ratio, SNR) is greater than 90 decibels.
[0084] Based on the above steps, the microphone array can capture sample microphone array data, and then the microphone array sends the sample microphone array data into the NPU for training of the sound effect adjustment model.
[0085] Specifically, the NPU obtains sample sound power data of a sound source corresponding to sample microphone array data, and inputs the sample sound power data into a sound wave attenuation model to obtain sample sound pressure data output by the sound wave attenuation model. Then, the NPU obtains an initial environmental space information prediction model, the sample sound pressure data, and label environmental space information of the sample sound pressure data, inputs the sample sound pressure data into the initial environmental space information prediction model to obtain predicted environmental space information output by the initial environmental space information prediction model. The label environmental space information includes a label space size and a label space type, and the predicted environmental space information includes a predicted space size and a predicted space type. Accordingly, a first loss can be determined based on the label space size and the predicted space size, a second loss can be determined based on the label space type and the predicted space type, and finally, a prediction loss can be determined based on the first loss and the second loss, and the initial environmental space information prediction model can be iterated based on the prediction loss to obtain an environmental space information prediction model. Further, after obtaining the environmental space information prediction model, the NPU can input the sample sound pressure data into the environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model. After obtaining the sample environmental space information, the NPU can return the sample environmental space information to the SOC. Finally, the SOC can train an initial sound effect adjustment model based on the sample environmental space information to obtain a sound effect adjustment model, that is, the sound effect adjustment model can set the AQ parameter by applying the sample environmental space information.
[0086] Based on the above embodiment, the training step of the sound wave attenuation model comprises: Step 410, obtaining label sound pressure data corresponding to the sample sound power data and an initial sound wave attenuation model; Step 420, inputting the sample sound power data into the initial sound wave attenuation model to obtain predicted sound pressure data output by the initial sound wave attenuation model; Step 430, iteratively updating parameters of the initial sound wave attenuation model based on the label sound pressure data and the predicted sound pressure data to obtain the sound wave attenuation model.
[0087] Specifically, first, label sound pressure data corresponding to sample sound power data and an initial sound wave attenuation model are obtained. Here, the parameters of the initial sound wave attenuation model can be pre-set or randomly generated, and the present embodiment does not make specific limitations thereon. The label sound pressure data refers to the sound pressure value directly obtained by an accurate acoustic measurement device in an actual measurement or experimental environment. The label sound pressure data is regarded as a "true" or "reference" sound pressure value for verifying and calibrating the predicted sound pressure data of the sound wave attenuation model.
[0088] After obtaining the sample sound power data, the sample sound power data can be input into the initial sound wave attenuation model to obtain predicted sound pressure data output by the initial sound wave attenuation model. Here, the predicted sound pressure data is used to reflect the sound pressure level after attenuation in the process of sound wave propagation from the loudspeaker to the microphone array.
[0089] After obtaining the predicted sound pressure data based on the initial sound wave attenuation model, the initial sound wave attenuation model can be iterated in parameters based on the label sound pressure data and the predicted sound pressure data, and the initial sound wave attenuation model after completing the parameter iteration is taken as the sound wave attenuation model.
[0090] It can be understood that the greater the difference between the label sound pressure data and the predicted sound pressure data, the greater the target loss; the smaller the difference between the label sound pressure data and the predicted sound pressure data, the smaller the target loss.
[0091] The method provided by the embodiment of the application obtains the label sound pressure data corresponding to the sample sound power data and the initial sound wave attenuation model, then inputs the sample sound power data into the initial sound wave attenuation model to obtain the predicted sound pressure data, and finally iterates the initial sound wave attenuation model in parameters based on the label sound pressure data and the predicted sound pressure data to obtain the optimized sound wave attenuation model. This process uses the label sound pressure data as a reference standard, can intuitively evaluate the accuracy of the prediction of the sound wave attenuation model, and thus adjust the parameters of the sound wave attenuation model in a targeted manner to make it closer to the actual sound field situation. Meanwhile, the parameter iteration process can continuously optimize the sound wave attenuation model, so that the sound wave attenuation model can maintain high prediction accuracy in different scenarios.
[0092] Based on the above embodiment, step 430 comprises: Step 431, determining a similarity loss based on the difference between the label sound pressure data and the predicted sound pressure data; Step 432, determining a target loss based on the similarity loss, and iterating the initial sound wave attenuation model in parameters based on the target loss.
[0093] Specifically, the similarity loss can be determined based on the difference between the label sound pressure data and the predicted sound pressure data. The similarity loss is used to reflect the deviation between the label sound pressure data and the predicted sound pressure data.
[0094] It can be understood that the greater the difference between the label sound pressure data and the predicted sound pressure data, the greater the similarity loss; the smaller the difference between the label sound pressure data and the predicted sound pressure data, the smaller the similarity loss.
[0095] After obtaining the similarity loss, the target loss can be determined based on the similarity loss, and the initial sound wave attenuation model is iterated in parameters based on the target loss, and the initial sound wave attenuation model after completing the parameter iteration is taken as the sound wave attenuation model.
[0096] Based on the above embodiments, the determining, in step 432, the target loss based on the similarity loss comprises: Step 4321, obtaining a positive sample pair and a negative sample pair; the positive sample pair comprises a first sound power data pair of a sound source corresponding to first microphone array data, the first microphone array data being collected based on a same sound source or a similar sound source; the negative sample pair comprises a second sound power data pair of a sound source corresponding to second microphone array data, the second microphone array data being collected based on a different sound source or a dissimilar sound source; Step 4322, inputting the first sound power data pair into the initial sound wave attenuation model to obtain first predicted sound pressure data and second predicted sound pressure data output by the initial sound wave attenuation model; Step 4323, inputting the second sound power data pair into the initial sound wave attenuation model to obtain third predicted sound pressure data and fourth predicted sound pressure data output by the initial sound wave attenuation model; Step 4324, determining a comparison loss based on the first predicted sound pressure data and the second predicted sound pressure data, and the third predicted sound pressure data and the fourth predicted sound pressure data; Step 4325, determining the target loss based on the similarity loss and the comparison loss.
[0097] Specifically, first, a positive sample pair and a negative sample pair are obtained, wherein the positive sample pair comprises a first sound power data pair of a sound source corresponding to first microphone array data, the first microphone array data being collected based on a same sound source or a similar sound source; the negative sample pair comprises a second sound power data pair of a sound source corresponding to second microphone array data, the second microphone array data being collected based on a different sound source or a dissimilar sound source.
[0098] After obtaining the first sound power data pair and the second sound power data pair, the first sound power data pair can be input into the initial sound wave attenuation model to obtain first predicted sound pressure data and second predicted sound pressure data output by the initial sound wave attenuation model. Then, the second sound power data pair is input into the initial sound wave attenuation model to obtain third predicted sound pressure data and fourth predicted sound pressure data output by the initial sound wave attenuation model. Here, the first predicted sound pressure data and the second predicted sound pressure data are used to reflect the first sound power data pair of the sound source corresponding to the first microphone array data collected based on the same sound source or the similar sound source, the sound pressure level under different conditions after processing by the initial sound wave attenuation model. The third predicted sound pressure data and the fourth predicted sound pressure data are used to reflect the second sound power data pair of the sound source corresponding to the second microphone array data collected based on the different sound source or the dissimilar sound source, the sound pressure level under different conditions after processing by the initial sound wave attenuation model.
[0099] After obtaining the first predicted sound pressure data, the second predicted sound pressure data, the third predicted sound pressure data and the fourth predicted sound pressure data, the contrast loss can be determined based on the first predicted sound pressure data and the second predicted sound pressure data, and the third predicted sound pressure data and the fourth predicted sound pressure data.
[0100] Since the goal of the contrast loss is to minimize the difference between the same class samples and maximize the difference between the different class samples, the greater the difference between the first predicted sound pressure data and the second predicted sound pressure data, the greater the contrast loss; the smaller the difference between the first predicted sound pressure data and the second predicted sound pressure data, the smaller the contrast loss.
[0101] It can be understood that the greater the difference between the third predicted sound pressure data and the fourth predicted sound pressure data, the smaller the contrast loss; the smaller the difference between the third predicted sound pressure data and the fourth predicted sound pressure data, the greater the contrast loss.
[0102] Finally, after obtaining the contrast loss, the target loss is determined based on the similarity loss and the contrast loss. Here, the target loss can be determined based on the sum of the similarity loss and the contrast loss, or the weighted sum of the similarity loss and the contrast loss.
[0103] The method provided by the embodiments of the present application improves the performance of the model in predicting the sound pressure data attenuation by using positive sample pairs and negative sample pairs in the process of optimizing the sound wave attenuation model. The positive sample pairs are composed of the first sound power data pairs of the sound sources corresponding to the first microphone array data, which are collected based on the same sound source or similar sound sources. The negative sample pairs are composed of the second sound power data pairs of the sound sources corresponding to the second microphone array data, which are collected based on different sound sources or dissimilar sound sources. By minimizing the difference between the positive sample pairs, the sound wave attenuation model can learn the common characteristics of the same sound source under different conditions, thereby significantly improving the accuracy of the sound power data attenuation prediction of the same sound source. At the same time, by maximizing the difference between the negative sample pairs, the model can more accurately distinguish the characteristics of different sound sources, thereby enhancing the distinguishing ability of the sound power data of different sound sources. On this basis, the initial sound wave attenuation model is iterated according to the target loss. After completing the parameter iteration, the optimized initial sound wave attenuation model is determined as the final sound wave attenuation model to achieve more accurate sound wave attenuation prediction.
[0104] Based on the above embodiments, step 4325 includes: Step 4325-1, determining the impulse response of the current room corresponding to the sample microphone array data, and the label impulse response corresponding to the impulse response; Step 4325-2: Determine an impulse response loss based on the impulse response and the label impulse response, and determine the target loss based on the impulse response loss, the similarity loss, and the contrast loss.
[0105] Specifically, during the process of optimizing the acoustic attenuation model, the impulse response of the current room corresponding to the sample microphone array data can be determined. An impulse response refers to the time-domain (time-amplitude) response characteristics exhibited by the system under test when receiving an impulse excitation signal. Simultaneously, a label impulse response corresponding to this impulse response is also obtained. Label impulse responses are standard impulse responses, either precisely measured or predefined, and serve as a reference benchmark.
[0106] After the impulse response and the label impulse response are obtained, the impulse response loss may be determined based on the impulse response and the label impulse response, and the target loss may be determined based on the impulse response loss, the similarity loss, and the contrast loss.
[0107] It's understandable that the greater the difference between the current room's impulse response corresponding to the sample microphone array data and the labeled impulse response, the greater the impulse response loss; the smaller the difference between the current room's impulse response corresponding to the sample microphone array data and the labeled impulse response, the smaller the impulse response loss. By comparing the current room's impulse response with the labeled impulse response, we can evaluate the model's accuracy in predicting the impulse response and, in turn, optimize the performance of the sound wave attenuation model.
[0108] Here, the impulse response can be expressed as The label impulse response can be expressed as Therefore, the impulse response loss is expressed as Accordingly, based on impulse response loss, similarity loss and contrast loss, the formula for determining the target loss is as follows: in, represents the target loss, represents the label impulse response, represents the impulse response, represents the similarity loss, represents the contrast loss, 、 and Both represent weight coefficients.
[0109] Based on the above embodiment, determining the impulse response loss based on the impulse response and the tag impulse response in step 4325-2 includes: Step 4325-21, determining a phase loss based on the first phase corresponding to the impulse response and the second phase corresponding to the tag impulse response; Step 4325-22, determining an amplitude loss based on the amplitude corresponding to the impulse response and the tag amplitude corresponding to the tag impulse response; Step 4325-23, determining the impulse response loss based on the phase loss and the amplitude loss.
[0110] Specifically, the impulse response is a time-domain (time-amplitude) response characteristic of a measured system when a pulse excitation signal is input. The impulse response is a signal varying with time in the time domain, and the amplitude information describes the intensity of the response, while the phase information is implied in the shape and time delay of the response. That is, the impulse response includes phase information and amplitude information.
[0111] Correspondingly, the phase loss can be determined based on the first phase corresponding to the impulse response and the second phase corresponding to the tag impulse response, the amplitude loss can be determined based on the amplitude corresponding to the impulse response and the tag amplitude corresponding to the tag impulse response, and finally, the impulse response loss can be determined based on the phase loss and the amplitude loss. The phase is the position of the pulse signal in the time axis in the impulse response, and is usually expressed in angle (degree or radian). The phase describes the position of the starting point of the signal waveform relative to the reference point. For a sinusoidal signal, the phase can be intuitively understood as the horizontal movement of the waveform.
[0112] It can be understood that the greater the difference between the first phase corresponding to the impulse response and the second phase corresponding to the tag impulse response, the greater the phase loss; the smaller the difference between the first phase corresponding to the impulse response and the second phase corresponding to the tag impulse response, the smaller the phase loss.
[0113] It can be understood that the greater the difference between the amplitude corresponding to the impulse response and the tag amplitude corresponding to the tag impulse response, the greater the amplitude loss; the smaller the difference between the amplitude corresponding to the impulse response and the tag amplitude corresponding to the tag impulse response, the smaller the amplitude loss.
[0114] Here, the impulse response loss can be determined based on the sum of the phase loss and the amplitude loss, or the weighted sum of the phase loss and the amplitude loss.
[0115] Based on any of the above embodiments, the present embodiment provides a method for training an audio effect adjustment model, Figure 5 is a fourth flowchart of the method for training the audio effect adjustment model provided by the present application, as shown in Figure 5 The training steps of the audio effect adjustment model are as follows: At step 510, the labeled sound pressure data corresponding to the sample sound power data is obtained, and an initial sound wave attenuation model is obtained. The sample sound power data is input into the initial sound wave attenuation model to obtain predicted sound pressure data output by the initial sound wave attenuation model. A similarity loss is determined based on a difference between the labeled sound pressure data and the predicted sound pressure data.
[0116] At step 520, a positive sample pair and a negative sample pair are obtained. The positive sample pair includes a first sound power data pair of a sound source corresponding to a first microphone array data, and the first microphone array data is collected based on a same sound source or a similar sound source. The negative sample pair includes a second sound power data pair of a sound source corresponding to a second microphone array data, and the second microphone array data is collected based on a different sound source or a dissimilar sound source.
[0117] At step 530, the first sound power data pair is input into the initial sound wave attenuation model to obtain first predicted sound pressure data and second predicted sound pressure data output by the initial sound wave attenuation model. The second sound power data pair is input into the initial sound wave attenuation model to obtain third predicted sound pressure data and fourth predicted sound pressure data output by the initial sound wave attenuation model. A contrast loss is determined based on the first predicted sound pressure data and the second predicted sound pressure data, and the third predicted sound pressure data and the fourth predicted sound pressure data.
[0118] At step 540, an impulse response of a current room corresponding to sample microphone array data is determined, and a labeled impulse response corresponding to the impulse response is determined. A phase loss is determined based on a first phase corresponding to the impulse response and a second phase corresponding to the labeled impulse response. An amplitude loss is determined based on an amplitude corresponding to the impulse response and a labeled amplitude corresponding to the labeled impulse response.
[0119] At step 550, an impulse response loss is determined based on the phase loss and the amplitude loss. A target loss is determined based on the impulse response loss, the similarity loss, and the contrast loss. A parameter iteration is performed on the initial sound wave attenuation model based on the target loss to obtain a sound wave attenuation model.
[0120] At step 560, sample sound power data of a sound source corresponding to sample microphone array data is obtained. The sample sound power data is input into the sound wave attenuation model to obtain sample sound pressure data output by the sound wave attenuation model.
[0121] At step 570, an initial environmental space information prediction model, sample sound pressure data, and labeled environmental space information of the sample sound pressure data are obtained. The sample sound pressure data is input into the initial environmental space information prediction model to obtain predicted environmental space information output by the initial environmental space information prediction model. The labeled environmental space information includes a labeled space size and a labeled space type. The predicted environmental space information includes a predicted space size and a predicted space type.
[0122] At step 580, a first loss is determined based on the label space size and the prediction space size, a second loss is determined based on the label space type and the prediction space type, a prediction loss is determined based on the first loss and the second loss, and parameter iteration is performed on the initial environmental space information prediction model based on the prediction loss to obtain an environmental space information prediction model. The sample sound pressure data is input into the environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model, and the initial sound effect adjustment model is trained based on the sample environmental space information to obtain a sound effect adjustment model.
[0123] Here, the prediction loss and the target loss can use a mean square error loss function, can use a mean absolute error loss function, can use a root mean square error loss function, a cross-entropy loss function, or a KL divergence loss, etc., and the embodiments of the present application do not make specific limitations.
[0124] After the sound effect adjustment model is trained, the sound effect adjustment model can be applied to perform sound effect adjustment, based on any of the above embodiments, the present application provides a sound effect adjustment method, Figure 6 is a flowchart of the sound effect adjustment method provided by the present application, as Figure 6 shown, the method comprises: At step 610, environmental space information corresponding to the sound effect to be adjusted is obtained. At step 620, the environmental space information is input into the sound effect adjustment model to obtain sound effect adjustment parameters output by the sound effect adjustment model, and the sound effect adjustment parameters are applied to perform sound effect adjustment on the sound effect to be adjusted. The sound effect adjustment model is obtained based on the training method of the sound effect adjustment model.
[0125] Specifically, first, environmental space information corresponding to the sound effect to be adjusted is obtained, which refers to environmental space information that needs to be adjusted in the future. The environmental space information can be determined based on microphone array data, which is obtained by playing audio of a specific frequency range through a loudspeaker and simultaneously collecting the audio by a microphone array. Here, the number of loudspeakers can be one or multiple, and the number of loudspeakers varies for different space environments. The number of microphone arrays can be 6, 8, or 1, etc., and the embodiments of the present application do not make specific limitations. The specific frequency range refers to 120-10khz, or 20 Hz-20 kHz, etc., and the embodiments of the present application do not make specific limitations.
[0126] After obtaining the microphone array data, the microphone array data can be input to the NPU for AI algorithm calculation, that is, the sound power data of the sound source corresponding to the microphone array data is input to the sound wave attenuation model to obtain sound pressure data, and then the sound pressure data is input to the environmental space information prediction model to obtain environmental space information, and the environmental space information is input to the sound effect adjustment model to obtain the sound effect adjustment parameters output by the sound effect adjustment model, and the sound effect adjustment parameters are applied for sound effect adjustment.
[0127] The training steps of the sound effect adjustment model are as follows: First, obtain the label sound pressure data corresponding to the sample sound power data, and an initial sound wave attenuation model. The sample sound power data is input to the initial sound wave attenuation model to obtain the predicted sound pressure data output by the initial sound wave attenuation model. Based on the difference between the label sound pressure data and the predicted sound pressure data, the similarity loss is determined.
[0128] Second, obtain a positive sample pair and a negative sample pair. The positive sample pair includes a first sound power data pair of a sound source corresponding to a first microphone array data, and the first microphone array data is collected based on the same sound source or similar sound source. The negative sample pair includes a second sound power data pair of a sound source corresponding to a second microphone array data, and the second microphone array data is collected based on different sound sources or dissimilar sound sources. Third, input the first sound power data pair to the initial sound wave attenuation model to obtain the first predicted sound pressure data and the second predicted sound pressure data output by the initial sound wave attenuation model. Input the second sound power data pair to the initial sound wave attenuation model to obtain the third predicted sound pressure data and the fourth predicted sound pressure data output by the initial sound wave attenuation model. Based on the first predicted sound pressure data and the second predicted sound pressure data, and the third predicted sound pressure data and the fourth predicted sound pressure data, the contrast loss is determined. Fourth, determine the impulse response of the current room corresponding to the sample microphone array data, and the label impulse response corresponding to the impulse response. Based on the first phase corresponding to the impulse response and the second phase corresponding to the label impulse response, the phase loss is determined, and based on the amplitude corresponding to the impulse response and the label amplitude corresponding to the label impulse response, the amplitude loss is determined.
[0129] Fifth, based on the phase loss and the amplitude loss, the impulse response loss is determined, based on the impulse response loss, the similarity loss and the contrast loss, the target loss is determined, and based on the target loss, the parameter iteration of the initial sound wave attenuation model is performed to obtain the sound wave attenuation model.
[0130] In the sixth step, sample sound power data of a sound source corresponding to the sample microphone array data is obtained, and the sample sound power data is input into the sound wave attenuation model to obtain sample sound pressure data output by the sound wave attenuation model. Here, the sample sound power data of the sound source corresponding to the sample microphone array data is obtained by using a signal processing algorithm to locate a sound source position and reconstruct a sound pressure distribution of a sound source surface for the microphone array data, estimating a sound intensity distribution through a sound pressure gradient of the sound pressure distribution, and performing numerical integration on the reconstructed sound intensity distribution on a virtual closed curved surface (such as a hemispherical surface or a measurement grid) to obtain sound power data. The signal processing algorithm can be beamforming, acoustic holography, or an inverse method, and embodiments of the present application do not make specific limitations thereto.
[0131] In the seventh step, initial environment space information prediction model, sample sound pressure data, and label environment space information of the sample sound pressure data are obtained, the sample sound pressure data is input into the initial environment space information prediction model to obtain predicted environment space information output by the initial environment space information prediction model, the label environment space information includes a label space size and a label space type, and the predicted environment space information includes a predicted space size and a predicted space type. In the eighth step, a first loss is determined based on the label space size and the predicted space size, a second loss is determined based on the label space type and the predicted space type, a prediction loss is determined based on the first loss and the second loss, and parameter iteration is performed on the initial environment space information prediction model based on the prediction loss to obtain an environment space information prediction model.
[0132] In the ninth step, the sample sound pressure data is input into the environment space information prediction model to obtain sample environment space information output by the environment space information prediction model, and an initial sound effect adjustment model is trained based on the sample environment space information to obtain a sound effect adjustment model.
[0133] Here, the initial sound effect adjustment model can be a multi-layer convolutional neural network with a cascade structure, a deep neural network, a long short-term memory network, or a combination structure of a CNN and a DNN, and embodiments of the present application do not make specific limitations thereto.
[0134] Here, the parameters of the initial sound effect adjustment model can be pre-set or randomly generated, and embodiments of the present application do not make specific limitations thereto.
[0135] Here, the space size and the space type in the sample environment space information can be stored using the following pseudo code: {"room_dim": [x,y,z] ±5cm, "room_typec":classroom, "reflectivity_map": "base64 encoded 10cm precision voxel grid"}.
[0136] room_dim: [x, y, z], precision: ±5cm, room_type: classroom, reflectivity_map: base64 encoded 10cm precision voxel grid.
[0137] It is worth noting that the sound effect parameters obtained in the embodiments of the present application are equivalent to "AI sound effects", which solve the confusion of traditional user self-selection, allow users to choose AI sound effects that are more suitable for the current environment, and improve user experience.
[0138] The method provided in the embodiments of the present application obtains the environmental space information corresponding to the sound effect to be adjusted, then inputs the environmental space information into the sound effect adjustment model to obtain the sound effect adjustment parameter output by the sound effect adjustment model, and applies the sound effect adjustment parameter to the sound effect adjustment of the sound effect to be adjusted; the sound effect adjustment model is obtained based on the training method of the sound effect adjustment model described above, thereby improving the accuracy and reliability of the sound effect adjustment.
[0139] The training device of the sound effect adjustment model provided in the present application is described below, and the training device of the sound effect adjustment model described below can be correspondingly referred to the training method of the sound effect adjustment model described above.
[0140] Based on any of the above embodiments, the present application provides a training device of a sound effect adjustment model, Figure 7 is a structural schematic diagram of the training device of the sound effect adjustment model provided in the present application, as Figure 7 shown, the device comprises: The first acquisition unit 710 is configured to input the sound pressure data into the environmental space information prediction model to obtain the sample environmental space information output by the environmental space information prediction model; the sample environmental space information comprises a space size and a space type. The training unit 720 is configured to train the initial sound effect adjustment model based on the sample environmental space information to obtain the sound effect adjustment model.
[0141] The device provided by the embodiment of the present application can input the sound pressure data into the environmental space information prediction model to obtain sample environmental space information, and then train an initial sound effect adjustment model based on the sample environmental space information to obtain a sound effect adjustment model. The sound effect adjustment model can learn the optimal sound effect parameters in different environments, thereby improving the accuracy of sound effect adjustment. In the prior art, the user needs to manually select an audio parameter mode matching the current environment according to his / her own judgment of the current environment, in order to achieve the best sound effect performance. In this method, the sound effect adjustment model can automatically adjust the sound effect parameters according to the change of the environment, without the need for the user to manually select the audio parameter mode, thereby improving the user experience.
[0142] Based on any of the above embodiments, further comprising a first training unit, specifically comprising: The acquisition sample unit is configured to acquire an initial environmental space information prediction model, sample sound pressure data, and label environmental space information of the sample sound pressure data. The prediction unit is configured to input the sample sound pressure data into the initial environmental space information prediction model to obtain predicted environmental space information output by the initial environmental space information prediction model. The parameter iteration unit is configured to perform parameter iteration on the initial environmental space information prediction model based on the predicted environmental space information and the label environmental space information, to obtain the environmental space information prediction model.
[0143] Based on any of the above embodiments, the label environmental space information includes a label space size and a label space type, and the predicted environmental space information includes a predicted space size and a predicted space type. The parameter iteration unit is specifically configured to: determine a first loss based on the label space size and the predicted space size; determine a second loss based on the label space type and the predicted space type; determine a prediction loss based on the first loss and the second loss, and perform parameter iteration on the initial environmental space information prediction model based on the prediction loss.
[0144] Based on any of the above embodiments, further comprising a sample sound pressure data acquisition unit, specifically configured to: acquire sample sound power data of a sound source corresponding to sample microphone array data; input the sample sound power data into the sound wave attenuation model to obtain the sample sound pressure data output by the sound wave attenuation model.
[0145] Based on any of the above embodiments, further comprising a second training unit, specifically comprising: An initial model unit is configured to obtain label sound pressure data corresponding to the sample sound power data and an initial sound wave attenuation model; A second input unit is configured to input the sample sound power data into the initial sound wave attenuation model to obtain predicted sound pressure data output by the initial sound wave attenuation model; A parameter iteration sub-unit is configured to perform parameter iteration on the initial sound wave attenuation model based on the label sound pressure data and the predicted sound pressure data to obtain the sound wave attenuation model.
[0146] Based on any of the above embodiments, the parameter iteration sub-unit specifically includes: A similarity loss determination unit is configured to determine a similarity loss based on a difference between the label sound pressure data and the predicted sound pressure data; A target loss determination unit is configured to determine a target loss based on the similarity loss and perform parameter iteration on the initial sound wave attenuation model based on the target loss.
[0147] Based on any of the above embodiments, the target loss determination unit specifically includes: A sample pair acquisition unit is configured to acquire positive sample pairs and negative sample pairs; the positive sample pairs include first sound power data pairs of sound sources corresponding to first microphone array data, the first microphone array data being acquired based on the same sound source or similar sound sources; the negative sample pairs include second sound power data pairs of sound sources corresponding to second microphone array data, the second microphone array data being acquired based on different sound sources or dissimilar sound sources; A third input unit is configured to input the first sound power data pairs into the initial sound wave attenuation model to obtain first predicted sound pressure data and second predicted sound pressure data output by the initial sound wave attenuation model; A fourth input unit is configured to input the second sound power data pairs into the initial sound wave attenuation model to obtain third predicted sound pressure data and fourth predicted sound pressure data output by the initial sound wave attenuation model; A contrast loss determination unit is configured to determine a contrast loss based on the first predicted sound pressure data and the second predicted sound pressure data, and the third predicted sound pressure data and the fourth predicted sound pressure data; A target loss determination sub-unit is configured to determine the target loss based on the similarity loss and the contrast loss.
[0148] Based on any of the above embodiments, the target loss determination sub-unit specifically includes: A pulse response determination unit is configured to determine a pulse response of a current room corresponding to the sample microphone array data and a label pulse response corresponding to the pulse response. The comprehensive loss unit is configured to determine a pulse response loss based on the pulse response and the label pulse response, and determine the target loss based on the pulse response loss, the similarity loss, and the contrast loss.
[0149] According to any one of the above embodiments, the comprehensive loss unit is specifically configured to: determine a phase loss based on a first phase corresponding to the pulse response and a second phase corresponding to the label pulse response; determine an amplitude loss based on an amplitude corresponding to the pulse response and a label amplitude corresponding to the label pulse response; determine the pulse response loss based on the phase loss and the amplitude loss.
[0150] The sound effect adjustment device provided by the present application is described below. The sound effect adjustment device described below can be referred to in correspondence with the sound effect adjustment method described above.
[0151] According to any one of the above embodiments, the present application provides a sound effect adjustment device, Figure 8 is a structural schematic diagram of the sound effect adjustment device provided by the present application, as Figure 8 shown, the device comprises: A second acquisition unit 810 is configured to acquire environmental space information corresponding to a sound effect to be adjusted. A first input unit 820 is configured to input the environmental space information into a sound effect adjustment model to obtain sound effect adjustment parameters output by the sound effect adjustment model, and apply the sound effect adjustment parameters to perform sound effect adjustment on the sound effect to be adjusted. The sound effect adjustment model is obtained based on the training method of the sound effect adjustment model.
[0152] The device provided by the present application acquires environmental space information corresponding to a sound effect to be adjusted, inputs the environmental space information into a sound effect adjustment model to obtain sound effect adjustment parameters output by the sound effect adjustment model, and applies the sound effect adjustment parameters to perform sound effect adjustment on the sound effect to be adjusted. The sound effect adjustment model is obtained based on the training method of the sound effect adjustment model, thereby improving the accuracy and reliability of sound effect adjustment.
[0153] Figure 9 is a structural schematic diagram of the electronic device provided by the present application, as Figure 9As shown, the electronic device can include a processor 910, a communications interface 920, a memory 930, and a communications bus 940, wherein the processor 910, the communications interface 920, and the memory 930 complete mutual communication through the communications bus 940. The processor 910 can invoke a logical instruction in the memory 930 to execute a training method of an audio effect adjustment model, the method including: inputting sound pressure data to an environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model; the sample environmental space information including a space size and a space type; training an initial audio effect adjustment model based on the sample environmental space information to obtain the audio effect adjustment model.
[0154] The processor 910 can also invoke a logical instruction in the memory 930 to execute an audio effect adjustment method, the method including: obtaining environmental space information corresponding to an audio effect to be adjusted; inputting the environmental space information into an audio effect adjustment model to obtain audio effect adjustment parameters output by the audio effect adjustment model, and applying the audio effect adjustment parameters to the audio effect to be adjusted for audio effect adjustment; the audio effect adjustment model being obtained based on the training method of the audio effect adjustment model.
[0155] In addition, the logical instruction in the memory 930 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0156] In another aspect, the present application also provides a computer program product comprising a computer program, which is stored in a non-transitory computer-readable storage medium and executable by a processor to cause a computer to perform the method for training an audio effect adjustment model provided by any of the above methods, the method comprising: inputting sound pressure data to an environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model; the sample environmental space information comprising a space size and a space type; and training an initial audio effect adjustment model based on the sample environmental space information to obtain the audio effect adjustment model.
[0157] The computer program is executable by the processor to cause the computer to perform the method for adjusting an audio effect provided by any of the above methods, the method comprising: obtaining environmental space information corresponding to an audio effect to be adjusted; inputting the environmental space information to an audio effect adjustment model to obtain audio effect adjustment parameters output by the audio effect adjustment model, and applying the audio effect adjustment parameters to the audio effect to be adjusted; the audio effect adjustment model being obtained based on the method for training an audio effect adjustment model.
[0158] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is executable by a processor to implement the method for training an audio effect adjustment model provided by any of the above methods, the method comprising: inputting sound pressure data to an environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model; the sample environmental space information comprising a space size and a space type; and training an initial audio effect adjustment model based on the sample environmental space information to obtain the audio effect adjustment model.
[0159] The computer program is executable by the processor to implement the method for adjusting an audio effect provided by any of the above methods, the method comprising: obtaining environmental space information corresponding to an audio effect to be adjusted; inputting the environmental space information to an audio effect adjustment model to obtain audio effect adjustment parameters output by the audio effect adjustment model, and applying the audio effect adjustment parameters to the audio effect to be adjusted; the audio effect adjustment model being obtained based on the method for training an audio effect adjustment model.
[0160] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0161] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0162] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A training method for a sound effect adjustment model, characterized in that: The method comprises: inputting the sound pressure data into an environmental space information prediction model to obtain sample environmental space information output by the environmental space information prediction model; the sample environmental space information comprises a space size and a space type; training an initial sound effect adjustment model based on the sample environmental space information to obtain the sound effect adjustment model.
2. The method of claim 1, wherein, The training step of the environmental space information prediction model comprises: obtaining an initial environmental space information prediction model, sample sound pressure data, and label environmental space information of the sample sound pressure data; inputting the sample sound pressure data into the initial environmental space information prediction model to obtain predicted environmental space information output by the initial environmental space information prediction model; performing parameter iteration on the initial environmental space information prediction model based on the predicted environmental space information and the label environmental space information to obtain the environmental space information prediction model.
3. The method of claim 2, wherein, The label environmental space information comprises a label space size and a label space type; the predicted environmental space information comprises a predicted space size and a predicted space type; The parameter iteration on the initial environmental space information prediction model based on the predicted environmental space information and the label environmental space information comprises: determining a first loss based on the label space size and the predicted space size; determining a second loss based on the label space type and the predicted space type; determining a prediction loss based on the first loss and the second loss, and performing parameter iteration on the initial environmental space information prediction model based on the prediction loss.
4. The method of claim 2, wherein the audio effect adjustment model is trained based on a plurality of audio effect adjustment models. The obtaining step of the sample sound pressure data comprises: obtaining sample sound power data of a sound source corresponding to sample microphone array data; inputting the sample sound power data into the sound wave attenuation model to obtain the sample sound pressure data output by the sound wave attenuation model.
5. The method of claim 4, wherein, The training step of the sound wave attenuation model comprises: obtaining label sound pressure data corresponding to the sample sound power data, and an initial sound wave attenuation model; inputting the sample sound power data into the initial sound wave attenuation model to obtain predicted sound pressure data output by the initial sound wave attenuation model; performing parameter iteration on the initial sound wave attenuation model based on the label sound pressure data and the predicted sound pressure data to obtain the sound wave attenuation model.
6. The method of claim 5, wherein, The parameter iteration on the initial sound wave attenuation model based on the label sound pressure data and the predicted sound pressure data comprises: determining a similarity loss based on the difference between the label sound pressure data and the predicted sound pressure data; determining a target loss based on the similarity loss, and performing parameter iteration on the initial sound wave attenuation model based on the target loss.
7. The method of claim 6, wherein, The determination of the target loss based on the similarity loss comprises: obtaining a positive sample pair and a negative sample pair; the positive sample pair comprises a first sound power data pair of a sound source corresponding to a first microphone array data, the first microphone array data being collected based on a same sound source or a similar sound source; the negative sample pair comprises a second sound power data pair of a sound source corresponding to a second microphone array data, the second microphone array data being collected based on a different sound source or a dissimilar sound source; inputting the first sound power data pair into the initial sound wave attenuation model to obtain first predicted sound pressure data and second predicted sound pressure data output by the initial sound wave attenuation model; inputting the second sound power data pair into the initial sound wave attenuation model to obtain third predicted sound pressure data and fourth predicted sound pressure data output by the initial sound wave attenuation model; determining a contrast loss based on the first predicted sound pressure data and the second predicted sound pressure data, and the third predicted sound pressure data and the fourth predicted sound pressure data; determining the target loss based on the similarity loss and the contrast loss.
8. The method of claim 7, wherein, The determining of the target loss based on the similarity loss and the contrast loss comprises: determining an impulse response of a current room corresponding to the sample microphone array data, and a label impulse response corresponding to the impulse response; determining an impulse response loss based on the impulse response and the label impulse response, and determining the target loss based on the impulse response loss, the similarity loss and the contrast loss.
9. The method of claim 8, wherein, The determining of the impulse response loss based on the impulse response and the label impulse response comprises: determining a phase loss based on a first phase corresponding to the impulse response and a second phase corresponding to the label impulse response; determining an amplitude loss based on an amplitude corresponding to the impulse response and a label amplitude corresponding to the label impulse response; determining the impulse response loss based on the phase loss and the amplitude loss.
10. A sound effect adjustment method characterized by comprising: comprising: obtaining environment space information corresponding to a sound effect to be adjusted; inputting the environment space information into a sound effect adjustment model to obtain sound effect adjustment parameters output by the sound effect adjustment model, and applying the sound effect adjustment parameters to the sound effect to be adjusted for sound effect adjustment; the sound effect adjustment model is obtained based on the training method of the sound effect adjustment model in any one of claims 1 to 9.
11. A training device for a sound effect adjustment model, characterized in that: comprising: a first obtaining unit configured to input sound pressure data into an environment space information prediction model to obtain sample environment space information output by the environment space information prediction model; the sample environment space information comprises a space size and a space type; a training unit configured to train an initial sound effect adjustment model based on the sample environment space information to obtain the sound effect adjustment model.
12. A sound effect adjusting apparatus characterized by comprising: comprising: a second obtaining unit configured to obtain environment space information corresponding to a sound effect to be adjusted; a first input unit configured to input the environment space information into a sound effect adjustment model to obtain sound effect adjustment parameters output by the sound effect adjustment model, and apply the sound effect adjustment parameters to the sound effect to be adjusted for sound effect adjustment; The sound effect adjustment model is obtained based on the training method of the sound effect adjustment model in any one of claims 1 to 9.
13. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor implements the training method of the sound effect adjustment model in any one of claims 1 to 9 or the sound effect adjustment method in claim 10 when executing the computer program.
14. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the training method of the sound effect adjustment model in any one of claims 1 to 9 or the sound effect adjustment method in claim 10.