Data processing method and apparatus, device, and storage medium
By extracting the loudness and amplitude characteristics of audio data and combining them with preset equal loudness curves and critical frequency bands, the problem of inaccurate volume adjustment in existing technologies has been solved, achieving more precise volume control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TCL TECHNOLOGY GROUP CORPORATION
- Filing Date
- 2022-12-30
- Publication Date
- 2026-07-21
AI Technical Summary
The accuracy and adjustability of volume adjustment in existing multimedia audio are poor. Existing technologies only measure sound amplitude through sound pressure level, which cannot fully describe the sound volume.
By extracting the loudness and amplitude features of audio data, and combining them with preset equal loudness curves and critical frequency bands, the target volume value of the audio data is calculated. A dual-channel amplitude extraction model and feature extraction channel are used to calculate the volume value by integrating subjective loudness and objective amplitude features.
It improves the accuracy and comprehensiveness of volume adjustment, achieving more precise volume control.
Smart Images

Figure CN118280376B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio data technology, specifically to a data processing method, apparatus, device, and storage medium. Background Technology
[0002] Currently, with the rapid development of smart terminals, their multimedia functions are becoming increasingly rich. Multimedia audio data, such as video, music, and voice, as well as noise audio data, are all measured by volume. However, existing multimedia audio volume values are typically represented by the objective quantity sound pressure level (SPL), meaning the greater the sound amplitude, the higher the SPL. This representation only describes the sound amplitude and cannot comprehensively describe the volume level of multimedia audio, resulting in poor accuracy and adjustability of multimedia audio volume adjustment. Summary of the Invention
[0003] This application provides a data processing method, apparatus, device, and storage medium, aiming to solve the technical problem of poor volume adjustment accuracy of audio data in the prior art.
[0004] On one hand, embodiments of this application provide a data processing method, which includes the following steps:
[0005] Extract loudness and amplitude features from the audio data to be processed;
[0006] The target volume value of the audio data is determined based on the loudness feature and the amplitude feature.
[0007] In one possible implementation of this application, the extraction of loudness and amplitude features from the audio data to be processed includes:
[0008] The loudness features of the audio data are extracted based on the preset equal loudness curves and critical frequency bands.
[0009] Based on the acoustic wave feature data and audio feature data in the audio data, the amplitude features of the audio data are extracted.
[0010] In one possible implementation of this application, the step of extracting loudness features from the audio data based on a preset equal loudness curve and a critical frequency band to obtain the loudness features of the audio data includes:
[0011] Query the preset equal loudness curves to obtain the hearing threshold sound level values corresponding to each preset critical frequency band;
[0012] The audio data is frequency domain converted to obtain an audio frequency domain signal. Based on the audio frequency domain signal and each of the critical frequency bands, the critical frequency band energy of each of the critical frequency bands is calculated.
[0013] For each critical frequency band, the loudness characteristics of the audio data are determined based on the hearing threshold loudness level value and the critical frequency band energy of the critical frequency band.
[0014] In one possible implementation of this application, determining the loudness characteristics of the audio data for each critical frequency band based on the hearing threshold loudness level value and the critical frequency band energy includes:
[0015] For each critical frequency band, the hearing threshold loudness level value and the critical frequency band energy of the critical frequency band are input into a preset loudness calculation model to obtain the characteristic loudness value of each critical frequency band.
[0016] The loudness values of each critical frequency band are statistically analyzed to generate the loudness features of the audio data.
[0017] In one possible implementation of this application, the step of extracting the amplitude features of the audio data based on the sound wave feature data and audio feature data in the audio data includes:
[0018] The audio data is converted into a sound wave image, and features are extracted from the sound wave image to obtain sound wave feature data in the sound wave image.
[0019] The acoustic wave feature data is pooled to obtain the image statistical features of the acoustic wave image;
[0020] Integrate the image statistical features and the sound wave feature data to generate the first amplitude feature of the audio data;
[0021] Obtain the volume feature data of the audio data, and generate the amplitude feature of the audio data based on the volume feature data and the first amplitude feature.
[0022] In one possible implementation of this application, the step of obtaining the volume feature data of the audio data and generating the amplitude feature of the audio data based on the volume feature data and the first amplitude feature includes:
[0023] Extract the volume feature data of the audio data, and perform statistical feature extraction on the volume feature data to obtain the statistical feature data of the volume feature data;
[0024] The statistical feature data is subjected to a nonlinear transformation to obtain low-dimensional statistical data. The second amplitude feature of the audio data is generated based on the volume feature data and the low-dimensional statistical data.
[0025] The amplitude features of the audio data are generated based on the first amplitude feature and the second amplitude feature.
[0026] In one possible implementation of this application, calculating the target volume value of the audio data based on the loudness feature and the amplitude feature includes:
[0027] Extract the target loudness weight of the loudness feature and the target amplitude weight of the amplitude feature;
[0028] The loudness features are weighted according to the target loudness weight to obtain weighted loudness features;
[0029] The amplitude features are weighted according to the target amplitude weight to obtain weighted amplitude features;
[0030] The target volume value of the audio data is calculated based on the weighted loudness feature and the weighted amplitude feature.
[0031] In one possible implementation of this application, before extracting the target loudness weight of the loudness feature and the target amplitude weight of the amplitude feature, the method further includes:
[0032] Obtain the initial loudness weight and initial amplitude weight of the preset feature weight model;
[0033] The preset training loudness features are weighted according to the initial loudness weight, and the preset training amplitude features are weighted according to the initial amplitude weight, and the training volume value is obtained by summing them up.
[0034] Calculate the training loss of the training volume value and the predicted volume value, and adjust the initial loudness weight and the initial amplitude weight according to the training loss to obtain the target loudness weight of the training loudness feature and the target amplitude weight of the training amplitude feature.
[0035] In one possible implementation of this application, after determining the target volume value of the audio data based on the loudness feature and the amplitude feature, the method further includes:
[0036] In response to an audio playback request, obtain the target audio data to be played in the audio playback request;
[0037] Play the target audio data in a preset audio application according to the target volume value of the target audio data.
[0038] On the other hand, this application provides a data processing apparatus, the data processing apparatus comprising:
[0039] The audio acquisition module is configured to respond to a volume adjustment request and acquire the audio data to be processed in the volume adjustment request.
[0040] The loudness extraction module is configured to extract loudness features from the audio data to obtain the loudness features of the audio data.
[0041] The amplitude extraction module is configured to acquire volume feature data and sound wave feature data in the audio data, and extract the amplitude feature of the audio data based on the volume feature data and sound wave feature data;
[0042] The volume calculation module is configured to calculate the target volume value of the audio data based on the loudness feature and the amplitude feature.
[0043] On the other hand, this application also provides a data processing device, the data processing device comprising:
[0044] One or more processors;
[0045] Memory; and
[0046] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the data processing method.
[0047] On the other hand, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps in the data processing method.
[0048] This application, in response to a volume adjustment request, acquires the audio data to be processed within that request; extracts loudness features from the audio data to obtain its loudness characteristics; furthermore, it extracts amplitude features from the audio data based on sound wave feature data and volume feature data; after acquiring the loudness and amplitude features of the audio data, it calculates the target volume value of the audio data using both loudness and amplitude features. This achieves the calculation of the audio data's volume value from a combination of subjective loudness and objective amplitude dimensions, thereby improving the accuracy and comprehensiveness of volume adjustment. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a schematic diagram illustrating a scenario of the data processing method in an embodiment of this application.
[0051] Figure 2This is a flowchart illustrating one embodiment of the data processing method in this application.
[0052] Figure 3 This is a schematic diagram of the model structure of the amplitude feature extraction model in this application;
[0053] Figure 4 A flowchart illustrating an embodiment of the data processing method for extracting loudness features from audio data provided in this application;
[0054] Figure 5 A flowchart illustrating an embodiment of the data processing method for obtaining amplitude characteristics of audio data provided in this application;
[0055] Figure 6 A flowchart illustrating an embodiment of the data processing method provided in this application for determining the calculation weight of a target volume value;
[0056] Figure 7 This is a schematic diagram of the structure of one embodiment of the data processing apparatus provided in this application;
[0057] Figure 8 This is a schematic diagram of one embodiment of the data processing device provided in this application. Detailed Implementation
[0058] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0059] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0060] In this application, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0061] Currently, with the rapid development of smart terminals, their multimedia functions are becoming increasingly rich. Multimedia audio data, such as video, music, and voice, as well as noise audio data, are all measured by volume. However, existing multimedia audio volume values are typically represented by the objective quantity sound pressure level (SPL), meaning the greater the sound amplitude, the higher the SPL. This representation only describes the sound amplitude and cannot comprehensively describe the volume level of multimedia audio, resulting in poor accuracy and adjustability of multimedia audio volume adjustment.
[0062] Based on this, this application proposes a data processing method, apparatus, device, and computer-readable storage medium to solve the technical problem of poor volume adjustment accuracy of audio data in the prior art.
[0063] The data processing method in this embodiment of the invention is applied to a data processing device, which is disposed in a data processing equipment. The data processing equipment includes one or more processors, a memory, and one or more application programs. The one or more application programs are stored in the memory and configured to be executed by the processor to implement the data processing method. The data processing equipment can be a smart terminal, such as a mobile phone, tablet computer, smart TV, network device, and smart computer. Optionally, the data processing equipment can also be a server or a service cluster composed of multiple servers.
[0064] like Figure 1 As shown, Figure 1 This is a schematic diagram of a data processing method according to an embodiment of the present application. The data processing scenario in this embodiment includes a data processing device 100 (the data processing device 100 integrates a data processing unit), and a computer-readable storage medium corresponding to the data processing method is running in the data processing device 100 to execute the steps of the data processing method.
[0065] Understandable Figure 1The data processing device in the data processing method scenario shown, or the device included in the data processing device, does not constitute a limitation on the embodiments of the present invention. That is, the number or type of data processing device in the data processing method scenario, or the number or type of device included in each device, does not affect the overall implementation of the technical solution in the embodiments of the present invention, and can all be considered as equivalent substitutions or derivatives of the technical solutions claimed in the embodiments of the present invention.
[0066] In this embodiment of the invention, the data processing device 100 is mainly used for:
[0067] Extract loudness and amplitude features from the audio data to be processed;
[0068] The target volume value of the audio data is determined based on the loudness feature and the amplitude feature.
[0069] The data processing device 100 in this embodiment of the invention can be an independent data processing device, such as a mobile phone, tablet computer, smart TV, network device, server and smart computer, or a data processing network or data processing cluster composed of multiple data processing devices.
[0070] This application provides a data processing method, apparatus, device, and computer-readable storage medium, which will be described in detail below.
[0071] It will be understood by those skilled in the art that Figure 1 The application environment shown is only one application scenario related to the solution of this application and does not constitute a limitation on the application scenario of this application. Other application environments may include more than one application scenario. Figure 1 The number of data processing devices shown, or the data processing network connections, for example Figure 1 Only one data processing device is shown in the figure. It is understood that the scenario of the data processing method may also include one or more data processing devices, which are not limited here. The data processing device may also include a memory for storing audio data and other data.
[0072] It should be noted that, Figure 1 The schematic diagram of the data processing method shown is merely an example. The scenarios of the data processing method described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided in the embodiments of the present invention.
[0073] Based on the scenarios described above, various embodiments of the data processing method disclosed in this invention are proposed.
[0074] like Figure 2 As shown, Figure 2This is a flowchart illustrating one embodiment of the data processing method in this application. The data processing method includes the following steps 201 to 202:
[0075] 201. Extract the loudness and amplitude features of the audio data to be processed;
[0076] The data processing method in this embodiment is applied to a data processing device. The type and number of data processing devices are not specifically limited. That is, the data processing device can be one or more smart terminals or servers. In a specific embodiment, the data processing device is a smart computer.
[0077] Specifically, the data processing device is configured to process audio data to obtain the loudness and amplitude characteristics of the audio data, and to determine the target volume value of the audio data through the loudness and amplitude characteristics, thereby improving the accuracy of the volume calculation of the audio data.
[0078] Specifically, during operation, the data processing device responds to volume adjustment requests. These requests carry audio data to be processed and request the data processing device to process the audio data to obtain a target volume value. Optionally, the triggering method for this volume adjustment request is not specifically limited; that is, the request can be user-initiated. For example, a user inputs the audio data to be processed into a defect detection device, actively triggering a defect detection request and driving the data processing device to adjust the target volume value of the audio data. Optionally, the volume adjustment request can also be automatically triggered by the data processing device. For example, the data processing device may have a pre-set audio detection process that automatically triggers the volume adjustment request upon detecting the presence of audio data to be processed.
[0079] Upon receiving the volume adjustment request, the data processing device parses the request and obtains the audio data to be processed carried within it. This audio data includes multimedia data or noise data capable of adjusting the volume, such as voice data, music data, video data, and noise data.
[0080] After acquiring the audio data to be processed carried in the volume adjustment request, the data processing device also acquires the loudness characteristics and amplitude characteristics of the audio data, and calculates the target volume value of the audio data using the loudness characteristics and amplitude characteristics.
[0081] After acquiring the audio data to be processed, the data processing device extracts the loudness features of the audio data based on preset equal-loudness curves and their critical frequency bands, thereby obtaining the loudness features of the audio data. The loudness feature represents the intensity of sound energy and is a subjective scale characteristic for judging sound volume. The preset equal-loudness curves are a cluster of curves whose subjective loudness perception (loudness level) is equal, obtained through subjective measurement. These preset equal-loudness curves correspond to the loudness of different standard tones. When the loudness of a sound is the same as the loudness of a standard tone, the sound intensity level of that standard tone is the loudness level of that sound.
[0082] Specifically, the data processing device divides the frequency range of 20 Hz to 16 kHz into several critical frequency bands based on the audible frequency range of the human ear. In one specific embodiment, there are 24 critical frequency bands. The characteristic loudness values of these 24 critical frequency bands, along with the nonlinear resolution characteristics of human hearing, reflect the characteristic information of the audio signal and can characterize human subjective auditory perception. The audio processing device determines the loudness characteristics of the audio data through preset equal loudness curves and each critical frequency band. The critical frequency band is the auditory frequency band that characterizes changes in human hearing.
[0083] Specifically, the data processing equipment obtains the hearing threshold sound level value corresponding to each critical frequency band and the critical frequency band energy corresponding to each audio signal through preset equal loudness curves and preset critical frequency bands. Based on the hearing threshold sound level value and critical frequency band energy, the loudness feature of the audio signal is extracted, thereby obtaining the loudness feature of the audio signal.
[0084] After acquiring the audio data to be processed, the data processing device also extracts the amplitude features from the audio data based on the sound wave feature data and volume feature data. The amplitude feature is an objective characteristic quantity used to measure the intensity and volume of sound.
[0085] like Figure 3 As shown, Figure 3 This is a schematic diagram of the amplitude feature extraction model in this application. Specifically, the data processing device pre-configures a dual-channel amplitude extraction model for extracting amplitude features from audio data. This amplitude extraction model includes a first feature extraction channel, a second feature extraction channel, and several fully connected layers. The first feature extraction channel includes a convolutional neural network module, a recurrent neural network module, a pooling layer, and a fully connected layer; the second feature extraction channel includes a low-level feature extraction module, a high-level feature extraction module, and several fully connected layers. Optionally, in other practical application scenarios, this amplitude extraction model can also be of other model structures.
[0086] Specifically, the data processing device acquires the acoustic wave feature data and volume feature data of the audio data, respectively. The acoustic wave feature data is input into the first feature extraction channel of the amplitude extraction model for neural network feature extraction, outputting the first amplitude feature of the audio data. The volume feature data is input into the second feature extraction channel of the amplitude extraction model for statistical amplitude feature extraction, outputting the second amplitude feature of the audio data. The first amplitude feature is the amplitude feature obtained by feature extraction from the acoustic wave image using the first feature extraction channel. The second amplitude feature is the amplitude feature obtained by statistical feature extraction and nonlinear transformation of the volume feature data using the second feature extraction channel, representing the dynamic changes in the audio data. The first and second amplitude features describe the sound state in the audio data from different perspectives and exist in different feature spaces, and they are complementary.
[0087] After acquiring the first amplitude feature and the second amplitude feature, the data processing device concatenates the first amplitude feature and the second amplitude feature to obtain the concatenated amplitude feature, and inputs the concatenated amplitude feature into the fully connected layer of the amplitude extraction model to obtain the amplitude feature of the audio data, wherein the amplitude feature is a feature parameter characterizing the strength of the sound pressure level.
[0088] 202. Determine the target volume value of the audio data based on the loudness characteristics and the amplitude characteristics.
[0089] After acquiring the loudness characteristics that represent the subjective perception of sound intensity by the human ear and the amplitude characteristics that represent the sound pressure level in the audio data, the data processing device calculates the target volume value of the audio data based on the loudness and amplitude characteristics.
[0090] Specifically, after acquiring the loudness and amplitude features of the audio data, the data processing device also acquires the target loudness weight of the loudness feature and the target amplitude weight of the amplitude feature. The target loudness weight and target amplitude weight are loudness weight values and amplitude weight values obtained through learning and training a preset feature weight model.
[0091] After acquiring the target loudness weight and target amplitude weight, the data processing device weights the loudness feature according to the target loudness weight to obtain a weighted loudness feature; it also weights the amplitude feature according to the target amplitude weight to obtain a weighted amplitude feature. After acquiring the weighted loudness feature and weighted amplitude feature, the data processing device calculates the target volume value of the audio data based on these features. Specifically, the calculation formula is as follows:
[0092] y = ω1||x1||1+ω2||x2||1+b
[0093] Where y is the target volume value of the audio data, ω1 is the target loudness weight, x1 is the loudness feature, ω2 is the target amplitude weight, x2 is the amplitude feature, and b is the bias constant.
[0094] After obtaining the target volume value of the audio data through loudness and amplitude characteristics, the data processing device also drives the smart terminal to adjust the playback status of the audio data according to the target volume value. Specifically, the data processing device receives an audio playback request, obtains the target audio data to be played corresponding to the audio playback request, obtains the target volume value of the target audio data according to the loudness and amplitude characteristics of the target audio data, and plays the target audio data in a preset audio application according to the target volume value.
[0095] In this embodiment, the data processing device, in response to a volume adjustment request, acquires the audio data to be processed within the request; extracts loudness features from the audio data to obtain its loudness characteristics; and further extracts amplitude features from the audio data based on sound wave feature data and volume feature data. After acquiring the loudness and amplitude features of the audio data, the device calculates the target volume value of the audio data using both loudness and amplitude features. This achieves the calculation of the audio data's volume value from a combination of subjective loudness and objective amplitude dimensions, thereby improving the accuracy and comprehensiveness of volume adjustment.
[0096] like Figure 4 As shown, Figure 4 A flowchart illustrating an embodiment of the data processing method for extracting loudness features from audio data provided in this application, specifically including steps 301 to 303:
[0097] 301. Query the preset equal loudness curves to obtain the hearing threshold sound level values corresponding to each preset critical frequency band;
[0098] 302. Perform frequency domain conversion on the audio data to obtain an audio frequency domain signal, and calculate the critical band energy of each critical band based on the audio frequency domain signal and each critical band;
[0099] 303. For each critical frequency band, determine the loudness characteristics of the audio data based on the hearing threshold loudness level value of the critical frequency band and the critical frequency band energy.
[0100] Based on the above embodiments, in this embodiment, after acquiring the audio data to be processed, the data processing device obtains the loudness characteristics of the audio data according to the preset equal loudness curve and each critical frequency band.
[0101] Specifically, the data processing device acquires the center frequency of each critical frequency band, queries a preset equal-loudness curve, obtains the hearing threshold sound level value corresponding to the center frequency of the critical frequency band in the preset equal-loudness curve, and determines the hearing threshold sound level value as the hearing threshold sound level value of the critical frequency band. Here, the center frequency of the critical frequency band is the middle frequency of the critical frequency band used to characterize its acoustic features. The hearing threshold sound level value is a parameter representing the perceived loudness of the sound in the human ear corresponding to the center frequency of the critical frequency band in the preset equal-loudness curve.
[0102] Specifically, after acquiring the hearing threshold sound level values corresponding to each critical frequency band, the data processing equipment performs a Fast Fourier Transform on the audio signal, thereby converting the audio signal from a time-domain signal to a frequency-domain signal, thus obtaining the audio frequency-domain signal. After acquiring the audio frequency-domain signal, the data processing equipment performs calculations based on the audio frequency-domain signal and each critical frequency band to obtain the energy of the audio signal in each critical frequency band, thus obtaining the critical frequency band energy of each critical frequency band.
[0103] After acquiring the hearing threshold sound level value and the critical band energy of the audio signal in each critical band, the data processing equipment determines the loudness characteristics of the audio data for each critical band based on the hearing threshold sound level value and the critical band energy.
[0104] Specifically, the data processing device pre-sets a loudness calculation model for calculating the characteristic loudness value of the critical frequency band. The hearing threshold loudness level value of each critical frequency band and the critical frequency band energy corresponding to the audio data are input into the pre-set loudness calculation model to calculate the characteristic loudness value of the critical frequency band. The pre-set loudness calculation model is as follows:
[0105]
[0106] Where N is the characteristic loudness value, Y is the hearing threshold loudness level value, E is the critical frequency band energy, and E0 is the energy corresponding to the reference sound intensity.
[0107] After the data processing device calculates the characteristic loudness value of each critical frequency band through a preset loudness calculation model, it statistically analyzes the characteristic loudness value of each critical frequency band, combines the characteristic loudness values of each critical frequency band, and generates the loudness feature of the audio data. In a specific embodiment, the loudness feature is a 1*24-dimensional loudness feature vector, which is a feature parameter characterizing the intensity of sound energy of the audio signal.
[0108] After acquiring the loudness characteristics of the audio data, the data processing device also acquires the amplitude data of the audio data, and calculates the target volume value of the audio data using the loudness characteristic data and the amplitude data.
[0109] In this embodiment, the data processing device obtains the hearing threshold sound level values corresponding to each preset critical frequency band by querying preset equal-loudness curves; performs frequency domain conversion on the audio data to obtain an audio frequency domain signal; calculates the critical frequency band energy of each critical frequency band based on the audio frequency domain signal and each critical frequency band; and determines the loudness characteristics of the audio data for each critical frequency band based on the hearing threshold sound level value and the critical frequency band energy. This achieves accurate acquisition of the loudness characteristics of the audio data, providing a subjective calculation scale for subsequently calculating the target volume value of the audio data.
[0110] like Figure 5 As shown, Figure 5 This is a flowchart illustrating an embodiment of the data processing method for obtaining amplitude characteristics of audio data provided in this application, specifically including steps 401 to 404:
[0111] 401. Convert the audio data into a sound wave image, extract features from the sound wave image, and obtain sound wave feature data in the sound wave image;
[0112] 402. Perform pooling processing on the acoustic wave feature data to obtain the image statistical features of the acoustic wave image;
[0113] 403. Integrate the image statistical features and the sound wave feature data to generate the first amplitude feature of the audio data;
[0114] 404. Obtain the volume feature data of the audio data, and generate the amplitude feature of the audio data based on the volume feature data and the first amplitude feature.
[0115] Based on the above embodiments, in this embodiment, after acquiring the audio data, the data processing device also extracts the amplitude features of the audio data according to the sound wave feature data and audio feature data in the audio data.
[0116] Specifically, since the convolutional neural network in the first feature extraction channel cannot recognize audio signals, the data processing device converts the audio signal into a sound wave image after acquiring it. This sound wave image is a sound wave function graph where the horizontal axis represents time and the vertical axis represents amplitude. After acquiring the sound wave image corresponding to the audio signal, the data processing device inputs this sound wave image into the first feature extraction channel. The first feature extraction channel then extracts features from the sound wave image to obtain sound wave feature data. This sound wave feature data consists of feature parameters characterizing the acoustic properties of the audio signal.
[0117] Specifically, the data processing device inputs the acoustic wave image into the convolutional neural network module in the first feature extraction channel for feature extraction, obtaining acoustic wave feature data from the acoustic wave image. After acquiring the acoustic wave feature data, the data processing device inputs it into the recurrent neural network module in the first feature extraction channel. The recurrent neural network module performs pooling processing on the acoustic wave feature data to obtain the image statistical features of the acoustic wave image. These image statistical features are feature vectors representing the sound wave image, including fundamental frequency, energy, and zero-crossing rate, which are related to volume.
[0118] The data processing device inputs acoustic wave feature data and image statistical features into a fully connected layer in the first feature extraction channel. The fully connected layer integrates the image statistical features and acoustic wave feature data to generate a first amplitude feature of the audio data. This first amplitude feature is the amplitude feature obtained by feature extraction from the acoustic wave image using the first feature extraction channel. The first feature extraction channel is a CRNN model, which includes a convolutional recurrent network, a recurrent neural network, and pooling layers.
[0119] Specifically, the data processing device learns the amplitude features of audio data through sound wave images, and also learns the volume-related statistical features of audio data through a second feature extraction channel to obtain the second amplitude feature. That is, the data processing device acquires the volume feature data of the audio data and generates the amplitude feature of the audio data based on the volume feature data and the first amplitude feature.
[0120] Specifically, the data processing device inputs the audio data into the low-level feature extraction module in the second feature extraction channel. The low-level feature extraction module extracts the volume features of the audio data to obtain the volume feature data of the audio data.
[0121] After acquiring the volume feature data of the audio data, the data processing device inputs this volume feature data into the advanced feature extraction module in the second feature extraction channel. The advanced feature extraction module performs statistical feature extraction on the volume feature data, thereby obtaining the statistical feature data of the volume feature data. This statistical feature data includes the average, maximum, minimum, variance, kurtosis, and skewness of the audio feature data. This statistical feature data represents parameters characterizing the global dynamic changes of the volume feature data.
[0122] After acquiring the statistical feature data of the volume feature data, the data processing device inputs this statistical feature data into the fully connected layer of the second feature extraction channel for nonlinear transformation, thereby mapping the high-dimensional statistical feature data sequentially to a low-dimensional feature space to obtain low-dimensional statistical data. The second amplitude feature of the audio data is generated from this volume feature data and the low-dimensional statistical data. The second amplitude feature is obtained by the second feature extraction channel through statistical feature extraction and nonlinear transformation of the volume feature data, and it characterizes the amplitude of dynamic changes in the audio data. The first amplitude feature and the second amplitude feature describe the sound state in the audio data from different perspectives and exist in different feature spaces, and they are complementary.
[0123] After acquiring the first amplitude feature and the second amplitude feature of the audio data, the data processing device concatenates the first amplitude feature and the second amplitude feature to obtain a concatenated amplitude feature, and then inputs the concatenated amplitude feature into the fully connected layer of the amplitude extraction model to obtain the amplitude feature of the audio data. In a specific embodiment, the amplitude feature is a 1*128 dimensional amplitude feature vector.
[0124] After acquiring the amplitude characteristics of the audio data, the data processing device also acquires the loudness data of the audio data, and calculates the target volume value of the audio data using the loudness characteristic data and the amplitude data.
[0125] In this embodiment, the data processing device converts the audio data into a sound wave image, extracts features from the sound wave image to obtain sound wave feature data; performs pooling processing on the sound wave feature data to obtain image statistical features of the sound wave image; integrates the image statistical features and the sound wave feature data to generate a first amplitude feature of the audio data; acquires the volume feature data of the audio data, and generates the amplitude feature of the audio data based on the volume feature data and the first amplitude feature. This achieves the acquisition of amplitude features of audio data from multiple angles, improving the accuracy of subsequent volume value calculations.
[0126] like Figure 6 As shown, Figure 6 This is a flowchart illustrating an embodiment of the data processing method provided in this application for determining the calculation weight of the target volume value, specifically including steps 501 to 503:
[0127] 501. Obtain the initial loudness weight and initial amplitude weight of the preset feature weight model;
[0128] 502. The preset training loudness features are weighted according to the initial loudness weight, and the preset training amplitude features are weighted according to the initial amplitude weight, and the training volume value is obtained by summing them up.
[0129] 503. Calculate the training loss of the training volume value and the predicted volume value, and adjust the initial loudness weight and the initial amplitude weight according to the training loss to obtain the target loudness weight of the training loudness feature and the target amplitude weight of the training amplitude feature.
[0130] Based on the above embodiments, in this embodiment, before acquiring audio data, the data processing device pre-generates a feature weight model for training loudness weight and amplitude weight, and trains the calculation weight of the target volume value through the feature weight model, wherein the calculation weight includes the target loudness weight and the target amplitude weight.
[0131] Specifically, the data processing device acquires the initial loudness weight and initial amplitude weight in the preset feature weight model, as well as the training audio data, and acquires the training loudness feature and training amplitude feature in the training audio data. The initial loudness weight and initial amplitude weight are then trained using the training loudness feature and training amplitude feature.
[0132] Specifically, after acquiring the training loudness features and training amplitude features, the data processing device inputs the training loudness features and training amplitude features into a preset feature weight model. The training loudness features are weighted using the initial loudness weight to obtain the training weighted loudness; the training amplitude features are weighted using the initial amplitude weight to obtain the training weighted amplitude.
[0133] After obtaining the training weighted loudness and training weighted amplitude, the data processing device obtains the training volume value of the training audio data through the training weighted loudness and training weighted amplitude.
[0134] After acquiring the training volume value, the data processing device acquires the predicted volume value of the training audio data, calculates the training loss of the training volume value and the predicted volume value in the preset feature weight model, and compares the training loss with the preset loss threshold to obtain the training result.
[0135] Optionally, if the training loss is greater than the preset loss threshold, the data processing device adjusts the initial loudness weight and the initial amplitude weight according to the training loss, and continues training according to the adjusted initial loudness weight and the adjusted initial amplitude weight until the training converges, thereby obtaining the target loudness weight and the target amplitude weight corresponding to the training audio data.
[0136] Optionally, if the training loss is less than the preset loss threshold, the data processing device determines the initial loudness weight and initial amplitude weight as the target loudness weight and target amplitude weight of the training audio data.
[0137] Specifically, after acquiring the loudness and amplitude features of the audio data, the data processing device inputs these features into a preset feature weighting model to obtain the target loudness weight, target amplitude weight, and target bias constant output by the model. The loudness and amplitude features are then weighted and summed using these target loudness weights, target amplitude weights, and target bias constant to obtain the target volume value of the audio data.
[0138] In this embodiment, the data processing device acquires the initial loudness weights and initial amplitude weights of a preset feature weight model; it then weights preset training loudness features according to the initial loudness weights and weights preset training amplitude features according to the initial amplitude weights, summing the results to obtain training volume values; it calculates the training loss between the training volume values and the predicted volume values, and adjusts the initial loudness weights and initial amplitude weights according to the training loss to obtain the target loudness weights and target amplitude weights of the training loudness features. This achieves effective acquisition of the target loudness weights and target amplitude weights of audio data.
[0139] To better implement the data processing method in the embodiments of this application, based on the data processing method, the embodiments of this application also provide a data processing apparatus, such as... Figure 7 As shown, Figure 7 This is a schematic diagram of one embodiment of the data processing apparatus in this application. The data processing apparatus 600 includes:
[0140] The feature extraction module 501 is configured to extract loudness features and amplitude features of the audio data to be processed.
[0141] The volume calculation module 504 is configured to calculate the target volume value of the audio data based on the loudness feature and the amplitude feature.
[0142] In some embodiments of this application, the data processing apparatus extracts loudness and amplitude features of the audio data to be processed, including:
[0143] The loudness features of the audio data are extracted based on the preset equal loudness curves and critical frequency bands.
[0144] Based on the acoustic wave feature data and audio feature data in the audio data, the amplitude features of the audio data are extracted.
[0145] In some embodiments of this application, the data processing device extracts loudness features from the audio data based on a preset equal loudness curve and a critical frequency band to obtain the loudness features of the audio data, including:
[0146] Query the preset equal loudness curves to obtain the hearing threshold sound level values corresponding to each preset critical frequency band;
[0147] The audio data is frequency domain converted to obtain an audio frequency domain signal. Based on the audio frequency domain signal and each of the critical frequency bands, the critical frequency band energy of each of the critical frequency bands is calculated.
[0148] For each critical frequency band, the loudness characteristics of the audio data are determined based on the hearing threshold loudness level value and the critical frequency band energy of the critical frequency band.
[0149] In some embodiments of this application, the data processing device determines the loudness characteristics of the audio data for each critical frequency band based on the hearing threshold loudness level value and the critical frequency band energy, including:
[0150] For each critical frequency band, the hearing threshold loudness level value and the critical frequency band energy of the critical frequency band are input into a preset loudness calculation model to obtain the characteristic loudness value of each critical frequency band.
[0151] The loudness values of each critical frequency band are statistically analyzed to generate the loudness features of the audio data.
[0152] In some embodiments of this application, the data processing device extracts the amplitude features of the audio data based on the sound wave feature data and audio feature data in the audio data, including:
[0153] The audio data is converted into a sound wave image, and features are extracted from the sound wave image to obtain sound wave feature data in the sound wave image.
[0154] The acoustic wave feature data is pooled to obtain the image statistical features of the acoustic wave image;
[0155] Integrate the image statistical features and the sound wave feature data to generate the first amplitude feature of the audio data;
[0156] Obtain the volume feature data of the audio data, and generate the amplitude feature of the audio data based on the volume feature data and the first amplitude feature.
[0157] In some embodiments of this application, the data processing device acquires volume feature data of the audio data and generates amplitude features of the audio data based on the volume feature data and the first amplitude feature, including:
[0158] Extract the volume feature data of the audio data, and perform statistical feature extraction on the volume feature data to obtain the statistical feature data of the volume feature data;
[0159] The statistical feature data is subjected to a nonlinear transformation to obtain low-dimensional statistical data. The second amplitude feature of the audio data is generated based on the volume feature data and the low-dimensional statistical data.
[0160] The amplitude features of the audio data are generated based on the first amplitude feature and the second amplitude feature.
[0161] In some embodiments of this application, the data processing device calculates the target volume value of the audio data based on the loudness feature and the amplitude feature, including:
[0162] Extract the target loudness weight of the loudness feature and the target amplitude weight of the amplitude feature;
[0163] The loudness features are weighted according to the target loudness weight to obtain weighted loudness features;
[0164] The amplitude features are weighted according to the target amplitude weight to obtain weighted amplitude features;
[0165] The target volume value of the audio data is calculated based on the weighted loudness feature and the weighted amplitude feature.
[0166] In some embodiments of this application, before the data processing apparatus extracts the target loudness weight of the loudness feature and the target amplitude weight of the amplitude feature, it further includes:
[0167] Obtain the initial loudness weight and initial amplitude weight of the preset feature weight model;
[0168] The preset training loudness features are weighted according to the initial loudness weight, and the preset training amplitude features are weighted according to the initial amplitude weight, and the training volume value is obtained by summing them up.
[0169] Calculate the training loss of the training volume value and the predicted volume value, and adjust the initial loudness weight and the initial amplitude weight according to the training loss to obtain the target loudness weight of the training loudness feature and the target amplitude weight of the training amplitude feature.
[0170] In this embodiment, the data processing device acquires the audio data to be processed in the volume adjustment request in response to the request; extracts loudness features from the audio data to obtain the loudness features; and extracts amplitude features from the audio data based on sound wave feature data and volume feature data. After acquiring the loudness and amplitude features of the audio data, the target volume value of the audio data is calculated using both loudness and amplitude features. This achieves the calculation of the audio data volume value from a combination of subjective loudness and objective amplitude dimensions, thereby improving the accuracy and comprehensiveness of volume adjustment.
[0171] This invention also provides a data processing device, such as... Figure 8 As shown, Figure 8 This is a schematic diagram of one embodiment of the data processing device provided in this application.
[0172] The data processing device integrates any of the data processing apparatuses provided in the embodiments of the present invention, and the data processing device includes:
[0173] One or more processors;
[0174] Memory; and
[0175] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor of the steps in the data processing method described in any of the embodiments of the above data processing method.
[0176] Specifically, the data processing device may include components such as a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, a power supply 703, and an input unit 704. Those skilled in the art will understand that... Figure 8 The data processing device structure shown does not constitute a limitation on the data processing device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0177] The processor 701 is the control center of the data processing device. It connects various parts of the data processing device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 702, and by calling data stored in the memory 702, it performs various functions of the data processing device and processes data, thereby providing overall monitoring of the data processing device. Optionally, the processor 701 may include one or more processing cores; preferably, the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 701.
[0178] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the data processing device, etc. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.
[0179] The data processing device also includes a power supply 703 that supplies power to the various components. Preferably, the power supply 703 can be logically connected to the processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0180] The data processing device may also include an input unit 704, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0181] Although not shown, the data processing device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 701 in the data processing device loads the executable files corresponding to the processes of one or more application programs into the memory 702 according to the following instructions, and the processor 701 runs the application programs stored in the memory 702 to realize various functions, as follows:
[0182] Extract loudness and amplitude features from the audio data to be processed;
[0183] The target volume value of the audio data is determined based on the loudness feature and the amplitude feature.
[0184] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.
[0185] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.
[0186] Optionally, embodiments of the present invention also provide a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, which is loaded by a processor to execute the steps in any of the data processing methods provided in the embodiments of the present invention. For example, the computer program loaded by the processor may execute the following steps:
[0187] Extract loudness and amplitude features from the audio data to be processed;
[0188] The target volume value of the audio data is determined based on the loudness feature and the amplitude feature.
[0189] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0190] The above provides a detailed description of a data processing method provided by the embodiments of this application. Specific embodiments have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A data processing method, characterized in that, The data processing method includes: Loudness features are extracted from audio data based on preset equal loudness curves and critical frequency bands to obtain loudness features of the audio data. The audio data is then converted into a sound wave image, and features are extracted from the sound wave image to obtain sound wave feature data. The sound wave feature data is then pooled to obtain image statistical features of the sound wave image. The image statistical features and the sound wave feature data are then integrated to generate the first amplitude feature of the audio data. The volume feature data of the audio data is then obtained, and the amplitude feature of the audio data is generated based on the volume feature data and the first amplitude feature. The target volume value of the audio data is determined based on the loudness feature and the amplitude feature.
2. The data processing method as described in claim 1, characterized in that, The step of extracting loudness features from the audio data based on a preset equal loudness curve and critical frequency band to obtain the loudness features of the audio data includes: Query the preset equal loudness curves to obtain the hearing threshold sound level values corresponding to each preset critical frequency band; The audio data is frequency domain converted to obtain an audio frequency domain signal. Based on the audio frequency domain signal and each of the critical frequency bands, the critical frequency band energy of each of the critical frequency bands is calculated. For each critical frequency band, the loudness characteristics of the audio data are determined based on the hearing threshold loudness level value and the critical frequency band energy of the critical frequency band.
3. The data processing method as described in claim 2, characterized in that, For each critical frequency band, the loudness characteristics of the audio data are determined based on the hearing threshold loudness level and the critical frequency band energy, including: For each critical frequency band, the hearing threshold loudness level value and the critical frequency band energy of the critical frequency band are input into a preset loudness calculation model to obtain the characteristic loudness value of each critical frequency band. The loudness values of each critical frequency band are statistically analyzed to generate the loudness features of the audio data.
4. The data processing method as described in claim 1, characterized in that, The step of acquiring the volume feature data of the audio data and generating the amplitude feature of the audio data based on the volume feature data and the first amplitude feature includes: Extract the volume feature data of the audio data, and perform statistical feature extraction on the volume feature data to obtain the statistical feature data of the volume feature data; The statistical feature data is subjected to a nonlinear transformation to obtain low-dimensional statistical data. The second amplitude feature of the audio data is generated based on the volume feature data and the low-dimensional statistical data. The amplitude features of the audio data are generated based on the first amplitude feature and the second amplitude feature.
5. The data processing method as described in claim 1, characterized in that, Determining the target volume value of the audio data based on the loudness feature and the amplitude feature includes: Extract the target loudness weight of the loudness feature and the target amplitude weight of the amplitude feature; The loudness features are weighted according to the target loudness weight to obtain weighted loudness features; The amplitude features are weighted according to the target amplitude weight to obtain weighted amplitude features; The target volume value of the audio data is calculated based on the weighted loudness feature and the weighted amplitude feature.
6. The data processing method as described in claim 5, characterized in that, Before extracting the target loudness weight of the loudness feature and the target amplitude weight of the amplitude feature, the method further includes: Obtain the initial loudness weight and initial amplitude weight of the preset feature weight model; The preset training loudness features are weighted according to the initial loudness weight, and the preset training amplitude features are weighted according to the initial amplitude weight, and the training volume value is obtained by summing them up. Calculate the training loss of the training volume value and the predicted volume value, and adjust the initial loudness weight and the initial amplitude weight according to the training loss to obtain the target loudness weight of the training loudness feature and the target amplitude weight of the training amplitude feature.
7. The data processing method according to any one of claims 1-6, characterized in that, After determining the target volume value of the audio data based on the loudness feature and the amplitude feature, the method further includes: In response to an audio playback request, obtain the target audio data to be played in the audio playback request; Play the target audio data in a preset audio application according to the target volume value of the target audio data.
8. A data processing apparatus, characterized in that, The data processing device includes: The feature extraction module is configured to extract loudness features from audio data based on a preset equal loudness curve and critical frequency band to obtain the loudness features of the audio data; convert the audio data into a sound wave image; extract features from the sound wave image to obtain sound wave feature data in the sound wave image; perform pooling processing on the sound wave feature data to obtain the image statistical features of the sound wave image; integrate the image statistical features and the sound wave feature data to generate the first amplitude feature of the audio data; obtain the volume feature data of the audio data; and generate the amplitude feature of the audio data based on the volume feature data and the first amplitude feature. The volume calculation module is configured to calculate the target volume value of the audio data based on the loudness feature and the amplitude feature.
9. A data processing device, characterized in that, The data processing device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to execute the steps of the data processing method according to any one of claims 1 to 7.