Audio classification method and device, electronic equipment, storage medium and program product

By extracting segmented audio and using feature extraction parameters and dimensional parameters to predict its classification probability, the inefficiency of audio classification caused by multiple inferences is solved, and efficient audio multi-dimensional classification is achieved.

CN120020950APending Publication Date: 2025-05-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311537625.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

With the increase of classification dimensions, the prior art requires inference of the same audio multiple times, resulting in inefficient audio classification.

Method used

By extracting segmented audio, the feature extraction parameters are used to determine the segmented audio characteristics, and the prediction probability of segmented audio in each classification dimension is predicted based on the dimension parameters of multiple classification dimensions. One reasoning can be achieved to determine the classification results of audio in multiple classification dimensions.

Benefits of technology

It improves the efficiency of audio classification, reduces data processing, avoids repeated inference, and improves the accuracy of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020950A_ABST
    Figure CN120020950A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio classification method and device, electronic equipment, a storage medium and a program product. According to the method, at least one segmented audio can be extracted from the to-be-classified audio; predicting a plurality of prediction probabilities of each segmented audio under each classification dimension according to dimension parameters corresponding to the plurality of classification dimensions and each segmented audio feature; for each classification dimension, determining a target type corresponding to each classification dimension by using the prediction probability of each segmented audio in a preset type corresponding to the classification dimension; and determining a classification result of the to-be-classified audio according to the target type. According to the method, feature extraction is performed on the segmented audio based on the same feature extraction parameter, and the prediction probability is calculated by using different dimension parameters, so that the target type of the to-be-classified audio under multiple classification dimensions can be determined through one-time reasoning, repeated reasoning is avoided, the classification data processing amount is reduced, and the audio classification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to an audio classification method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] Audio classification refers to determining the predetermined category to which an audio belongs based on the characteristics of the audio. When classifying audio, different classification dimensions can produce different classification results. Since the preset types under different classification dimensions are different, if the classification results of the audio under different classification dimensions are to be obtained, targeted processing of the audio needs to be performed under each classification dimension.

[0003] However, with the increase in the number of classification dimensions, it is necessary to perform inference on the same audio multiple times to determine the types of the audio under multiple classification dimensions. This method requires repeated inference on the same audio, increasing the data processing volume of classification and resulting in low efficiency of audio classification. Summary of the Invention

[0004] Embodiments of this application provide an audio classification method, apparatus, electronic device, storage medium, and program product, which can infer the classification results of an audio under multiple classification dimensions at one time, improving the efficiency of audio classification.

[0005] Embodiments of this application provide an audio classification method, which includes:

[0006] Extracting at least one segmented audio from the audio to be classified;

[0007] Determining the segmented audio features corresponding to each of the segmented audios based on the feature extraction parameters;

[0008] Predicting multiple prediction probabilities of each of the segmented audios under each of the classification dimensions according to the dimension parameters corresponding to multiple classification dimensions and each of the segmented audio features, where each of the classification dimensions corresponds to multiple preset types, and the prediction probabilities correspond one-to-one to the preset types;

[0009] For each of the classification dimensions, using the prediction probabilities of each of the segmented audios under the preset types corresponding to the classification dimension, determining the target type corresponding to each classification dimension;

[0010] Determining the classification result of the audio to be classified based on the target type corresponding to each classification dimension.

[0011] Embodiments of this application also provide an audio classification apparatus, which includes:

[0012] A segmentation unit, configured to extract at least one segmented audio from the audio to be classified;

[0013] A feature extraction unit, configured to determine the segment audio features corresponding to each of the segment audios based on feature extraction parameters;

[0014] A probability prediction unit, configured to predict multiple prediction probabilities of each of the segment audios under each of the classification dimensions according to the dimension parameters corresponding to multiple classification dimensions and each of the segment audio features, wherein each of the classification dimensions corresponds to multiple preset types, and the prediction probabilities correspond one-to-one to the preset types;

[0015] A type determination unit, configured to, for each of the classification dimensions, determine the target type corresponding to each classification dimension by using the prediction probabilities of each of the segment audios under the preset types corresponding to the classification dimensions;

[0016] A classification unit, configured to determine the classification result of the audio to be classified based on the target type corresponding to each classification dimension.

[0017] In some embodiments, the feature extraction parameters include depth convolution parameters and pointwise convolution parameters, and the feature extraction unit further includes:

[0018] A depth convolution subunit, configured to, for each of the segment audios, perform depth convolution processing on the segment audio by using the depth convolution parameters to obtain the depth audio features corresponding to the segment audio;

[0019] A pointwise convolution subunit, configured to perform pointwise convolution processing on the depth audio features by using the pointwise convolution parameters to obtain the segment audio features corresponding to the segment audio.

[0020] In some embodiments, the probability prediction unit further includes:

[0021] An acquisition subunit, configured to acquire the multiple preset types corresponding to each of the classification dimensions;

[0022] A linear transformation subunit, configured to perform linear transformation on each of the segment audio features by using the dimension parameters corresponding to multiple classification dimensions to obtain the dimension features corresponding to each of the classification dimensions;

[0023] A normalization subunit, configured to perform normalization processing on the dimension features corresponding to each classification dimension to obtain the multiple prediction probabilities of each of the segment audios under each of the classification dimensions.

[0024] In some embodiments, the type determination unit further includes:

[0025] A probability acquisition subunit, configured to, for each of the classification dimensions, acquire the multiple prediction probabilities corresponding to each of the preset types, wherein one prediction probability corresponds to one segment audio;

[0026] A mean calculation subunit, configured to calculate the mean of multiple prediction probabilities corresponding to each of the preset types, to obtain a mean probability corresponding to each of the preset types;

[0027] A determination subunit, configured to determine the preset type corresponding to the maximum mean probability as the target type corresponding to the classification dimension.

[0028] In some embodiments, the segmentation unit further includes:

[0029] A splitting subunit, configured to perform splitting processing on the audio to be classified based on a preset duration, to obtain at least one sub-audio, where the duration of the sub-audio is the preset duration;

[0030] A first determination subunit, configured to determine the preset number of segmented audios from the at least one sub-audio if the number of the sub-audios is greater than a preset number;

[0031] A second determination subunit, configured to determine each of the sub-audios as a segmented audio if the number of the sub-audios is less than or equal to the preset number.

[0032] In some embodiments, the audio classification device further includes a training unit, configured to:

[0033] Obtain a training sample set, where the training sample set includes multiple sample audios and a sample type corresponding to each of the sample audios, and where the sample type is one of the preset types under the multiple classification dimensions;

[0034] Determine, from the training sample set, the sample audio to be used corresponding to each of the classification dimensions;

[0035] Extract features from the sample audio to be used based on initial feature extraction parameters, to obtain sample audio features corresponding to the classification dimension;

[0036] Calculate a predicted type of the sample audio to be used under the classification dimension based on the initial dimension parameters corresponding to the classification dimension and the sample audio features corresponding to the classification dimension;

[0037] Adjust the initial feature extraction parameters and the initial dimension parameters according to the predicted type of the sample audio to be used and the sample type corresponding to the sample audio to be used, to obtain feature extraction parameters and dimension parameters.

[0038] In some embodiments, the training unit is further configured to:

[0039] For each classification dimension, construct a dimension loss function by using the predicted type of the sample audio to be used and the sample type corresponding to the sample audio to be used;

[0040] Construct a target loss function based on the dimension weights corresponding to the classification dimensions and the dimension loss functions corresponding to the classification dimensions;

[0041] According to the target loss function, adjust the initial feature extraction parameters and the initial dimension parameters until the target loss function meets the preset conditions, and then obtain the feature extraction parameters and the dimension parameters.

[0042] In some embodiments, the training unit is further configured to:

[0043] Obtain an initial sample audio according to the multimedia information corresponding to the preset multimedia;

[0044] Perform enhancement processing on the initial sample audio in the time domain to obtain a time-domain enhanced sample audio;

[0045] Perform enhancement processing on the initial sample audio in the frequency domain to obtain a frequency-domain enhanced sample audio;

[0046] Determine the initial sample audio, the time-domain enhanced sample audio, and the frequency-domain enhanced sample audio as the sample audio of the training sample set;

[0047] Determine the sample type corresponding to the sample audio according to the preset types corresponding to multiple classification dimensions.

[0048] In some embodiments, the multimedia information includes multimedia content and multimedia description information, and the training unit is further configured to:

[0049] Perform entity recognition on the multimedia description information to extract audio entities in the multimedia content;

[0050] Perform audio content recognition on the multimedia content to extract audio content features in the multimedia content;

[0051] Use the audio entities and the audio content features to determine the initial sample audio from multiple preset audios in a preset audio library.

[0052] In some embodiments, the training unit is further configured to:

[0053] Extract the time-domain features of the initial sample audio in the time domain;

[0054] Perform transformation processing on the time-domain features to obtain transformed time-domain features;

[0055] Based on the transformed time-domain features, obtain a time-domain enhanced sample audio.

[0056] In some embodiments, the training unit is further configured to:

[0057] Convert the initial sample audio into a spectrogram feature map;

[0058] Perform perturbation processing on the spectrogram feature map according to the mask matrix to obtain a perturbed feature map;

[0059] Perform linear interpolation processing on the spectrogram feature maps corresponding to different initial sample audios to obtain interpolated feature maps;

[0060] Generate the frequency-domain enhanced sample audio according to the perturbed feature map and the interpolated feature map.

[0061] An embodiment of the present application further provides an electronic device, including a memory storing multiple instructions; the processor loads the instructions from the memory to execute the steps in any one of the audio classification methods provided by the embodiments of the present application.

[0062] An embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the audio classification methods provided by the embodiments of the present application.

[0063] An embodiment of the present application further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps in any one of the audio classification methods provided by the embodiments of the present application are implemented.

[0064] An embodiment of the present application can extract at least one segmented audio from the audio to be classified, and use the feature extraction parameters to determine the segmented audio features corresponding to each segmented audio; according to the dimension parameters of multiple classification dimensions and each segmented audio feature, predict multiple prediction probabilities of each segmented audio under each classification dimension; for each classification dimension, use the prediction probabilities of each segmented audio under the preset type corresponding to the classification dimension to determine the target type corresponding to each classification dimension; according to all the target types, determine the classification result of the audio to be classified. Based on the same feature extraction parameters to extract features from the segmented audio, and then use different dimension parameters to calculate the prediction probabilities, it is possible to determine the target types of the audio to be classified under multiple classification dimensions with one inference, avoiding repeated inferences, reducing the data processing volume of classification, and thus improving the efficiency of audio classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0066] Figure 1aIt is a schematic diagram of the application scenario of the audio classification method provided by the embodiment of the present application;

[0067] Figure 1b It is a schematic flowchart of the audio classification method provided by the embodiment of the present application;

[0068] Figure 1c It is a schematic diagram of the structure of the classification model provided by the embodiment of the present application;

[0069] Figure 1d It is a schematic diagram of training the classification model provided by the embodiment of the present application;

[0070] Figure 1e It is a schematic diagram of obtaining the initial sample audio provided by the embodiment of the present application;

[0071] Figure 1f It is a schematic diagram of perturbing the spectral feature map provided by the embodiment of the present application;

[0072] Figure 2a It is a schematic flowchart of the audio classification method provided by another embodiment of the present application;

[0073] Figure 2b It is a schematic diagram of the overall process of the audio classification method provided by the embodiment of the present application;

[0074] Figure 3 It is a schematic diagram of the structure of the audio classification device provided by the embodiment of the present application;

[0075] Figure 4 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0076] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0077] The embodiment of the present application provides an audio classification method, device, electronic device, storage medium and program product.

[0078] Among them, the audio classification device can be specifically integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can be a device such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC), a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.; the server can be a single server or a server cluster composed of multiple servers.

[0079] In some embodiments, the audio classification device may also be integrated in multiple electronic devices. For example, the audio classification device may be integrated in multiple servers, and the audio classification method of the present application may be implemented by multiple servers.

[0080] In some embodiments, the server may also be implemented in the form of a terminal.

[0081] For example, referring to Figure 1a , a schematic diagram of an application scenario of the audio classification method is shown. This application scenario may include a server 101 and a terminal 102. Among them, the terminal 102 can access a video platform so that users can share videos on the video platform to provide better recommendation and search services. The videos shared by users on the terminal 102 can be sent to the server 101 for classification processing. Among them, the server 101 can extract audio from the video to obtain the audio to be classified. Of course, the terminal 102 can also extract the audio to be classified from the video and send it to the server 101.

[0082] The server 101 can extract at least one segmented audio from the audio to be classified; determine the segmented audio features corresponding to each segmented audio based on the feature extraction parameters; according to the dimension parameters corresponding to multiple classification dimensions and each segmented audio feature, predict multiple prediction probabilities of each segmented audio under each classification dimension, where each classification dimension corresponds to multiple preset types, and the prediction probabilities correspond one-to-one to the preset types; for each classification dimension, use the prediction probabilities of each segmented audio under the preset types corresponding to the classification dimension to determine the target type corresponding to each classification dimension; based on the target type corresponding to each classification dimension, determine the classification result of the audio to be classified.

[0083] After the server 101 classifies the audio, it can implement recommendation and search services according to the classification result of the audio. For example, recommend to the user the types of audio that the user often watches. Another example is that when the user searches for a certain type of audio, all audio of that type can be quickly displayed. It can be understood that in the specific implementation of the present application, it involves user-related data. For example, when making recommendations, it is necessary to obtain the data of the types that the user often watches. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0084] The following will be described in detail respectively. It should be noted that the order of the following embodiments does not limit the preferred order of the embodiments.

[0085] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0086] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0087] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.

[0088] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, Artificial Intelligence Generated Content (AIGC), conversational interaction, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. In the embodiments of this application, through machine learning, the neural network model is enabled to have the ability to infer the type of audio, and then the neural network model is used to analyze the audio to infer the type of audio.

[0089] In this embodiment, an audio classification method involving artificial intelligence is provided. As Figure 1b shown, the specific process of this audio classification method can be as follows:

[0090] 110. Extract at least one segmented audio from the audio to be classified.

[0091] The audio to be classified refers to the audio that needs to be classified, and the segmented audio refers to a part of the audio to be classified or the entire audio to be classified. Among them, the audio to be classified can refer to all the audio in a certain multimedia interaction platform. Usually, users can share multimedia through an interaction account in the multimedia interaction platform, and the multimedia can be video, audio, etc. The audio to be classified can be the shared audio or the audio extracted from the shared video.

[0092] In some embodiments, in order to facilitate the processing of the audio to be classified, the audio to be classified can be transcoded so that the audio to be classified has a preset sampling rate, a preset number of channels, and a preset format. Among them, the preset sampling rate and preset format can be set according to actual needs. In the embodiments of the present application, the audio to be classified can be transcoded into an audio with a sampling rate of 8000 Hz, a single channel, and a wav format.

[0093] The segmented audio can be extracted from the audio to be classified. As an implementation, it can be to perform segmentation processing on the audio to be classified based on a preset duration to obtain at least one sub-audio, and the duration of the sub-audio is the preset duration; if the number of sub-audios is greater than the preset number, determine the preset number of segmented audios from the at least one sub-audio; if the number of sub-audios is less than or equal to the preset number, each sub-audio is determined as a segmented audio.

[0094] The preset duration refers to the duration set in advance, and the preset duration can be set according to actual needs and is not specifically limited here. For example, it can be 10 s, 15 s, etc. The sub-audio is the audio segmented from the audio to be classified.

[0095] When performing segmentation processing on the audio to be classified based on the preset duration, it can be to calculate the segmentation points according to the total duration of the audio to be classified and the preset duration, and perform segmentation processing on the audio to be classified based on the segmentation points to obtain at least one intermediate audio; fill the duration of each intermediate audio to the preset duration to obtain a sub-audio.

[0096] Among them, the total duration of the audio to be classified refers to the duration of the entire audio to be classified. After obtaining the total duration of the audio to be classified, the ratio of the total duration to the preset duration can be rounded down to obtain a specified number, and then the preset duration is used to determine a specified number of segmentation points from the audio to be classified.

[0097] Among them, the specified quantity is the number of segmentation points. For example, if the total duration is 50s, the preset duration is 15s, and the ratio is 10 / 3, which is rounded down to 3 after rounding, then the specified quantity is 3. Then, according to the specified quantity and the preset duration, the segmentation time points can be determined. For example, if the total duration is 50s, the first segmentation point is at 15s, the second segmentation point is at 30s, and the third segmentation point is at 45s. Then, by segmenting the audio to be classified at the segmentation points, at least one intermediate audio can be obtained.

[0098] After obtaining at least one intermediate audio, the duration of the intermediate audio may be less than the preset duration. Therefore, the duration of each intermediate audio can be filled to the preset duration to obtain sub-audios of the preset duration. As an implementation manner, it may be to first select an intermediate audio with a duration less than the preset duration according to the preset duration, and only fill this intermediate audio. For example, it may be to fill a silent audio segment at the end of the intermediate audio to fill the duration of the intermediate audio to the preset duration. Another example is to copy the content of the intermediate audio to the end of the intermediate audio for filling until the preset duration is satisfied.

[0099] After segmenting the audio to be classified to obtain at least one sub-audio, the number of sub-audios can be obtained. In the aforementioned example, 4 sub-audios are segmented. Among them, the preset quantity can be a pre-set quantity, and the preset quantity is related to the number of classification dimensions. For example, if classification is required under n classification dimensions, the preset quantity can be set to n.

[0100] If the number of sub-audios is greater than the preset quantity, only the preset quantity of sub-audios need to be selected from at least one sub-audio as the segmented audio. At this time, the segmented audio is a partial audio of the audio to be classified. As an implementation manner, it may be to determine the preset quantity of segmented audio from at least one sub-audio based on the start time and end time of each sub-audio in the audio to be classified.

[0101] As an implementation, a preset number of sub-audios can be randomly selected from at least one sub-audio as segmented audios. As another implementation, in order to ensure that the preset number of segmented audios can better represent the entire audio to be classified, the total duration of the audio to be classified can be divided by the preset number to obtain the auxiliary audio duration. Within each auxiliary audio duration of the audio to be classified, one sub-audio can be selected as the segmented audio. For example, if the preset number is 2 and the auxiliary audio duration is 25s, among the aforementioned 4 sub-audios, the first segment is from the 1st to the 15th second, the second segment is from the 16th to the 30th second, the third segment is from the 31st to the 45th second, and the fourth segment is from the 46th to the 50th second. One sub-audio can be selected from the first auxiliary audio duration, i.e., from the 1st to the 25th second, and another sub-audio can be selected from the 26th to the 50th second. Then, the first sub-audio and the fourth sub-audio can be determined as the segmented sub-audios. When the total duration of the audio to be classified is relatively long, reasonably selecting some audio for processing can not only reduce the data processing volume but also ensure that the segmented audio can represent the entire audio to be classified, avoiding affecting the subsequent classification accuracy.

[0102] If the number of sub-audios is less than or equal to the preset number, each sub-audio can be directly determined as the segmented audio. At this time, the segmented audio is the entirety of the audio to be classified. For example, if the total duration of the audio to be classified is 15s, only one sub-audio is obtained after segmentation. When the preset number is 3, then this sub-audio is the segmented audio.

[0103] 120. Based on the feature extraction parameters, determine the segmented audio features corresponding to each of the segmented audios.

[0104] The feature extraction parameters refer to the model parameters for feature extraction obtained after pre-training the classification model. The at least one segmented audio obtained above can be used as the input data of the classification model.

[0105] In some embodiments, the feature extraction parameters can be standard convolution parameters, that is, the feature extraction layer in the classification model is a standard convolution layer. Based on the standard convolution parameters, a standard convolution operation can be performed on each segmented audio, that is, convolution processing is performed on each channel of at least one segmented audio using the same convolution kernel to extract the segmented audio features corresponding to each segmented audio.

[0106] In other embodiments, in order to reduce the computational complexity of feature extraction, in the embodiments of the present application, the feature extraction layer can include a depth convolution layer and a pointwise convolution layer, and the feature extraction parameters can include depth convolution parameters and pointwise convolution parameters. For example, reference can be made to Figure 1c , which shows the structural schematic diagram of the classification model.

[0107] When determining the segment audio features corresponding to the segmented audio, the deep convolutional parameters can be used to perform deep convolutional processing on the segmented audio to obtain the deep audio features corresponding to the segmented audio; the pointwise convolutional parameters can be used to perform pointwise convolutional processing on the deep audio features to obtain the segment audio features corresponding to the segmented audio.

[0108] Among them, the deep convolutional parameters refer to the parameters related to the deep convolutional layer. At least one segmented audio can be input into the deep convolutional layer so as to perform deep convolutional processing on it using the deep convolutional parameters. That is, the data of one channel of at least one segmented audio is convolved by a convolutional kernel to obtain the convolutional result of this channel, and then the convolutional results of each channel are superimposed to obtain the deep audio features corresponding to the segmented audio. Among them, the dimension of the deep audio features is consistent with the dimension of the input data.

[0109] The pointwise convolutional parameters refer to the parameters related to the pointwise convolutional layer. Its input data is the deep audio features. Using the pointwise convolutional parameters, the data of each channel in the deep audio features can be fused to obtain the final segment audio features. Among them, the number of segment audio features is consistent with the number of segmented audio.

[0110] Among them, the size of the convolutional kernel used for pointwise convolutional processing is 1*1. Pointwise convolutional processing can be understood as an operation of multiplying and adding the features of each channel in the deep audio features point by point. Finally, the convolutional results of each channel are superimposed to obtain the final output result. In this process, the number of convolutional kernels is the same as the number of channels of the finally obtained segment audio features. That is, each segmented audio uses the same feature extraction parameters for feature extraction to obtain the segment audio features corresponding to each segmented audio.

[0111] 130. Predict multiple prediction probabilities of each segmented audio under each classification dimension according to the dimension parameters corresponding to multiple classification dimensions and each of the segmented audio features.

[0112] The classification dimension refers to the dimension for classifying the audio to be classified. The same audio to be classified has different classification results under different classification dimensions. This classification dimension can include dimensions such as music genre, language, emotion, etc., and can be specifically set according to actual needs, and no specific limitation is made here.

[0113] The dimension parameters refer to the model parameters for classification obtained after training the classification model. The dimension parameters correspond to the classification dimensions, that is, different classification dimensions use different dimension parameters. That is, in the classification model, multiple classification layers can be included. One classification layer corresponds to one classification dimension, and one classification layer has its corresponding dimension parameters. For example, Figure 1cAmong them, it is assumed that classification is performed under 3 classification dimensions, then there are 3 classification layers.

[0114] Among them, each classification dimension corresponds to multiple preset types. The multiple prediction probabilities of the segmented audio under each classification dimension refer to the probabilities that the segmented audio belongs to each preset type under that classification dimension, and the preset probabilities and preset types correspond one by one. For example, it is assumed that there are 26 preset types under the music genre dimension, then the segmented audio can obtain 26 prediction probabilities under this music genre dimension.

[0115] In some embodiments, the classification layer may include a fully connected layer and a normalization layer. The fully connected layer can perform linear transformation processing on the features, and the normalization layer can perform normalization processing on the features. When predicting the multiple prediction probabilities of each segmented audio under each classification dimension, it may be to obtain the multiple preset types corresponding to each classification dimension; use the dimension parameters corresponding to the multiple classification dimensions to perform linear transformation on each segmented audio feature to obtain the dimension features corresponding to each classification dimension; perform normalization processing on the dimension features corresponding to each classification dimension to obtain the multiple prediction probabilities of each segmented audio under each classification dimension.

[0116] Among them, the multiple preset types corresponding to each classification dimension can be preset and stored at a specified location, and the multiple preset types corresponding to each classification dimension can be directly read from the specified location. Using the dimension parameters corresponding to the multiple classification dimensions to perform linear transformation processing on each segmented audio feature, the dimension features corresponding to each classification dimension can be obtained.

[0117] To facilitate the description of the above process, the following will take one classification dimension as an example for illustration. For one classification dimension, using the dimension parameters corresponding to this classification dimension to perform linear transformation on each segmented audio feature, the transformed features corresponding to each segmented audio feature can be obtained, and all the transformed features are used as the dimension features corresponding to this classification dimension. The above processing is performed for each classification dimension, then the dimension features corresponding to each classification dimension can be obtained.

[0118] In some embodiments, the dimension parameters corresponding to one classification dimension may include a dimension weight and a dimension bias. When performing linear transformation processing on the segmented audio feature, it may be to perform weighted processing on the segmented audio feature based on the dimension weight to obtain a weighted feature; obtain the transformed feature according to the weighted feature and the dimension bias. The transformed feature may include the scores of the segmented audio under each preset type corresponding to this classification dimension. Thus, the dimension feature may include the scores of each segmented audio under the preset type corresponding to this classification dimension.

[0119] To convert it into a probability distribution, the dimensional features corresponding to each classification dimension can be normalized to obtain multiple prediction probabilities of each segmented audio under each classification dimension.

[0120] The above-mentioned dimensional features corresponding to each classification dimension can convert the scores into a probability distribution through a normalization function, and the normalization function can be selected according to actual needs. For example, in the embodiments of the present application, the normalization function can be the softmax function. Since a dimensional feature contains the scores of each segmented audio under multiple preset types, when performing normalization processing, it can be a single normalization processing for the scores of each segmented audio under multiple preset types. For example, a dimensional feature contains the scores of 3 segmented audios under multiple preset types. If the preset types are 10, then in this dimensional feature, a segmented audio corresponds to 10 scores, and a single normalization processing is performed on these 10 scores. Since there are 3 segmented audios, 3 normalization processes are required to obtain the prediction probabilities of each segmented audio under this classification dimension. Performing the above processing on the dimensional features corresponding to each classification dimension can obtain multiple prediction probabilities of each segmented audio under each classification dimension.

[0121] For example, refer to Table 1, which shows the multiple prediction probabilities of each segmented audio under each classification dimension.

[0122] Table 1

[0123]

[0124] Among them, in Table 1 The superscript represents that the serial number of the segmented audio is 1, and the subscript represents the first preset type under this music genre dimension. The meanings of other parameters are similar. Among them, there are 3 segmented audios, namely segmented audio 1, segmented audio 2, and segmented audio 3. Assume that there are 26 preset types in the music genre dimension, 18 preset types in the language dimension, and 8 preset types in the emotion dimension. Then a segmented audio has 26 prediction probabilities under the music genre dimension, and the sum of these 26 prediction probabilities is 1; a segmented audio has 18 prediction probabilities under the language dimension, and the sum of these 18 prediction probabilities is 1; a segmented audio has 8 prediction probabilities under the emotion dimension, and the sum of the 8 prediction probabilities is 1.

[0125] 140. For each of the classification dimensions, using the prediction probabilities of each segmented audio under the preset types corresponding to the classification dimension, determine the target type corresponding to each classification dimension.

[0126] The foregoing prediction obtains multiple prediction probabilities of each segmented audio under each classification dimension. Then, for each classification dimension, the prediction probabilities of each segmented audio under the preset type corresponding to its classification dimension can be obtained. For example, in Table 1 above, under the music genre dimension, the 3 segmented audios each correspond to 26 prediction probabilities, and one prediction probability corresponds to one preset type.

[0127] In some embodiments, when determining the target type corresponding to each classification dimension, it can also be that for each of the classification dimensions, multiple prediction probabilities corresponding to each of the preset types are obtained, where one prediction probability corresponds to one segmented audio; the average probability corresponding to each of the preset types is calculated by calculating the average of the multiple prediction probabilities corresponding to each of the preset types; the preset type corresponding to the maximum average probability is determined as the target type corresponding to the classification dimension.

[0128] For each preset type of a classification dimension, multiple prediction probabilities corresponding to that preset type can be obtained. For example, in Table 1 above, for the first preset type in the music genre dimension, there can be 3 prediction probabilities, which are respectively where one prediction probability corresponds to one segmented audio. For each preset type, the multiple prediction probabilities corresponding to that preset type can be averaged to obtain the average probability. For example, the 3 prediction probabilities of the first preset type in the music genre dimension are Then the average probability is Thus, under one classification dimension, the average probability corresponding to each preset type can be calculated.

[0129] Then, the preset type with the maximum average probability can be determined as the target type corresponding to the classification dimension. Thus, the target type corresponding to each classification dimension can be obtained. Therefore, when calculating the target type, the prediction probabilities of each segmented audio can be taken into account to improve the accuracy of determining the target type.

[0130] In other embodiments, when determining the target type corresponding to each classification dimension based on the prediction probabilities of each segmented audio under the preset type corresponding to the classification dimension, it can be that for each classification dimension, using the multiple prediction probabilities of each segmented audio corresponding to that classification dimension, the prediction type corresponding to each segmented audio is determined; from the prediction types corresponding to each segmented audio, the target type corresponding to the classification dimension is determined.

[0131] Among them, when determining the prediction type corresponding to each segmented audio, it can be that the preset type corresponding to the maximum probability among the multiple prediction probabilities corresponding to that segmented audio is determined as the prediction type. For example, in Table 1 above, under the music genre dimension, the maximum prediction probability corresponding to segmented audio 1 is That is, the second preset type under the music genre dimension is the prediction type, and the maximum prediction probability corresponding to the segmented audio 2 is That is, the first preset type under the music genre dimension is the prediction type, and the maximum prediction probability corresponding to the segmented audio 3 is That is, the fifth preset type under the music genre dimension is the prediction type.

[0132] Then, the frequency of occurrence of the prediction type can be determined, and the prediction type with the highest frequency is determined as the target type. For example, the prediction types corresponding to each segmented audio are the first preset type, the first preset type, and the second preset type in sequence. Since the highest frequency of the first preset type is 2, the first preset type can be directly determined as the target type.

[0133] If the frequencies of occurrence of each prediction type are the same, the prediction type with the maximum prediction probability can be selected as the target type. For example, the prediction types corresponding to each segmented audio are the first preset type with a prediction probability of 30%, the second preset type with a prediction probability of 40%, and the third preset type with a prediction probability of 70%. Then, the third preset type can be directly determined as the target type. In this way, the target type corresponding to each classification dimension can be determined.

[0134] 150. Based on the target type corresponding to each classification dimension, determine the classification result of the audio to be classified.

[0135] After obtaining the target type corresponding to each classification dimension, the classification result of the audio to be classified can be determined based on the target type corresponding to each classification dimension. For example, all the target types can be directly determined as the classification result of the audio to be classified.

[0136] It should be noted that the above feature extraction parameters and dimension parameters are both obtained by training the classification model. Among them, the feature extraction parameters and dimension parameters can be obtained through the following steps: obtain a training sample set, where the training sample set includes multiple sample audios and the sample type corresponding to each sample audio, and the sample type is one of the preset types under multiple classification dimensions; determine the sample audio to be used corresponding to each classification dimension from the training sample set; based on the initial feature extraction parameters, perform feature extraction on the sample audio to be used to obtain the sample audio features corresponding to the classification dimension; based on the initial dimension parameters corresponding to the classification dimension and the sample audio features corresponding to the classification dimension, calculate the prediction type of the sample audio to be used under the classification dimension; according to the prediction type of the sample audio to be used and the sample type corresponding to the sample audio to be used, adjust the initial feature extraction parameters and the initial dimension parameters to obtain the feature extraction parameters and dimension parameters.

[0137] The training sample set refers to the data set used to train the classification model. The training sample set may include multiple sample audios and the sample types corresponding to each sample audio. Among them, the sample type may be one of the preset types under multiple classification dimensions.

[0138] After obtaining the training sample set, the training sample set can be used to train the classification model to obtain feature extraction parameters and dimension parameters. For example, refer to Figure 1d , which shows a schematic diagram of training the classification model. Among them, the classification model may include a feature extraction layer and multiple classification layers, where one classification layer corresponds to one classification dimension.

[0139] Among them, the input data of the classification model can be selected from the training sample set according to the classification dimension. The selected sample audio is used as the sample audio to be used, and the sample type corresponding to the sample audio to be used is used as the sample type to be used. Among them, in order to ensure that the trained classification model can infer the types under multiple classification dimensions at one time, the number of sample audios to be used can be the same as the number of classification dimensions. For example, if the number of classification dimensions is 3, then there are also 3 sample audios to be used, and the types to be used of the sample audios to be used correspond to the 3 classification dimensions respectively.

[0140] Optionally, the training sample set can be divided into multiple groups according to the classification dimension, one group corresponding to one classification dimension, and one sample audio is selected from the sample audios in each classification dimension as the sample audio to be used. Then 3 sample audios to be used can be obtained, and these 3 sample audios to be used are input into the classification model together.

[0141] The feature extraction layer of the classification model has initial feature extraction parameters, and each classification layer has corresponding initial dimension parameters. After the sample to be used is input into the feature extraction layer, the initial feature extraction parameters can perform feature extraction processing on the sample to be used, and the process of this feature extraction processing is similar to that in the aforementioned step 120. That is, the initial feature extraction parameters may include initial depth convolution parameters and initial pointwise convolution parameters, and finally the sample audio features corresponding to each sample audio to be used can be extracted.

[0142] Since each sample audio feature to be used corresponds to a classification dimension, and one classification dimension corresponds to a classification layer, the sample audio feature to be used enters the corresponding classification layer to calculate the predicted type of the sample audio to be used. For example, as Figure 1cAmong them, the sample audio features in the music style dimension enter the classification layer corresponding to the music style dimension, the sample audio features in the language dimension enter the classification layer corresponding to the language dimension, and the sample audio features in the emotion dimension enter the classification layer corresponding to the emotion dimension. The processing of a similar classification layer is similar to the processing process of the aforementioned step 130, and reference can be made to the description of the corresponding part above. The difference lies in that after calculating multiple prediction probabilities corresponding to the sample audio to be used in the corresponding classification dimension, the preset type corresponding to the largest prediction probability is directly selected as the prediction type.

[0143] Thus, a prediction type can be calculated for each sample audio to be used corresponding to each classification dimension. Then, based on the prediction type of the sample audio to be used and the sample type of the sample audio to be used, the initial feature extraction parameters and the initial dimension parameters can be adjusted to obtain the feature extraction parameters and the dimension parameters. Optionally, for each classification dimension, a dimension loss function can be constructed by using the prediction type of the sample audio to be used and the sample type corresponding to the sample audio to be used; based on the dimension weight corresponding to the classification dimension and the dimension loss function corresponding to the classification dimension, an objective loss function can be constructed; according to the objective loss function, the initial feature extraction parameters and the initial dimension parameters are adjusted until the objective loss function meets the preset conditions, and the feature extraction parameters and the dimension parameters are obtained.

[0144] For each classification dimension, a dimension loss function can be established according to the prediction type and the sample type of the sample audio to be used corresponding to it. This dimension loss function can be used to measure the difference between the prediction type and the sample type in this classification dimension, and this dimension loss function can be selected according to actual needs. For example, in the embodiments of the present application, it can be a cross-entropy loss function.

[0145] If there are the above three classification dimensions, three dimension loss functions can be correspondingly established, and an objective loss function for the entire training is constructed based on the dimension weight corresponding to the classification dimension and the dimension loss function corresponding to the classification dimension. This objective loss function can be expressed by the following formula:

[0146] Loss Toyal =α×Loss Genre +β×Loss Lang +γ×Loss Mood ;

[0147] Among them, Loss Total represents the objective loss function; Loss Genre represents the dimension loss function corresponding to the music style dimension; Loss Lang represents the dimension loss function corresponding to the language dimension; Loss MoodDenote the dimensional loss function corresponding to the emotion dimension; α denotes the dimensional weight corresponding to the music style dimension; β denotes the dimensional weight corresponding to the language dimension; γ denotes the dimensional weight corresponding to the emotion dimension.

[0148] Among them, the dimensional weights corresponding to each classification dimension are default to be the same. However, in the actual training process, the dimensional weights can be adjusted according to the difficulty levels of training for different classification dimensions. The initial feature extraction parameters and initial dimensional parameters can be continuously adjusted according to the value of the target loss function. When the target loss function converges, the feature extraction parameters and dimensional parameters can be obtained.

[0149] In some embodiments, in order to ensure that the data in the training sample set conforms to the actual classification scenario, multiple training sample sets can be determined using the data in the actual classification scenario. For example, the classification scenario can be to classify videos or audio in a live broadcast room. Then, the relevant information of the video or live broadcast room can be used to obtain the training sample set. As an implementation, in order to expand the number of sample audios for better training of the classification model, the initial sample audio can be obtained according to the multimedia information corresponding to the preset multimedia; the initial sample audio is enhanced in the time domain to obtain the time-domain enhanced sample audio; the initial sample audio is enhanced in the frequency domain to obtain the frequency-domain enhanced sample audio; the initial sample audio, the time-domain enhanced sample audio, and the frequency-domain enhanced sample audio are determined as the sample audios of the training sample set; the sample types corresponding to the sample audios are determined according to the preset types corresponding to multiple classification dimensions.

[0150] Among them, the preset multimedia refers to the multimedia used for transmitting audio, which can include videos, live videos in a live broadcast room, etc. Usually, the multimedia platform can provide some audios for users to use. However, users can also add the audios not provided by the platform to the preset multimedia according to actual needs. That is, the audios in the preset multimedia can be those provided by the platform in advance, and these audios can be directly obtained from the media library of the multimedia platform. For the audios not provided by the platform, they can be obtained through the multimedia information corresponding to the preset multimedia.

[0151] In some embodiments, the multimedia information can include multimedia description information and multimedia content. When obtaining the initial sample audio based on the multimedia information, entity recognition can be performed on the multimedia description information to extract the audio entities in the multimedia content; audio content recognition is performed on the multimedia content to extract the audio content features in the multimedia content; the initial sample audio is determined from multiple preset audios in the preset audio library using the audio entities and the audio content features.

[0152] Multimedia description information refers to information used to describe the preset multimedia. For example, the title of a video, the brief introduction of a video, the description of the account that publishes the video, etc. Multimedia content refers to the actual content of the preset multimedia. For example, the actual video content, the actual live video, etc.

[0153] Reference can be made to Figure 1e , which shows a schematic diagram of obtaining the initial sample audio. Named entity recognition can be performed on the multimedia description information to determine the audio entity therein, and based on this audio entity, the audio data corresponding to the audio entity can be obtained from the Internet. Among them, the audio entity refers to entity information related to music, including music names, singer names, lyrics, etc. The preset audio library can refer to the audio library composed of all the audio on the Internet, and the audio on the Internet is the preset audio. The initial sample audio can be determined from the preset audio using the audio entity.

[0154] For example, if the audio entity is the singer name, then all the preset audio of this singer can be obtained from the preset audio library as the initial sample audio. Similarly, when obtaining the audio corresponding to the music entity from the Internet, the initial sample audio can be obtained using both the multimedia content and the description information of the multimedia. After obtaining the initial sample audio from the preset audio library, the relevant information of the initial sample audio can be obtained at the same time. For example, the song name, singer, album, lyrics, release date, language, music style, etc. corresponding to the audio. The sample type of the initial sample audio can be obtained based on this information, and the preset type under the classification dimension can be sorted out using this information for subsequent use.

[0155] The multimedia content may include audio. Audio content recognition can be performed on the multimedia content to extract the audio content features in the multimedia content. Among them, the audio can be first extracted from the multimedia content, and the extracted audio is used as the initial sample audio. Then, audio content recognition is performed on the extracted audio to obtain the audio content features. Then, based on this audio content feature, the relevant information of the initial sample audio is obtained from the preset database. Optionally, the preset audio content features corresponding to each preset audio can be determined in the same content recognition manner, and the relevant information of the preset audio corresponding to the maximum similarity can be obtained by calculating the similarity between each preset audio content feature and the audio content feature. Of course, the preset audio corresponding to the maximum similarity can also be used as the initial sample audio.

[0156] In some embodiments, the initial sample audio can be processed by data augmentation to improve the richness of the training sample set, avoid overfitting of the model to a certain classification dimension, and thus improve the generalization performance of the trained model. Among them, the initial sample audio can be augmented in the time domain to obtain the time-domain augmented sample audio. This step may include: extracting the time-domain features of the initial sample audio in the time domain; performing a transformation process on the time-domain features to obtain the transformed time-domain features; and obtaining the time-domain augmented sample audio based on the transformed time-domain features.

[0157] Among them, the time-domain features may include amplitude, phase, frequency, background sound, noise, echo, etc. Amplitude refers to the amplitude size of the audio signal at different time points. In the time domain, the waveform height of the audio signal represents the amplitude, which can be calculated by sampling and quantifying the audio signal. Phase refers to the starting point position of the audio signal waveform, which can be determined by comparing the time offset of two waveforms. Frequency refers to the repeating pattern in the audio signal, and the frequency can be determined by calculating the period of the waveform. Background sound refers to other sounds in the environment, and noise refers to unwanted audio signals, such as the noise on the street. Echo refers to the delayed signal that reaches the listener's ear after the sound is reflected in the environment.

[0158] By performing a transformation process on the time-domain features, the transformed time-domain features can be obtained. For example, adjusting the amplitude size can simulate the sense of distance of the sound to simulate different near and far effects. Generally, reducing the amplitude can simulate the effect of the sound moving away, and increasing the amplitude can simulate the effect of the sound approaching. Adjusting the phase can simulate different spatial effects. For example, in a stereo sound field, by adjusting the left and right channels, the audio can sound like it is coming from different directions. Of course, the phase angle can also be adjusted to simulate the effect of the audio being emitted at different angles. Using filters can change the frequency components of the audio. For example, a low-pass filter allows lower-frequency signals to pass through and can be used to simulate the low-frequency attenuation effect in a distant environment; a band-pass filter allows frequency signals within a certain range to pass through and can be used to simulate specific spatial effects; a high-pass filter allows higher-frequency signals to pass through and can be used to simulate the low-frequency attenuation in a close-range effect.

[0159] To make the audio more close to the actual audio classification scenario, background sound, noise, echo, etc. can also be added to the initial sample audio. By adding background sound, the realism of the audio can be enhanced. By adding noise, the audio in different environments can be simulated. By adding echo, the reverberation effect in different spaces can be simulated, making the audio sound more stereo and surround.

[0160] After transforming the time-domain features, the transformed time-domain features can be obtained. Based on the transformed time-domain features, time-domain enhanced samples can be obtained. That is, by changing one or more of the amplitude, phase, frequency, background sound, noise, and echo of the initial sample audio, the corresponding time-domain enhanced sample audio can be generated. Performing data augmentation in the time domain can simulate a more realistic scenario, but the data processing is time-consuming. To improve the efficiency of data augmentation, the steps of performing data augmentation processing in the time domain can be executed by a GPU.

[0161] Perform enhancement processing on the initial sample audio in the frequency domain to obtain frequency-domain enhanced sample audio. Optionally, when obtaining the frequency-domain enhanced sample audio, the initial sample audio can be transformed into a spectral feature map; the spectral feature map is perturbed according to a mask matrix to obtain a perturbed feature map; linear interpolation processing is performed on the spectral feature maps corresponding to different initial sample audios to obtain an interpolated feature map; the frequency-domain enhanced sample audio is generated according to the perturbed feature map and the interpolated feature map.

[0162] Among them, the spectral feature map is a two-dimensional matrix evenly distributed in time and frequency, usually used to represent the frequency-domain features of an audio signal. The spectral feature map can be a linear spectrogram, a Mel spectrogram, etc. In the embodiments of the present application, the spectral feature map is a Mel spectrogram. Optionally, when transforming the initial sample audio into a spectral feature map, the initial sample audio signal corresponding to the initial sample audio can be obtained; the initial sample audio signal is windowed to obtain a plurality of processing windows; for each processing window, the audio signal in the processing window is subjected to a short-time Fourier transform to obtain a window spectral feature map corresponding to the processing window; the window spectral feature map is compressed to obtain a compressed window spectral feature map; the compressed window spectral feature map is logarithmically amplitude-normalized to obtain a normalized spectral feature map; the normalized spectral feature maps of each processing window are spliced to obtain the spectral feature map.

[0163] Among them, when compressing the window spectral feature map, a set of filters can be used to compress continuous frequency intervals into several fixed frequency intervals. For example, the filters can be Mel filters, and the fixed frequency intervals can be Mel frequency intervals. The set of filters can be designed according to the frequency distribution of the Mel scale, with wider filters for lower frequencies and narrower filters for higher frequencies. During the compression process, each frequency point of the window spectral feature map is convolved with the set of filters to obtain the compressed window spectral feature map. Then, the amplitude in the window spectral feature map is normalized using a logarithmic function to obtain the normalized spectral feature map. Finally, the normalized spectral feature maps of each processing window are spliced to convert the initial sample audio into a spectral feature map.

[0164] The mask matrix can be used to occlude part of the information in the spectral feature map. Thus, a random mask matrix can be generated on the spectral feature map to obtain a perturbed feature map, where the size and generation position of the mask matrix are both random. For example, refer to Figure 1f , which shows a schematic diagram of the perturbed spectral feature map. Among them, the first one is the spectral feature map, and the remaining two are both perturbed feature maps. Compared with the spectral feature map, part of the information in the perturbed feature map is occluded.

[0165] The aforementioned spectral feature map corresponding to the initial sample audio can be obtained. The spectral feature maps corresponding to at least two different initial sample audios can be obtained, and then linear interpolation processing is performed on these spectral feature maps to generate an interpolated feature map. For example, a corresponding number of weight coefficients can be randomly generated according to the number of spectral feature maps, and the sum of these weight coefficients is 1. The spectral feature maps are weighted and summed according to the weight coefficients to obtain an interpolated feature map. If two spectral feature maps are denoted as x and y, and their corresponding weight coefficients are a and b, then the interpolated feature map is a*x + b*y. Finally, the interpolated feature map can be converted into audio, and the perturbed feature map can be converted into audio, so as to obtain the frequency-domain enhanced sample audio.

[0166] The above initial sample audio, time-domain enhanced sample audio, and frequency-domain enhanced sample audio can all be used as sample audios in the training sample set. Then, by setting the corresponding sample type for each sample audio in the training sample set, the training sample set can be obtained.

[0167] The preset type corresponding to the classification dimension is obtained based on the analysis of the relevant information of the initial sample audio. Then, when obtaining the initial sample audio, the relevant information of the initial sample audio can be directly used to determine the sample type corresponding to the initial sample audio. As for the sample type of the time-domain enhanced sample audio, the sample type of the corresponding initial sample audio can be directly used. For the frequency-domain enhanced sample audio, if only the mask matrix is used for enhancement, its sample type is the same as that of the corresponding initial sample audio; if it is enhanced by interpolation, its sample type can be determined according to the weight coefficients, that is, the sample type of the enhanced sample audio is the same as that of the initial sample audio with a larger weight coefficient. Thus, the sample type corresponding to each sample audio can be accurately determined. Finally, the training sample set can be used to train the classification model to obtain the feature extraction parameters and dimension parameters.

[0168] The audio classification solution provided by the embodiments of the present application can be applied to various scenarios of multi-classifying audio. For example, taking the audio in 3 classification dimensions as an example, by adopting the solution provided by the embodiments of the present application, the types in 3 classification dimensions can be obtained by one inference of the audio, reducing the data processing volume and improving the audio classification efficiency.

[0169] Through the audio classification method provided by the embodiments of the present application, at least one segmented audio is extracted from the audio to be classified, and the segmented audio features corresponding to each segmented audio are determined using the same feature extraction parameters. Then, the target types of the segmented audio in multiple classification dimensions are predicted using the dimension parameters of multiple classification dimensions, and the classification result of the audio to be classified is obtained. Extracting at least one segmented audio can ensure accurate classification while reducing the data processing volume. Using the same feature extraction parameters to extract features from the segmented audio and then calculating the prediction probabilities using different dimension parameters can determine the target types of the audio to be classified in multiple classification dimensions with a single inference, avoiding repeated inferences, reducing the data processing volume for classification, and thus improving the efficiency of audio classification. At the same time, when training the classification model as described above, the training sample set is expanded through data augmentation, enabling the classification model to have good generalization ability. Moreover, a lightweight model is selected for the classification model, which can balance the accuracy and efficiency of classification.

[0170] The method described in the above embodiments will be further described in detail below.

[0171] In this embodiment, taking the example of using a classification model to infer the types of the audio to be classified in multiple classification dimensions in one go, the method of the embodiments of the present application will be described in detail.

[0172] As Figure 2a shown, the specific process of an audio classification method is as follows:

[0173] 210. Extract at least one segmented audio from the audio to be classified.

[0174] 220. Input at least one segmented audio into the feature extraction layer to obtain the segmented audio features corresponding to each segmented audio.

[0175] 230. Input each segmented audio feature into multiple classification layers respectively to predict multiple prediction probabilities of each segmented audio under each classification layer.

[0176] 240. For each classification layer, use the multiple prediction probabilities of each segmented audio under the classification layer to determine the target type under each classification layer;

[0177] 250. Determine the classification result of the audio to be classified according to the target types output by each classification layer.

[0178] The content in the above steps 210 to 250 can refer to the corresponding descriptions in the foregoing embodiments. To describe the audio classification method in more detail, the process from training the classification model to using the classification model for audio classification will be described in detail below. This process can refer to Figure 2b , which shows the overall process schematic diagram of the audio classification method.

[0179] Among them, the construction of the classification system can be obtained by statistically analyzing the audio reported by video accounts every day. A total of three classification dimensions are sorted out, namely the music genre dimension, the language dimension, and the emotion dimension. The music genre dimension includes 26 preset types, namely: Country, New Age, Metal, Classical, Children's Music, Punk, Hip-Hop Rap, Reggae, Dance, Alternative, Rock, Easy Listening, Jazz, Latin, Rhythm and Blues, Acoustic, World Music, Electronic, Blues, Accompaniment, Chinese Folk Music, Pop, Folk, Ancient Style, Grassland, Chinese Style. The emotion dimension includes 8 preset types, namely: Sweet, Lonely, Happy, Healing, Missing, Sad, Quiet, Other. The language dimension includes 18 preset types such as English, Pure Music, Other, etc.

[0180] In the data cleaning and processing stage, a training sample set needs to be created, and each sample audio in the training sample set needs to be labeled with the corresponding sample type. The sample audio can be obtained by collecting the audio in the video account. At the same time, analyze the titles, descriptions, etc. of the videos published by the video account to identify the music entities in them, such as singers, song names, genres, etc., in order to obtain more sample audio. And in order to facilitate subsequent training of the model, data augmentation processing can be carried out in various ways to obtain more sample audio. Since ordinary people cannot accurately identify the type of audio based on the audio, when obtaining the sample audio, relevant information of the audio can be obtained at the same time, such as song name, singer, album, lyrics, release date, language, genre, etc., so as to determine the sample type of the sample audio.

[0181] Among them, when performing data augmentation, it can be in the time dimension, simulating music at different distances, in different spaces, and under noisy environments, mainly including amplitude change, phase change, angle change, background sound, noise, low-pass, band-pass, high-pass filters, echo, etc. Performing data augmentation in the time domain can simulate a more realistic scenario. In order to improve the efficiency of performing data augmentation processing in the time domain, the data augmentation in the time domain can be transferred to the GPU for processing, and the processing speed can be rapidly improved. In addition to performing data augmentation in the time domain, the spectral feature map of the audio can be extracted, and the masking method can be used, mainly masking the time dimension and the frequency dimension, or linearly interpolating different spectral feature maps to generate a new spectral feature map. In this way, the sample audio in the obtained training sample set is more abundant, which is beneficial to improving the generalization performance of the model.

[0182] When training the classification model, the sample audio under the music genre dimension, the language dimension, and the emotion dimension, that is, the three sample audio, are simultaneously input into the network for training. Among them, the sample audio data is uniformly processed into audio segments with a sampling rate of 8000Hz, single channel, a duration of 15s, and a format of wav.

[0183] Due to the huge demand for audio classification and the large number of calls to the audio model, in order to reduce the computing power for audio classification, a lightweight model can be selected and modified according to actual training needs. For example, MobileNet can be selected as the classification model. To make the classification model adapt to data with multiple different classification dimensions, the fully connected layer of MobileNet can be removed, and three new fully connected layers can be added, corresponding to the classification of music genre dimension, language dimension, and emotion dimension respectively, with outputs of 26 dimensions, 18 dimensions, and 7 dimensions respectively. These three fully connected layers are all connected to the output of the feature extraction layer, so that they share the features of a backbone network, and the model does not need to repeat the inference during reasoning. The output of the feature extraction layer is dimensionally reduced, and the final model feature dimension is 64 dimensions, which greatly improves the inference speed compared to the default 1024 dimensions.

[0184] Since the classification in three classification dimensions is trained simultaneously, the objective function for training optimization is to optimize the classification losses of the three classification dimensions simultaneously. During training, the weights of the loss functions for different classification dimensions are default to be the same. During the training process, the weights of each loss function can be adjusted according to the difficulty of different task training. Through the above process, the classification model is trained to obtain the trained classification model.

[0185] Finally, the trained classification model can be directly used for inference to obtain the types of the audio to be classified in three classification dimensions. Specifically, the audio to be classified can be transcoded into audio data with a sampling rate of 8000Hz, single channel, and wav format, and then at least one segmented audio can be extracted from it for prediction. When the total duration of the audio to be classified is greater than 45s, the audio to be classified can be split into multiple sub-audios according to 15s, and then three sub-audios at the front, middle, and back positions in the entire audio to be classified can be selected as input data and output to the classification model for classification.

[0186] To further improve the classification speed, the classification model can be converted into the onnx format for inference. Compared with the traditional pytorch model, the inference speed is greatly improved. Since the classification model can parallelly predict the types in three classification dimensions, finally the classification model can output the target types in 3 classification dimensions as the final classification result of the audio to be classified.

[0187] As can be seen from the above, when training the classification model, the training sample set is expanded through data augmentation, enabling the classification model to have better generalization ability. Moreover, the classification model selects a lightweight model, which can balance the accuracy and efficiency of classification. When using the classification model for classification, the same feature extraction parameters are used to extract features from the segmented audio, and then different dimension parameters are used to calculate the prediction probability, so that the target type of the audio to be classified under multiple classification dimensions can be determined with a single inference, avoiding repeated inferences, reducing the data processing volume of classification, and thus improving the efficiency of audio classification.

[0188] To better implement the above method, an embodiment of the present application further provides an audio classification device. This audio classification device can be specifically integrated in an electronic device, which can be a device such as a terminal, a server, etc. Among them, the terminal can be a device such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.; the server can be a single server or a server cluster composed of multiple servers.

[0189] For example, in this embodiment, taking the audio classification device being specifically integrated in the server as an example, the method of the embodiment of the present application will be described in detail.

[0190] For example, as Figure 3 shown, the audio classification device may include a segmentation unit 310, a feature extraction unit 320, a probability prediction unit 330, a type determination unit 340, and a classification unit 350, as follows:

[0191] (1) Segmentation unit 310

[0192] It is used to extract at least one segmented audio from the audio to be classified.

[0193] In some embodiments, the segmentation unit 310 further includes:

[0194] A splitting sub-unit, which is used to perform splitting processing on the audio to be classified based on a preset duration to obtain at least one sub-audio, and the duration of the sub-audio is the preset duration;

[0195] A first determination sub-unit, which is used to determine the preset number of segmented audios from the at least one sub-audio if the number of sub-audios is greater than the preset number;

[0196] A second determination sub-unit, which is used to determine each sub-audio as a segmented audio if the number of sub-audios is less than or equal to the preset number.

[0197] (2) Feature extraction unit 320

[0198] To determine the segment audio features corresponding to each of the segment audios based on feature extraction parameters.

[0199] In some embodiments, the feature extraction parameters include depth convolution parameters and pointwise convolution parameters, and the feature extraction unit 320 further includes:

[0200] A depth convolution subunit, configured to perform depth convolution processing on each of the segment audios by using the depth convolution parameters to obtain depth audio features corresponding to the segment audios;

[0201] A pointwise convolution subunit, configured to perform pointwise convolution processing on the depth audio features to obtain segment audio features corresponding to the segment audios.

[0202] (III) Probability prediction unit 330

[0203] To predict multiple prediction probabilities of each of the segment audios under each of the classification dimensions according to dimension parameters corresponding to multiple classification dimensions and each of the segment audio features, where each of the classification dimensions corresponds to multiple preset types, and the prediction probabilities correspond to the preset types one by one.

[0204] In some embodiments, the probability prediction unit 330 further includes:

[0205] An acquisition subunit, configured to acquire multiple preset types corresponding to each of the classification dimensions;

[0206] A linear transformation subunit, configured to perform linear transformation on each of the segment audio features by using dimension parameters corresponding to multiple classification dimensions to obtain dimension features corresponding to each of the classification dimensions;

[0207] A normalization subunit, configured to perform normalization processing on the dimension features corresponding to each classification dimension to obtain multiple prediction probabilities of each of the segment audios under each of the classification dimensions.

[0208] (IV) Type determination unit 340

[0209] To determine, for each of the classification dimensions, a target type corresponding to each classification dimension by using the prediction probabilities of each of the segment audios under the preset types corresponding to the classification dimensions.

[0210] In some embodiments, the type determination unit 340 further includes:

[0211] A probability acquisition subunit, configured to acquire, for each of the classification dimensions, multiple prediction probabilities corresponding to each of the preset types, where one prediction probability corresponds to one segment audio;

[0212] A mean calculation subunit, configured to calculate the mean of multiple prediction probabilities corresponding to each of the preset types, so as to obtain a mean probability corresponding to each of the preset types;

[0213] A determination subunit, configured to determine the preset type corresponding to the maximum mean probability as the target type corresponding to the classification dimension.

[0214] (V) Classification unit 350

[0215] configured to determine the classification result of the audio to be classified based on the target type corresponding to each classification dimension.

[0216] In some embodiments, the audio classification device further includes a training unit, configured to:

[0217] Obtain a training sample set, where the training sample set includes multiple sample audios and a sample type corresponding to each sample audio, and the sample type is one of the preset types under the multiple classification dimensions;

[0218] Determine, from the training sample set, the sample audio to be used corresponding to each classification dimension;

[0219] Extract features from the sample audio to be used based on initial feature extraction parameters, so as to obtain sample audio features corresponding to the classification dimension;

[0220] Calculate the predicted type of the sample audio to be used under the classification dimension based on the initial dimension parameters corresponding to the classification dimension and the sample audio features corresponding to the classification dimension;

[0221] Adjust the initial feature extraction parameters and the initial dimension parameters according to the predicted type of the sample audio to be used and the sample type corresponding to the sample audio to be used, so as to obtain feature extraction parameters and dimension parameters.

[0222] In some embodiments, the training unit is further configured to:

[0223] For each classification dimension, construct a dimension loss function by using the predicted type of the sample audio to be used and the sample type corresponding to the sample audio to be used;

[0224] Construct an objective loss function based on the dimension weight corresponding to the classification dimension and the dimension loss function corresponding to the classification dimension;

[0225] Adjust the initial feature extraction parameters and the initial dimension parameters according to the objective loss function until the objective loss function meets a preset condition, so as to obtain the feature extraction parameters and the dimension parameters.

[0226] In some embodiments, the training unit is further configured to:

[0227] Obtain an initial sample audio according to the multimedia information corresponding to the preset multimedia;

[0228] Perform enhancement processing on the initial sample audio in the time domain to obtain a time-domain enhanced sample audio;

[0229] Perform enhancement processing on the initial sample audio in the frequency domain to obtain a frequency-domain enhanced sample audio;

[0230] Determine the initial sample audio, the time-domain enhanced sample audio, and the frequency-domain enhanced sample audio as the sample audio of the training sample set;

[0231] Determine the sample type corresponding to the sample audio according to the preset types corresponding to multiple classification dimensions.

[0232] In some embodiments, the multimedia information includes multimedia content and multimedia description information, and the training unit is further configured to:

[0233] Perform entity recognition on the multimedia description information to extract audio entities in the multimedia content;

[0234] Perform audio content recognition on the multimedia content to extract audio content features in the multimedia content;

[0235] Use the audio entity and the audio content features to determine the initial sample audio from multiple preset audios in a preset audio library.

[0236] In some embodiments, the training unit is further configured to:

[0237] Extract time-domain features of the initial sample audio in the time domain;

[0238] Perform transformation processing on the time-domain features to obtain transformed time-domain features;

[0239] Obtain a time-domain enhanced sample audio based on the transformed time-domain features.

[0240] In some embodiments, the training unit is further configured to:

[0241] Convert the initial sample audio into a spectral feature map;

[0242] Perform perturbation processing on the spectral feature map according to a mask matrix to obtain a perturbed feature map;

[0243] Perform linear interpolation processing on the spectral feature maps corresponding to different initial sample audios to obtain an interpolated feature map;

[0244] Generate the frequency-domain enhanced sample audio according to the perturbed feature map and the interpolated feature map.

[0245] In specific implementation, each of the above units can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments and will not be elaborated herein.

[0246] As can be seen from the above, the audio classification device of this embodiment can extract features from segmented audio based on the same feature extraction parameters, and then calculate the prediction probability using different dimensional parameters, so as to determine the target type of the audio to be classified under multiple classification dimensions with one inference, avoiding repeated inference, reducing the data processing volume of classification, and thus improving the efficiency of audio classification.

[0247] The embodiment of the present application also provides an electronic device, which can be a device such as a terminal or a server. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a notebook computer, a personal computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.

[0248] In some embodiments, the audio classification device can also be integrated in multiple electronic devices. For example, the audio classification device can be integrated in multiple servers, and the audio classification method of the present application can be implemented by multiple servers.

[0249] In this embodiment, the electronic device of this embodiment being a server will be described in detail as an example. For example, as Figure 4 shown, it shows a schematic structural diagram of the electronic device involved in the embodiment of the present application. Specifically:

[0250] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input module 404, a communication module 405 and other components. Those skilled in the art can understand that Figure 4 the structural diagram of the electronic device shown in

[0251] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and by invoking the data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby performing an overall detection of the electronic device. In some embodiments, the processor 401 may include one or more processing cores; in some embodiments, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.

[0252] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0253] The electronic device further includes a power supply 403 that powers each component. In some embodiments, the power supply 403 may be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0254] The electronic device may further include an input module 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0255] The electronic device may further include a communication module 405. In some embodiments, the communication module 405 may include a wireless module. The electronic device can perform short-range wireless transmission through the wireless module of the communication module 405, thereby providing users with wireless broadband Internet access. For example, the communication module 405 can be used to help users send and receive emails, browse web pages, and access streaming media, etc.

[0256] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0257] Extract at least one segmented audio from the audio to be classified;

[0258] Determine the segmented audio features corresponding to each of the segmented audios based on the feature extraction parameters;

[0259] Predict multiple prediction probabilities of each of the segmented audios under each of the classification dimensions according to the dimension parameters corresponding to multiple classification dimensions and each of the segmented audio features, where each of the classification dimensions corresponds to multiple preset types, and the prediction probabilities correspond one-to-one with the preset types;

[0260] For each of the classification dimensions, use the prediction probabilities of each of the segmented audios under the preset types corresponding to the classification dimensions to determine the target type corresponding to each classification dimension;

[0261] Determine the classification result of the audio to be classified based on the target type corresponding to each classification dimension.

[0262] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0263] As can be seen from the above, the electronic device provided by the embodiments of the present application can extract features from segmented audios based on the same feature extraction parameters, and then calculate prediction probabilities using different dimension parameters, so as to determine the target types of the audio to be classified under multiple classification dimensions with one inference, avoiding repeated inferences, reducing the data processing volume of classification, and thus improving the efficiency of audio classification.

[0264] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0265] For this reason, the embodiments of the present application provide a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any of the audio classification methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:

[0266] Extract at least one segmented audio from the audio to be classified;

[0267] Determine the segment audio features corresponding to each of the segment audios based on the feature extraction parameters;

[0268] According to the dimension parameters corresponding to multiple classification dimensions and each of the segment audio features, predict multiple prediction probabilities of each of the segment audios under each of the classification dimensions, where each of the classification dimensions corresponds to multiple preset types, and the prediction probabilities correspond one-to-one to the preset types;

[0269] For each of the classification dimensions, use the prediction probabilities of each of the segment audios under the preset types corresponding to the classification dimensions to determine the target type corresponding to each classification dimension;

[0270] Determine the classification result of the audio to be classified based on the target type corresponding to each classification dimension.

[0271] Wherein, the storage medium may include: read only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0272] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in various alternative implementations in the aspects of audio classification or model training provided in the above embodiments.

[0273] Since the instructions stored in the storage medium can execute the steps in any of the audio classification methods provided in the embodiments of the present application, the beneficial effects achievable by any of the audio classification methods provided in the embodiments of the present application can be realized. For details, see the previous embodiments and will not be elaborated herein.

[0274] The above has introduced in detail an audio classification method, device, electronic device, storage medium and program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An audio classification method, characterized in that: The method comprises: Extracting at least one audio segment from the audio to be classified; Determine, based on the feature extraction parameters, a segmented audio feature corresponding to each of the segmented audios; According to the dimension parameters corresponding to the multiple classification dimensions and each of the segmented audio features, predict multiple prediction probabilities of each of the segmented audio under each of the classification dimensions, wherein each of the classification dimensions corresponds to multiple preset types, and the prediction probabilities correspond to the preset types one by one; For each of the classification dimensions, using the predicted probability of each of the segmented audio under the preset type corresponding to the classification dimension, determine the target type corresponding to each of the classification dimensions; Based on the target type corresponding to each classification dimension, a classification result of the audio to be classified is determined.

2. The method according to claim 1, characterized in that The feature extraction parameters include a depth convolution parameter and a point-by-point convolution parameter, and determining the segmented audio feature corresponding to each segmented audio based on the feature extraction parameters includes: For each of the segmented audios, using the deep convolution parameters, performing deep convolution processing on the segmented audios to obtain deep audio features corresponding to the segmented audios; The deep audio features are subjected to point-by-point convolution processing using the point-by-point convolution parameters to obtain segmented audio features corresponding to the segmented audio.

3. The method according to claim 1, characterized in that The step of predicting multiple prediction probabilities of each of the audio segments under each of the classification dimensions according to the dimension parameters corresponding to the multiple classification dimensions and each of the audio segment features comprises: Obtaining multiple preset types corresponding to each of the classification dimensions; Using dimensional parameters corresponding to multiple classification dimensions, linearly transform each of the segmented audio features to obtain dimensional features corresponding to each of the classification dimensions; The dimensional features corresponding to each classification dimension are normalized to obtain multiple prediction probabilities of each segmented audio under each classification dimension.

4. The method according to claim 1, characterized in that The step of determining the target type corresponding to each classification dimension by using the predicted probability of each segmented audio under the preset type corresponding to the classification dimension for each classification dimension includes: For each of the classification dimensions, obtaining a plurality of prediction probabilities corresponding to each of the preset types, wherein one prediction probability corresponds to one segmented audio; Calculate the mean of multiple predicted probabilities corresponding to each of the preset types to obtain the mean probability corresponding to each of the preset types; The preset type corresponding to the maximum mean probability is determined as the target type corresponding to the classification dimension.

5. The method according to claim 1, characterized in that The step of extracting at least one audio segment from the audio to be classified comprises: Segmenting the audio to be classified based on a preset duration to obtain at least one sub-audio, where the duration of the sub-audio is the preset duration; If the number of the sub-audios is greater than a preset number, determining the preset number of segmented audios from the at least one sub-audio; If the number of the sub-audios is less than or equal to the preset number, each of the sub-audios is determined as a segmented audio.

6. The method according to claim 1, characterized in that The method further comprises: Acquire a training sample set, wherein the training sample set includes a plurality of sample audios and a sample type corresponding to each of the sample audios, wherein the sample type is one of the preset types under the plurality of classification dimensions; Determine, from the training sample set, sample audio to be used corresponding to each classification dimension; Based on the initial feature extraction parameters, feature extraction is performed on the sample audio to be used to obtain sample audio features corresponding to the classification dimension; Calculate the predicted type of the sample audio to be used under the classification dimension based on the initial dimension parameter corresponding to the classification dimension and the sample audio feature corresponding to the classification dimension; According to the predicted type of the sample audio to be used and the sample type corresponding to the sample audio to be used, the initial feature extraction parameters and the initial dimension parameters are adjusted to obtain feature extraction parameters and dimension parameters.

7. The method according to claim 6, characterized in that The adjusting the initial feature extraction parameters and the initial dimension parameters according to the predicted type of the sample audio to be used and the sample type corresponding to the sample audio to be used to obtain the feature extraction parameters and the dimension parameters includes: For each classification dimension, construct a dimension loss function using the predicted type of the sample audio to be used and the sample type corresponding to the sample audio to be used; Construct a target loss function based on the dimension weight corresponding to the classification dimension and the dimension loss function corresponding to the classification dimension; According to the target loss function, the initial feature extraction parameters and the initial dimension parameters are adjusted until the target loss function meets a preset condition, thereby obtaining the feature extraction parameters and the dimension parameters.

8. The method according to claim 6, characterized in that The step of obtaining a training sample set includes: Obtaining initial sample audio according to multimedia information corresponding to the preset multimedia; Performing enhancement processing on the initial sample audio in the time domain to obtain a time domain enhanced sample audio; Performing enhancement processing on the initial sample audio in the frequency domain to obtain frequency domain enhanced sample audio; Determine the initial sample audio, the time domain enhanced sample audio, and the frequency domain enhanced sample audio as sample audio of the training sample set; According to preset types corresponding to multiple classification dimensions, a sample type corresponding to the sample audio is determined.

9. The method according to claim 8, characterized in that The multimedia information includes multimedia content and multimedia description information, and the obtaining of the initial sample audio according to the multimedia information corresponding to the preset multimedia includes: Performing entity recognition on the multimedia description information to extract audio entities in the multimedia content; Performing audio content recognition on the multimedia content to extract audio content features in the multimedia content; The initial sample audio is determined from a plurality of preset audios in a preset audio library by using the audio entity and the audio content feature.

10. The method according to claim 8, characterized in that The step of performing enhancement processing on the initial sample audio in the time domain to obtain the time domain enhanced sample audio includes: Extracting time domain features of the initial sample audio in the time domain; Performing transformation processing on the time domain features to obtain transformed time domain features; Based on the transformed time domain features, a time domain enhanced sample audio is obtained.

11. The method according to claim 8, characterized in that The step of performing enhancement processing on the initial sample audio in the frequency domain to obtain frequency domain enhanced sample audio includes: Converting the initial sample audio into a frequency spectrum feature graph; Performing a perturbation process on the frequency spectrum feature map according to a mask matrix to obtain a perturbation feature map; Performing linear interpolation processing on the frequency spectrum feature graphs corresponding to different initial sample audios to obtain interpolation feature graphs; The frequency domain enhanced sample audio is generated according to the disturbance feature map and the interpolation feature map.

12. An audio classification device, characterized in that: The device comprises: A segmentation unit, used for extracting at least one segmented audio from the audio to be classified; A feature extraction unit, configured to determine a segmented audio feature corresponding to each of the segmented audios based on a feature extraction parameter; A probability prediction unit, configured to predict a plurality of prediction probabilities of each of the segmented audios under each of the classification dimensions according to the dimension parameters corresponding to the plurality of classification dimensions and the features of each of the segmented audios, wherein each of the classification dimensions corresponds to a plurality of preset types, and the prediction probabilities correspond to the preset types one by one; A type determination unit, configured to determine, for each of the classification dimensions, a target type corresponding to each of the classification dimensions by using a predicted probability of each of the segmented audios under a preset type corresponding to the classification dimension; The classification unit is used to determine the classification result of the audio to be classified based on the target type corresponding to each classification dimension.

13. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps in the audio classification method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the audio classification method according to any one of claims 1 to 11.

15. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the steps of the audio classification method according to any one of claims 1 to 11.