Method, apparatus, electronic device and storage medium for audio classification

By grouping the sound spectrum characteristics of audio files and processing neural network models, the feature labels of audio clips are extracted, and the problem of low audio classification accuracy in the prior art is solved, and higher audio classification accuracy is achieved.

CN115240707BActive Publication Date: 2025-06-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210895273.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2025-06-13
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

When the existing audio classification algorithm processes rich audio clip types in audio data, it only focuses on the global information of the entire audio segment, resulting in a low classification accuracy.

Method used

By obtaining the sound spectrum characteristics of the audio file, the first time domain characteristics are extracted using the neural network model and grouped them to obtain the second time domain characteristics. The second time domain feature is input into the classification model to obtain the feature tags for each audio clip, thereby determining the type of the audio file.

Benefits of technology

By paying attention to the local information of each audio clip in the audio file, the accuracy of audio classification is improved, and the diverse audio clips in the audio data can be processed more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240707B_ABST
    Figure CN115240707B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, apparatus, device, and storage medium for audio classification, which relates to the field of artificial intelligence technology. The method for audio classification includes: obtaining the spectrogram features of an audio file; using a neural network model to obtain the first time-domain features of the spectrogram features; grouping the first time-domain features to obtain the second time-domain features of the spectrogram features, where the feature vector of each dimension of the second time-domain features represents the time-domain feature vector of each audio segment in the audio file; inputting the second time-domain features into a classification model to obtain the first feature labels of each audio segment in the audio file; and determining the type of the audio file according to the first feature labels of each audio segment. According to the local information of the audio segments in the audio file, the embodiment of the present application classifies the audio file, can pay attention to finer-grained local information, and helps to improve the accuracy of audio file classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more particularly, to methods, devices, equipment, and storage media for audio classification. Background Art

[0002] Audio classification is one of the most widespread applications of audio deep learning. It classifies sounds through deep learning and predicts the categories of sounds. Audio classification can be applied to many practical scenarios, such as classifying music segments to identify music genres, or classifying short utterances through a set of speakers to identify people based on their voices.

[0003] Common audio classification algorithms use an entire audio segment as the input to the model and predict the global label of the entire audio segment. However, when the types of audio segments in the audio data are rich, only focusing on the global information of the entire audio segment will result in a low accuracy of audio classification. Summary of the Invention

[0004] Embodiments of this application provide a method, device, equipment, and storage medium for audio classification, which can help improve the accuracy of audio classification.

[0005] In a first aspect, embodiments of this application provide a method for audio classification, including:

[0006] Obtain the spectrogram features of an audio file;

[0007] Using a neural network model, obtain the first time-domain features of the spectrogram features;

[0008] Group the first time-domain features to obtain the second time-domain features of the spectrogram features. The feature vectors of each dimension of the second time-domain features represent the time-domain feature vectors of each audio segment in the audio file;

[0009] Input the second time-domain features into a classification model to obtain the first feature labels of each audio segment in the audio file;

[0010] Determine the type of the audio file according to the first feature labels of each audio segment.

[0011] In a second aspect, embodiments of this application provide a method for training an audio classification model, including:

[0012] Obtain a training sample set, where the training sample set includes audio file samples and classification labels of the audio file samples;

[0013] Copy the classification labels of the audio file samples M times to obtain the classification labels of each audio segment in the audio file samples, where the value of M is determined according to the number of audio segments in the audio file samples;

[0014] Obtain the spectrogram features of the audio file sample;

[0015] Using a neural network model, obtain the third time-domain features of the spectrogram features;

[0016] Group the third time-domain features to obtain the fourth time-domain features of the spectrogram features, where the feature vector of each dimension of the fourth time-domain features represents the time-domain feature vector of each audio segment in the audio file sample; and

[0017] Input the fourth time-domain features into a classification model to obtain the third feature labels of each audio segment in the audio file sample;

[0018] Determine a first loss according to the third feature labels of each audio segment in the audio file sample and the classification labels of each audio segment;

[0019] Update the parameters of the classification model according to the first loss.

[0020] In a third aspect, an embodiment of the present application provides a device for predicting object risks, including:

[0021] A device for audio classification, characterized by including:

[0022] A first acquisition unit for acquiring the spectrogram features of an audio file;

[0023] A neural network model for obtaining the first time-domain features of the spectrogram features;

[0024] A grouping unit for grouping the first time-domain features to obtain the second time-domain features of the spectrogram features, where the feature vector of each dimension of the second time-domain features represents the time-domain feature vector of each audio segment in the audio file;

[0025] A classification model for inputting the second time-domain features to obtain the first feature labels of each audio segment in the audio file;

[0026] A determination unit for determining the type of the audio file according to the first feature labels of each audio segment.

[0027] In a fourth aspect, an embodiment of the present application provides a device for training an audio classification model, including:

[0028] A first acquisition unit for acquiring a training sample set, where the training sample set includes audio file samples and classification labels of the audio file samples;

[0029] The first acquisition unit is further configured to copy the classification label of the audio file sample M times, and obtain the classification label of each audio segment in the audio file sample, where the value of M is determined according to the number of audio segments in the audio file sample;

[0030] A second acquisition unit, configured to acquire the spectrogram feature of the audio file sample;

[0031] A neural network model, configured to obtain a third time-domain feature of the spectrogram feature;

[0032] A grouping unit, configured to group the third time-domain feature to obtain a fourth time-domain feature of the spectrogram feature, where the feature vector of each dimension of the fourth time-domain feature represents the time-domain feature vector of each audio segment in the audio file sample; and

[0033] A classification model, configured to input the fourth time-domain feature into the classification model to obtain a third feature label of each audio segment in the audio file sample;

[0034] A determination unit, configured to determine a first loss according to the third feature label of each audio segment in the audio file sample and the classification label of each audio segment;

[0035] An update unit, configured to update the parameters of the classification model according to the first loss.

[0036] In a fifth aspect, an embodiment of the present application provides an electronic device, including:

[0037] A processor, adapted to implement computer instructions; and,

[0038] A memory, storing computer instructions, where the computer instructions are adapted to be loaded and executed by the processor to perform the method in the first aspect or the method in the second aspect above.

[0039] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are read and executed by a processor of a computer device, the computer device is caused to execute the method in the first aspect or the method in the second aspect above.

[0040] In a seventh aspect, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in the first aspect or the method in the second aspect above.

[0041] In the embodiments of the present application, by grouping the first time-domain features of the spectrogram features of an audio file, the second time-domain features of the spectrogram features are obtained. Since the feature vectors of each dimension in the second time-domain features represent the time-domain feature vectors of each audio segment in the audio file, the feature labels of each audio segment in the audio file can be obtained according to the second time-domain features. Furthermore, the type of the audio file can be determined according to the feature labels of each audio segment in the audio file. The embodiments of the present application classify the audio file according to the local information of the audio segments in the audio file, can focus on finer-grained local information, and help improve the accuracy of audio file classification. Description of the Drawings

[0042] Figure 1 It is a schematic diagram of a system architecture related to the embodiments of the present application;

[0043] Figure 2 It is a schematic flowchart of a method for audio classification provided by the embodiments of the present application;

[0044] Figure 3 It is a schematic diagram of the network architecture of an audio classification model according to the embodiments of the present application;

[0045] Figure 4 It is a schematic flowchart of another method for audio classification according to the embodiments of the present application;

[0046] Figure 5 It is Figure 4 a specific example of the first transformation operator in

[0047] Figure 6 It is a schematic flowchart of a method for training a model provided by the embodiments of the present application;

[0048] Figure 7 It is Figure 4 a specific example of the second transformation operator in

[0049] Figure 8 It is a schematic block diagram of an apparatus for audio classification provided by the embodiments of the present application;

[0050] Figure 9 It is a schematic block diagram of an apparatus for training a model provided by the embodiments of the present application;

[0051] Figure 10 It is a schematic block diagram of an electronic device provided by the embodiments of the present application. Detailed Embodiments

[0052] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0053] It should be understood that in the embodiment of the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based only on A, and B can also be determined based on A and / or other information.

[0054] In the description of the present application, unless otherwise specified, "at least one" means one or more, and "plurality" means two or more than two. In addition, "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0055] It should also be understood that the first, second, etc. descriptions appearing in the embodiments of the present application are only used for illustration and distinction of the description objects, without any distinction of order, nor do they indicate any special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.

[0056] It should also be understood that the specific features, structures or characteristics related to the embodiments in the specification are included in at least one embodiment of the present application. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.

[0057] In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or device.

[0058] The embodiments of the present application are applied to the field of artificial intelligence technology.

[0059] Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0060] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0061] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0062] The embodiments of this application may be related to speech technology in artificial intelligence technology (Speech Technology). The key technologies of speech technology include Automatic Speech Recognition (ASR), speech synthesis technology, and voiceprint recognition technology. Enabling the computer to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.

[0063] The embodiments of this application may also be related to Machine Learning (ML) in artificial intelligence technology. ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0064] Figure 1 This is a schematic diagram of a system architecture related to an embodiment of the present application. As Figure 1 shown, the system architecture may include a user device 101, a data acquisition device 102, a training device 103, an execution device 104, a database 105, and a content library 106.

[0065] Among them, the data acquisition device 102 is used to read training data from the content library 106 and store the read training data in the database 105. The training data related to the embodiments of the present application includes audio file samples and classification labels of the audio file samples.

[0066] The training device 103 trains a machine learning model based on the training data maintained in the database 105, so that the trained machine learning model can effectively classify audio files. The machine learning model obtained by the training device 103 can be applied to different systems or devices.

[0067] In addition, referring to Figure 1 , the execution device 104 is configured with an I / O interface 107 to interact with external devices. For example, it receives an audio file to be classified sent by the user device 101 through the I / O interface. The computing module 109 in the execution device 104 uses the trained classification model to classify at least one input audio file and outputs the type of the audio file. The classification model can send the corresponding result to the user device 101 through the I / O interface.

[0068] Among them, the user device 101 may include a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted terminal, a mobile internet device (MID), or other terminal devices.

[0069] The execution device 104 may be a server.

[0070] Exemplarily, the server may be a computing device such as a rack server, a blade server, a tower server, or a cabinet server. The server may be an independent test server or a test server cluster composed of multiple test servers.

[0071] The server may be one or more. When there are multiple servers, at least two servers are used to provide different services, and / or at least two servers are used to provide the same service. For example, the same service is provided in a load balancing manner. The embodiments of the present application do not limit this.

[0072] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can also become a node of the blockchain.

[0073] In this embodiment, the execution device 104 is connected to the user device 101 through a network. The network can be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, or a call network.

[0074] It should be noted that Figure 1 is only a schematic diagram of a system architecture provided by the embodiments of the present application. The positional relationships among the devices, components, modules, etc. shown in the figure do not constitute any limitation. In some embodiments, the above data acquisition device 102, user device 101, training device 103, and execution device 104 can be the same device. The above database 105 can be distributed on one server or multiple servers, and the above content library 106 can be distributed on one server or multiple servers.

[0075] Exemplarily, the application scenarios of the embodiments of the present application include, but are not limited to, classifying music segments to identify music types, or classifying short utterances through a set of speakers to identify people based on voices.

[0076] Common audio classification algorithms use an entire audio segment as the model input to predict the global label of the entire audio segment. At this time, the model tends to focus on the global label of the entire audio segment, that is, on the global information of the entire audio segment, and lacks attention to the local information of a certain segment of the audio. When the types of audio segments in the audio data are rich, for example, when the audio data is spliced and fused by audio segments with rich variations, only focusing on the global information of the entire audio segment will result in a low accuracy rate of audio classification.

[0077] In view of this, embodiments of the present application provide a method, apparatus, electronic device, and storage medium for audio classification. By classifying an audio file based on the local information of audio segments in the audio file, it is possible to focus on finer-grained local information, which helps to improve the accuracy of audio file classification.

[0078] Specifically, obtain the spectrogram features of the audio file, and use a convolutional neural network model to obtain the first time-domain features of the spectrogram features. Then, group the first time-domain features to obtain the second time-domain features of the spectrogram features. The feature vectors of each dimension of the second time-domain features represent the time-domain feature vectors of each audio segment in the audio file. And input the second time-domain features into a classification model to obtain the first feature labels of each audio segment in the audio file. Finally, determine the type of the audio file according to the first feature labels of each audio segment.

[0079] Therefore, in embodiments of the present application, by grouping the first time-domain features of the spectrogram features of the audio file to obtain the second time-domain features of the spectrogram features, since the feature vectors of each dimension in the second time-domain features represent the time-domain feature vectors of each audio segment in the audio file, the feature labels of each audio segment in the audio file can be obtained according to the second time-domain features. Furthermore, the type of the audio file can be determined according to the feature labels of each audio segment in the audio file. Embodiments of the present application classify an audio file based on the local information of audio segments in the audio file, which can focus on finer-grained local information and helps to improve the accuracy of audio file classification.

[0080] Figure 2 FIG. 200 shows a schematic flowchart of a method 200 for audio classification provided by an embodiment of the present application. The method 200 can be executed by any electronic device with data processing capabilities. For example, the electronic device can be implemented as a server or a terminal device. For another example, the electronic device can be implemented as the Figure 1 computing module 109 therein, and the present application does not make any limitations in this regard.

[0081] In some embodiments, a machine learning model may be included (such as deployed) in the electronic device. The machine learning model can be a deep learning model, a neural network model, or other models, without limitation. In some embodiments, the machine learning model can be a classification model and can be used to classify audio files. Figure 3FIG. 0 is a schematic diagram of a network architecture 300 of an audio classification model provided by an embodiment of the present application, which includes a spectrogram extraction module 310, a Convolutional Neural Networks (CNN) 320, a second time-frequency fusion module 330, and a fully connected layer 340. Among them, the second time-frequency fusion module 330 includes a first transformation operator 3301. Optionally, the network architecture 300 may further include a first time-frequency fusion module 360, a first fully connected layer 370, and a second transformation operator 350. Hereinafter, the method 200 for audio classification will be described in conjunction with Figure 3 the network architecture in

[0082] 210, obtain the spectrogram features of the audio file.

[0083] Exemplarily, the audio file may be a music file, such as a song segment. The audio file may be a fixed-length audio segment, such as a 15s or 30s audio segment. Exemplarily, the audio file may be in formats such as mp3, aac, wav, wma, etc., without limitation.

[0084] Referring to Figure 3 , the spectrogram features of the audio file can be obtained through Figure 3 the spectrogram extraction module 310 in

[0085] In some embodiments, the feature vector may be represented as a feature tensor of [N, C, T, F], where N represents the number of batches, which can also be denoted as batch_size, C represents the number of channels, which can also be denoted as channel, T represents the number of time intervals, which can also be denoted as time_bins, and F represents the number of frequency intervals, which can also be denoted as freq_bins. Among them, T can also be referred to as the time-domain feature, and F can also be referred to as the frequency-domain feature.

[0086] As a specific example, the spectrogram features may be [N, C0, T0, F0] = [25, 1, 1500, 64], and the spectrogram features include a time-domain feature vector with a time dimension of 1500 and a frequency-domain feature vector with a frequency dimension of 64.

[0087] 220, use a neural network model to obtain the first time-domain feature of the spectrogram features.

[0088] The neural network model can perform feature extraction on the input spectrogram features to obtain the first time-domain feature. The neural network model may be a CNN, and the present application does not limit this.

[0089] In some embodiments, the neural network model can also extract features from the input acoustic spectrum features to obtain the spectral features and the first number of channels of the acoustic spectrum features.

[0090] Referring to Figure 3 , the acoustic spectrum features can be used as the Figure 3 input of the CNN 320 in . The CNN 320 extracts features from the acoustic spectrum features to obtain CNN features. The CNN 320 can extract features from the time-domain features in the acoustic spectrum features to obtain the first time-domain features of the acoustic spectrum features. Optionally, the CNN 320 can also extract features from the frequency-domain features in the acoustic spectrum features to obtain the frequency-domain features of the acoustic spectrum features. Optionally, the number of channels of the acoustic spectrum features obtained by the CNN 320 is the first number of channels.

[0091] Continuing with the above example, after the acoustic spectrum features [N, C0, T0, F0] = [25, 1, 1500, 64] are input into the CNN 320, CNN features [N, C1, T1, F1] = [25, 1024, 45, 2] can be obtained, where the batch number N remains unchanged, the number of channels increases from 1 to 1024, the time dimension of the time-domain feature vector is reduced to 45, and the frequency of the frequency-domain feature vector is reduced to 2. That is to say, after passing through the CNN 320, an audio file with a fixed duration (such as 15s) is mapped to 45 feature vectors in the time dimension and 2 feature vectors in the frequency dimension.

[0092] 230 groups the first time-domain features to obtain the second time-domain features of the acoustic spectrum features. Each dimensional feature vector of the second time-domain features represents the time-domain feature vector of each audio segment in the audio file.

[0093] That is to say, the second time-domain features are obtained by grouping the feature vectors of each time dimension of the first time-domain features. For example, when the first time-domain features include time-domain feature vectors with a time dimension of 45, the 45 time-dimensional feature vectors can be grouped, and each group of feature vectors corresponds to a time dimension, representing the time-domain feature vector of each audio segment. Here, the time dimension of the second time-domain features is the same as the number of audio segments in the audio file.

[0094] In some embodiments, the audio segment can be an audio segment corresponding to a unit time, such as an audio segment of 1s.

[0095] In some embodiments, a grouping coefficient of an audio file can be obtained, where the grouping coefficient represents the number of sampling points of the audio file per unit time. Then, the first time-domain features are grouped according to the grouping coefficient to obtain the second time-domain features described above. According to the grouping coefficient, uniform grouping of the first time-domain features can be achieved, such that the time dimension corresponding to each group of feature vectors is the same as the number of sampling points of the audio file per unit time.

[0096] Taking the example that there are 3 sampling points in a unit time of 1s, the grouping coefficient is 3 at this time. Each group of feature vectors in the second time-domain features obtained by grouping the first time-domain features can include feature vectors of 3 time dimensions. Continuing with the above example, when the first time-domain features include a time-domain feature vector with a time dimension of 45, the obtained second time-domain features include time-domain feature vectors with a time dimension of (45÷3 = 15).

[0097] In some embodiments, the second number of channels input to the classification model can also be determined according to the first number of channels output by the convolutional neural network model, the frequency-domain features of the spectrogram features, and the grouping coefficient. That is to say, the frequency-domain features and the grouping coefficient of the spectrogram features can be incorporated into the second number of channels, such that the second number of channels input to the classification model can fuse the frequency-domain features and the time-domain features of the spectrogram features.

[0098] In some embodiments, the second time-domain features and the second number of channels described above can represent the time-frequency features of the audio file. Refer to Figure 3 , the CNN features output by CNN320 can be input to the first transformation operator 331 in the second time-frequency fusion module 330 to obtain time-frequency feature 2, which includes the second time-domain features described above, and the number of channels of time-frequency feature 2 is the second number of channels. As a specific example, the second time-frequency fusion module 330 can be a Glance time-frequency fusion module, and time-frequency feature 2 can be referred to as Glance time-frequency.

[0099] As an implementation manner, refer to Figure 4 , the second number of channels can be obtained through steps 401 to 404.

[0100] 401. According to the first feature tensor, a second feature tensor is obtained, and the second feature tensor sequentially includes the first time-domain features, the first number of channels, and the frequency-domain features.

[0101] Exemplarily, the convolutional neural network model can output a first feature tensor, which sequentially includes the first number of channels, the first time-domain features, and the frequency-domain features described above. Then, by exchanging the order of the first number of channels and the first time-domain features in the first feature tensor, a second feature vector can be obtained.

[0102] For example, refer to Figure 5For a CNN feature in the form of [N, C1, T1, F1] = [25, 1024, 45, 2], the C1 axis and the T1 axis can be swapped, that is, the number of channels 1024 and the frequency domain feature 45 are swapped, to obtain a CNN feature in the form of [N, T1, C1, F1] = [25, 45, 1024, 2].

[0103] 402. According to the second feature tensor and the grouping coefficient, a third feature tensor is obtained. The third feature tensor sequentially includes a second time domain feature, a grouping coefficient, and a third number of channels, and the third number of channels is obtained by fusing the first number of channels and the frequency domain feature. Among them, the second time domain feature is obtained according to the first time domain feature and the grouping coefficient in the second feature tensor. Specifically, the process of obtaining the second time domain feature can refer to the description in the above text and will not be elaborated here.

[0104] Continue to refer to Figure 5 For a CNN feature in the form of [N, T1, C1, F1] = [25, 45, 1024, 2], the C1 axis and the F1 axis can be fused, that is, the number of channels and the frequency domain feature are fused, to obtain a CNN feature in the form of [N, T1, C1*F1] = [25, 45, 2048], and the T1 axis (i.e., the time domain feature) is evenly grouped according to the grouping coefficient g to obtain a CNN feature in the form of [N, T1 / / g, g, C1*F1] = [25, 15, 3, 2048], where g represents the grouping coefficient and / / represents the division operation.

[0105] 403. According to the third feature tensor, a fourth feature tensor is obtained. The fourth feature tensor sequentially includes the second time domain feature and the second number of channels, where the second number of channels is obtained by fusing the grouping coefficient and the third number of channels.

[0106] Continue to refer to. For a CNN feature in the form of [N, T1 / / g, g, C1*F1] = [25, 15, 3, 2048], the g axis and the C1*F1 axis are fused, that is, the grouping coefficient and the number of channels are fused, and finally a time-frequency feature in the form of [N, T2, C2] = [25, 15, 6144] can be obtained. Here, fusing the g axis and the C1*F1 axis can also be referred to as performing a feature transformation on g, C1, and F1, and the present application does not limit this.

[0107] 240. The second time domain feature is input into the classification model to obtain the first feature label of each audio segment in the audio file.

[0108] Specifically, the classification model can determine the probabilities of the respective feature labels corresponding to each audio segment in the audio file based on the feature vectors of each dimension of the second time-domain feature, and determine the first feature label corresponding to each audio segment. For example, the first several feature labels with the highest probabilities corresponding to each audio segment can be used as the first feature label of the audio segment.

[0109] Since the feature vector of each dimension in the second time-domain feature represents the time-domain feature vector of each audio segment in the audio file, the first feature label of each audio segment in the audio file can be obtained based on this second time-domain feature. That is to say, the first feature label of each audio segment is obtained based on the local information of the audio segment.

[0110] In some embodiments, when the audio segment is a music segment, the feature labels of the music file or music segment can include labels classified according to the style of the music, such as pop, rock, folk, electronic, jazz, rap, classical, etc. The feature labels of the music file or music segment can also be labels classified according to the emotion of the music, such as happy, sad, angry, relaxed, etc. The feature labels can also be called classification labels or others, and the present application does not make any limitations in this regard.

[0111] In some examples, a music file or music segment can correspond to one or more feature labels, and the present application does not make any limitations in this regard.

[0112] In some embodiments, the second time-domain feature can also be input into the classification model according to the second number of channels to obtain the first feature label of each audio segment in the audio file. Since the second number of channels combines the frequency-domain feature and the time-domain feature of the spectrogram feature, the first feature label of each audio segment in the audio file is made more accurate.

[0113] In some embodiments, the classification model can be a fully connected layer. Refer to Figure 3 , for the time-frequency feature 2 obtained by the second time-frequency fusion module 330, it can be input into the second fully connected layer 340 to obtain the first feature label of each audio segment in the audio file, that is, the output result 2. Exemplarily, the output result 2 can be a feature tensor in the form of [N, T2, class_num] = [25, 15, 12].

[0114] 250, determine the type of the audio file according to the first feature label of each audio segment.

[0115] Exemplarily, the probabilities of the first feature tags corresponding to each audio segment in the audio file that belong to the same category can be added up to obtain the sum of the probabilities of the first feature tags of each category, and the category with the largest sum of probabilities can be determined as the type of the audio file, so that the audio file can be classified according to the local information of the audio segments.

[0116] Therefore, in the embodiment of the present application, by grouping the first time-domain features of the spectrogram features of the audio file, the second time-domain features of the spectrogram features are obtained. Since the feature vectors of each dimension in the second time-domain features represent the time-domain feature vectors of each audio segment in the audio file, the feature tags of each audio segment in the audio file can be obtained according to the second time-domain features, and then the type of the audio file can be determined according to the feature tags of each audio segment in the audio file. The embodiment of the present application classifies the audio file according to the local information of the audio segments in the audio file, can pay attention to finer-grained local information, and helps to improve the accuracy of audio file classification.

[0117] In some embodiments, the second feature tag of the audio file can also be obtained according to the first time-domain feature, and the type of the audio file can be determined according to the second feature tag and the first feature tag of each audio segment.

[0118] That is to say, in step 220, after obtaining the first time-domain features of the spectrogram features, the second feature tag corresponding to the entire audio file can be directly obtained according to the first time-domain features. That is to say, the second feature tag is obtained according to the global information of the audio file.

[0119] Exemplarily, the probabilities of the first feature tags corresponding to each audio segment in the audio file and the second feature tag corresponding to the entire audio file that belong to the same category can be added up to obtain the sum of the probabilities of the feature tags of each category, and the category with the largest sum of probabilities can be determined as the type of the audio file, so that the audio file can be classified according to the local information of the audio segments and the global information of the entire audio file.

[0120] See Figure 3, the CNN features output by the CNN 320 can be input into the time-frequency fusion operator 361 in the first time-frequency fusion module 360 to obtain time-frequency features 1. Further, the time-frequency features 1 can be input into the first fully connected layer 370 to obtain the second feature label of the audio file, that is, the output result 1. Exemplarily, the output result 1 can be a feature tensor in the form of [N, class_num] = [25, 12]. Then, according to the output result 1 and the output result 2, for the first feature label of each audio segment and the second feature label of the audio file, the classification of the audio file is obtained. As a specific example, the first time-frequency fusion module 360 can be called the valinna time-frequency fusion module, and the time-frequency features 1 can be called the valinna time-frequency features. The output result 1 can also be called valinnalogits, and the output result 2 can also be called glancelogits.

[0121] Therefore, the embodiment of the present application classifies the audio file according to the local information of the audio segments in the audio file and the global information of the whole audio file, which can achieve both attention to the global information and more fine-grained local information, and helps to further improve the accuracy of audio file classification.

[0122] In some embodiments, before using the classification model to classify the audio file, the classification model can also be trained to obtain a trained classification model.

[0123] Figure 6 shows a schematic flowchart of a method 600 for training a model provided by an embodiment of the present application. The method 600 can be executed by any electronic device with data processing capabilities. For example, the electronic device can be implemented as a server or a terminal device. For another example, the electronic device can be implemented as Figure 1 the training module 103 in, and the present application does not make any limitations on this.

[0124] 610, obtain a training sample set, where the training sample set includes audio file samples and classification labels of the audio file samples.

[0125] The training sample set can include a large number of audio file samples and the corresponding classification labels for each audio file sample. As an example, when the audio file is a music file, the data of the training sample set can come from the songs in the Million Song Dataset (MillionSongDataset) and the classification labels of each song. Exemplarily, the classification labels of the songs can be the labels obtained by classifying according to the music style, such as pop, rock, folk, electronic, jazz, rap, classical, etc. The classification labels of the songs can also be the labels obtained by classifying according to the emotion of the music, such as happy, sad, angry, relaxed, etc.

[0126] In some embodiments, the audio file sample may be a fixed-length audio segment, such as a 15s or 30s song segment. Exemplarily, the audio file sample may be in formats such as mp3, aac, wav, wma, etc., without limitation.

[0127] 620. Duplicate the classification label of the audio file sample M times to obtain the classification label of each audio segment in the audio file sample, where the value of M is determined according to the number of audio segments in the audio file sample.

[0128] Specifically, the M duplicated classification labels can be used as the classification labels of each audio segment in the audio file sample. By duplicating the target label, the target label can be extended in the time dimension to obtain the classification label corresponding to each audio segment. Exemplarily, the classification label of the audio file sample can be duplicated by Figure 3 the second transformation operator 350 in to obtain the classification label of each audio segment in the audio file sample.

[0129] Refer to Figure 7 . The copy coefficient M = 15 can be determined according to the sampling time unit (such as 1s) and the duration of the audio file (such as 15s). For the target label 1, a new time domain feature dimension (i.e., the T dimension) can be added, and the label can be copied M (such as 15) times in the T dimension to obtain the target label 2. As a specific example, for the target label 1 in the form of [N, class_num] = [25, 12], by adding the T dimension and copying the label M times in the T dimension, the target label 2 in the form of [N, T, class_num] = [25, 15, 12] can be obtained.

[0130] 630. Obtain the spectrogram features of the audio file sample.

[0131] 640. Use a neural network model to obtain the third time domain features of the spectrogram features.

[0132] 650. Group the third time domain features to obtain the fourth time domain features of the spectrogram features, and the feature vector of each dimension of the fourth time domain features represents the time domain feature vector of each audio segment in the audio file sample.

[0133] 660. Input the fourth time domain features into a classification model to obtain the third feature label of each audio segment in the audio file sample.

[0134] Specifically, the process of obtaining the third feature label of each audio segment in the audio file sample through steps 630 to 660 is Figure 2 similar to the process of obtaining the first feature label of each audio segment in the audio file through step 210 in, and can be referred toFigure 2 The description in [reference] is not repeated here.

[0135] 670. Determine a first loss according to the third feature label of each audio segment in the audio file sample and the classification label of each audio segment.

[0136] See Figure 3 , when the output result 2 includes the third feature label of each audio segment in the audio file sample, the third feature label of each audio segment of the audio file sample in the output result 2 can be compared with the labels of each time dimension of the target label 2 to determine the first loss between the third feature label of each audio segment of the audio file sample and the label of the corresponding time dimension of the target label 2, that is, loss 2. Loss 2 can also be called glanceloss. Exemplarily, the first loss can be cross-entropy loss, and the present application does not limit this.

[0137] In some embodiments, a fourth feature label of the audio file sample may also be obtained, and a second loss may be determined according to the fourth feature label of the audio file sample and the classification label of the audio file sample. Exemplarily, the second loss can be cross-entropy loss, and the present application does not limit this.

[0138] Specifically, the process of obtaining the fourth feature label of the audio file sample is similar to the process of obtaining the second feature label of the audio file above, and reference can be made to the description above, which is not repeated here.

[0139] See Figure 3 , when the output result 1 includes the fourth feature label of the audio file sample, the fourth feature label of the audio file sample in the output result 1 can be compared with the target label 1 to determine the second loss between the fourth feature label of the audio file sample and the classification label of the audio file sample, that is, loss 1. Loss 1 can also be called valinnaloss.

[0140] 680. Update the parameters of the classification model according to the first loss.

[0141] Exemplarily, the model parameters of each layer of the classification model can be iteratively updated according to the first loss to gradually improve the accuracy of the classification model. In some embodiments, the model parameters of each layer of the neural network model for extracting the third time-domain features of the spectrogram features can also be iteratively updated according to the first loss.

[0142] In some embodiments, the parameters of the classification model can also be updated according to the first loss and the second loss, or the model parameters of each layer of the classification model and the neural network model for extracting the third time-domain features of the spectrogram features can be iteratively updated.

[0143] Therefore, in the present application, by copying the classification labels of the audio file samples, the feature labels of each audio segment in the audio file samples are obtained. After obtaining the third feature labels of each audio segment in the audio file samples according to the classification model, the parameters of the classification model are updated according to the third feature labels and the feature labels of each audio segment, and a trained classification model is obtained. The classification model of the embodiment of the present application can classify the audio file by using the local information of the audio segments of the audio file, can pay attention to finer-grained local information, and helps to improve the accuracy of audio file classification.

[0144] The specific embodiments of the present application have been described in detail above with reference to the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all belong to the protection scope of the present application. For example, in the various specific technical features described in the above specific embodiments, they can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the present application will not separately describe various possible combination methods. Again, for example, any combination can be made between various different embodiments of the present application, as long as it does not violate the idea of the present application, it should also be regarded as the content disclosed by the present application.

[0145] It should also be understood that in the various method embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. It should be understood that these serial numbers can be interchanged under appropriate circumstances, so that the embodiments of the present application described can be implemented in an order other than those illustrated or described.

[0146] The method embodiments of the present application have been described in detail above. Below, in combination with Figures 8 to 10 , the device embodiments of the present application are described in detail.

[0147] Figure 8 FIG. is a schematic block diagram of an audio classification device 800 provided by an embodiment of the present application. As Figure 8 shown, the audio classification device 800 may include a first acquisition unit 810, a neural network model 820, a grouping unit 830, a classification model 840, and a determination unit 850.

[0148] The first acquisition unit 810 is configured to acquire the spectrogram features of an audio file;

[0149] The neural network model 820 is configured to obtain the first time-domain features of the spectrogram features;

[0150] A grouping unit 830 is configured to group the first time-domain features to obtain second time-domain features of the spectrogram features, and a feature vector of each dimension of the second time-domain features represents a time-domain feature vector of each audio segment in the audio file;

[0151] A classification model 840 is configured to input the second time-domain features to obtain a first feature label of each audio segment in the audio file;

[0152] A determination unit 850 is configured to determine the type of the audio file according to the first feature label of each audio segment.

[0153] In some embodiments, the grouping unit 830 is specifically configured to:

[0154] Obtain a grouping coefficient of the audio file, where the grouping coefficient represents the number of sampling points of the audio file per unit time;

[0155] Group the first time-domain features according to the grouping coefficient to obtain the second time-domain features.

[0156] In some embodiments, the convolutional neural network model 820 is further configured to obtain frequency-domain features of the spectrogram features;

[0157] The determination unit 850 is further configured to determine a second number of channels input to the classification model 830 according to a first number of channels output by the convolutional neural network model 820, the frequency-domain features, and the grouping coefficient;

[0158] The classification model 830 is specifically configured to input the second time-domain features into the classification model according to the second number of channels to obtain a first feature label of each audio segment in the audio file.

[0159] In some embodiments, the convolutional neural network model outputs a first feature tensor, and the first feature tensor sequentially includes the first number of channels, the first time-domain features, and the frequency-domain features.

[0160] Wherein, the determination unit 850 is specifically configured to:

[0161] Obtain a second feature tensor according to the first feature tensor, where the second feature tensor sequentially includes the first time-domain features, the first number of channels, and the frequency-domain features;

[0162] Obtain a third feature tensor according to the second feature tensor and the grouping coefficient, where the third feature tensor sequentially includes the second time-domain features, the grouping coefficient, and a third channel coefficient, and the third number of channels is obtained by fusing the first number of channels and the frequency-domain features;

[0163] According to the third feature tensor, a fourth feature tensor is obtained, and the fourth feature tensor sequentially includes the second time-domain feature and the second number of channels, where the second number of channels is obtained by fusing the grouping coefficient and the third number of channels.

[0164] In some embodiments, the apparatus 800 further includes a second obtaining unit, configured to:

[0165] Obtain a training sample set, where the training sample set includes audio file samples and classification labels of the audio file samples;

[0166] Duplicate the classification labels of the audio file samples M times to obtain the classification labels of each audio segment in the audio file samples, where the value of M is determined according to the number of audio segments in the audio file samples.

[0167] In some embodiments, the apparatus 800 further includes a third obtaining unit, configured to:

[0168] Obtain a second feature label of the audio file according to the first time-domain feature;

[0169] Wherein, the determining unit 850 is specifically configured to determine the type of the audio file according to the first feature label of each audio segment and the second feature label of the audio file.

[0170] In some embodiments, the spectrogram feature includes Mel feature.

[0171] In some embodiments, the audio file includes a music file.

[0172] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, details are not described herein again. Specifically, in this embodiment, the apparatus 800 for audio classification may correspond to the corresponding main body that executes the method 200 of the embodiments of the present application, and the foregoing and other operations and / or functions of each module in the apparatus 800 respectively implement the corresponding processes in the method 200 described above. For the sake of brevity, details are not described herein again.

[0173] Figure 9 is a schematic block diagram of an apparatus 900 for training a model according to an embodiment of the present application. As Figure 9 shown, the apparatus 900 for training a model may include a first obtaining unit 910, a second obtaining unit 920, a neural network model 930, a grouping unit 940, a classification model 950, a determining unit 960, and an updating unit 970.

[0174] The first obtaining unit 910 is configured to obtain a training sample set, where the training sample set includes audio file samples and classification labels of the audio file samples;

[0175] The first acquisition unit 910 is further configured to copy the classification label of the audio file sample M times, and obtain the classification label of each audio segment in the audio file sample, where the value of M is determined according to the number of audio segments in the audio file sample;

[0176] A second acquisition unit 920, configured to acquire the spectrogram feature of the audio file sample;

[0177] A neural network model 930, configured to obtain a third time-domain feature of the spectrogram feature;

[0178] A grouping unit 940, configured to group the third time-domain feature to obtain a fourth time-domain feature of the spectrogram feature, where the feature vector of each dimension of the fourth time-domain feature represents the time-domain feature vector of each audio segment in the audio file sample; and

[0179] A classification model 950, configured to input the fourth time-domain feature into the classification model to obtain a third feature label of each audio segment in the audio file sample;

[0180] A determination unit 960, configured to determine a first loss according to the third feature label of each audio segment in the audio file sample and the classification label of each audio segment;

[0181] An update unit 970, configured to update the parameters of the classification model according to the first loss.

[0182] In some embodiments, it may further include a third acquisition unit, configured to acquire a fourth feature label of the audio file sample. The determination unit 960 may specifically determine a second loss according to the fourth feature label of the audio file sample and the classification label of the audio file sample. The update unit 970 may update the parameters of the classification model according to the first loss and the second loss.

[0183] It should be understood that the device embodiments and the method embodiments may correspond to each other, and similar descriptions may refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, in this embodiment, the device 900 for training the model may correspond to the corresponding main body that executes the method 600 of the embodiment of the present application, and the foregoing and other operations and / or functions of each module in the device 900 respectively implement the corresponding processes in the method 600 described above. For the sake of brevity, it will not be elaborated here.

[0184] In the foregoing, the apparatuses and systems of the embodiments of the present application have been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the embodiments of the present application can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0185] As Figure 10 is a schematic block diagram of an electronic device 1100 provided by an embodiment of the present application.

[0186] As Figure 10 shown, the electronic device 1100 may include:

[0187] A memory 1110 and a processor 1120. The memory 1110 is used to store a computer program and transmit the program code to the processor 1120. In other words, the processor 1120 can call and run the computer program from the memory 1110 to implement the method in the embodiments of the present application.

[0188] For example, the processor 1120 can be used to execute the steps in the above method 200 or 600 according to the instructions in the computer program.

[0189] In some embodiments of the present application, the processor 1120 may include, but is not limited to:

[0190] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0191] In some embodiments of the present application, the memory 1110 includes, but is not limited to:

[0192] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0193] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 1110 and executed by the processor 1120 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 1100.

[0194] Optionally, as Figure 10 shown, the electronic device 1100 may further include:

[0195] A transceiver 1130, which may be connected to the processor 1120 or the memory 1110.

[0196] Among them, the processor 1120 may control the transceiver 1130 to communicate with other devices. Specifically, it may send information or data to other devices, or receive information or data sent by other devices. The transceiver 1130 may include a transmitter and a receiver. The transceiver 1130 may further include an antenna, and the number of antennas may be one or more.

[0197] It should be understood that each component in the electronic device 1100 is connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0198] According to one aspect of the present application, a communication device is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory, so that the encoder executes the method of the above method embodiment.

[0199] According to one aspect of the present application, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method of the above method embodiment. Or rather, the embodiment of the present application further provides a computer program product including instructions. When the instructions are executed by the computer, the computer executes the method of the above method embodiment.

[0200] According to another aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of the above method embodiment.

[0201] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).

[0202] It is understandable that in the specific embodiments of the present application, data related to user information and the like may be involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0203] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0204] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other form.

[0205] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0206] The above is only the specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for audio classification, characterized in that, comprising: Obtaining the spectrogram features of an audio file; Using a neural network model to obtain the first time-domain features of the spectrogram features; Grouping the first time-domain features to obtain the second time-domain features of the spectrogram features, where the feature vector of each dimension of the second time-domain features represents the time-domain feature vector of each audio segment in the audio file; Inputting the second time-domain features into a classification model to obtain the first feature labels of each audio segment in the audio file; Determining the type of the audio file according to the first feature labels of each audio segment; Among them, the grouping of the first time-domain features to obtain the second time-domain features of the spectrogram features includes: Obtaining the grouping coefficient of the audio file, where the grouping coefficient represents the number of sampling points of the audio file per unit time; Grouping the first time-domain features according to the grouping coefficient to obtain the second time-domain features.

2. The method according to claim 1, characterized in that, further comprising: Using a convolutional neural network model to obtain the frequency-domain features of the spectrogram features; Determining the second number of channels input to the classification model according to the first number of channels output by the convolutional neural network model, the frequency-domain features, and the grouping coefficient; Among them, the inputting the second time-domain features into a classification model to obtain the first feature labels of each audio segment in the audio file includes: Inputting the second time-domain features into the classification model according to the second number of channels to obtain the first feature labels of each audio segment in the audio file.

3. The method according to claim 2, characterized in that, The convolutional neural network model outputs a first feature tensor, and the first feature tensor sequentially includes the first number of channels, the first time-domain features, and the frequency-domain features; Among them, the determining the second number of channels according to the first number of channels output by the convolutional neural network model, the frequency-domain features, and the grouping coefficient includes: Obtaining a second feature tensor according to the first feature tensor, where the second feature tensor sequentially includes the first time-domain features, the first number of channels, and the frequency-domain features; Obtaining a third feature tensor according to the second feature tensor and the grouping coefficient, where the third feature tensor sequentially includes the second time-domain features, the grouping coefficient, and a third channel coefficient, and the third number of channels is obtained by fusing the first number of channels and the frequency-domain features; Obtaining a fourth feature tensor according to the third feature tensor, where the fourth feature tensor sequentially includes the second time-domain features and the second number of channels, and the second number of channels is obtained by fusing the grouping coefficient and the third number of channels.

4. The method according to claim 1, characterized in that, further comprising: Obtaining a training sample set, where the training sample set includes audio file samples and classification labels of the audio file samples; Copy the classification label of the audio file sample M times to obtain the classification label of each audio segment in the audio file sample, where the value of M is determined according to the number of audio segments in the audio file sample.

5. The method according to claim 1, wherein, further comprising: Obtaining a second feature label of the audio file according to the first time-domain feature; wherein, determining the type of the audio file according to the first feature label of each audio segment includes: Determining the type of the audio file according to the first feature label of each audio segment and the second feature label of the audio file.

6. The method according to claim 1, wherein, The spectrogram feature includes Mel feature.

7. The method according to claim 1, wherein, The audio file includes a music file.

8. A method for training an audio classification model, wherein, comprising: Obtaining a training sample set, the training sample set including an audio file sample and a classification label of the audio file sample; Copying the classification label of the audio file sample M times to obtain the classification label of each audio segment in the audio file sample, where the value of M is determined according to the number of audio segments in the audio file sample; Obtaining the spectrogram feature of the audio file sample; Using a neural network model to obtain a third time-domain feature of the spectrogram feature; Grouping the third time-domain feature to obtain a fourth time-domain feature of the spectrogram feature, and the feature vector of each dimension of the fourth time-domain feature represents the time-domain feature vector of each audio segment in the audio file sample; and Inputting the fourth time-domain feature into a classification model to obtain a third feature label of each audio segment in the audio file sample; Determining a first loss according to the third feature label of each audio segment in the audio file sample and the classification label of each audio segment; Updating the parameters of the classification model according to the first loss; wherein, grouping the third time-domain feature to obtain a fourth time-domain feature of the spectrogram feature includes: Obtaining a grouping coefficient of the audio file sample, and the grouping coefficient represents the number of sampling points of the audio file sample per unit time; Grouping the third time-domain feature according to the grouping coefficient to obtain the fourth time-domain feature.

9. An audio classification device, wherein, comprising: A first obtaining unit for obtaining the spectrogram feature of the audio file; A neural network model for obtaining a first time-domain feature of the spectrogram feature; A grouping unit for grouping the first time-domain feature to obtain a second time-domain feature of the spectrogram feature, and the feature vector of each dimension of the second time-domain feature represents the time-domain feature vector of each audio segment in the audio file; A classification model for inputting the second time-domain feature to obtain a first feature label of each audio segment in the audio file; A determining unit for determining the type of the audio file according to the first feature label of each audio segment; The grouping unit is specifically used for: Obtain the grouping coefficient of the audio file, where the grouping coefficient represents the number of sampling points for the audio file per unit time; Group the first time-domain feature according to the grouping coefficient to obtain the second time-domain feature.

10. An apparatus for training an audio classification model, characterized in that, it includes: A first acquisition unit for acquiring a training sample set, where the training sample set includes audio file samples and classification labels of the audio file samples; The first acquisition unit is further configured to copy the classification labels of the audio file samples M times to obtain the classification labels of each audio segment in the audio file samples, where the value of M is determined according to the number of audio segments in the audio file samples; A second acquisition unit for acquiring the spectrogram features of the audio file samples; A neural network model for obtaining the third time-domain feature of the spectrogram features; A grouping unit for grouping the third time-domain feature to obtain the fourth time-domain feature of the spectrogram features, where the feature vector of each dimension of the fourth time-domain feature represents the time-domain feature vector of each audio segment in the audio file samples; and A classification model for inputting the fourth time-domain feature into the classification model to obtain the third feature label of each audio segment in the audio file samples; A determination unit for determining a first loss according to the third feature label of each audio segment in the audio file samples and the classification label of each audio segment; An update unit for updating the parameters of the classification model according to the first loss; The grouping unit is specifically configured to: Obtain the grouping coefficient of the audio file sample, where the grouping coefficient represents the number of sampling points for the audio file sample per unit time; Group the third time-domain feature according to the grouping coefficient to obtain the fourth time-domain feature.

11. An electronic device, characterized in that, it includes a processor and a memory, and instructions are stored in the memory. When the processor executes the instructions, the processor executes the method according to any one of claims 1-8.

12. A computer storage medium, characterized in that, it is used to store a computer program, and the computer program includes the method for executing any one of claims 1-8.

13. A computer program product, characterized in that, it includes computer program code, and when the computer program code is run on an electronic device, the electronic device executes the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Audio recognition model training method and device and abnormal audio recognition method and device

    CN112259078A

  • Music genre identification method and device, equipment and storage medium

    CN113450828A

  • Music style identification method and device, storage medium and electronic equipment

    CN114141270A