An Event Audio Detection Method and System Based on Few-Shot Metric Learning
By adopting a method based on few-sample metric learning in event audio detection, the similarity is calculated through feature extraction and mask filling, the problem of degradation of detection accuracy caused by insufficient sample number is solved, and the accuracy of event audio detection is improved.
Patent Information
- Application Number
- CN202210906275.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-07-29
AI Technical Summary
Existing event audio detection methods reduce detection accuracy when the sample size is insufficient, and the small sample method based on metric learning ignores the potential connections between samples supported by the support of centralization, resulting in a decrease in detection accuracy.
A method of event audio detection based on few-sample metric learning is proposed. By obtaining the sample data set of event audio, the samples are divided into support sets and query sets, feature extraction is performed separately, and the feature set is filled based on masks through a convolutional network to calculate the similarity to train the detection model.
By learning the commonalities between samples among various types, the accuracy of event audio detection is improved and the model's detection ability of rare event audio is enhanced.
Smart Images

Figure CN115221350B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an event audio detection method and system based on few-shot metric learning. Background Art
[0002] Event audio detection is to determine the presence or absence of an event and the occurrence and termination times in a piece of audio. It has important application values in many fields, which can greatly save labor costs and is particularly widely used in fields such as security and monitoring. Event audio detection is more difficult than speech recognition because the temporal structure of event audio is complex and the frequency is uncertain, unlike human speech which contains fixed structures and rhythms. Therefore, the accuracy rate of event audio detection in the industry has always been relatively low. In recent years, methods based on deep learning have greatly improved the accuracy of event audio detection, but deep learning requires a large number of labeled training samples. For some rare event audio (with few samples), its detection accuracy will drop significantly due to insufficient training samples. At the same time, few-shot methods based on metric learning have largely alleviated the problem of poor performance caused by a small number of samples in other deep learning tasks. These methods calculate the similarity between each sample in the support set and the query set samples during training, and rely on the learned similarity calculation to infer the labels of the test samples during testing. However, this method ignores the potential connections between the samples in the support set, which will cause the deep learning model to be unable to learn some important features of this support set, such as the common characteristics of all samples and the unique characteristics of some samples. This phenomenon will greatly reduce the detection accuracy. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose an event audio detection method and system based on few-shot metric learning, which can learn the commonalities between samples of various types and improve the accuracy of event audio detection.
[0004] To achieve the above object, the first aspect of the embodiments of this application proposes an event audio detection method based on few-shot metric learning, and the method includes:
[0005] Obtain a sample data set of event audio, and divide multiple samples in the sample data set into a support set and a query set according to the sample categories;
[0006] Extract features from the support set and the query set respectively to obtain a first support sample feature set and a first query sample feature set;
[0007] Obtain a mask according to the first support sample feature set;
[0008] Padding the first support sample feature set and the first query sample feature set based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set;
[0009] Calculating the similarity according to the mask, the second support sample feature set and the second query sample feature set;
[0010] Determining the value of the loss function according to the similarity, and training the event audio detection model according to the value of the loss function until the value of the loss function meets a preset threshold to obtain a trained event audio detection model;
[0011] Inputting the event audio to be detected into the trained event audio detection model to obtain a detection result.
[0012] In some embodiments, the obtaining the sample data set of the event audio and dividing the multiple samples in the sample data set into a support set and a query set according to the sample categories includes:
[0013] Obtaining the sample data set of the event audio, the sample data set includes samples of multiple sample categories, and the number of samples corresponding to each sample category is multiple;
[0014] For each sample category, respectively obtaining K samples corresponding to the sample category, where K is an integer greater than 1;
[0015] Dividing the K samples corresponding to each sample category into support samples and query samples according to a preset ratio value to obtain the support set and the query set.
[0016] In some embodiments, the event audio detection method based on few-shot metric learning further includes:
[0017] Reducing the dimension of the first support sample feature set by using a convolutional layer to obtain a dimension-reduced sample feature set;
[0018] Calculating the mean value of the samples of the same category in the first dimension in the dimension-reduced sample feature set to obtain an average dimension-reduced sample feature set;
[0019] Regarding each feature in the average dimension-reduced sample feature set as the common representation of the corresponding category.
[0020] In some embodiments, the obtaining the mask according to the first support sample feature set includes:
[0021] Deforming the dimension-reduced sample feature set to obtain a deformed dimension-reduced sample feature set;
[0022] Deforming the deformed dimension-reduced sample feature set through a convolutional layer to obtain a mask;
[0023] Activate the mask using a normalization function.
[0024] In some embodiments, the convolutional network fills the first support sample feature set and the first query sample feature set based on the mask to obtain a second support sample feature set and a second query sample feature set, including:
[0025] Use a convolutional network to fill the first support sample feature set so that the dimension of the first support sample feature set matches the dimension of the mask, obtaining a second support sample feature set;
[0026] Use a convolutional network to fill the first query sample feature set so that the dimension of the first query sample feature set matches the dimension of the mask, obtaining a second query sample feature set.
[0027] In some embodiments, calculating the similarity according to the mask, the second support sample feature set, and the second query sample feature set includes:
[0028] Obtain the mask for each category;
[0029] Measure the Euclidean distance between the mask of each category and the second query sample feature set. The Euclidean distance represents the similarity, and the similarity calculation formula is as follows:
[0030]
[0031] where S represents the support set, Q represents the query set, represents the feature extraction network, θ represents the parameters of the prototype network, P represents the mask of each category, r(S) represents the second support sample feature set, r(Q) represents the second query sample feature set, and ⊙ represents the sequence mask calculation;
[0032] Use the Euclidean distance as the basis for determining the label of the second query sample feature set.
[0033] In some embodiments, the loss function is:
[0034]
[0035] where, represents the j-th sample of the i-th category in the support set, x q represents the q-th sample in the query set, and k represents the number of samples.
[0036] To achieve the above object, a second aspect of the embodiments of the present application proposes an event audio detection system based on few-shot metric learning. The system includes:
[0037] A data division module, configured to obtain a sample data set of event audio, and divide a plurality of samples in the sample data set into a support set and a query set according to sample categories;
[0038] A feature acquisition module, configured to perform feature extraction on the support set and the query set respectively to obtain a first support sample feature set and a first query sample feature set;
[0039] A mask acquisition module, configured to obtain a mask according to the first support sample feature set;
[0040] A feature filling module, configured to fill the first support sample feature set and the first query sample feature set based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set;
[0041] A similarity calculation module, configured to calculate a similarity according to the mask, the second support sample feature set, and the second query sample feature set;
[0042] A model training module, configured to determine a value of a loss function according to the similarity, and train an event audio detection model according to the value of the loss function until the value of the loss function meets a preset threshold to obtain a trained event audio detection model;
[0043] An event audio detection module, configured to input the event audio to be detected into the trained event audio detection model to obtain a detection result.
[0044] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, including:
[0045] At least one memory;
[0046] At least one processor;
[0047] At least one computer program;
[0048] The at least one computer program is stored in the at least one memory, and the at least one processor executes the at least one computer program to implement an event audio detection method based on few-shot metric learning described in the first aspect above.
[0049] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, the storage medium is a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the computer program is used to cause a computer to execute an event audio detection method based on few-shot metric learning described in the first aspect above.
[0050] An event audio detection method and system based on few-shot metric learning proposed in this application obtain a sample data set of event audio, and divide multiple samples in the sample data set into a support set and a query set according to the sample categories; in order to facilitate subsequent learning and generalization, feature extraction is performed on the support set and the query set respectively to obtain a first support sample feature set and a first query sample feature set. In order to learn the commonality between samples of various categories and learn the potential connections between samples in the support set, a mask is obtained according to the first support sample feature set; the first support sample feature set and the first query sample feature set are filled based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set. In order to improve the accuracy of event audio detection, the similarity is calculated according to the mask, the second support sample feature set and the second query sample feature set; the value of the loss function is determined according to the similarity, and the event audio detection model is trained according to the value of the loss function until the value of the loss function meets the preset threshold, and a trained event audio detection model is obtained. This application obtains a mask according to the first support sample feature set, and calculates the similarity according to the mask, the second support sample feature set and the second query sample feature set, which can learn the commonality between samples of various categories and improve the accuracy of event audio detection. Description of the Drawings
[0051] Figure 1 is a flowchart of an event audio detection method based on few-shot metric learning provided by an embodiment of this application;
[0052] Figure 2 is another flowchart of an event audio detection method based on few-shot metric learning provided by an embodiment of this application;
[0053] Figure 3 is Figure 1 a flowchart of step S130 in
[0054] Figure 4 is Figure 1 a flowchart of step S140 in
[0055] Figure 5 is a schematic structural diagram of an event audio detection system based on few-shot metric learning provided by an embodiment of this application;
[0056] Figure 6 is a schematic hardware structure diagram of an electronic device provided by an embodiment of this application. Detailed Embodiments
[0057] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0058] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the device or a different order from that in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0060] First, several terms involved in this application are analyzed:
[0061] Deep Learning (DL): It is a new research direction in the field of Machine Learning (ML). It is introduced into machine learning to make it closer to the original goal - Artificial Intelligence (AI).
[0062] Deep learning is to learn the internal laws and representation levels of sample data. The information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies.
[0063] Convolutional Neural Networks (CNN): It is a multi-layer supervised learning neural network. The convolutional layer and pooling layer in the hidden layer are the core modules for implementing the feature extraction function of the convolutional neural network. This network model adjusts the weight parameters in the network layer by layer in reverse by using the gradient descent method to minimize the loss function, and improves the accuracy of the network through frequent iterative training.
[0064] Euclidean distance: The Euclidean distance or Euclidean metric is the "ordinary" (i.e., straight-line) distance between two points in Euclidean space. Using this distance, Euclidean space becomes a metric space. The associated norm is called the Euclidean norm. Earlier literature referred to it as the Pythagorean metric.
[0065] Metric Learning: Metric Learning is a traditional machine learning method commonly used in face recognition. It was proposed by Eric Xing at NIPS 2002 and can be divided into two types: one is metric learning through linear transformation, and the other is metric learning through non-linear transformation. Its basic principle is to autonomously learn a metric distance function for a specific task according to different tasks. Later, metric learning was migrated to the field of text classification, especially for text processing of high-dimensional data, and metric learning has good classification effects.
[0066] Event audio detection is to determine the presence or absence of an event and the occurrence and termination times in a piece of audio. It has important application values in many fields, which can greatly save labor costs and is particularly widely used in the fields of security, monitoring, etc. Event audio detection is more difficult than speech recognition because the temporal structure of event audio is complex and the frequency is uncertain, unlike human speech which contains fixed structures and rhythms. Therefore, the accuracy rate of event audio detection in the industry has always been relatively low. In recent years, methods based on deep learning have greatly improved the accuracy of event audio detection, but deep learning requires a large number of labeled training samples. For some rare event audio (with few samples), its detection accuracy will drop significantly due to insufficient training samples. At the same time, the few-shot method based on metric learning has largely alleviated the problem of poor performance caused by a small number of samples in other deep learning tasks. This type of method calculates the similarity between each sample in the support set and the query set samples during training, and relies on the learned similarity calculation to infer the labels of the test samples during testing. However, this method ignores the potential connections between the samples in the support set, which will cause the deep learning model to be unable to learn some important features of this support set, such as the common characteristics of all samples and the unique characteristics of some samples. This phenomenon will greatly reduce the detection accuracy.
[0067] Based on this, the embodiments of the present application provide an event audio detection method and system based on few-shot metric learning. By obtaining a sample data set of event audio, multiple samples in the sample data set are divided into a support set and a query set according to the sample categories; in order to facilitate subsequent learning and generalization, feature extraction is performed on the support set and the query set respectively to obtain a first support sample feature set and a first query sample feature set. In order to learn the commonalities between samples of different classes and learn the potential connections between samples in the support set, a mask is obtained according to the first support sample feature set; the first support sample feature set and the first query sample feature set are filled based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set. In order to improve the accuracy of event audio detection, the similarity is calculated according to the mask, the second support sample feature set and the second query sample feature set; the value of the loss function is determined according to the similarity, and the event audio detection model is trained according to the value of the loss function until the value of the loss function meets a preset threshold to obtain a trained event audio detection model. The present application obtains a mask according to the first support sample feature set and calculates the similarity according to the mask, the second support sample feature set and the second query sample feature set, which can learn the commonalities between samples of different classes and improve the accuracy of event audio detection.
[0068] The event audio detection method based on few-shot metric learning, the event audio detection system based on few-shot metric learning, the electronic device and the storage medium provided by the embodiments of the present application will be specifically described through the following embodiments. First, an event audio detection method based on few-shot metric learning in the embodiments of the present application will be described.
[0069] An event audio detection method based on few-shot metric learning provided by the embodiments of the present application relates to the field of artificial intelligence technology. The event audio detection method based on few-shot metric learning provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the event audio detection method based on few-shot metric learning, etc., but is not limited to the above forms.
[0070] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0071] Please refer to Figure 1 , Figure 1 FIG. is an alternative flowchart of an event audio detection method based on few-shot metric learning provided by an embodiment of this application. Figure 1 The method in can specifically include but is not limited to steps S110 to S170.
[0072] Step S110: Obtain a sample data set of event audio, and divide multiple samples in the sample data set into a support set and a query set according to the sample categories.
[0073] Step S120: Respectively perform feature extraction on the support set and the query set to obtain a first support sample feature set and a first query sample feature set.
[0074] Step S130: Obtain a mask according to the first support sample feature set.
[0075] Step S140: Fill the first support sample feature set and the first query sample feature set based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set.
[0076] Step S150: Calculate the similarity according to the mask, the second support sample feature set, and the second query sample feature set.
[0077] Step S160: Determine the value of the loss function according to the similarity, and train the event audio detection model according to the value of the loss function until the value of the loss function meets a preset threshold to obtain a trained event audio detection model.
[0078] Step S170: Input the event audio to be detected into the trained event audio detection model to obtain a detection result.
[0079] In steps S110 to S170 of some embodiments, by obtaining a sample data set of event audio, a plurality of samples in the sample data set are divided into a support set and a query set according to sample categories; in order to facilitate subsequent learning and generalization, feature extraction is performed on the support set and the query set respectively to obtain a first support sample feature set and a first query sample feature set. In order to learn the commonalities between samples of various categories and learn the potential connections between samples in the support set, a mask is obtained according to the first support sample feature set; the first support sample feature set and the first query sample feature set are filled based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set. In order to improve the accuracy of event audio detection, the similarity is calculated according to the mask, the second support sample feature set, and the second query sample feature set; the value of the loss function is determined according to the similarity, and the event audio detection model is trained according to the value of the loss function until the value of the loss function meets a preset threshold, and a trained event audio detection model is obtained. In this application, a mask is obtained according to the first support sample feature set, and the similarity is calculated according to the mask, the second support sample feature set, and the second query sample feature set, which can learn the commonalities between samples of various categories and improve the accuracy of event audio detection.
[0080] In step S110 of some embodiments, in order to implement event audio detection of few-shot metric learning, first, a sample data set of event audio is obtained, and a plurality of samples in the sample data set are divided into a support set and a query set according to sample categories. Specifically, a sample data set of event audio is obtained from DCASE task5. The sample data set includes samples of multiple sample categories, and the number of samples corresponding to each sample category is multiple; for each sample category, K samples corresponding to the sample category are obtained respectively, where K is an integer greater than 1; according to a preset ratio value, the K samples corresponding to each sample category are divided into support samples and query samples to obtain a support set and a query set.
[0081] It should be noted that DCASE task5 is a public data set, which will not be described in detail in this embodiment, and this embodiment does not limit the use of only the DCASE task5 data set.
[0082] It should be noted that according to the preset ratio value, the K samples corresponding to each sample category are divided into support samples and query samples. The preset ratio value in this embodiment can be changed as needed, and this embodiment does not make a limitation.
[0083] In step S120 of some embodiments, in order to facilitate subsequent learning and generalization, feature extraction is performed on the support set and the query set respectively to obtain a first support sample feature set and a first query sample feature set. Specifically, by using a convolutional network f φ to extract the features of the samples, this convolutional network f φIt contains five layers. The first two layers are convolutional layers with a 3x3 convolutional kernel, the third layer is a batch normalization layer, the fourth layer is a max pooling layer with a size of 4x4, and the fifth layer is a max pooling layer with a size of 1x1. Through the convolutional network Feature extraction is respectively performed on the support set S and the query set Q in step S110 to obtain the first support sample feature set and the first query sample feature set
[0084] In step S130 of some embodiments, in order to learn the commonalities between samples of different classes and learn the potential connections between samples in the support set, a mask is obtained according to the first support sample feature set. Specifically, a mask is obtained according to the first support sample feature set, and the mask is used as the commonality between samples of each class to learn the commonalities between samples of different classes and learn the potential connections between samples in the support set.
[0085] In step S140 of some embodiments, in order to learn the commonalities between samples of different classes and learn the potential connections between samples in the support set, the convolutional network fills the first support sample feature set and the first query sample feature set based on the mask to obtain the second support sample feature set and the second query sample feature set. Specifically, the convolutional network fills the first support sample feature set and the first query sample feature set based on the mask so that the dimension of the mask matches the dimensions of the first support sample feature set and the first query sample feature set, thereby learning the commonalities between samples of different classes and learning the potential connections between samples in the support set.
[0086] In step S150 of some embodiments, in order to improve the accuracy of event audio detection, the similarity is calculated according to the mask, the second support sample feature set, and the second query sample feature set. Specifically, the mask of each class is obtained, and the Euclidean distance between the mask of each class and the second query sample feature set is measured. The Euclidean distance represents the similarity, and the similarity calculation formula is as follows:
[0087]
[0088] where S represents the support set, Q represents the query set, represents the feature extraction network, θ represents the parameters of the prototype network, P represents the mask of each class, r(S) represents the second support sample feature set, r(Q) represents the second query sample feature set, and ⊙ represents the sequence mask calculation;
[0089] The Euclidean distance is used as the basis for discriminating the label of the second query sample feature set. The second query sample feature set is discriminated as the class with the smallest distance.
[0090] It should be noted that the prototype network in this embodiment is a classic few-shot method based on metric learning. The principle of the prototype network is to take the average of all sample representations of each class in the support set as the prototype of this class of samples, and then measure the Euclidean distance between the query set samples and each class prototype as the basis for discriminating the prototype labels of each class, which will not be elaborated here.
[0091] In step S160 of some embodiments, after calculating the similarity, determine the value of the loss function according to the similarity, and train the event audio detection model according to the value of the loss function until the value of the loss function meets the preset threshold, and obtain the trained event audio detection model. Specifically, the loss function is:
[0092]
[0093] Wherein, represents the j-th sample of the i-th class in the support set, x q represents the q-th sample in the query set, and k represents the number of samples;
[0094] After obtaining the similarity, determine the value L of the loss function according to the similarity through the above formula, and train the event audio detection model according to the value of the loss function until the value of the loss function meets the preset threshold, and obtain the trained event audio detection model.
[0095] It should be noted that the preset threshold of this embodiment can be adjusted according to actual needs, which will not be elaborated here.
[0096] In step S170 of some embodiments, after obtaining the trained event audio detection model, input the event audio to be detected into the trained event audio detection model to obtain the detection result. Specifically, using the trained event audio detection model to detect the event audio to be detected can obtain the detection result. The embodiment of the present application can learn the commonalities between various samples through the event audio detection method based on few-shot metric learning, and improve the accuracy of event audio detection.
[0097] Please refer to Figure 2 , Figure 2 which is another optional flowchart of the event audio detection method based on few-shot metric learning provided by the embodiment of the present application. Figure 2 The method in Figure 2 includes but is not limited to step S210 and step S230. The following will introduce these three steps in detail with reference to
[0098] Step S210, use a convolutional layer to reduce the dimension of the first support sample feature set to obtain a dimension-reduced sample feature set;
[0099] Step S220: Calculate the mean of samples of the same category in the first dimension in the dimensionality-reduced sample feature set to obtain the average dimensionality-reduced sample feature set.
[0100] Step S230: Take each feature in the average dimensionality-reduced sample feature set as the common representation of the corresponding category.
[0101] In steps S210 and S230 of some embodiments, in order to reduce the dimension and computational amount, a convolutional layer is used to reduce the dimension of the first support sample feature set to obtain the dimensionality-reduced sample feature set; calculate the mean of samples of the same category in the first dimension in the dimensionality-reduced sample feature set to obtain the average dimensionality-reduced sample feature set; take each feature in the average dimensionality-reduced sample feature set as the common representation of the corresponding category. Specifically, take the first support sample feature set as the input, with the aim of discovering the common features of K samples of each category in the support set S. Assume the output of the support set S after passing through the feature extraction network has a dimension of (N×K, m1, w1, h1), where m1, w1, and h1 respectively represent the number of channels, spatial width, and spatial height. Use a convolutional layer to reduce the dimension of the first support sample feature set to obtain the dimensionality-reduced sample feature set, calculate the mean of samples of the same category in the first dimension in the dimensionality-reduced sample feature set to obtain the average dimensionality-reduced sample feature set, and take each feature o: (N, m2, w2, h2) in the average dimensionality-reduced sample feature set as the common representation of the corresponding category. In the embodiments of the present application, the convolutional layer is used to reduce the dimension of the first support sample feature set, reducing the computational amount, and by taking each feature in the average dimensionality-reduced sample feature set as the common representation of the corresponding category, the commonalities between samples of different categories can be learned.
[0102] Please refer to Figure 3 , Figure 3 which is a flowchart of the specific method of step S130 in some embodiments of the present application. In some embodiments of the present application, step S130 specifically includes but is not limited to steps S310 and S330. The following combines Figure 3 to introduce these three steps in detail.
[0103] Step S310: Deform the dimensionality-reduced sample feature set to obtain the deformed dimensionality-reduced sample feature set.
[0104] Step S320: Use a convolutional layer to deform the deformed dimensionality-reduced sample feature set to obtain a mask.
[0105] Step S330: Activate the mask using a normalization function.
[0106] In step S310 and step S330 of some embodiments, in order to learn the commonalities between samples of different categories and learn the potential connections between samples in the support set, the reduced-dimensionality sample feature set is deformed to obtain a deformed reduced-dimensionality sample feature set, and the deformed reduced-dimensionality sample feature set is deformed by a convolution layer to obtain a mask, and the mask is activated by a normalization function. Specifically, each feature o:(N,m2,w2,h2) in the reduced-dimensionality sample feature set is used as input, with the purpose of discovering the characteristics of each category. First, o:(N,m2,w2,h2) is deformed into Using convolutional layers The mask is transformed into a mask P: (1, m3, w3, h3), and a normalized exponential function is used to act on the channel dimension of the mask P to activate the mask. In the embodiment of the present application, the deformed reduced-dimensional sample feature set is deformed by a convolution layer to obtain a mask, so as to learn the commonalities between samples of different categories and learn the potential connections between samples in the support set.
[0107] See also Figure 4 , Figure 4 is a flowchart of a specific method of step S140 in some embodiments of the present application. In some embodiments of the present application, step S140 specifically includes but is not limited to step S410 and step S420. Figure 4 These two steps are introduced in detail.
[0108] Step S410, using a convolutional network to fill the first support sample feature set so that the dimension of the first support sample feature set matches the dimension of the mask, and obtaining a second support sample feature set;
[0109] Step S420: Use a convolutional network to fill the first query sample feature set so that the dimension of the first query sample feature set matches the dimension of the mask, and obtain a second query sample feature set.
[0110] In step S410 and step S420 of some embodiments, in order to learn the commonalities between samples of different classes and learn the potential connections between samples in the support set, a convolutional network is used to fill the first support sample feature set so that the dimension of the first support sample feature set matches the dimension of the mask, and a second support sample feature set is obtained. A convolutional network is used to fill the first query sample feature set so that the dimension of the first query sample feature set matches the dimension of the mask, and a second query sample feature set is obtained. Specifically, in order to make the dimension of the mask P match the dimension of the first support sample feature set and the first query sample feature set The dimensions of and On the top, a convolutional network-based filling module is used to and The dimension alignment of r(.) is (N×K, m3, w3, h3). The mask P performs sequence mask calculation on r(S) and r(Q) as the representation of the commonality within each category and the characteristics between categories. In the embodiment of the present application, the first support sample feature set and the first query sample feature set are filled based on the mask, so that the dimension of the mask matches the dimensions of the first support sample feature set and the first query sample feature set, thereby learning the commonality between samples of different categories, learning the potential connections between samples in the support set, and improving the accuracy of event audio detection.
[0111] It should be noted that sequence mask calculation is a conventional technique in the art, and this embodiment will not be described in detail.
[0112] An event audio detection method based on few-shot metric learning provided by an embodiment of the present application includes: obtaining a sample data set of event audio, and dividing multiple samples in the sample data set into a support set and a query set according to the sample category; for promoting subsequent learning and generalization, feature extraction is respectively performed on the support set and the query set to obtain a first support sample feature set and a first query sample feature set. To reduce the amount of calculation, the first support sample feature set is dimension-reduced through a convolutional layer. To learn the commonality between samples of different categories and the potential connections between samples in the support set, a mask is obtained according to the first support sample feature set; the first support sample feature set and the first query sample feature set are filled based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set. To improve the accuracy of event audio detection, the similarity is calculated according to the mask, the second support sample feature set, and the second query sample feature set; the value of the loss function is determined according to the similarity, and the event audio detection model is trained according to the value of the loss function until the value of the loss function meets a preset threshold, and a trained event audio detection model is obtained. The present application obtains a mask according to the first support sample feature set and calculates the similarity according to the mask, the second support sample feature set, and the second query sample feature set, which can learn the commonality between various samples and improve the accuracy of event audio detection.
[0113] Please refer to Figure 5 In addition, an event audio detection system based on few-shot metric learning is provided in an embodiment of the present application, which can implement the above-mentioned event audio detection method based on few-shot metric learning. The system includes a data division module 510, a feature acquisition module 520, a mask acquisition module 530, a feature filling module 540, a similarity calculation module 550, a model training module 560, and an event audio detection module 570.
[0114] The data division module 510 is configured to obtain a sample data set of event audio and divide multiple samples in the sample data set into a support set and a query set according to the sample category;
[0115] A feature acquisition module 520, configured to perform feature extraction on a support set and a query set respectively, to obtain a first support sample feature set and a first query sample feature set;
[0116] A mask acquisition module 530, configured to obtain a mask according to the first support sample feature set;
[0117] A feature filling module 540, configured to fill the first support sample feature set and the first query sample feature set based on the mask through a convolutional network, to obtain a second support sample feature set and a second query sample feature set;
[0118] A similarity calculation module 550, configured to calculate a similarity according to the mask, the second support sample feature set, and the second query sample feature set;
[0119] A model training module 560, configured to determine a value of a loss function according to the similarity, and train an event audio detection model according to the value of the loss function until the value of the loss function meets a preset threshold, to obtain a trained event audio detection model;
[0120] An event audio detection module 570, configured to input an event audio to be detected into the trained event audio detection model, to obtain a detection result.
[0121] It should be noted that the event audio detection based on few-shot metric learning in the embodiments of the present application is used to implement the above-mentioned event audio detection method based on few-shot metric learning. The event audio detection system based on few-shot metric learning in the embodiments of the present application corresponds to the foregoing event audio detection method based on few-shot metric learning. For the specific processing process, please refer to the foregoing event audio detection method based on few-shot metric learning, which will not be elaborated here.
[0122] An event audio detection system based on few-shot metric learning provided by an embodiment of the present application can implement the above-mentioned event audio detection method based on few-shot metric learning. By obtaining a sample data set of event audio, multiple samples in the sample data set are divided into a support set and a query set according to the sample categories; to facilitate subsequent learning and generalization, feature extraction is performed on the support set and the query set respectively to obtain a first support sample feature set and a first query sample feature set. To reduce the computational amount, the first support sample feature set is dimensionally reduced through a convolutional layer. To learn the commonality between samples of various categories and learn the potential connection between samples in the support set, a mask is obtained according to the first support sample feature set; the first support sample feature set and the first query sample feature set are filled based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set. To improve the accuracy of event audio detection, the similarity is calculated according to the mask, the second support sample feature set, and the second query sample feature set; the value of the loss function is determined according to the similarity, and the event audio detection model is trained according to the value of the loss function until the value of the loss function meets a preset threshold to obtain a trained event audio detection model. According to the present application, a mask is obtained according to the first support sample feature set, and the similarity is calculated according to the mask, the second support sample feature set, and the second query sample feature set, which can learn the commonality between samples of various categories and improve the accuracy of event audio detection.
[0123] An embodiment of the present application further provides an electronic device, which includes: at least one memory, at least one processor, at least one computer program, at least one computer program is stored in at least one memory, and at least one processor executes at least one computer program to implement the event audio detection method based on few-shot metric learning in any one of the above embodiments. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0124] Please refer to Figure 6 , Figure 6 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0125] A processor 610, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0126] The memory 620 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 620 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 620 and are called by the processor 610 to execute a method for event audio detection based on few-shot metric learning according to an embodiment of the present application;
[0127] The input / output interface 630 is used to implement information input and output;
[0128] The communication interface 640 is used to implement communication and interaction between this device and other devices. It can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0129] The bus 650 transmits information between the various components of the device (such as the processor 610, the memory 620, the input / output interface 630, and the communication interface 640);
[0130] Among them, the processor 610, the memory 620, the input / output interface 630, and the communication interface 640 are communicatively connected to each other inside the device through the bus 650.
[0131] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is used to make a computer execute the method for event audio detection based on few-shot metric learning according to any one of the above embodiments.
[0132] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0133] The embodiments described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. As those skilled in the art know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0134] Those skilled in the art can understand that Figures 1 to 4 the technical solutions shown in do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0135] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0136] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0137] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0138] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one)" or its similar expression below refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0139] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0140] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0142] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0143] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.
Claims
1. An event audio detection method based on few-shot metric learning, characterized in that, The method includes: Obtain a sample data set of event audio, and divide multiple samples in the sample data set into a support set and a query set according to sample categories; Extract features from the support set and the query set respectively to obtain a first support sample feature set and a first query sample feature set; Obtain a mask according to the first support sample feature set; Fill the first support sample feature set and the first query sample feature set based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set; Calculate the similarity according to the mask, the second support sample feature set and the second query sample feature set; where: Obtain the mask for each category; Measure the Euclidean distance between the mask of each category and the second query sample feature set, and the Euclidean distance represents the similarity. The similarity calculation formula is as follows: Among them, represents the support set, represents the query set, represents the feature extraction network, represents the parameters of the prototype network, represents the mask for each of the said categories, represents the second support sample feature set, represents the second query sample feature set, represents the sequence mask calculation, represents the Euclidean distance calculation function; Use the Euclidean distance as the basis for determining the label of the second query sample feature set; Determine the value of the loss function according to the similarity, and train the event audio detection model according to the value of the loss function until the value of the loss function meets a preset threshold to obtain a trained event audio detection model; where the loss function is: in, Express support for the concentration Class samples, Represents the first samples, represents the number of samples; Input the event audio to be detected into the trained event audio detection model to obtain a detection result.
2. The event audio detection method based on few-shot metric learning according to claim 1, characterized in that, The obtaining of the sample data set of event audio and dividing multiple samples in the sample data set into a support set and a query set according to sample categories includes: Obtain a sample data set of event audio, where the sample data set includes samples of multiple sample categories, and the number of samples corresponding to each sample category is multiple; For each sample category, obtain K samples corresponding to the sample category respectively, where K is an integer greater than 1; Divide the K samples corresponding to each sample category into support samples and query samples according to a preset ratio value to obtain the support set and the query set.
3. The event audio detection method based on few-shot metric learning according to claim 1, characterized in that, The event audio detection method based on few-shot metric learning further includes: Use a convolutional layer to reduce the dimension of the first support sample feature set to obtain a reduced-dimension sample feature set; Calculate the mean value of samples of the same category in the first dimension in the reduced-dimension sample feature set to obtain an average reduced-dimension sample feature set; Use each feature in the average reduced-dimension sample feature set as the common representation of the corresponding category.
4. The event audio detection method based on few-shot metric learning according to claim 3, characterized in that, The obtaining of the mask according to the first support sample feature set includes: Deform the reduced-dimension sample feature set to obtain a deformed reduced-dimension sample feature set; Deform the deformed reduced-dimension sample feature set through a convolutional layer to obtain a mask; Activate the mask using a normalization function.
5. The event audio detection method based on few-shot metric learning according to claim 1, characterized in that, The filling of the first support sample feature set and the first query sample feature set based on the mask through a convolutional network to obtain a second support sample feature set and a second query sample feature set includes: Use a convolutional network to fill the first support sample feature set to make the dimension of the first support sample feature set match the dimension of the mask to obtain a second support sample feature set; Using a convolutional network, pad the first query sample feature set so that the dimension of the first query sample feature set matches the dimension of the mask, obtaining a second query sample feature set.
6. An event audio detection system based on few-shot metric learning, characterized in that, The system includes: A data partitioning module, configured to obtain a sample data set of event audio, and partition multiple samples in the sample data set into a support set and a query set according to sample categories; A feature acquisition module, configured to perform feature extraction on the support set and the query set respectively, obtaining a first support sample feature set and a first query sample feature set; A mask acquisition module, configured to obtain a mask according to the first support sample feature set; A feature padding module, configured to pad the first support sample feature set and the first query sample feature set based on the mask through a convolutional network, obtaining a second support sample feature set and a second query sample feature set; A similarity calculation module, configured to calculate a similarity according to the mask, the second support sample feature set, and the second query sample feature set; where: Obtain the mask for each category; Measure the Euclidean distance between the mask of each category and the second query sample feature set, the Euclidean distance represents the similarity, and the similarity calculation formula is as follows: Among them, represents the support set, represents the query set, represents the feature extraction network, represents the parameters of the prototype network, represents the mask for each of the said categories, represents the second support sample feature set, represents the second query sample feature set, represents the sequence mask calculation, represents the Euclidean distance calculation function; Use the Euclidean distance as the basis for discriminating the label of the second query sample feature set; A model training module, configured to determine the value of a loss function according to the similarity, and train an event audio detection model according to the value of the loss function until the value of the loss function meets a preset threshold, obtaining a trained event audio detection model; where the loss function is: Among them, represents the th sample of the th class in the support set, represents the th sample in the query set, and represents the number of samples. An event audio detection module, configured to input the event audio to be detected into the trained event audio detection model, obtaining a detection result.
7. An electronic device, characterized in that, Includes: At least one memory; At least one processor; At least one computer program; The at least one computer program is stored in the at least one memory, and the at least one processor executes the at least one computer program to implement: An event audio detection method based on few-shot metric learning according to any one of claims 1 to 5.
8. A storage medium, the storage medium being a computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is used to cause a computer to execute: An event audio detection method based on few-shot metric learning according to any one of claims 1 to 5.
Citation Information
Patent Citations
Classification model training and classification method and device, computer equipment and storage medium
CN113299346A
Speech classification network training method and device, computing equipment and storage medium
CN113593611A