Environmental Audio Open-Set Recognition Method, Model and Electronic Device

Through closed-set classification and generative adversarial network combined with kernel density estimation methods, edge samples are generated and sample distribution is controlled, which solves the problem of unknown categories recognition in open environments, and achieves good in-class aggregation and identification accuracy under open space risks.

CN119862454BActive Publication Date: 2025-07-01NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510347006.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-01
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The prior art cannot effectively identify unknown categories of sound events in open environments, resulting in false alarms and false recognition, especially in autonomous driving and audio monitoring, and the existing open set recognition algorithms do not fully utilize generative models and prototype point learning.

Method used

The closed-set classification algorithm is used for preliminary classification, combining the generative adversarial network and kernel density estimation, and edge samples are generated by the generation of the loss function of the adversarial network, the density loss and offset loss training of the kernel density estimation, and edge samples are generated, and sample distribution is controlled using the attraction-mutual reversal point loss to identify unknown categories, while maintaining intra-class aggregation.

Benefits of technology

Effectively identify unknown categories under open space risks, maintain good in-class aggregation, and improve the accuracy and security of environmental audio identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862454B_ABST
    Figure CN119862454B_ABST
Patent Text Reader

Abstract

The present invention discloses an environmental audio open-set recognition method, model, and electronic device. The method includes: preliminary classification, obtaining a first output vector, dimensionality reduction into a second output vector, training using the loss function of a generative adversarial network, density loss, and offset loss between them, dividing into C mutually distinct subsets according to classification labels, assigning reciprocal points to and maximizing the difference between and, assigning attractors to to control the sample distribution, and using to control the open-set data in a low-magnitude region, and calculating the attractor-reciprocal point loss. According to the environmental audio open-set recognition method of the present invention, unknown categories can be recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of environmental audio recognition, and more particularly, to an environmental audio open-set recognition method, an environmental audio open-set recognition model, and an electronic device. Background Art

[0002] Under the closed-set assumption, significant achievements have been made in acoustic target recognition using various traditional machine learning algorithms and deep learning methods. However, the closed-set assumption has limitations in real-world scenarios. For example, in the field of autonomous driving, it is unrealistic to collect acoustic event samples of all possible classes for model training. Even in a relatively controlled indoor environment, it is inevitable to encounter objects outside the training set. When directly applying models trained based on the closed-set paradigm to an open environment, they tend to confidently classify unknown classes as one of the known classes in the training set. This is because the classification problem often generates probabilities through the Softmax layer in the output layer, and the dimension of this layer is fixed. This idealized closed-set assumption, especially in practical applications such as audio surveillance and autonomous driving, may lead to false alarms or misidentifications, causing serious safety problems. Although open-set recognition algorithms have been studied in non-audio application fields, they do not fully utilize generative models, and the prototype point learning is relatively single, without considering the risks in open spaces, and cannot guarantee good intra-class aggregation. Summary of the Invention

[0003] The present invention aims to at least solve one of the technical problems existing in the prior art. To this end, the present invention proposes an environmental audio open-set recognition method, which can identify unknown classes and maintain good intra-class aggregation while considering the risks in open spaces.

[0004] The present invention also proposes an environmental audio open-set recognition model.

[0005] The present invention also proposes an electronic device.

[0006] According to an embodiment of the first aspect of the present invention, the environmental audio open-set recognition method includes:

[0007] Preliminarily classifying the environmental audio using a closed-set classification algorithm and obtaining a first output vector and a second output vector after dimensionality reduction ;

[0008] Taking the as a generated training set, and using kernel density estimation to act on a generative adversarial network to generate first marginal samples , and obtaining second marginal samples after dimensionality reduction ;

[0009] Using the loss function of the generative adversarial network , the density loss of the kernel density estimation and the with the offset loss between are used for training:

[0010]

[0011] In the formula, , are constant coefficients; represents the training of the generative adversarial network; G , D represent the generator and the discriminator respectively;

[0012] According to C classification labels, the is divided into C distinct subsets , , reciprocal points are assigned to the , and the difference between the and the with the is maximized. Attractors are assigned to each to control the sample distribution of the , and the is used to control the open-set data in the low-magnitude region. The attractor-reciprocal point loss is calculated according to the following formula : :

[0013]

[0014] In the formula, represents the loss for constraining the ; represents the loss for constraining the ; is the distance cross-entropy loss; is the loss for constraining the data distribution range; , are hyperparameters used to control the proportion of the loss function.

[0015] According to the environmental audio open-set recognition method of the embodiments of the present invention, unknown classes can be recognized, and good intra-class aggregation can be maintained while considering the risks in the open space.

[0016] In addition, the environmental audio open-set recognition method according to the embodiments of the present invention further has the following additional technical features:

[0017] According to some embodiments of the present invention, the environmental audio open-set recognition method further includes:

[0018] Obtain each of the The set composed of the center points ;

[0019] Calculate the in the th data point and the center point The distance between , ;

[0020] Use the distribution function and kernel function to fit the probability density function , and calculate the belongs to the interval , probability:

[0021]

[0022] Calculate the according to the following formula:

[0023]

[0024] In the formula, is the number of the second marginal samples; is the first label encoding; is the in the data point and the The distance between; .

[0025] In some embodiments of the present invention, calculate the according to the following formula:

[0026]

[0027] In the formula, belongs to the class sample mean.

[0028] In some specific embodiments of the present invention, the environmental audio open set recognition method further includes:

[0029] Obtain the in each data point feature vector ;

[0030] Calculate the and each pair of the and the The distance between, and according to the following formula, the is classified as the class:

[0031]

[0032] In the formula, represents the category corresponding to the maximum value among C results in the brackets.

[0033] In some embodiments of the present invention, the is calculated according to the following formula:

[0034]

[0035] In the formula, is the number of batches of the current training; is the second label encoding; is the belongs to probability of class; is a hyperparameter for controlling the distance-probability conversion; is a hyperparameter for controlling that the denominator is not zero.

[0036] In some specific embodiments of the present invention, the , the is calculated according to the following formula:

[0037]

[0038] In the formula, is a hyperparameter for controlling the proportion of the Euclidean distance and the dot product.

[0039] In some embodiments of the present invention, the is calculated according to the following formula:

[0040]

[0041] In the formula, is the boundary value.

[0042] In some specific embodiments of the present invention, the is calculated according to the following formula:

[0043]

[0044] In the formula, represents calculating the two-norm.

[0045] The environmental audio open-set recognition model according to the second aspect embodiment of the present invention includes: a first classifier, the first classifier includes a feature extraction layer, a convolutional layer, a pre-logit layer, a first logit layer and a first kernel density estimation layer, the first classifier uses a closed-set classification algorithm to preliminarily classify environmental audio, and the pre-logit layer is used to output a first output vector , the first logit layer is used to reduce the dimension and output a second output vector , the first kernel density estimation layer is used to generate a probability density curve for each category; a generative adversarial network, the generative adversarial network includes a generator, a discriminator, a second logit layer, and a second kernel density estimation layer, the generator uses the kernel density estimation function of the second kernel density estimation layer to generate a first marginal sample , the second logit layer is used to reduce the dimension and output a second marginal sample to the second kernel density estimation layer, the discriminator is used to determine whether the belongs to the , the generative adversarial network uses the loss function of the generative adversarial network , the density loss of the kernel density estimation and the and the offset loss between for training; a second classifier, the second classifier includes a plurality of convolutional layers, the second classifier is used to perform attractor-reciprocal point training on the , the , divide the into mutually distinct subsets according to classification labels , , assign reciprocal points to the and maximize the difference between the and the , assign attractors to each to control the sample distribution of the , and use the to control the open set data in a low-magnitude region, and calculate the attractor-reciprocal point loss according to the loss constraining the , the loss constraining the , the loss constraining the , the distance cross-entropy loss and the loss constraining the data distribution range .

[0046] The environmental audio open set recognition model according to the embodiment of the present invention can identify unknown categories and maintain good intra-class aggregation while considering the open space risk.

[0047] An electronic device according to an embodiment of the third aspect of the present invention includes a processor and a memory. The processor is connected to the memory, and the memory is used to store a computer program. When the computer program is executed by the processor, the environmental audio open-set recognition method described in the embodiment of the first aspect of the present invention is implemented.

[0048] The electronic device according to the embodiment of the present invention can identify unknown categories and maintain good intra-class aggregation while considering the risks in open spaces.

[0049] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0050] Figure 1 It is a schematic structural diagram of an environmental audio open-set recognition model according to an embodiment of the present invention;

[0051] Figure 2 It is a schematic structural diagram of a first classifier according to an embodiment of the present invention;

[0052] Figure 3 It is a schematic structural diagram of a generative adversarial network according to an embodiment of the present invention;

[0053] Figure 4 It is a schematic structural diagram of a second classifier according to an embodiment of the present invention;

[0054] Figure 5 It is a schematic diagram of sample distribution according to an embodiment of the present invention; among them, (a) is the original sample distribution; (b) is the original GAN-generated sample distribution; (c) is the generated sample distribution without Offset Loss constraint; (d) is the generated sample distribution of Offset Loss + Density Loss;

[0055] Figure 6 It is a schematic diagram of the training result according to an embodiment of the present invention. Detailed Embodiments

[0056] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0057] The environmental audio open-set recognition method and the environmental audio open-set recognition model according to the embodiments of the present invention will be described below with reference to the drawings.

[0058] As Figures 1-6As shown in the figure, the environmental audio open-set recognition model according to an embodiment of the present invention includes: a first classifier, a generative adversarial network, and a second classifier.

[0059] Specifically, the first classifier includes a feature extraction layer, a two-dimensional convolutional layer, a pre-logit layer, a first logit layer, and a first kernel density estimation layer. The first classifier uses a closed-set classification algorithm to preliminarily classify environmental audio, and the environmental audio is a closed-set data set. The feature extraction layer uses MFCC (Mel-scale Frequency Cepstral Coefficients) for feature extraction. After passing through the two-dimensional convolutional layer, the pre-logit layer outputs a first output vector , which is used as the training set for sample generation. Since the samples at the edge of the pre-logit layer may not be aligned with the edge samples of the first logit layer, constraints need to be imposed on the first logit layer. The first logit layer is used to reduce the dimension and output a second output vector . The first kernel density estimation layer is used to generate a probability density curve for each category.

[0060] Since open-set samples are unknown, even if the closed-set classification algorithm has excellent performance, it is difficult to apply to open-set recognition problems without open-set samples as a reference. And interference samples generally surround the sample clusters, bringing great difficulties to the model's judgment. Therefore, the present invention is based on a kernel density estimation-constrained generative adversarial network (GAN) to generate edge samples for data augmentation.

[0061] Obtain and then the GAN can be trained. The generative adversarial network includes a generator, a discriminator, a second logit layer, and a second kernel density estimation layer. Through noise sampling, the generator receives a random noise distribution , and tries to generate forged samples close to , that is, the first edge samples , to deceive the discriminator. The discriminator judges whether the input sample belongs to a real sample or a forged sample, and reaches the training goal through adversarial training between the two. The second logit layer is used to reduce the dimension and output second edge samples to the second kernel density estimation layer. Calculate the loss function of the generative adversarial network according to the following formula

[0062] (1)

[0063] In the formula, Indicates the training of the generative adversarial network; G 、 D respectively represent the generator and the discriminator; Indicates the probability that the sample comes from ; Indicates the probability that the sample comes from . This optimization objective function can make the distribution of forged samples approximate the original sample distribution as much as possible, as shown in Figure 5 (a) and Figure 5 (b).

[0064] Since belongs to high-dimensional vectors, and the size of the environmental audio dataset is usually much smaller than that of image and text datasets, it is difficult to fit its distribution. In addition, we need to generate its marginal data for each type of sample. If the original data is understood as several hypersphere distributions, the marginal samples are wrapped around the periphery of these hyperspheres, and it is also difficult in terms of visualization. To simplify the modeling, is divided into subsets . After calculating the center points for each type of sample, a set of center points is obtained, and the corresponding distances are calculated to obtain a new one-dimensional data set :

[0065] (2)

[0066] In the formula, is the distance set of the th class in the known classes, and is the distance function. For example, when calculating the distance between the th data point in and the center point , . Converting to distance set operations can reduce high-dimensional data points to one-dimensional data points, making it more convenient for modeling. When plotting it as a histogram, the marginal samples to be generated are equivalent to the samples at the tail of the histogram. Similarly, represents converting the data in into the corresponding distance set.

[0067] Next, use the kernel density estimation of the second kernel density estimation layer to model . Assume that the cumulative distribution function of is , and the probability density function is , then:

[0068] (3)

[0069] Introduce the empirical distribution function :

[0070] (4)

[0071] is an unbiased estimator of, and is approximated by using the ratio of the number of occurrences of in observations to to describe . Substitute it into and introduce the kernel function:

[0072] (5)

[0073] In the formula, is the kernel function, is the bandwidth of the kernel function.

[0074] Thus, use the distribution function of and the kernel function to fit the probability density function , and calculate the probability that belongs to the interval , according to the following formula:

[0075] (6)

[0076] Since the generated data points are around each class of samples, that is, the data points farther from the corresponding center point, therefore, calculate according to the following formula:

[0077] (7)

[0078] In the formula, is the number of second marginal samples; is the first label encoding; is the distance between the data points in and .

[0079] In addition, to make the generated samples as evenly distributed as possible around the corresponding class hypersphere and avoid the network output from deviating in a certain direction towards the hypersphere (such as Figure 5(as shown in (c)), it is necessary to apply the Offset Loss constraint to make the network converge in the expected direction. In fact, the forged samples to be generated are understood as the spherical shell of a hypersphere, and this spherical shell should share the same center with the hypersphere corresponding to the category. The uniform distribution of the generated forged samples on the periphery of the hypersphere can be ensured by constraining the distance between the center of the spherical shell and the center of the hypersphere. Therefore, the offset loss is calculated according to the following formula :

[0080] (8)

[0081] In the formula, is the sample mean belonging to the th category.

[0082] Thus, the generator uses the loss function of the generative adversarial network , the density loss of kernel density estimation and and the offset loss between for training:

[0083] (9)

[0084] In the formula, , are constant coefficients; represents the training of the generative adversarial network; G , D represent the generator and the discriminator respectively.

[0085] The second classifier includes multiple convolutional layers. The second classifier is used to perform attractor-reciprocal point training on , . According to classification labels, is divided into distinct subsets , . Assign reciprocal points to and maximize the difference between and . The role of the reciprocal point is to push the samples apart as much as possible. Assign attractors to each to control the sample distribution of . The role of the attractor is to ensure that the distribution of one class of samples is as small as possible while the reciprocal point pushes them apart, and use to control the open-set data in the low-magnitude region. Calculate the attractor-reciprocal point loss according to the following formula :

[0086] (10)

[0087] In the formula, Express Loss of restraint; Express Loss of restraint; is the distance cross entropy loss; It is the loss of constraining the data distribution range; , is a hyperparameter used to control the weight of the loss function. and Conduct training, The data in is pushed away from the low-density region by the dual point, while maintaining strong intra-class consistency under the influence of the attractor; at the same time, The data in is constrained to remain in low-density regions.

[0088] Specifically, obtain Each data point in The eigenvector of ;

[0089] calculate With each pair and The distance between them is calculated according to the following formula Classified as kind:

[0090] (11)

[0091] In the formula, Indicates the category corresponding to the maximum value of the C results in brackets.

[0092] The attractor should be located at the center of a type of sample points as much as possible, and the difference between the reciprocal point and the sample points of this type should be as large as possible. Therefore, the position and angle direction of the reciprocal point and the attractor in the feature space should be relative, such as Figure 6 As shown. Calculated according to the following formula :

[0093] (12)

[0094] In the formula, is the number of batches currently being trained; is the second label encoding; yes belong Probability of class; is a hyperparameter used to control the distance-to-probability conversion; is a hyperparameter used to control the denominator from being zero.

[0095] In some embodiments of the present invention, it is calculated according to the following formula :

[0096] (13)

[0097] It is calculated according to the following formula 、 :

[0098] (14)

[0099] In the formula, is a hyperparameter used to control the proportion of the Euclidean distance and the dot product.

[0100] It is calculated according to the following formula :

[0101] (15)

[0102] In the formula, is the boundary value.

[0103] In order to keep the unknown samples in the low-magnitude region and push the known samples away, the second norm of the samples in is used for constraint to make it approach 0, and it is calculated according to the following formula :

[0104] (16)

[0105] In the formula, represents calculating the second norm.

[0106] After dimensionality reduction visualization, most of the open-set samples are distributed at the tail of the closed-set samples. Kernel density estimation is used to constrain the generated samples so that the forged samples are as close as possible to the tail of the closed-set samples, and these forged samples are used to approximate the open-set samples to complete data augmentation. In addition, the present invention proposes to use attractor-reciprocal point learning to maintain intra-class aggregation while considering the risk of the open space. UrbanSound8K, AudioEventDataset, and TUT AcousticScenes 2017 are selected as the closed-set datasets for the experiment, and the classes that do not overlap with these datasets are selected from ESC-50 as the open-set samples. Finally, the results are compared with the open-set recognition algorithms in other fields, showing that the method proposed in the present invention performs excellently in the open-set recognition task of environmental audio. Specifically, for the UrbanSound8K dataset, the AUROC and OSCR are 0.9251 and 0.8743 respectively; for the AudioEventDataset, the AUROC and OSCR are 0.7921 and 0.7135 respectively; and for the TUT Acoustic Scenes 2017 dataset, the AUROC and OSCR are 0.8209 and 0.6262 respectively. The openness of the above three datasets exceeds 90%, indicating that the proposed algorithm still shows strong performance in highly open scenarios.

[0107] The environmental audio open-set recognition model according to an embodiment of the present invention can recognize unknown classes and maintain good intra-class aggregation while considering the risk of the open space.

[0108] The environmental audio open-set recognition method according to an embodiment of the present invention can recognize unknown classes and maintain good intra-class aggregation while considering the risk of the open space.

[0109] The electronic device according to an embodiment of the present invention includes a processor and a memory. The processor is connected to the memory, and the memory is used to store a computer program. When the computer program is executed by the processor, the environmental audio open-set recognition method according to the embodiment of the first aspect of the present invention is implemented.

[0110] The electronic device according to an embodiment of the present invention can recognize unknown classes and maintain good intra-class aggregation while considering the risk of the open space.

[0111] The other configurations and operations of the electronic device according to an embodiment of the present invention are known to those of ordinary skill in the art and will not be described in detail here.

[0112] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0113] In the description of the present invention, the "first feature" and "second feature" may include one or more of such features. The first feature being "above" or "below" the second feature may include the first and second features being in direct contact, or may include the first and second features not being in direct contact but in contact through additional features therebetween. The first feature being "above", "over" and "on top of" the second feature includes the first feature being directly above and obliquely above the second feature, or merely indicating that the first feature has a higher level height than the second feature.

[0114] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "specific embodiments", "examples" or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0115] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for identifying an open set of ambient audio, characterized in that: include: Use the closed set classification algorithm to preliminarily classify the ambient audio and obtain the first output vector and the second output vector after dimensionality reduction ; The As a generated training set, and using the kernel density estimation to generate adversarial networks to generate the first edge samples , after dimensionality reduction, we get the second edge sample ; Using the loss function of the generative adversarial network , the density loss of the kernel density estimate and the With the The offset loss between To train: In the formula, , is a constant coefficient; Representing the training of the generative adversarial network; G , D Represent the generator and discriminator respectively; According to C classification labels, Divide into C distinct subsets , , for the Assign reciprocal points and maximize the With the The difference, for each of the Distribution Attractor To control the The sample distribution of The open set data is controlled in the low-level region, and the attractor-reciprocal point loss is calculated according to the following formula : In the formula, Expressing the Loss of restraint; Expressing the Loss of restraint; is the distance cross entropy loss; It is the loss of constraining the data distribution range; , It is a hyperparameter used to control the weight of the loss function.

2. The method for identifying an open set of ambient audio according to claim 1, characterized in that: Also includes: Get each of the The set of centers of ; Calculate the Middle Data points With center point The distance between , ; Using the The distribution function and kernel function fitting probability density function , calculated according to the following formula Belongs to the interval [ , ] probability: According to the following formula, the : In the formula, is the number of the second edge samples; Encode the first label; For the The data points in The distance between .

3. The method for identifying an open set of ambient audio according to claim 2, characterized in that: According to the following formula, the : In the formula, For the The sample mean of the class.

4. The method for identifying an open set of ambient audio according to claim 2, characterized in that: Also includes: Get the Each data point in The eigenvector of ; Calculate the With each pair of and stated The distance between them is calculated according to the following formula: Classified as kind: In the formula, Indicates the category corresponding to the maximum value of the C results in brackets.

5. The method for identifying an open set of ambient audio according to claim 4, characterized in that: According to the following formula, the : In the formula, is the number of batches currently being trained; is the second label encoding; is the belong Probability of class; is a hyperparameter used to control the distance-to-probability conversion; is a hyperparameter used to control the denominator from being zero.

6. The method for identifying an open set of ambient audio according to claim 5, characterized in that: According to the following formula, the , : In the formula, is a hyperparameter used to control the weight of Euclidean distance and dot product.

7. The method for identifying an open set of ambient audio according to claim 4, characterized in that: According to the following formula, the : In the formula, is the boundary value.

8. The method for identifying an open set of ambient audio according to claim 4, characterized in that: According to the following formula, the : In the formula, Represents the calculation of the bi-norm.

9. An open set recognition model for ambient audio, characterized in that: include: The first classifier includes a feature extraction layer, a convolution layer, a pre-logit layer, a first logit layer and a first kernel density estimation layer. The first classifier uses a closed set classification algorithm to preliminarily classify the ambient audio. The pre-logit layer is used to output a first output vector The first logit layer is used to Reduce dimension and output the second output vector , the first kernel density estimation layer is used to generate a probability density curve for each category; Generate an adversarial device, the generative adversarial device includes a generator, a discriminator, a second logit layer and a second kernel density estimation layer, the generator uses the kernel density estimation effect of the second kernel density estimation layer to generate a first edge sample The second logit layer is used to Reduce dimension and output the second edge sample To the second kernel density estimation layer, the discriminator is used to determine the Does it belong to the , the generative adversarial machine uses the loss function of the generative adversarial network , the density loss of the kernel density estimate and the With the The offset loss between Conduct training; The second classifier includes a plurality of convolutional layers, and the second classifier is used to classify the , Perform attractor-reciprocal point training, according to The classification labels will be described Divided into distinct subsets , , for the Assign reciprocal points and maximize the With the The difference, for each of the Distribution Attractor To control the The sample distribution of The open set data is controlled in the low-level area, according to the Constraint loss , Constraint loss , distance cross entropy loss And the loss of constraining the data distribution range Calculate attractor-reciprocal point loss .

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the processor is connected to the memory, and the memory is used to store a computer program. When the computer program is executed by the processor, the method for identifying an open set of ambient audio according to any one of claims 1 to 8 is implemented.