Sound classification system

Through multi-layer neural network processing audio signals, combined with Mel frequency cepspectrum and filter groups, the problem of low classification efficiency of sound FX database in the prior art is solved, and automated efficient sound classification and tactile feedback is realized, suitable for movies and video games.

CN112912897BActive Publication Date: 2025-08-29SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201980061832.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-28
Filing Date
2019-09-23
Publication Date
2025-08-29
Estimated Expiration
2040-02-10

AI Technical Summary

Technical Problem

Existing FX database classification methods for movie and video game sounds rely on manual classification, making it difficult for content creators to find the sounds they need efficiently, and machine learning systems cannot classify them reliably based on sound characteristics.

Method used

Multi-layer neural network is used to classify sound, combine Mel frequency cepspectrum and filter groups to process audio signals, use convolutional neural network and recurrent neural network for feature extraction and training, and combine cross entropy loss and metric learning loss function optimization model to achieve hierarchical classification of sound.

Benefits of technology

It improves the search efficiency and accuracy of the sound database, can automatically classify and classify sounds, supports tactile feedback and assists users with disabilities, and is suitable for large sound FX databases and video game location determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112912897B_ABST
    Figure CN112912897B_ABST
Patent Text Reader

Abstract

A system method and computer program product for hierarchical classification of sounds includes one or more neural networks implemented on one or more processors. The one or more neural networks are configured to classify sounds into two or more hierarchical procedural and fine-grained levels of classification in a hierarchy. The classified sounds can be used to search a database for similar or contextually related sounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the classification of sounds. Specifically, the present disclosure relates to multi-layer classification of sounds using neural networks. Background Art

[0002] The growing popularity of films and video games, which rely heavily on computer-generated special effects (FX), has led to the creation of vast databases containing sound files. These databases are categorized to provide film and video game creators with better access to the sound files. While categorization aids accessibility, using the databases still requires familiarity with the database's contents and the categorization scheme. Content creators without knowledge of available sounds will struggle to find the sounds they desire. Consequently, new content creators may waste time and resources creating sounds that already exist.

[0003] Despite these large accessible sound databases, movies and video games often create new custom sounds. This requires a lot of time and people familiar with the classification schemes to add new sounds to these databases.

[0004] It is against this background that embodiments of the present disclosure are presented. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Aspects of the present disclosure may be readily understood by considering the following detailed description taken in conjunction with the accompanying drawings, in which:

[0007] Figure 1 is a block diagram illustrating a method of sound classification using a trained sound classification and classification neural network according to aspects of the present disclosure.

[0008] Figure 2A is a simplified node graph of a recurrent neural network for use in a sound classification system according to aspects of the present disclosure.

[0009] Figure 2B is a simplified node graph of an unfolded recurrent neural network used in a sound classification system according to aspects of the present disclosure.

[0010] Figure 2C is a simplified node diagram of a convolutional neural network for use in a sound classification system according to aspects of the present disclosure.

[0011] Figure 2D is a block diagram of a method for training a neural network in a sound classification system according to aspects of the present disclosure.

[0012] Figure 3 is a block diagram illustrating a combined metric learning and cross-entropy loss function learning method for training a neural network in a sound classification system according to aspects of the present disclosure.

[0013] Figure 4 A block diagram of a system implementing a sound classification method using a trained sound classification and categorization neural network according to aspects of the present disclosure is shown.

[0014] Detailed description of specific implementation plan

[0015] Although the following detailed description contains many specific details for the purpose of illustration, anyone skilled in the art will appreciate that many variations and modifications to the following details are within the scope of the present disclosure. Therefore, the examples of embodiments of the present disclosure described below are set forth without losing the generality of the claimed disclosure and without implying limitations on the claimed disclosure.

[0016] Although numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure, those skilled in the art will appreciate that other embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail in order to avoid obscuring the present disclosure. Some portions of the description herein are presented in terms of algorithms and symbolic representations of operations on data bits or binary digital signals within a computer memory. These algorithmic descriptions and representations can be a technique used by those skilled in the art of data processing to convey the substance of their work to others skilled in the art.

[0017] As used herein, an algorithm is a self-consistent sequence of acts or operations leading to a desired result. These involve physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0018] Unless explicitly stated or otherwise apparent from the following discussion, it should be understood that throughout the description, discussions utilizing terms such as "process," "compute," "convert," "coordinate," "determine," or "identify" refer to actions and processes of a computer platform, which is an electronic computing device that includes a processor that manipulates and converts data represented as physical (e.g., electronic) quantities within the processor's registers and accessible platform memory to other data similarly represented as physical quantities within the computer platform memory, processor registers, or display screens.

[0019] A computer program may be stored in a computer-readable storage medium, such as, but not limited to: any type of magnetic disk, including floppy disks, optical disks (e.g., compact disk read-only memory (CD-ROM), digital video disks (DVD), Blu-Ray Discs™, etc.), and magneto-optical disks; read-only memory (ROM); random access memory (RAM); EPROM; EEPROM; magnetic or optical cards; flash memory; or any other type of non-transitory medium suitable for storing electronic instructions.

[0020] The terms "coupled" and "connected," and their derivatives, may be used herein to describe structural relationships between components of devices for performing the operations herein. It should be understood that these terms are not intended to be synonymous with each other. Rather, in particular embodiments, "connected" may be used to indicate that two or more elements are in direct physical or electrical contact with each other. In some cases, "connected," "connected," and their derivatives are used to indicate a logical relationship, such as between layers of nodes in a neural network (NN). "Coupled" may be used to indicate that two or more elements are in physical or electrical contact with each other, directly or indirectly (with other intervening elements between them), and / or that two or more elements cooperate or communicate with each other (e.g., as if in a causal relationship).

[0021] Sound classification system

[0022] Currently, there are large databases of sound effects for movies and video games. These large databases are manually classified using a non-uniform multi-layer classification scheme. In one example scheme, the database has many categories, and each category has one or more sub-subcategories, with actual sounds listed under each subcategory. Machine learning has been used to train neural networks to cluster and classify datasets. Previous datasets often consist of objects that already have inherent classifications based on their design. For example, a previous clustering problem involves determining whether a car is a sedan or a coupe. The automotive industry explicitly manufactures cars as either sedans or coupe designs, and therefore, the differences between these two types of vehicles are inherent.

[0023] When a human classifies sounds in the Sound FX database, factors other than the acoustic properties of the sound are used to determine which classification the sound requires. For example, a human classifier will know things about the origin of the sound, such as whether the sound was produced by an explosion or a gunshot. This information cannot be obtained from the sound data alone. The insight of the present disclosure is that sounds that have no inherent classification can have a classification that can be applied in a reliable manner based solely on the characteristics of the sound. In addition, by learning similarities between sound samples based on their acoustic similarity, a machine learning system can learn new sound classes. Therefore, searching and classifying these large databases can be facilitated by using machine learning. It should be understood that although the present disclosure discusses two-tier classification, it is not limited to this structure and the teachings can be applied to databases with any number of tiers.

[0024] A sound classification system can facilitate searching for specific sounds or types of sounds within a large database. Classification refers to the general concept of hierarchical classification, i.e., from coarse to fine categories (e.g., nodes along a sound hierarchy tree). By way of example and not limitation, there are coarse and fine-grained classifications, with multiple coarse-grained categories or classes preceding the finest category or class. In some implementations, the system can classify and categorize onomatopoeic sounds uttered by a user. As used herein, the term onomatopoeic sound refers to a word or vocalization that describes or suggests a specific sound. The system can use the classification of onomatopoeic sounds to search the SoundFX database for actual recorded sounds that match the classification of onomatopoeic sounds. One advantage of a hierarchical database is that the system can present multiple similar sounds with varying degrees of dissimilarity based on both subcategories and categories. In some embodiments, the search may, for example, provide the top three most similar sounds. In other embodiments, the search may provide a selection of similar sounds from the same subcategory and class. The selection of similar sounds may be sounds that are contextually relevant to the input sound. In other implementations, the system may simply classify the actual recorded sounds and store them in the SoundFX database according to the classification. In some alternative embodiments, the actual recorded sounds may be onomatopeic sounds uttered by the user, as in some cases the user may wish to use spoken sounds as actual sound FX, such as in comic book media where onomatopeic sounds may be used to emphasize actions.

[0025] While a large sound FX database is one application of aspects of this disclosure, other applications are also contemplated. In an alternative implementation, the large sound FX database can be replaced with a time-stamped or location-organized database of sounds for a specific video game. A trained classification neural network determines the player's location within the game based on the classified sounds played to the player. This particular implementation is beneficial for non-localized simulators because no information from the simulation server is required beyond the sounds already sent to the user.

[0026] In another implementation, a hierarchical database of sounds can be associated with haptic events, such as vibrations from a controller or pressure on a joystick. A sound classification NN can be used to add self-organized haptic feedback to games that lack actual haptic information. The classification NN can simply receive game audio, classify the audio, and use the classification to determine the haptic events to provide to the user. Furthermore, the classification system can provide additional information to users with disabilities. For example, but not limited to, an application can provide a higher-level category sound to control a left-hand vibration pattern, and use a lower-level subcategory sound to control a right-hand high-frequency vibration.

[0027] By way of example and not limitation, some output modalities can assist hearing-impaired individuals, for example, using a mobile phone or glasses with a built-in microphone, a processor, and a visual display. A neural network running on the processor can classify the sounds picked up by the microphone. The display can present visual feedback of the identified sound category, for example in a text format such as closed captioning.

[0028] Current sound FX databases rely on a two-tier approach. In a two-tier database, the goal is to build a classifier that can correctly classify both categories and subcategories. A two-tier database according to aspects of the present disclosure has labels that are divided into two categories: "coarse" and "fine." The coarse class label is the category label. The fine class label is the category + subcategory label. Although the methods and systems described herein discuss a two-tier database, it should be understood that the teachings can be applied to databases with any number of category layers. According to aspects of the present disclosure and without limitation, the database can be a two-tier, unevenly hierarchical database with 183 categories and 4721 subcategories. The database can be searched by known methods such as index lookups. Examples of suitable search programs include commercially available digital audio workstation (DAW) software and other software, such as Soundminer software from Soundminer Ltd. of Toronto, Ontario, Canada.

[0029] The operating scheme of the sound classification system 100 begins with a sound clip 101. A plurality of filters are applied 102 to the sound clip 101 to produce a windowed sound and generate a representation of the sound in a mel-frequency cepstrum 103. The mel-frequency cepstrum representation is provided to a trained sound classification neural network 104. The trained sound classification NN outputs a vector 105 representing the class and subclass of the sound and a vector representing the finest level class of the sound, i.e., the finest level classification 106. This classification can then be used to search a database 110 for similar sounds as discussed above.

[0030] Filter the samples

[0031] Before classification and clustering, the sound FX can be processed to aid in classification. In some implementations, Mel-cepstral spectrogram features are extracted from the audio file. To extract the Mel-cepstral spectrogram features, the audio signal is divided into multiple time windows, and each window is converted into a frequency domain signal, for example, using a fast Fourier transform (FFT). The frequency or spectral domain signal is then compressed by taking the logarithm of the spectral domain signal and then performing another FFT. The cepstrum of the time domain signal S(t) can be mathematically represented as FT(log(FT(S(t)))+j2πq), where q is the integer required to correctly unwrap the angle or imaginary part of the complex logarithm function. Algorithmically, the cepstrum can be generated via the following sequence of operations: signal → FT → logarithm → phase unwrapping → FT → cepstrum. The cepstrum can be viewed as information about the rate of change of different spectral bands within the sound window. The spectrum is first transformed using a Mel filter bank (MFB), which differs from Mel-frequency cepstral coefficients (MFCCs) by having a fewer final processing step, a discrete cosine transform (DCT). The frequency f in Hertz (cycles per second) can be converted to the Mel frequency m according to the following formula: m = (1127.01048 Hz) log e (1+f / 700). Similarly, the Mel frequency m can be converted to a frequency f in Hertz using the following formula: f = (700 Hz) (e m / 1127.01048 -1). For example, but not limitation, the sound FX may be converted into a 64-dimensional Mel-Cepstral spectrogram with a moving window 42.67ms length and 10.67ms shift.

[0032] Batch training is used for NN training. Random windows of mel-frequency-converted sound, i.e., feature frames, are generated and the dimensional samples are fed into the model for training. By way of example and not limitation, 100 random feature frames can be selected, where each feature frame is a 64-dimensional sample.

[0033] Neural network training

[0034] The neural network 104 that implements the classification of the sound FX may include one or more of several different types of neural networks and may have many different layers. By way of example and not limitation, the classification neural network may be composed of one or more convolutional neural networks (CNNs), recurrent neural networks (RNNs), and / or dynamic neural networks (DNNs).

[0035] Figure 2AThe basic form of an RNN with a layer of nodes 220 is shown, each node being characterized by an activation function S, an input weight U, a recurrent hidden node transition weight W, and an output transition weight V. It should be noted that the activation function S can be any nonlinear function known in the art and is not limited to (hyperbolic tangent (tanh) function). For example, the activation function S can be a Sigmoid or ReLu function. Unlike other types of neural networks, RNNs have a set of activation functions and weights throughout the layer. As shown in FIG. Figure 2B As shown, the RNN can be thought of as a series of nodes 220 with the same activation function moving through time T and T + 1. Thus, the RNN maintains historical information by feeding the results from the previous time T to the current time T + 1.

[0036] There are many ways to configure the weights U, W, and V. Input weights U can be applied based on the mel-spectrogram. Weights for these different inputs can be stored in a lookup table and applied as needed. There may be default values ​​that the system initially applies. These values ​​can then be modified manually by the user or automatically through machine learning.

[0037] In some embodiments, a convolutional RNN can be used. Another type of RNN that can be used is a long short-term memory (LSTM) neural network, which adds an input gate activation function, an output gate activation function, and a forget gate activation function to the memory block in the RNN node, thereby generating a gated memory that allows the network to retain some information for a longer period of time, as described in Hochreiter & Schmidhuber "Long Short-term memory" Neural Computation 9(8):1735-1780 (1997), which is incorporated herein by reference.

[0038] Figure 2C An example layout of a convolutional neural network (such as a CRNN) according to aspects of the present disclosure is shown. In this figure, a convolutional neural network is generated for an image 232, the size of which is 4 units in height and 4 units in width, resulting in a total area of ​​16 units. The convolutional neural network shown has a filter 233 size of 2 units in height and 2 units in width with a skip value of 1 and a channel 236 size of 9. In order to Figure 2C For clarity in FIG, only the connections 234 between the first column of channels and their filter windows are shown. However, aspects of the present disclosure are not limited to such implementations. According to aspects of the present disclosure, the convolutional neural network implementing classification 229 can have any number of additional neural network node layers 231 and can include any size of such layer types as additional convolutional layers, fully connected layers, pooling layers, max pooling layers, local contrast normalization layers, etc.

[0039] As in Figure 2D As seen in , training a neural network (NN) begins with initialization 241 of the NN's weights. Typically, the initial weights should be randomly distributed. For example, a NN with a tanh activation function should have weights distributed over and A random value between , where n is the number of inputs to the node.

[0040] After initialization, the activation function and optimizer are defined. The NN is then provided with a feature or input dataset 242. Each of the different feature vectors can be provided with an input having a known label. Similarly, the classification NN can be provided with a feature vector corresponding to an input having a known label or classification. The NN then predicts the label or classification of the feature or input 243. The predicted label or class is compared to the known label or class (also known as the ground truth), and a loss function measures the total error between the prediction and the ground truth for all training samples 244. By way of example and not limitation, the loss function can be a cross-entropy loss function, a quadratic cost, a triple contrast function, an exponential cost, etc. A variety of different loss functions can be used depending on the purpose. The result of the loss function is then used to optimize and train the NN 245 using known training methods for neural networks (such as backpropagation using stochastic gradient descent, etc.). At each training epoch, the optimizer attempts to select model parameters (i.e., weights) that minimize the training loss function (i.e., the total error). The data is divided into training, validation, and test samples.

[0041] During training, the optimizer minimizes the loss function on the training examples. After each training epoch, the model is evaluated on the validation examples by calculating the validation loss and accuracy. If there is no significant change, training can be stopped. The trained model can then be used to predict labels for test data.

[0042] Thus, a classification neural network can be trained on audio inputs with known labels or classifications to identify and classify those audio inputs by minimizing the cross-entropy loss given the known target labels.

[0043] Layer-preserving representation learning

[0044] In addition to simple cross entropy loss, training NNs according to aspects of the present disclosure can employ metric learning. Metric learning via Siamese or triplet loss has an inherent ability to learn complex manifolds or representations. For SFX classification, metric learning improves clustering in the embedding space compared to using only cross entropy. The overall joint loss function according to aspects of the present disclosure is given by:

[0045] L 总 =L CE +L 度量(1)

[0046] Among them L CE is the cross entropy loss for classification, and L 度量 is the metric learning loss, which is L 三元组 or L 四元组 , as described below.

[0047] Uniform triplet training does not consider the hierarchical label structure, and it is possible that most of the negative samples come from categories different from the anchor category. To cope with this situation, triplet negative mining can be performed probabilistically so that the model encounters some triplets when the negative samples come from the same category but different subcategories. Given M samples, the triplets The anchor samples, positive samples and negative samples are selected as follows:

[0048] 1. Select anchors from category C and subcategory S

[0049] 2. Choose positive Make

[0050] 3. Select negative

[0051] and

[0052] Here, r is a Bernoulli random variable r~Ber(0.5). I(.) is the indicator function. Then, we obtain the triplet loss on a batch of N examples.

[0053]

[0054] Here, m is a non-negative margin parameter.

[0055] Another improvement in metric learning networks can be achieved by using a quadruple loss instead of a triplet. The quadruple loss attempts to preserve the embedding structure and provides more information about the hierarchy used to classify the input. The quadruple loss is given by:

[0056]

[0057] Select the four-tuple tuple as follows:

[0058] 1. Select anchors from category C and subcategory S

[0059] 2. Choose strong positive Make

[0060] 3. Choose weak positive Make and

[0061] 4. Select negative Make

[0062] like Figure 2D As shown, during training 245, a quadruple or triplet loss function and a cross entropy loss function are used in the optimization process, as discussed above. In addition, the combination of the cross entropy loss and the quadruple / triplet loss function used in the optimization process can be improved by adding a weight factor λ to compensate for the trade-off between the two types of losses. Therefore, the new combined loss equation is expressed as:

[0063] L 总 =(λ)L ce +(1-λ)L 度量

[0064] To extend this system to hierarchies with an arbitrary number of layers, the metric learning loss function is simply modified by recursively applying the method to all nodes of the hierarchy.

[0065] This combined training allows for improved classification of the input.

[0066] Combined training

[0067] Figure 3A schematic diagram illustrating training a sound FX classification system 300 using a combined cross-entropy loss function and a metric learning loss function 309 is shown. During training, samples representing anchors 301, strong positives 302, weak positives 303, and negatives 304 are provided to a neural network 305. Note that in implementations using triplet learning, only anchor, positive, and negative samples are provided. Neural network 305 may include one or more neural networks with any number of layers. By way of example and not limitation, in a two-layer network, there are four networks that share parameters during training. These networks represent (f(anchor), f(strong+), f(weak+), f(-)). An L2 normalization layer 306 is used on the output layer of neural network 305 to generate embedding distances 307. The output from L2 normalization 306 is a "normalized" vector called "embedding," which is converted 308 to a class vector but 307 is used to calculate distances between three pairs of embeddings corresponding to inputs 301 to 304. These distances are then used in the metric learning loss function. The labels of the anchors 311 are passed along for use in the loss function. The metric learning loss function is then applied to the embedding distance 307. In addition, the result of f(anchor) is also used to provide the finest level classification, which can be in the form of a vector representing the finest level sub-category 308. As discussed above, during training, the losses of the metric learning function and the cross entropy function are calculated and added together 309. The combined metric learning loss and cross entropy loss are then used for optimization in a mini-batch backpropagation using a stochastic gradient descent algorithm 310.

[0068] Implementation

[0069] Figure 4 A sound classification system according to aspects of the present disclosure is shown. The system may include a computing device 400 coupled to a user input device 402. The user input device 402 may be a controller, a touch screen, a microphone, a keyboard, a mouse, a joystick, or other device that allows a user to input information, including sound data, into the system. The user input device may be coupled to a tactile feedback device 421. The tactile feedback device 421 may be, for example, a vibration motor, a force feedback system, an ultrasonic feedback system, or a pneumatic feedback system.

[0070] The computing device 400 may include one or more processor units 403, which may be configured according to well-known architectures, such as, for example, single-core, dual-core, quad-core, multi-core, processor-coprocessor, cell processor, etc. The computing device may also include one or more memory units 404 (e.g., random access memory (RAM), dynamic random access memory (DRAM), read-only memory (ROM), etc.).

[0071] The processor unit 403 may execute one or more programs, portions of which may be stored in the memory 404, and the processor 403 may be operatively coupled to the memory, for example, by accessing the memory via the data bus 405. The program may be configured to implement the sound filter 408 to convert the sound into a mel-frequency cepstrum. In addition, the memory 404 may contain programs that implement the training of the sound classification NN 421. The memory 404 may also contain software modules such as the sound filter 408, the multi-layer sound database 422, and the sound classification NN module 421. The sound database 422 may be stored as data 418 in a mass storage device 418 or at a server coupled to the network 420 accessed via the network interface 414.

[0072] The overall structure and probabilities of the NN may also be stored as data 418 in the mass storage device 415. The processor unit 403 is further configured to execute one or more programs 617 stored in the mass storage device 415 or memory 404, which cause the processor to perform the method 300 of training a sound classification NN 421 from a sound database 422. As part of the NN training process, the system may generate neural networks. These neural networks may be stored in the memory 404 in the sound classification NN module 421. The complete NN may be stored in the memory 404 or in the mass storage device 415 as data 418. The program 417 (or portions thereof) may also be configured, for example through appropriate programming, to apply appropriate filters 408 to the sounds input by the user, classify the filtered sounds using the sound classification NN 421, and search the sound category database 422 for similar or identical sounds. In addition, the program 417 may be configured to utilize the results of the sound classification to create haptic feedback events using the haptic feedback device 421.

[0073] The computing device 400 may also include well-known support circuits such as input / output (I / O) 407, circuitry, power supply (P / S) 411, clock (CLK) 412, and cache 413, which may communicate with other components of the system, for example, via bus 405. The computing device may include a network interface 414. The processor unit 403 and the network interface 414 may be configured to implement a local area network (LAN) or a personal area network (PAN) via a suitable network protocol suitable for the PAN (e.g., Bluetooth). The computing device may optionally include a mass storage device 415, such as a disk drive, CD-ROM drive, tape drive, flash memory, etc., and the mass storage device may store programs and / or data. The computing device may also include a user interface 616 for facilitating interaction between the system and a user. The user interface may include a monitor, a television screen, speakers, headphones, or other device that conveys information to the user.

[0074] Computing device 400 may include a network interface 414 to facilitate communication via an electronic communications network 420. Network interface 414 may be configured to enable wired or wireless communication over a local area network and a wide area network such as the Internet. System 400 may send and receive data and / or file requests via one or more message packets over network 420. Message packets sent over network 420 may be temporarily stored in buffer 409 in memory 404. The classified sound database may be obtained via network 420 and partially stored in memory 404 for use.

[0075] Although the above is a complete description of the preferred embodiments of the present disclosure, it is possible to use various alternatives, modifications and equivalents. It should be understood that the above description is intended to be illustrative and not restrictive. For example, although the flowcharts in the accompanying drawings show a particular order of operations performed by certain embodiments of the present disclosure, it should be understood that such order is not required (for example, alternative embodiments may perform operations in a different order, combine certain operations, overlap certain operations, etc.). In addition, after reading and understanding the above description, many other embodiments will be apparent to those skilled in the art. Although the present disclosure has been described with reference to specific exemplary embodiments, it will be recognized that the present disclosure is not limited to the embodiments described, but may be practiced with modification and variation within the spirit and scope of the appended claims. Therefore, the scope of the present disclosure should be determined by reference to the appended claims and the full scope of equivalents to which such claims are entitled. Any feature described herein (whether preferred or not) may be combined with any other feature described herein (whether preferred or not). In the appended claims, unless otherwise expressly stated, The indefinite article "a" or "a kind" refers to the quantity of one or more of the item following the article. The following claims are not to be understood as including means-plus-function limitations unless such a limitation is explicitly recited in a given claim using the phrase "means for..."

Claims

1. A system for hierarchical classification of sounds, the system comprising: one or more processors; one or more neural networks implemented on the one or more processors, the one or more neural networks configured to classify sounds into two or more hierarchical levels of coarse classification and finest level classification in a hierarchy, wherein the one or more neural networks are trained using a combination of a metric learning function and a cross-entropy loss function, wherein a loss from the metric learning function and a loss from the cross-entropy loss function are added together; and One or more haptic feedback devices, wherein a first haptic event of the one or more haptic feedback devices is determined by a coarse classification in the hierarchical classification of the sound, and wherein a second haptic event of the one or more haptic feedback devices is determined by a finer classification in the hierarchical classification of the sound.

2. The system of claim 1, wherein the metric learning function is a triplet loss function.

3. The system of claim 1, wherein the metric learning function is a quadruple loss function.

4. The system of claim 1 , wherein the one or more neural networks are further configured to classify and categorize onomatopoeic sounds uttered by the user.

5. The system of claim 1 , wherein the one or more neural networks are executable instructions stored in a non-transitory computer-readable medium that, when executed, cause the processor to perform neural network calculations.

6. The system of claim 5, further comprising a database, and wherein the executable instructions further comprise searching the database for results of the classification from the neural network.

7. The system of claim 5, wherein the executable instructions further comprise storing sound data in a hierarchical database based on the classification performed by the one or more neural networks.

8. The system of claim 5, wherein the executable instructions further comprise determining hierarchical synchronization of audio events within a video game based on the classification performed by the one or more neural networks.

9. The system of claim 5, wherein the instructions further comprise executable instructions for using results from the classification of the one or more neural networks to find context-relevant sounds in a database.

10. The system of claim 1, wherein the sound is digitized and converted to a Mel-frequency cepstrum before classification.

11. Computer-executable instructions embodied in a non-transitory computer-readable medium, the computer-executable instructions, when executed, implementing one or more neural networks configured to classify sounds into a two or more level hierarchy, the sounds being classified according to a finest level in the hierarchy. wherein the one or more neural networks are trained using a combination of a metric learning function and a cross entropy loss function, wherein the loss of the metric learning function and the loss of the cross entropy loss function are added together; and A coarse classification in the hierarchical classification of the sounds determines a first haptic event of one or more haptic feedback devices, and a fine classification in the hierarchical classification of the sounds determines a second haptic event of one or more haptic feedback devices.

12. The computer-executable instructions of claim 11, wherein the one or more neural networks are further configured to classify and categorize onomatopoeic sounds uttered by a user.

13. The computer-executable instructions of claim 11, wherein the instructions further comprise searching a database for results of the classification from the neural network.

14. A method for hierarchical classification of sounds, the method comprising: Use neural networks to classify sounds into two or more layers of hierarchy, coarse classification and finest classification, wherein the neural network is trained using a combination of a metric learning function and a cross entropy loss function, wherein the loss of the metric learning function and the loss of the cross entropy loss function are added together; and A coarse classification in the hierarchical classification of the sounds determines a first haptic event of one or more haptic feedback devices, and a fine classification in the hierarchical classification of the sounds determines a second haptic event of one or more haptic feedback devices. The method of claim 14 , further comprising classifying and categorizing the onomatopoeic sounds emitted by the user.

16. The method of claim 14, further comprising searching a database for results of the classification from the neural network.

Citation Information

Patent Citations

  • Terminal device and controlling method thereof

    CN106375547A

  • Method and device for training acoustic characteristic extraction model, equipment and computer storage medium

    CN107221320A

  • Binaural recording for processing audio signals to enable alerts

    US20160192073A1

  • Identifying And Extracting Video Game Highlights Based On Audio Analysis

    US20170065889A1