Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Sound classification" patented technology

Classification of Sounds. THE ALPHABET. ORTHOGRAPHY. 2. The simple Vowels are a, e, i, o, u, y. The Diphthongs are ae, au, ei, eu, oe, ui, and, in early Latin, ai, oi, ou. In the diphthongs both vowel sounds are heard, one following the other in the same syllable.

Regional bird chirp classification method for small sample registration and related equipment

PendingCN121331146ASpeech analysisCosine similaritySound classification
The invention discloses a small sample registered regional bird chirp classification method and related equipment, and the method comprises the steps: inputting a first to-be-classified Fbank feature into a bird voiceprint feature extraction model, generating a to-be-classified bird voiceprint template feature, and carrying out the updating of a small sample registration template library according to the to-be-classified bird voiceprint template feature; inputting the second to-be-classified Fbank feature into a bird voice voiceprint feature extraction model, and extracting a to-be-classified voiceprint embedding feature; and performing cosine similarity calculation on the to-be-classified voiceprint embedded features and the voiceprint features in the updated small sample registration template library to obtain a to-be-classified feature matching score set, and selecting a bird corresponding to the highest matching score as a regional bird buzzing classification result. The method can improve the feature extraction robustness in a field complex noise environment, improves the accuracy of regional bird recognition, does not need to train a model again, greatly reduces the extension cost, and can be widely applied to the technical field of sound signal recognition.
Owner:GUANGZHOU UNIVERSITY +1

Double-path CNN heart sound classification method based on time-frequency and double-spectrum fusion features

PendingCN121054045AStethoscopeSpeech analysisBispectral analysisNerve network
The invention relates to the technical field of audio signal processing and biomedical signal analysis, and still has a further optimized space for the recognition of anti-noise requirements, signal individual differences and complex pathological modes in a noise environment. The invention provides a double-path CNN heart sound classification method based on time-frequency and double-spectrum fusion features, and the method comprises the steps: carrying out the preprocessing of an original heart sound signal of a data set which is classified into a normal heart sound and an abnormal heart sound, and obtaining a to-be-recognized heart sound signal; based on dynamic continuous wavelet transform, adaptively selecting parameters to extract time-frequency characteristics, introducing bispectrum analysis, capturing nonlinear characteristics, generating a dual-channel characteristic pattern, and efficiently storing the dual-channel characteristic pattern in an HDF5 format; and based on a designed double-path convolutional neural network structure, respectively processing the extracted time-frequency and double-spectrum features, performing classification after fusion, and training a model in combination with category weighted loss and an optimization strategy to obtain a heart sound classification result. The heart sound recognition accuracy can be improved.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Method for analyzing sound data for use in an anti-snoring system and apparatus

ActiveUS12502002B2SofasAuscultation instrumentsNerve networkSound classification
An anti-snoring system comprising an adjustable bed having a sleeping surface that may be mechanically raised or lowered, and a control module adapted to receive commands from a source external to the adjustable bed; a mobile device in direct or indirect communication with the control module in the adjustable bed, and wherein the mobile device has sound recording capabilities; a mobile application resident on the mobile device, wherein the mobile application includes a sound classification machine learning model that includes an artificial intelligence or neural network operative to determine whether or not a person on the sleep surface is snoring, and wherein upon a determination that the person is snoring and has been snoring for a predetermined period of time, the mobile application instructs the control module to raise or adjust the sleeping surface to a height or position that will discourage the person from snoring.
Owner:SKY BACON TECH HLDG LLC

Method for analyzing sound data for use in an Anti-snoring system and apparatus

PendingUS20260108074A1SofasAuscultation instrumentsNerve networkSound classification
An anti-snoring system comprising an adjustable bed having a sleeping surface that may be mechanically raised or lowered, and a control module adapted to receive commands from a source external to the adjustable bed; a mobile device in direct or indirect communication with the control module in the adjustable bed, and wherein the mobile device has sound recording capabilities; a mobile application resident on the mobile device, wherein the mobile application includes a sound classification machine learning model that includes an artificial intelligence or neural network operative to determine whether or not a person on the sleep surface is snoring, and wherein upon a determination that the person is snoring and has been snoring for a predetermined period of time, the mobile application instructs the control module to raise or adjust the sleeping surface to a height or position that will discourage the person from snoring.
Owner:SKY BACON TECH HLDG LLC

Micro-motor abnormal sound classification method and device based on multi-scale feature fusion and attention mechanism

ActiveCN120048285BSustainable transportationSpeech analysisFeature vectorSound classification
The application provides a micro-motor abnormal sound classification method and device based on multi-scale feature fusion and attention mechanism, which comprises the following steps: first, sound signal data acquisition; second, sound signal preprocessing; third, sound data feature extraction; fourth, multi-scale feature fusion: a convolutional neural network with three different scale convolution kernels is used to extract multi-level feature information of the comprehensive feature vector to obtain a feature map; a channel attention and a spatial attention mechanism are integrated in each convolutional neural network, the importance of each channel feature vector is weighted through the channel attention mechanism, and the attention area of the feature vector is dynamically adjusted in the spatial dimension through the spatial attention mechanism; fifth, feature classification and judgment of the running state or fault type of the current micro-motor. The method introduces the fusion of multi-scale convolution and attention mechanism, enhances the robustness of the model in a complex environment, and improves the accuracy of fault classification.
Owner:MINZHUO ELECTRIC CO LTD

Machine learning (ML) algorithm for sound classification and cancellation

ActiveUS12718788B2NoiseSound classification
This disclosure provides systems, methods, and devices for audio signal processing that support noise cancellation. In a first aspect, a method of signal processing includes determining a location of the apparatus; receiving an audio signal including sounds at the location of the apparatus; determining, based on a machine learning (ML) model, to reduce a presence of the one or more sounds in the audio signal based on the location; and determining an output audio signal by reducing the presence of the one or more sounds in the audio signal. Other aspects and features are also claimed and described.
Owner:QUALCOMM INC

A method, system, device, medium, and product for classifying animal sounds

The application relates to the technical field of machine learning, and provides an animal sound classification method, system, device, medium and product. The application trains a classification model through a staged strategy, including pre-training the classification model based on a general audio dataset, end-to-end training the classification model based on an animal sound dataset, and deploying the trained classification model to an intelligent device; the intelligent device collects original audio signals in real time and extracts log-mel spectrum features as feature inputs of the classification model, and the model finally outputs prediction probabilities for each target sound category, so as to determine an animal sound classification result. The staged training method significantly reduces the model training difficulty, the log-mel spectrum features are used as inputs to improve the robustness of the model to background noise, and the generalization ability and overall accuracy of the classification model are effectively improved.
Owner:VERISILICON MICROELECTRONICS (NANJING) CO LTD +2

Air conditioner production line abnormal sound detection method and system

The application provides an air conditioner production line abnormal sound detection method and system, comprising: acquiring sound signals in real time through a microphone spherical array; processing the sound signals based on an HOA-SHT domain analysis framework to generate a panorama sound image; acquiring an air conditioner image through a panorama camera; locking an air conditioner area according to the air conditioner image; performing multi-modal positioning on the air conditioner area through the panorama sound image to obtain a target area; extracting sound signals from the target area through a multi-scale mel spectrum, and inputting the extracted sound signals into a beamformer of a neural network to obtain noise signals; detecting and identifying the noise signals based on a convolutional neural network-based anomaly classifier to obtain air conditioner abnormal sound classification results; efficiently and accurately detecting air conditioner hanging machine abnormal sounds, effectively overcoming the limitations of traditional manual detection methods, significantly improving production efficiency, reducing the false detection rate and the missed detection rate, and providing strong quality guarantee for the production line.
Owner:BEIJING FRYHUIER TECHNOLOGY CO LTD

A bird chirping sound recognition method based on a combination of voiceprints and spatial distribution

PendingCN122392544AData setSound classification
The application discloses a bird chirp sound recognition method based on a combination of voiceprints and spatial distribution, and belongs to the technical field of intelligent sound classification and recognition. In view of the problems of ignoring geographical distribution prior knowledge and sample imbalance in the prior art, the application firstly constructs a voiceprint recognition model: a training data set is constructed by audio preprocessing, logarithmic mel spectrum and dynamic difference feature extraction, a model is trained based on DenseNet-121 by adopting a two-stage training strategy, and recognition confidence of each species is obtained; meanwhile, a spatial distribution model is constructed: based on public observation data, an average observer ability index is used to correct an original encounter rate, and spatial distribution probability of the species in a specific city is obtained; finally, a Sigmoid function is used to perform nonlinear fusion on the two, a joint recognition probability is calculated, and a classification result is output. The application introduces ecological spatial constraints into the recognition decision, effectively reduces false positive misjudgment, improves rare species monitoring capability, and makes the recognition result have ecological interpretability.
Owner:INST OF URBAN ENVIRONMENT CHINESE ACAD OF SCI

Self-learning musical instrument pitch correction device that eliminates environmental noise

PendingCN122637737AEnvironmental noiseNoise
The present application relates to the technical field of musical instrument tuning, and particularly relates to a self-learning musical instrument pitch correction device capable of eliminating environmental noise. The present application comprises an image acquisition module, a vibration feature acquisition module, a fundamental frequency acquisition module, a correction indication module and a self-learning recording module. The image acquisition module is used to collect string vibration images to physically isolate noise, the vibration feature acquisition module is used to track string displacement to generate vibration feature data, the fundamental frequency acquisition module is used to perform frequency domain transformation and interpolation on the data to obtain the actual vibration fundamental frequency, the correction indication module is used to compare with the standard pitch and output a graded color indication and a tuning direction, and the self-learning recording module is used to analyze previous deviation data to generate stability evaluation and maintenance suggestions. The present application realizes sound classification level pitch detection and correction without sound collection, and improves the tuning reliability in a noisy environment.
Owner:HUIZHOU UNIV +1

Brain-like low-power-consumption audio classification method and device suitable for edge device, equipment, medium and product

The invention relates to the field of computer technology and audio signal processing, and provides a brain-like low-power-consumption audio classification method and device suitable for edge equipment, equipment, a medium and a product. Decomposing the target audio by using an audio decomposition module of the audio classification model to obtain a first sub-band and a second sub-band; performing feature extraction on the first sub-band and the second sub-band by using a feature extraction module of the audio classification model to obtain a first feature corresponding to the first sub-band and a second feature corresponding to the second sub-band; and predicting a sound classification of the target audio based on the first feature and the second feature by using a classification module of the audio classification model. According to the method and the device, the problem of redundant calculation caused by homogenization processing of the audio can be solved, the problem that high-frequency information and low-frequency information cannot be distinguished is avoided, targeted feature extraction can be performed on different frequency components, so that redundant calculation can be avoided, and the calculation efficiency of audio classification is improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Model parameter determination method, vehicle abnormal sound classification method, device and equipment

This application discloses a method for determining model parameters, a method for classifying vehicle abnormal noises, an apparatus, and a device. The method includes: acquiring a training dataset of vehicle abnormal noise features; generating an initial population based on the training dataset; the population containing multiple individuals; constructing a kernel principal component analysis (KPC) model based on preset model parameters; the KPC model is used to extract target abnormal noise feature vectors from the vehicle's abnormal noise features for vehicle abnormal noise classification; determining an objective function based on the KPC model; iteratively updating the initial population based on the objective function to obtain an updated population; determining the optimal individual from the updated population; and determining the target model parameters of the KPC model based on the optimal individual and a preset parameter range. This method can determine the target model parameters of the KPC model most suitable for the current abnormal noise environment, thereby improving the classification accuracy of vehicle abnormal noise classification.
Owner:CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD

Sound production object multi-classification method and device based on multiple modes and computer equipment

PendingCN120612947ASpeech analysisNeural learning methodsNoiseSound classification
The invention discloses a sound production object multi-classification method and device based on multiple modes and computer equipment. The method comprises the following steps: acquiring audios and videos of sounding objects to be classified to obtain audio information and video information; inputting the audio information and the video information into a classification model for classification to obtain a classification result; obtaining the classification result; wherein the classification model is obtained by taking a plurality of pieces of audio and video information with category labels as a sample set to train a deep learning model. By implementing the method provided by the invention, the multi-class sound classification problem can be effectively solved, the classification accuracy can be enhanced in combination with multi-modal information, and the method has strong noise filtering capability to improve the recognition precision in a noisy environment; the problem that sound classification recognition accuracy is not high only from a pure voice mode due to strong noise interference is solved.
Owner:WUXI UNIV

Construction method of insect sound recognition algorithm model

PendingCN120656462ASpeech analysisAlgorithmWoodworm
The invention discloses a construction method of an insect sound recognition algorithm model, and belongs to the field of moth detection. Comprising the following steps: S1: data processing: carrying out background noise separation processing on collected original insect sound audio data to obtain audio data only containing insect sound segments, and then carrying out data enhancement and feature extraction; s2, model training: performing model training on the processed data; and S3, testing the model: constructing a test set, and evaluating the accuracy of insect sound classification. Compared with the prior art, the method is based on VAD and speaker recognition and other deep learning technologies, data enhancement is combined, an efficient insect sound recognition model is achieved, insect sound can be accurately recognized at the second-level precision, the model size is remarkably reduced while the recognition effect is kept through the model optimization and compression technology, and the recognition efficiency is improved. And efficient insect sound identification in a laboratory scene can be realized.
Owner:SHENZHEN CUSTOMS ANIMAL & PLANT INSPECTION & QUARANTINE TECH CENT

Pasture decision management system and method

The embodiment of the invention provides a pasture decision management system and method, and belongs to the technical field of pasture management. The system comprises an equipment end used for collecting multi-modal data and sending the multi-modal data to an edge end; the edge end is used for performing behavior recognition by adopting a behavior recognition model based on the dairy cow image data to obtain dairy cow behavior category data; based on the cow audio data, performing sound classification by adopting a sound classification model to obtain cow barking category data; performing abnormal value filtering and aggregation processing based on the sensor data to obtain aggregated data; the multi-modal data, the dairy cow behavior category data, the dairy cow sound category data and the aggregated data are sent to a cloud end; the cloud is used for storing the knowledge graph; updating the knowledge graph based on the data sent by the edge end; and generating a management decision based on the updated knowledge graph. The system is used for overcoming the defects in existing pasture intelligent decision making.
Owner:INNER MONGOLIA UNIV OF TECH

Mask auto-identification via breathing sound classification

PCT designated stageWO2025201927A1Mechanical/radiation/invasive therapiesRespiratory masksInhalationSound classification
A system and associated method for automatically identifying a mask (9) used in a pressure support system (2) for delivering a flow of breathing gas to the airway of a patient. The system includes a controller (20) implementing a trained machine learning model (22). The controller is structured and configured to receive a sound signal, the sound signal being indicative of breathing sounds (e.g., exhalation and / or inhalation sounds) captured from the patient during use of the mask in the pressure support system, generate acoustic spectrum data indicative of an acoustic spectrum of the exhalation sounds based on the sound signal, provide the acoustic spectrum data to the trained machine learning model (22), and determine a brand, type and / or size of the mask (9) in the trained machine learning model based on the provided acoustic spectrum data.
Owner:KONINKLIJKE PHILIPS NV

Snoring sound recognition intervention system and method

PendingCN122455017ASound classificationSound recognition
The application relates to a snoring sound identification intervention system and method, wherein the snoring sound identification intervention system comprises a data acquisition module, a snoring sound identification module and a snoring sound intervention module; the data acquisition module is used for acquiring heart impact snoring sound vibration data of a user; the snoring sound identification module is used for generating a corresponding snoring sound feature vector based on the heart impact snoring sound vibration data, inputting the snoring sound feature vector into a trained snoring sound classification model, and outputting a snoring sound judgment result and a snoring sound intensity level; the snoring sound intervention module is used for generating a corresponding intervention scheme based on the snoring sound intensity level when the snoring sound judgment result is that snoring sound is identified, and adjusting an intelligent pillow used by the user based on the intervention scheme. Through the application, the problem of high snoring sound misjudgment rate is solved.
Owner:HANGZHOU SHENGWEI INNOVATION TECHNOLOGY CO LTD

Smart classroom noise monitoring device

ActivePH22025051235U1MicrocontrollerSound detection
The present utility model relates to a classroom noise monitoring device that detects ambient sound levels and provides immediate visual feedback to regulate classroom behavior. The device comprises a sound detection module configured to capture noise, a microcontroller board programmed to classify the detected sound into predefined threshold ranges, and a visual output module that displays indicators corresponding to acceptable, moderately high, and excessive noise levels. In one embodiment, the visual output module employs colored light indicators, while in another embodiment it utilizes a graphic display presenting emoticon icons. An optional wireless communication module may transmit noise data to a remote server or mobile device for monitoring and record-keeping. The device is enclosed in a wall-mountable or desktop casing and powered by a standard low-voltage supply. By providing real-time, intuitive feedback, the utility model offers an affordable and effective tool for promoting discipline and self-awareness in educational settings.

An Automatic Classification Method for Indoor Ambient Sound Based on a Lightweight ECAPA-TDNN Neural Network

ActiveCN116013276BSpeech recognitionFeature extractionSound classification
This invention discloses an automatic classification method for indoor ambient sound based on a lightweight ECAPA-TDNN neural network, belonging to the field of ambient sound classification technology. The method includes the following steps: First, augmenting the initial ambient sound data through time masking, frequency masking, and audio data shifting. Second, extracting Mel spectrogram features from the indoor ambient sound data through pre-emphasis, short-time Fourier transform, and Mel filtering; dividing the obtained ambient sound Mel spectrogram features into training and testing sets. Third, constructing an ECAPA-TDNN network model, optimizing the neuron parameters of the ECAPA-TDNN network using the training set; and then using the trained neural network for classifying the ambient sound test set. Compared to ambient sound classification methods using traditional training classification frameworks, the method proposed in this invention has higher accuracy, wider applicability, and consumes fewer computational resources.
Owner:GUANGDONG UNIV OF TECH

Method, device and equipment for detecting water pipe leakage point based on vision and sound

ActiveCN116907742BAlgorithmAnomaly detection
The application relates to the technical field of artificial intelligence, and provides a method, device and equipment for jointly detecting a water pipe leakage point based on vision and sound, to solve the problem that there is no leakage anomaly detection method with high accuracy and good universality in related technologies. First, an image sample of a water pipe is taken as input of a vision network model to obtain a positioning result and a positioning confidence for representing a leakage point position in the image sample, and a vision classification result and a vision classification confidence for representing a leakage point category in the image sample; then, a sound sample of the water pipe is taken as input of a sound network model to obtain a sound classification result and a sound classification confidence for representing a sound sample category; the position of the leakage point is comprehensively obtained according to the positioning confidence and the sound classification confidence, and then the category information of the leakage point is comprehensively obtained according to the vision classification confidence and the sound classification confidence, so that the position and type of the leakage point are finally obtained.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Sound classification method based on time-frequency multi-aggregation and cross Gaussian attention

PendingCN120954447ASpeech analysisTime domainSound classification
The invention provides a sound classification method based on time-frequency multi-aggregation and cross Gaussian attention, which can realize stronger model characterization capability based on lower parameter quantity, realize long-range modeling, better integrate global context information and effectively improve the accuracy of sound classification and recognition. A time-frequency multi-aggregation network is designed, then a cross Gaussian attention mechanism is introduced, a time-frequency multi-aggregation and cross Gaussian attention network is constructed, and an acoustic model is obtained. In a time-frequency multi-aggregation convolution block in the acoustic model, targeted frequency domain feature and time domain feature extraction is realized; then, constructing a time-frequency double-channel attention mechanism, respectively strengthening effective time-frequency information weights, and aggregating effective information for multiple times; and a cross Gaussian attention mechanism is introduced between adjacent TFMA blocks for feature fusion, so that hierarchical fusion from shallow to deep is realized, stronger global association is established, and the classification accuracy is effectively improved.
Owner:JIANGNAN UNIV

Urban sound event marking and identifying method based on saliency judgment

The invention discloses a city sound event labeling and recognition method based on significance judgment, and belongs to the technical field of city sound environment monitoring and intelligent recognition. The method comprises the following steps: continuously collecting real environment audio; segmenting the audio data and extracting samples, and performing de-identification preprocessing to reduce privacy risks; screening reliable annotators through classification capability averaging and annotation consistency, and achieving a significance judgment consensus; performing significant sound event labeling based on a significance judgment consensus, and constructing a data set; the judgment rule of human on acoustic significance is analyzed and concluded through labeling consistency, and then the model is guided to learn significance standards with cognitive fitness; a deep learning model integrated with a channel attention mechanism is adopted to extract spectrum features, and intelligent judgment of a significant sound event is realized under supervised classification guidance of a significant sound event data set. According to the method, redundant data can be remarkably compressed while sound classification accuracy is kept, and the method is suitable for identification and analysis of multi-source and key sound events in a real urban environment.
Owner:DALIAN UNIV OF TECH

Method, system and production line for detecting abnormal sound of automobile seat driver

The invention discloses an abnormal sound detection method and system for an automobile seat driver and a production line. The method comprises the following steps: taking collected audio signal data as reference data; extracting multi-dimensional features of the reference data, and constructing an abnormal sound detection model based on an SVM algorithm; a multi-dimensional feature contribution vector is calculated by adopting an SHAP value analysis method, and feature contribution templates of different abnormal sound subdivision types in the reference data are obtained; calculating a multi-dimensional feature contribution vector of each to-be-detected abnormal sound sample through an SHAP value analysis method; and carrying out similarity comparison on the obtained multi-dimensional feature contribution vector and the feature contribution templates of different abnormal sound subdivision types, and carrying out abnormal sound subdivision classification on an abnormal sound sample to be detected. The system adopts the method, and the system is deployed on a production line. The method has the advantages that manual intervention is not needed from audio collection, feature extraction, model judgment and abnormal sound classification, the detection efficiency is improved, and the long-term cost of mute room facilities and professionals is reduced.
Owner:NINGBO SHUANGLIN AUTO PARTS CO LTD

Bird chirp sound classification and recognition method and device

ActiveCN115762533BSpeech analysisFrequency spectrumSound classification
The application discloses a bird chirp sound classification and recognition method and device, comprising the following steps: acquiring bird chirp sound audio data; pre-processing the bird chirp sound audio data to obtain pre-processed audio data; performing Fourier transform on the pre-processed audio data to obtain a spectrogram of the bird chirp sound; obtaining an MFCC hybrid feature vector of the pre-processed audio data based on a mel-frequency cepstrum coefficient and a difference operation; processing the spectrogram by using a CNN network to obtain local fine-grained spectral features after training; processing the MFCC hybrid feature vector by using a Transformer encoder network to obtain global sequence features considering context after training; and obtaining a recognition classification result of the bird chirp sound by using a Softmax classifier after splicing and fusing the local fine-grained spectral features and the global sequence features. The application can improve the bird sound classification and recognition accuracy.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Apparatus for classifying sounds based on neural code in spiking neural network and method thereof

A method of classifying sounds based on a neural code in a spiking neural network includes: receiving sounds to be classified and digitally converting the received sounds into sound data; preprocessing the sound data using a multiple neural code-based encoding method including rate code encoding and synchrony code encoding; inputting the preprocessed sound data to a biological spiking neural network to extract features; performing biological spike timing-dependent plasticity (STDP) rule-based learning using the extracted features; and performing classification of the sounds according to neural code propagation characteristics using a test dataset according to a result of the performing of the learning.
Owner:KOREA UNIV RES & BUSINESS FOUND

Machine learning (ML) algorithm for sound classification and cancellation

PCT designated stageWO2025184109A1Speech analysisSound producing devicesNoiseSound classification
This disclosure provides systems, methods, and devices for audio signal processing that support noise cancellation. In a first aspect, a method of signal processing includes determining a location of the apparatus; receiving an audio signal including sounds at the location of the apparatus; determining, based on a machine learning (ML) model, to reduce a presence of the one or more sounds in the audio signal based on the location; and determining an output audio signal by reducing the presence of the one or more sounds in the audio signal. Other aspects and features are also claimed and described.
Owner:QUALCOMM INC

Sound classification method and system based on channel attention and multi-scale mel spectrogram

ActiveCN117854546BFrequency spectrumSound classification
The application discloses a sound classification method and system based on channel attention and a multi-scale mel spectrum diagram, and the method comprises the following steps: collecting cough audio data, and performing noise reduction processing on the audio; performing cough event detection on long-time audio and removing the mute section, and segmenting out short-time audio signals containing cough events; performing adaptive scale audio feature extraction on the uniformly processed short-time audio signals, generating multi-channel mel spectrum data of the audio, obtaining a mel spectrum feature matrix set K of the audio; building a convolutional neural network model based on channel attention, extracting features of a three-channel mel spectrum diagram; taking the mel spectrum feature matrix set K of the audio as the input of the feature model M of the three-channel mel spectrum diagram, and generating a sound classification result. weight The application has the characteristics of low cost, high precision and fast identification of cough sound.
Owner:GUIZHOU UNIV

Elevator fault sound classification method and device based on self-knowledge distillation

PendingCN121483304ASpeech analysisBiological modelsAlgorithmSound classification
The invention discloses an elevator fault sound classification method and device based on self-knowledge distillation. The method comprises the following steps: S1, acquiring an audio signal, and generating a time-frequency diagram according to the audio signal; s2, inputting the time-frequency diagram into an acoustic classification model to obtain elevator fault sound classification; the acoustic classification model comprises a channel expansion module, a convolutional layer and a teacher network which are sequentially connected in series; the channel extension module adds channel dimensions to an input time-frequency graph and converts the time-frequency graph into a three-dimensional feature tensor; the convolution layer carries out convolution on the three-dimensional feature tensor to obtain bottom layer acoustic features; and the teacher network gradually extracts semantic features from the underlying acoustic features through n extraction modules connected in series, and sequentially processes the semantic features output by the nth extraction module through an attention-based time sequence pooling layer SAP, feature projection and a full-connection classification layer to obtain a prediction result of the teacher network, namely elevator fault sound classification. The problem of minority class event detection caused by multi-source signal mixed interference and class imbalance is solved.
Owner:ZHEJIANG NEW ZAILING TECH CO LTD +1

Method and system for categorizing musical sound according to emotions

A computer implemented method for analysing sounds, such as audio tracks, and automatically classifying the sounds in a space in which arousal is one axis and valence is another axis. The location of a sound or track in that arousal-valence space is automatically determined using a computer implemented system that analyses, measures or infers values for each of the following base feature parameters: harmonicity, turbulence, rhythmicity, sharpness, volume and linear harmonic cost, or any combination of two or more of those parameters.
Owner:X SYSTEM LTD

Methods and systems for detecting vomit-related events

PCT designated stageWO2025238616A1Medical data miningSpeech analysisSound classificationAcoustics
A method (300) includes receiving a sequence of acoustic frames (122) and segmenting the sequence of acoustic frames into a plurality of events of interest characterizing one or more audible sounds (106). Each event of interest is associated with a respective subset of sequential acoustic frames (122) segmented from the sequence of acoustic frames (122). For corresponding event of interest, the method also includes: processing, using a machine learning (ML) model (200), the respective subset of sequential acoustic frames (122) to generate a corresponding probability distribution over possible audible sound classifications (202); and labeling, using a labeler (130), based on the corresponding probability distribution over possible audible sound classifications, the particular event of interest as a vomit-related event or a non- vomit-related event. The method also includes determining quantitative bout information indicating a number of the plurality of events of interest that are labeled as vomit-related events.
Owner:TAKEDA PHARMA CO LTD