Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

32 results about "Sound event detection" patented technology

System and method for CLAP4sed

A method for real-time sound event detection on an embedded device includes pretraining a contrastive language-audio pretraining model as an audio foundation model and preparing offline multimodal query prototypes for sound events of interest. The pretrained model and query prototypes are deployed on an embedded device. The device receives an input audio stream and extracts audio embeddings using the pretrained model. Similarity scores are calculated between the extracted audio embeddings and the prepared query prototypes. The presence of a sound event is determined based on the calculated similarity scores, and a real-time sound event detection result is output. The system includes a memory storing the pretrained model and query prototypes, an audio input interface, and a processor configured to perform the extraction, calculation, determination, and output operations. A non-transitory computer-readable medium stores instructions that, when executed, cause a processor to perform the method.
Owner:ROBERT BOSCH GMBH

Acoustic sound event detection system

ActiveUS12505852B2Speech analysisSound event detectionSpeech sound
In general, the disclosure describes a computing system to automatically identify and classify audio input, including non-speech audio signals. The computing system may also add new classes based on only a limited number of examples of the new classes, to identify classes of sounds for which the system had not been trained.
Owner:SRI INTERNATIONAL

A sound event detection method, device, apparatus and storage medium

ActiveCN115985294BSpeech recognitionSound event detectionAudio frequency
The application discloses a sound event detection method, device and equipment and a storage medium. The method comprises the following steps: acquiring to-be-detected audio data, and extracting acoustic characteristics of the to-be-detected audio data; determining a sound event detection result of the to-be-detected audio data according to the acoustic characteristics and a pre-determined sound event relationship characteristic; wherein the sound event relationship characteristic is determined based on a pre-constructed sound event relationship graph; and the sound event relationship graph is determined according to a statistical result of sound events in an audio data set. The technical scheme solves the problems of limited application range and low accuracy of sound event detection, and can effectively expand the application range of the detection scene while improving the accuracy of sound event detection.
Owner:AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG

Human voice data collection method and device, equipment and storage medium

The invention discloses a human voice data collection method and device, equipment and a storage medium. The method comprises the following steps: inputting received audio data into a sound event detection model; a feature extraction module of the model performs feature extraction on the audio data to obtain audio features; respectively inputting the audio features into a strong label learning module, a multi-instance learning module and a CTC module of the model, correspondingly outputting a frame-level event probability, a frame-level boundary probability, an audio-level existence probability, a frame-level attention weight and a sound event prediction sequence, and fusing the output results to obtain a target detection result; and when the target detection result represents that the voice event exists in the audio data, extracting target voice data from the audio data based on the target detection result. According to the method, the audio data with different granularity labels are optimized through the multi-task learning framework of the model, different task output results are fused, the time sequence detection problem of multi-sound event overlapping is solved, and the accuracy of the target human voice data is guaranteed.
Owner:ANHUI ENBOLI ELECTRIC CO LTD

Systems, methods, and apparatuses

In one possible example aspect, the present disclosure provides a context-aware noise compensation system for a wearable device playing back media content to a user in a noisy environment, wherein the system is configured to: obtain a sound signal from the noisy environment; determine context information associated with the sound signal, wherein the context information comprises at least one of: vocal sound detection information indicative of presence of one or more vocal sounds of the user in the sound signal, or sound event detection information indicative of presence of one or more environmental sound events in the sound signal; and optimize, based on the context information, the media content for playback by the wearable device.
Owner:DOLBY LABORATORIES LICENSING CORP

Small sample sound event detection method based on multi-scale feature aggregation

The invention discloses a small sample sound event detection method based on multi-scale feature aggregation, and relates to the technical field of audio signal processing and machine learning. The method comprises the following steps: firstly, preprocessing an input audio signal and extracting Mel spectrogram features; then, the features are input into a specially designed multi-scale feature aggregation network, the network extracts a time dimension, a frequency dimension and a local time-frequency joint feature through three parallel convolution paths, and deep fusion is carried out on multiple paths of features; thirdly, training the network by adopting a model-independent meta-learning framework, and enabling the network to learn the capability of quickly adapting to new tasks by simulating a large number of small sample classification tasks; and finally, using the trained model to quickly identify a new sound event category only containing a small number of samples. According to the method, through effective combination of multi-scale feature extraction and a meta-learning strategy, the problems of incomplete feature extraction and poor model generalization ability of a traditional method in a data scarcity scene are solved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Sound signal periodic feature extraction method, network model training method, storage medium and equipment

The invention discloses a sound signal periodic feature extraction method, a network model training method, a storage medium and equipment, and belongs to the technical field of sound event detection. The objective of the invention is to solve the problems of high sensing difficulty and poor decoupling effect of overlapped acoustic events in the current acoustic detection process. The method comprises the following steps: for a sound signal i, mapping the sound signal i to a low-dimensional space through two different linear layers to obtain p and g, and respectively carrying out expansion convolution operation on p and g to obtain pconv and gconv; for p and g, feature coding is carried out based on a Fourier basis function and a gating mechanism to obtain Fourier features, for pconv and gconv, Fourier features are obtained in the same mode, and Hadamard product is carried out on the pconv and the gconv to obtain representation of periodic features. And in the training process of the corresponding model, performing reconstruction error on the sum and the original signal i, respectively calculating two norms of the sum, and adding the two obtained two norms to obtain a Fourier series regular term for training the model.
Owner:HARBIN UNIV OF SCI & TECH

Sound event detection method, apparatus, device, and storage medium

The application relates to the technical field of sound recognition, and provides a low-resource sound event detection method, device and equipment and a storage medium, wherein the method comprises the following steps: acquiring sound to be detected; inputting the sound to be detected into a trained sound group category neural network model to obtain sound group category information; inputting the sound to be detected into an encoder of a pre-trained sound event category judgment model to obtain fine-grained feature information, and splicing and fusing the group category information and the fine-grained feature information to obtain fusion feature information; inputting the fusion feature information into a decoder of the sound event category judgment model, and decoding the fusion feature information based on an attention mechanism and in combination with a pre-obtained group category representation matrix to obtain a sound event judgment result. The application forms large-class group sound group category information containing rich information by using the sound to be detected, and realizes auxiliary discrimination based on large-class sound event results through a self-attention mechanism.
Owner:PING AN TECH (SHENZHEN) CO LTD

An acoustic event detection method based on multi-scale spatial feature and coordinate attention fusion

This invention discloses a sound event detection method based on the fusion of multi-scale spatial features and coordinate attention, belonging to the field of sound event detection technology. It aims to address the problems in existing sound event detection methods, such as the limited receptive field of the network, which makes it difficult to capture multi-scale temporal span features, and the inability to accurately focus on target sounds and suppress noise in complex time-frequency spaces. This method innovatively introduces a multi-scale coordinate attention module into the sound event feature extraction network. This module first uses parallel dilated convolutions with an increasing time dilation rate to obtain multi-scale temporal features and performs fusion and dimensionality reduction. Then, it employs a dual-axis time-frequency decoupling mechanism, performing adaptive pooling aggregation along both the time and frequency dimensions to generate directional time attention weights and frequency attention weights. Finally, the dual-axis weights are used to recalibrate the multi-scale features. This method significantly improves the accuracy and robustness of sound event detection.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Systems and methods for CLAP4SED

PendingCN121687102ASpeech analysisEngineeringSound event detection
A method for real-time sound event detection on an embedded device includes pre-training a contrast language-audio pre-training model as an audio base model and preparing an offline multi-modal query prototype for a sound event of interest. The pre-trained model and the query prototype are deployed on an embedded device. The device receives an input audio stream and extracts an audio embedding using a pre-trained model. A similarity score between the extracted audio embedding and the prepared query prototype is calculated. The presence of the sound event is determined based on the calculated similarity score, and the real-time sound event detection result is output. The system includes a memory storing a pre-trained model and a query prototype, an audio input interface, and a processor configured to perform extraction, calculation, determination, and output operations. A non-transitory computer readable medium stores instructions that, when executed, cause a processor to implement the method.
Owner:ROBERT BOSCH GMBH

SOUND EVENT DETECTION SYSTEM

The invention relates to a sound event detection system (1) comprising a data processing unit (2) and a sound event recording unit (11). The data processing unit (2) is suitable for generating a user interface (5) that can be displayed on a data display device (4) that is data-connected to the data processing unit (2). Several user data entries (8) can be generated via the user interface (5) using a data input device that is data-connected to the data processing unit (2). These user data entries can be stored in a database (10) that is connected to the data processing unit (2) via at least one database interface. Functional data (9) can be generated in the data processing unit (2) by comparing and / or data-connected to the system data (7) stored in the data processing unit (2).The sound event recording device (11) is arranged in a transport vehicle (12), in particular a small van, which is connected to the database (10) via data processing. The sound event recording device (11) comprises at least one digital audio interface (15) arranged in a mobile and / or portable frame (14) that can be attached to the transport vehicle (12), in particular a small van, and having at least several microphone inputs for several microphones that can be stowed in the transport vehicle (12), in particular a small van. The transport vehicle (12), in particular a small van, includes a navigation and communication device (16) which is connected to the data processing device (2) via data processing.The transport vehicle (12), in particular a small van, includes a computing device (17), in particular a portable device, which is connected to the audio interface (15) and to the data processing device (2). When navigating the transport vehicle (12), in particular a small van, and / or when communicating with the communication device (16), and / or when exchanging data with the computing device (17), at least some of the functional data (9) can be queried electronically.
Owner:SCHIERENBERG KAYE +1

A multi-sound event detection and positioning method and device based on a neural network model

This invention relates to the field of artificial intelligence technology, and in particular to a method and apparatus for detecting and locating multiple sound events based on a neural network model. The method includes: innovatively designing time-frequency multi-scale residual convolutional blocks, which, together with a Conformer module and a cross-stitch unit module, form a network model to extract features at multiple scales, enhance long sequence modeling, and promote task-based collaborative optimization, thereby improving performance and accuracy; in terms of data processing, pre-emphasis and frame-by-frame windowing improve feature quality, audio channel swapping and spectrum enhancement increase data diversity and reduce overfitting, and SALSA-Lite features are used to enhance feature representation; in terms of training strategy, a multivariate loss function is used to accelerate convergence while considering task requirements, and hyperparameters are flexibly adjusted using a validation set. This makes the method highly efficient in training, has excellent practical performance, strong generalization ability on unknown data, and can accurately cope with complex and ever-changing real-world scenarios, effectively overcoming the shortcomings of traditional methods.
Owner:UNIV OF SCI & TECH BEIJING

Sound event detection method based on recursive gated convolution and self-attention mechanism

The invention discloses a sound event detection method based on recursive gating convolution and a self-attention mechanism, and the method comprises the steps: collecting a to-be-detected audio signal, and constructing a sound event detection model; inputting the audio signal into a sound event detection model, and extracting a time-frequency feature through a preprocessing module to obtain a logarithmic Mel spectrum feature; inputting the logarithmic Mel spectrum features into a convolution module, and performing spatial feature fusion through recursive gating convolution to obtain spatial fusion features; inputting the spatial fusion features into a time domain modeling module, and performing context information interaction through a self-attention mechanism to obtain global information features; and inputting the global information features into a KANLinear classifier, and carrying out frame-by-frame classification to obtain each sound event category in the audio and the occurrence time period thereof. According to the method, the accuracy of sound event detection can be remarkably improved on the premise that the number of parameters is small and the number of floating point operation times is low.
Owner:ANHUI UNIV

A smart audio non-invasive monitoring method and system for a home environment

The application discloses a smart audio non-inductive monitoring method and system for a home environment, and belongs to the technical field of smart home and health monitoring, which comprises the following steps: collecting audio data in a room and extracting features to obtain an input feature tensor; based on the input feature tensor, sound event detection and sound source positioning are performed through a SELD model based on a deep neural network to obtain local observation data, which is converted to a global coordinate system; the position of a target object in the global coordinate system is continuously tracked and state estimation is performed based on a particle filter to obtain a global motion trajectory; based on the local observation data and the global motion trajectory, combined with trigger information of infrared and radar sensors, decision-level fusion is performed by using Dempster-Shafer evidence theory to output a final activity state and an environmental abnormal event, thereby solving the problems that in the current monitoring scheme, normal life of a user is disturbed, privacy and self-esteem are highly invasive, and monitoring data is not comprehensive.
Owner:NORTH CHINA UNIVERSITY OF TECHNOLOGY

User profile generation methods, devices, electronic devices, and computer-readable media

This application discloses a user profile generation method, apparatus, electronic device, and computer-readable medium. An embodiment of the method includes: segmenting a user's call audio into an audio segment without human voice and an audio segment containing human voice; extracting a first audio feature from the audio segment without human voice, inputting the first audio feature into a pre-trained first ambient sound event detection model to obtain a first detection result; extracting a second audio feature from the audio segment containing human voice, inputting the second audio feature into a pre-trained second ambient sound event detection model to obtain a second detection result; and generating a user profile based on the first and second detection results. This implementation enriches the methods for generating user profiles. User profiles generated in this way can provide users with services related to their environment, thereby improving service quality.
Owner:CHINA TELECOM CORP LTD

Multi-source positioning and detection method based on global-local feature recalibration

The application discloses a multi-sound source positioning and detection method based on global-local feature recalibration, which comprises the following steps: calculating the short-time Fourier transform of a multi-channel spatial audio signal in a first-order stereo format, obtaining a log-linear spectrum and a normalized sound intensity vector as input features, and then performing data augmentation on the features of a training set; splicing the augmented spectrum and sound intensity vector as the input of a neural network model, training the neural network model, obtaining optimal network model parameters and saving them; preprocessing a test sample and inputting it into the trained model, outputting the predicted sound event category and position information, drawing a sound event detection graph, a direction angle and an azimuth angle trajectory curve graph according to the prediction result, and comparing them with the visualized image of the test sample real label to analyze the performance of the model. The application can achieve high sound source positioning and detection performance, and the model shows good generalization on real and synthetic data sets.
Owner:XINJIANG UNIVERSITY

Systems and methods for remote multi-directional bark deterrence

An apparatus is described that comprises a microphone array and a plurality of transducers. The microphone array and the plurality of transducers are communicatively coupled with at least one processor. The apparatus includes the microphone array for receiving at least one signal. Each transducer of the plurality of transducers is configured to deliver a correction signal along a transducer beam spread axis, wherein the plurality of transducers is positioned on the apparatus for providing a combined transducer beam spread coverage in the horizontal plane. One or more applications running on the at least one processor use information of the at least one signal to detect a sound event. The detecting the sound event includes selecting transducers from the plurality of transducers and instructing the selected transducers to deliver a correction signal.
Owner:RADIO SYST CORP

Artificial intelligence-based sound event detection method, apparatus, device, and medium

The present application relates to the technical field of digital medical treatment, and in particular to a sound event detection method, device and equipment based on artificial intelligence and a medium. The method separates mixed sound into independent sound, inputs the independent sound into an encoder to obtain sound features, inputs the sound features into a recurrent layer to obtain timing features, uses a label prediction model to process and predict pseudo-event labels according to the sound features and the timing features, queries the pseudo-event labels as reference labels, uses the reference labels and the independent sound to form training samples, trains an event detection model, and then obtains an event detection result, extracts timing information of the sound features, enriches input of event prediction, improves accuracy of event prediction, queries and filters the pseudo-event labels to determine the reference labels, so that the event detection model better adapts to a scene, improves accuracy of sound event detection, and can assist medical staff in timely discovering abnormal sound events of patients in a medical environment, so as to respond in time.
Owner:PING AN TECH (SHENZHEN) CO LTD

Animal sound event detection model training method, animal sound event detection method and animal sound event detection device

PendingCN121768404ASpeech analysisBiological modelsSmall sampleSound event detection
The invention provides a training method, a detection method and a device of an animal sound event detection model, and relates to the technical field of intelligent audio signal processing. The training method comprises the steps that a sample support set and a multi-class training set are obtained, the sample support set comprises first positive and negative samples and a query sample, and the multi-class training set comprises the first positive sample; extracting sample frame-level acoustic features, wherein the sample frame-level acoustic features comprise a sample Log-Mel feature and a sample PCEN feature; performing multi-task joint fine tuning on the pre-training detection model through the sample frame-level acoustic features to obtain an intermediate detection model; performing animal sound event detection on the query sample through the intermediate detection model to obtain a first detection result so as to determine a second positive sample, and updating the frame-level acoustic features of the sample; and iteratively executing multi-task joint fine tuning and query sample detection to finally obtain an animal sound event detection model. According to the invention, high-precision and high-robustness animal sound event detection can be realized under the condition of small samples.
Owner:HAINAN KEXUN ZHIAN TECHNOLOGY CO LTD +3

Open audio tracking system with audio base model and large language model

PendingCN122177165ASpeech analysisHearing aids signal processingFeedback loopSound event detection
A method of executing an open audio tracking system is disclosed. A feedback loop between an audio foundation model (AFM) and a large language model (LLM) enables both low-level sound event detection and high-level acoustic scene detection in real-time, which are then used to generate additional text-based event descriptions that are applied to subsequent iteration cycles of the system. The AFM can be similar to a contrastive language-audio pre-training (CLAP) model configured to detect sound events, while the LLM receives the detected given sound events and categorizes these events into acoustic sound classes that can interpret the sound events in the context of the environment.
Owner:ROBERT BOSCH GMBH

Systems and methods to predict aggression in surveillance camera video

Methods and systems for predicting aggressive behavior associated with a surveillance scene. Images are generated from a camera of a surveillance scene, and audio is also generated. A local computing system executes an object classification model on the images to predict one or more classes of objects in the scene. The local computing system also executes a sound-event detection model on the audio to predict one or more classes of events occurring in the scene. Metadata is generated associated with the image-based classes and audio-based classes. The metadata is transferred to a remote computing system that executes a knowledge graph on the metadata to implement knowledge graph-based reasoning to predict aggressive behavior occurring in the surveillance scene based on the metadata. The metadata associated with the predicted aggressive behavior is labeled as such, and control commands are output accordingly.
Owner:ROBERT BOSCH GMBH

Systems, methods, and apparatuses

In one possible example aspect, the present disclosure provides a context-aware noise compensation system for a wearable device playing back media content to a user in a noisy environment, wherein the system is configured to: obtain a sound signal from the noisy environment; determine context information associated with the sound signal, wherein the context information comprises at least one of: vocal sound detection information indicative of presence of one or more vocal sounds of the user in the sound signal, or sound event detection information indicative of presence of one or more environmental sound events in the sound signal; and optimize, based on the context information, the media content for playback by the wearable device.
Owner:DOLBY LABORATORIES LICENSING CORP

Pig abnormal sound monitoring and early warning method, system and equipment based on audio and video fusion

The invention discloses a live pig abnormal sound monitoring and early warning method, system and device based on audio and video fusion. The method comprises the steps of obtaining audio and video stream data; extracting audio frequency spectrum features and inputting the audio frequency spectrum features into a sound event detection and positioning model to obtain a prediction result; analyzing the video to obtain the real-time position and activity state of the live pig; performing multi-modal fusion judgment on the audio prediction result and the live pig state to generate a high-confidence identification result; and generating a structured log and carrying out early warning pushing. According to the method, audio classification, positioning and video analysis technologies are combined, the recognition accuracy is improved through multi-modal comparison, the problem that detection and positioning are difficult in a strong noise environment is effectively solved, and efficient and accurate monitoring of the abnormal state of the live pig is achieved.
Owner:SOUTH CHINA AGRICULTURAL UNIVERSITY

Sound event detection method and system with class distribution and temporal context collaborative cues

The application discloses a sound event detection method and system based on distribution and timing context cooperative prompting, converts an original audio signal into a signal frame sequence, extracts a mel filter bank feature through a pre-training model branch, and extracts a mel spectrum feature through a downstream model branch; a combination of a local timing prompt module and a global distribution prompt module is introduced into each layer of the pre-training model, the pre-training model outputs an audio sequence feature and the global distribution prompt module; the audio sequence and the output of the downstream model are fused in features, frame-level prediction probabilities of the downstream model are calculated, and an audio positioning task is realized; then, the frame-level prediction probabilities of the downstream model are used to generate sentence-level prediction probabilities; for the pre-training model, the global distribution prompt module is processed to obtain sentence-level prediction probabilities of the pre-training model; finally, the two sentence-level prediction probabilities are fused to obtain a classification result of a sound event. The application can significantly improve the audio classification and positioning performance of sound event detection.
Owner:JIANGSU UNIV

Sound event detection method based on time-frequency convolution and feature enhancement

The invention discloses a sound event detection model based on a time-frequency double-branch dynamic convolution and feature fusion enhancement module, and the model employs a time-frequency double-branch dynamic convolution structure to extract local structure features, a time domain branch is only convolved along a time dimension, and a frequency domain branch is only convolved along a frequency dimension. The two branches extract multi-scale features in parallel and implement gating fusion in a channel dimension to obtain compact time-frequency joint representation with stronger discrimination; and meanwhile, a feature fusion enhancement module is introduced to carry out explicit alignment on convolutional features and pre-training semantic embedding, so that the representation consistency of fusion features is improved. According to the method, the difference of sound events with different time lengths and frequency modes can be more accurately modeled under similar parameter quantities and calculated quantities, and the sound event boundary positioning stability and the category recognition accuracy are improved.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

A method and system for cetacean overlapping sound event detection based on adaptive multi-scale synthetic attention

The application discloses a kind of based on adaptive multi-scale synthetic attention whale overlapping sound event detection method and system, belong to ocean engineering and marine signal technical field, method includes: collection historical whale sound signal, pre-process and ACT transform historical whale sound signal to obtain transformed data set;Cross-shift feature extraction network is constructed to time-frequency perception, and cross-shift feature extraction network is trained using transformed data to time-frequency perception, to obtain adaptive multi-scale synthetic attention network;Adaptive window size prediction network is constructed, and according to the feature map output by adaptive multi-scale synthetic attention network, the window parameters required for constructing multi-scale attention fusion network are predicted;Multi-scale attention fusion network is constructed based on window parameter and is iterated, to obtain whale overlapping sound event detection model;Real-time whale sound signal is obtained, and real-time whale sound signal is detected using detection model, to obtain detection result.
Owner:ANHUI UNIV

Cross-branch feature interaction sound event detection method and system guided by category semantic prior module

The invention discloses a cross-branch feature interaction sound event detection method and system guided by a category semantic prior module, and the method comprises the steps: converting an original audio signal into a signal frame sequence with a continuous time sequence, extracting the features of a Mel filter bank through a pre-training model branch, and extracting the features of a Mel spectrum through a downstream model branch; the pre-training model branch adopts an audio pre-training model based on a Transform architecture, and the downstream model branch adopts a convolutional neural network; a category semantic prior module is introduced into a pre-training model branch to serve as a core medium of cross-branch feature interaction, and fine feature interaction with a downstream CNN branch is realized through a hierarchical and bidirectional multi-head cross attention mechanism, so that the problems of semantic dislocation and information redundancy in a traditional fusion method are solved; and finally, outputting a sound event positioning and classification result through feature splicing, time sequence modeling and decision-level fusion.
Owner:JIANGSU UNIV

A large-scale bird sound recognition method based on deep learning technology

The application relates to the cross field of artificial intelligence and bird ecology, and specifically discloses a large-scale bird sound recognition method based on a deep learning technology, which collects bird sound audio through a project area recording device, trains a bird sound event detection model by using a Bird audio detection challenge 2018 data set, and separates effective bird sound and background noise by using field recording data for category prediction; meanwhile, a bird sound recognition data set is constructed by combining a bird species directory of a Chinese bird watching recording center and an audio file of Xeno-Canto, a bird sound recognition model is trained by using the background noise and the bird sound recognition data set, and verification is carried out; finally, effective bird sound data is input into the model for label prediction to recognize bird sound. The application aims to save data processing cost and improve the recognition accuracy in a complex scene by carrying out extensive bird sound recognition in a project area through the deep learning technology.
Owner:GUANGZHOU UNIVERSITY

Sound event detection model training methods, electronic devices and storage media

This invention discloses a method for training a sound event detection model, an electronic device, and a storage medium. The method for training the sound event detection model includes: acquiring pairs of audio and text data from a preset dataset as training data, wherein sound events appearing in the text have corresponding frame-level labels, and the training data also has corresponding segment-level labels; calculating a segment-level weak loss based on the segment-level labels and the output of the sound event detection model to pre-train the sound detection model; calculating a frame-level strong loss based on the frame-level labels and the output of the sound event detection model; and fine-tuning the sound event detection model by combining the frame-level strong loss and the segment-level weak loss.
Owner:AISPEECH CO LTD

A sound event detection method and system based on group feature calibration

This invention provides a sound event detection method and system based on grouped feature calibration, comprising: acquiring audio feature data of the sound event to be detected; inputting the audio feature data into a time-frequency learning network, obtaining a time-spectrum graph through a convolutional neural network, performing grouped feature learning based on the intermediate representations of the time-spectrum graph from multiple dimensions to obtain grouped enhanced features, performing task-aware activation on the grouped enhanced features to obtain adaptive features; inputting the adaptive features into a context modeling network to obtain the time-domain correlation features of the audio signal, classifying the time-domain correlation features of the audio signal, and obtaining the sound event category detection result. This invention introduces a grouped feature calibration module based on the time-frequency characteristics of different types of audio in the sound event detection task, enhancing the feature representation capability of the sound event detection network for various types of audio, with a small number of parameters and strong versatility, and introducing it into existing mainstream sound event detection models with a relatively low computational cost while improving their performance.
Owner:WUHAN UNIV