Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

59 results about "Sound event detection" patented technology

Sound event detection using surface mounted vibration sensors

A method for detecting a sound field in an outside environment of a solid structure is provided. The solid structure comprises at least one sound transducing element, which is a part of the solid structure and forms a part of an outer surface of the solid structure, wherein the incident sound falls at least partly onto the sound transducing element. The method comprises obtaining sensor data from at least one vibration sensor, which is mounted onto a mounting surface of the at least one sound transducing element, wherein the at least one vibration sensor detects vibrations in the at least one sound transducing element, which are induced into the at least one sound transducing element by the incident sound field, processing the sensor data to generate output data, and providing the output data.
Owner:HARMAN INT IND INC

Sound event detection method and system for collaborative prompting of class distribution and time sequence context

The invention discloses a sound event detection method and system based on class distribution and time sequence context collaborative prompt, and the method comprises the steps: converting an original audio signal into a signal frame sequence, extracting the features of a Mel filter bank through a pre-training model branch, and extracting the features of a Mel spectrum through a downstream model branch; a combination of a local time sequence prompt module and a global distribution prompt module is introduced into each layer of the pre-training model, and the pre-training model outputs audio sequence features and the global distribution prompt module; performing feature fusion on the audio sequence and the output of the downstream model, calculating the frame level prediction probability of the downstream model, and realizing an audio positioning task; generating a sentence level prediction probability from the frame level prediction probability of the downstream model; for the pre-training model, processing the global distribution prompt module to obtain a sentence level prediction probability of the pre-training model; and finally, fusing the two sentence level prediction probabilities to obtain a classification result of the sound event. According to the method, the audio classification and positioning performance of sound event detection can be remarkably improved.
Owner:JIANGSU UNIV

System and method for CLAP4sed

A method for real-time sound event detection on an embedded device includes pretraining a contrastive language-audio pretraining model as an audio foundation model and preparing offline multimodal query prototypes for sound events of interest. The pretrained model and query prototypes are deployed on an embedded device. The device receives an input audio stream and extracts audio embeddings using the pretrained model. Similarity scores are calculated between the extracted audio embeddings and the prepared query prototypes. The presence of a sound event is determined based on the calculated similarity scores, and a real-time sound event detection result is output. The system includes a memory storing the pretrained model and query prototypes, an audio input interface, and a processor configured to perform the extraction, calculation, determination, and output operations. A non-transitory computer-readable medium stores instructions that, when executed, cause a processor to perform the method.
Owner:ROBERT BOSCH GMBH

Acoustic sound event detection system

In general, the disclosure describes a computing system to automatically identify and classify audio input, including non-speech audio signals. The computing system may also add new classes based on only a limited number of examples of the new classes, to identify classes of sounds for which the system had not been trained.
Owner:SRI INTERNATIONAL

Sound event detection data synthesis and sound event detection model training method

The invention discloses a semantic prompt-based sound event detection data synthesis and sound event detection model training method. The sound event detection task is converted into the semantic description information, and the structured semantic prompt instruction which accurately reflects the target sound event characteristics is generated in combination with the semantic constraint rule, so that automatic mapping from semantic description to instruction generation is realized, and the manual intervention cost is reduced. And inputting the structured semantic prompt instruction into the audio generation model, and synthesizing the audio data in a large scale, thereby reducing the data acquisition cost and improving the sample diversity and expandability. The structured semantic prompt instruction can guide the model to synthesize audios of various sound event types in batches, and is automatically generated by a large language model to ensure that the synthesized audios are strictly aligned with instruction semantics. When the label is generated, the sample event type can be obtained without manual labeling, and an efficient and reliable data source is provided for sound event detection model training.
Owner:SHANGHAI NORMAL UNIVERSITY +1

Multi-sound event detection positioning method and device based on neural network model

The invention relates to the technical field of artificial intelligence, in particular to a multi-sound event detection positioning method and device based on a neural network model. The method comprises the following steps: innovatively designing a time-frequency multi-scale residual error convolution block, forming a network model with a Conformer module and a cross-stitch unit module, extracting features in a multi-scale manner, enhancing long sequence modeling, promoting task collaborative optimization, and improving performance and accuracy; in the aspect of data processing, pre-emphasis and framing windowing improve feature quality, audio channel exchange and spectrum enhancement increase data diversity and reduce overfitting, and SALSA-Lite features are adopted to enhance feature expression. In the training strategy, a multivariate loss function is applied to accelerate convergence in consideration of task requirements, and hyper-parameters are flexibly adjusted by means of a verification set. Therefore, the method is high in training efficiency, excellent in actual performance, high in generalization ability on unknown data, capable of accurately coping with complex and changeable actual scenes and capable of effectively overcoming the defects of a traditional method.
Owner:UNIV OF SCI & TECH BEIJING

A sound event detection method, device, apparatus and storage medium

The application discloses a sound event detection method, device and equipment and a storage medium. The method comprises the following steps: acquiring to-be-detected audio data, and extracting acoustic characteristics of the to-be-detected audio data; determining a sound event detection result of the to-be-detected audio data according to the acoustic characteristics and a pre-determined sound event relationship characteristic; wherein the sound event relationship characteristic is determined based on a pre-constructed sound event relationship graph; and the sound event relationship graph is determined according to a statistical result of sound events in an audio data set. The technical scheme solves the problems of limited application range and low accuracy of sound event detection, and can effectively expand the application range of the detection scene while improving the accuracy of sound event detection.
Owner:AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG

Human voice data collection method and device, equipment and storage medium

The invention discloses a human voice data collection method and device, equipment and a storage medium. The method comprises the following steps: inputting received audio data into a sound event detection model; a feature extraction module of the model performs feature extraction on the audio data to obtain audio features; respectively inputting the audio features into a strong label learning module, a multi-instance learning module and a CTC module of the model, correspondingly outputting a frame-level event probability, a frame-level boundary probability, an audio-level existence probability, a frame-level attention weight and a sound event prediction sequence, and fusing the output results to obtain a target detection result; and when the target detection result represents that the voice event exists in the audio data, extracting target voice data from the audio data based on the target detection result. According to the method, the audio data with different granularity labels are optimized through the multi-task learning framework of the model, different task output results are fused, the time sequence detection problem of multi-sound event overlapping is solved, and the accuracy of the target human voice data is guaranteed.
Owner:ANHUI ENBOLI ELECTRIC CO LTD

Systems, methods, and apparatuses

In one possible example aspect, the present disclosure provides a context-aware noise compensation system for a wearable device playing back media content to a user in a noisy environment, wherein the system is configured to: obtain a sound signal from the noisy environment; determine context information associated with the sound signal, wherein the context information comprises at least one of: vocal sound detection information indicative of presence of one or more vocal sounds of the user in the sound signal, or sound event detection information indicative of presence of one or more environmental sound events in the sound signal; and optimize, based on the context information, the media content for playback by the wearable device.
Owner:DOLBY LABORATORIES LICENSING CORP

Small sample sound event detection method based on multi-scale feature aggregation

The invention discloses a small sample sound event detection method based on multi-scale feature aggregation, and relates to the technical field of audio signal processing and machine learning. The method comprises the following steps: firstly, preprocessing an input audio signal and extracting Mel spectrogram features; then, the features are input into a specially designed multi-scale feature aggregation network, the network extracts a time dimension, a frequency dimension and a local time-frequency joint feature through three parallel convolution paths, and deep fusion is carried out on multiple paths of features; thirdly, training the network by adopting a model-independent meta-learning framework, and enabling the network to learn the capability of quickly adapting to new tasks by simulating a large number of small sample classification tasks; and finally, using the trained model to quickly identify a new sound event category only containing a small number of samples. According to the method, through effective combination of multi-scale feature extraction and a meta-learning strategy, the problems of incomplete feature extraction and poor model generalization ability of a traditional method in a data scarcity scene are solved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A bird sound event detection method and system for monitoring state key protected birds

The application discloses a bird sound event detection method and system for monitoring state-protected birds, and comprises the following steps: collecting bird sound data and environmental noise data of state-protected birds, and constructing a bird sound spectrogram data set and a noisy bird sound data set; constructing a bird sound noise reduction model and a target detection model; training and jointly optimizing the bird sound noise reduction model and the target detection model according to the bird sound spectrogram data set and the noisy bird sound data set; constructing an image-audio mapping module for converting the image detection result of the target detection model into an audio sound event detection result; acquiring bird sound data to be detected, searching and sound source positioning the target bird sound according to the bird sound noise reduction model, the target detection model and the image-audio mapping module, and determining the species, time-frequency and direction information of the target bird sound. The application can improve the detection effect of field bird sound, and further improve the classification accuracy, and can be widely applied to the technical field of acoustic information processing.
Owner:GUANGZHOU UNIVERSITY +1

Sound signal periodic feature extraction method, network model training method, storage medium and equipment

The invention discloses a sound signal periodic feature extraction method, a network model training method, a storage medium and equipment, and belongs to the technical field of sound event detection. The objective of the invention is to solve the problems of high sensing difficulty and poor decoupling effect of overlapped acoustic events in the current acoustic detection process. The method comprises the following steps: for a sound signal i, mapping the sound signal i to a low-dimensional space through two different linear layers to obtain p and g, and respectively carrying out expansion convolution operation on p and g to obtain pconv and gconv; for p and g, feature coding is carried out based on a Fourier basis function and a gating mechanism to obtain Fourier features, for pconv and gconv, Fourier features are obtained in the same mode, and Hadamard product is carried out on the pconv and the gconv to obtain representation of periodic features. And in the training process of the corresponding model, performing reconstruction error on the sum and the original signal i, respectively calculating two norms of the sum, and adding the two obtained two norms to obtain a Fourier series regular term for training the model.
Owner:HARBIN UNIV OF SCI & TECH

Sound event detection model training method and device, equipment and storage medium

The invention relates to a sound event detection model training method and device, equipment and a storage medium. The method comprises the following steps: inputting a sample audio into a pre-trained sound event positioning model to obtain a prediction label of the sample audio; the sample audio is labeled with an initial label, and the initial label is used for representing whether the sample audio contains a sound event; the prediction label comprises probability distribution information that each audio frame in the sample audio comprises a sound event; cutting out an audio clip conforming to the target audio length from the sample audio, and marking a clip tag corresponding to the audio clip according to the initial tag and the prediction tag; the fragment label is used for representing whether the audio fragment contains a sound event; and according to the audio clip and the clip tag corresponding to the audio clip, training a to-be-trained sound event detection model to obtain a pre-trained sound event detection model. By adopting the method, the training efficiency of the sound event detection model can be improved.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

A data augmentation method for audio, a real-time sound event detection system and method

The present invention discloses a method for data augmentation of audio and a real-time sound event detection system and method, including: establishing two deep learning models, where the deep learning models include a pre-awakening model and a detection model; deforming audio data through a series of audio transformations to obtain deformed audio; extracting spectral features from the deformed audio; randomly masking the spectral features to obtain data after data augmentation; using the data after data augmentation to train the deep learning models in the sound event detection method and saving the trained models; using a microphone to record an audio stream in real time, slicing the audio stream and sending it into the trained models for detection to obtain detection results. According to the present invention, the sound events to be detected can be quickly and accurately fed back, and the number of model parameters and the memory occupancy in this method are small, meeting the conditions for mobile use.
Owner:UNIV OF SHANGHAI FOR SCI & TECH

Sound event detection method, apparatus, device, and storage medium

The application relates to the technical field of sound recognition, and provides a low-resource sound event detection method, device and equipment and a storage medium, wherein the method comprises the following steps: acquiring sound to be detected; inputting the sound to be detected into a trained sound group category neural network model to obtain sound group category information; inputting the sound to be detected into an encoder of a pre-trained sound event category judgment model to obtain fine-grained feature information, and splicing and fusing the group category information and the fine-grained feature information to obtain fusion feature information; inputting the fusion feature information into a decoder of the sound event category judgment model, and decoding the fusion feature information based on an attention mechanism and in combination with a pre-obtained group category representation matrix to obtain a sound event judgment result. The application forms large-class group sound group category information containing rich information by using the sound to be detected, and realizes auxiliary discrimination based on large-class sound event results through a self-attention mechanism.
Owner:PING AN TECH (SHENZHEN) CO LTD

Model generation methods, sound event detection methods, devices, media and equipment

This disclosure relates to a model generation method, a sound event detection method, an apparatus, a medium, and a device. The method includes: extracting audio features from a first sample audio file and inputting the audio features into a sound event detection model to obtain a first detection result; determining the target loss of the model based on the first annotation result of the first sample audio file and the first detection result, the target loss including cross-entropy loss and continuity error penalty loss; and updating the model parameters based on the target loss. Since continuity detection errors are the main cause of misidentification, the continuity error penalty term is included as part of the loss function during the sound event detection model training process. The model trained using this method can maintain high accuracy while keeping the sound event recall rate unchanged. That is, it can significantly reduce misidentifications that affect user experience without introducing more missed detections, thereby improving the accuracy of sound event detection and user experience.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

An acoustic event detection method based on multi-scale spatial feature and coordinate attention fusion

This invention discloses a sound event detection method based on the fusion of multi-scale spatial features and coordinate attention, belonging to the field of sound event detection technology. It aims to address the problems in existing sound event detection methods, such as the limited receptive field of the network, which makes it difficult to capture multi-scale temporal span features, and the inability to accurately focus on target sounds and suppress noise in complex time-frequency spaces. This method innovatively introduces a multi-scale coordinate attention module into the sound event feature extraction network. This module first uses parallel dilated convolutions with an increasing time dilation rate to obtain multi-scale temporal features and performs fusion and dimensionality reduction. Then, it employs a dual-axis time-frequency decoupling mechanism, performing adaptive pooling aggregation along both the time and frequency dimensions to generate directional time attention weights and frequency attention weights. Finally, the dual-axis weights are used to recalibrate the multi-scale features. This method significantly improves the accuracy and robustness of sound event detection.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Train water tank water level positioning method, device, equipment and storage medium

The present invention provides a method, device, equipment and storage medium for locating the water level of a train water tank, comprising: obtaining an audio segment to be tested; performing feature extraction on the audio segment to be tested to obtain multiple frames of feature vectors to be tested corresponding to the audio segment to be tested; and substituting the multiple frames of feature vectors to be tested corresponding to the audio segment to be tested into a pre-established sound event detection model to obtain the water level height magnitude corresponding to the audio segment to be tested. By establishing a sound event detection model, a correspondence between the feature vector corresponding to the audio segment and the water level height magnitude is constructed. After obtaining the audio segment to be tested, its features are first extracted to obtain the feature vector to be tested, and then the feature vector to be tested is substituted into the model. This allows for rapid and accurate determination of the water level magnitude corresponding to the feature vector to be tested, i.e., the water level magnitude corresponding to the audio segment to be tested. This facilitates timely judgment of the water level condition of the train water tank during the water filling process, improves water filling efficiency and avoids waste of water resources caused by water overflow.
Owner:CRRC ZHUZHOU ELECTRIC LOCOMOTIVE RESEARCH INSTITUTE CO LTD

Systems and methods for CLAP4SED

A method for real-time sound event detection on an embedded device includes pre-training a contrast language-audio pre-training model as an audio base model and preparing an offline multi-modal query prototype for a sound event of interest. The pre-trained model and the query prototype are deployed on an embedded device. The device receives an input audio stream and extracts an audio embedding using a pre-trained model. A similarity score between the extracted audio embedding and the prepared query prototype is calculated. The presence of the sound event is determined based on the calculated similarity score, and the real-time sound event detection result is output. The system includes a memory storing a pre-trained model and a query prototype, an audio input interface, and a processor configured to perform extraction, calculation, determination, and output operations. A non-transitory computer readable medium stores instructions that, when executed, cause a processor to implement the method.
Owner:ROBERT BOSCH GMBH

Audio data enhancement method, device and medium based on sound pickup environment factors

The present invention discloses an audio data enhancement method, device, and medium based on the collection of sound pickup environmental factors. The method comprises obtaining a sample training set of the original audio data to be enhanced; sequentially batching, verifying, labeling, and merging the sample training set; scheduling a microphone and a speaker to mix the entire audio data of each sample batch in a real environment with the sound pickup environmental factors; sequentially segmenting and labeling the entire recording data based on the batching and labeling to obtain an enhanced sample training set for the current sample batch; obtaining the enhanced sample training set for each sample batch and splicing them together to obtain the final enhanced sample training set for the original audio data. Advantages: The present invention simultaneously takes into account environmental factors such as environmental background noise, the distance between the microphone and the sound source, and interference generated within the microphone, thereby more effectively introducing environmental information, thereby improving the accuracy of the sound event detection model in a real environment and reducing performance degradation.
Owner:WUHAN UNIV

SOUND EVENT DETECTION SYSTEM

The invention relates to a sound event detection system (1) comprising a data processing unit (2) and a sound event recording unit (11). The data processing unit (2) is suitable for generating a user interface (5) that can be displayed on a data display device (4) that is data-connected to the data processing unit (2). Several user data entries (8) can be generated via the user interface (5) using a data input device that is data-connected to the data processing unit (2). These user data entries can be stored in a database (10) that is connected to the data processing unit (2) via at least one database interface. Functional data (9) can be generated in the data processing unit (2) by comparing and / or data-connected to the system data (7) stored in the data processing unit (2).The sound event recording device (11) is arranged in a transport vehicle (12), in particular a small van, which is connected to the database (10) via data processing. The sound event recording device (11) comprises at least one digital audio interface (15) arranged in a mobile and / or portable frame (14) that can be attached to the transport vehicle (12), in particular a small van, and having at least several microphone inputs for several microphones that can be stowed in the transport vehicle (12), in particular a small van. The transport vehicle (12), in particular a small van, includes a navigation and communication device (16) which is connected to the data processing device (2) via data processing.The transport vehicle (12), in particular a small van, includes a computing device (17), in particular a portable device, which is connected to the audio interface (15) and to the data processing device (2). When navigating the transport vehicle (12), in particular a small van, and / or when communicating with the communication device (16), and / or when exchanging data with the computing device (17), at least some of the functional data (9) can be queried electronically.
Owner:SCHIERENBERG KAYE +1

A multi-sound event detection and positioning method and device based on a neural network model

This invention relates to the field of artificial intelligence technology, and in particular to a method and apparatus for detecting and locating multiple sound events based on a neural network model. The method includes: innovatively designing time-frequency multi-scale residual convolutional blocks, which, together with a Conformer module and a cross-stitch unit module, form a network model to extract features at multiple scales, enhance long sequence modeling, and promote task-based collaborative optimization, thereby improving performance and accuracy; in terms of data processing, pre-emphasis and frame-by-frame windowing improve feature quality, audio channel swapping and spectrum enhancement increase data diversity and reduce overfitting, and SALSA-Lite features are used to enhance feature representation; in terms of training strategy, a multivariate loss function is used to accelerate convergence while considering task requirements, and hyperparameters are flexibly adjusted using a validation set. This makes the method highly efficient in training, has excellent practical performance, strong generalization ability on unknown data, and can accurately cope with complex and ever-changing real-world scenarios, effectively overcoming the shortcomings of traditional methods.
Owner:UNIV OF SCI & TECH BEIJING

ECAPA-TDNN bird sound identification method based on dynamic frequency band division

The invention relates to the technical field of bird sound recognition, and provides an ECAPA-TDNN bird sound recognition method based on dynamic frequency band division, and the method is characterized in that the method comprises the steps: recognition data set construction, which comprises the steps: downloading a target bird species audio file from an original bird species audio through a bird sound event detection model, constructing an identification data set containing the target bird species audio file; the construction of a dynamic sub-band ECAPA-TDNN model comprises the step of integrating a dynamic band segmentation module, a multi-band parallel processing architecture and a cross-band gating fusion unit on the basis of an existing ECAPA-TDNN model. And carrying out verification and post-processing on the dynamic sub-band ECAPA-TDNN model. According to the method, the dynamic sub-band ECAPA-TDNN model is constructed, so that the features of different frequency bands of the bird sound are effectively extracted.
Owner:GUANGZHOU UNIVERSITY

Three-dimensional sound event positioning and detection method based on audio-guided visual attention

The invention discloses a three-dimensional sound event positioning and detection method based on audio guidance visual attention, and belongs to the field of audio-visual fusion positioning detection, and the method comprises the steps: extracting audio features of audio data through an audio encoder, and outputting the audio features to an audio guidance visual attention module; extracting visual features of the video data through a visual encoder and outputting the visual features to an audio-guided visual attention module; an audio-guided visual attention module guides visual features through audio features to obtain attention visual features, and an output branch serves as a visual sound positioning result to be output; the audio features and the attention visual features are fused through the fusion module, and the obtained multi-modal features are input into the context network to obtain a sound event detection result and a sound source coordinate estimation result which are output through the two output branches. The method effectively improves the positioning and detection performance of sound events in a three-dimensional space through an audio-guided visual attention mechanism and audio-visual fusion in combination with a sound source coordinate estimation branch.
Owner:UNIV OF SCI & TECH OF CHINA

Sound event detection method based on recursive gated convolution and self-attention mechanism

The invention discloses a sound event detection method based on recursive gating convolution and a self-attention mechanism, and the method comprises the steps: collecting a to-be-detected audio signal, and constructing a sound event detection model; inputting the audio signal into a sound event detection model, and extracting a time-frequency feature through a preprocessing module to obtain a logarithmic Mel spectrum feature; inputting the logarithmic Mel spectrum features into a convolution module, and performing spatial feature fusion through recursive gating convolution to obtain spatial fusion features; inputting the spatial fusion features into a time domain modeling module, and performing context information interaction through a self-attention mechanism to obtain global information features; and inputting the global information features into a KANLinear classifier, and carrying out frame-by-frame classification to obtain each sound event category in the audio and the occurrence time period thereof. According to the method, the accuracy of sound event detection can be remarkably improved on the premise that the number of parameters is small and the number of floating point operation times is low.
Owner:ANHUI UNIV

Systems and methods for predicting aggression in surveillance camera video recordings

Methods and systems for predicting aggressive behavior associated with a surveillance scene. Images are generated by a camera of a surveillance scene, and audio is also generated. A local computing system executes an object classification model on the images to predict one or more classes of objects in the scene. The local computing system also executes a sound event detection model on the audio recordings to predict one or more classes of events occurring in the scene. Metadata associated with the image-based classes and the audio-based classes is generated. The metadata is transmitted to a remote computing system, which executes a knowledge graph on the metadata to implement a knowledge graph-based assessment to predict aggressive behavior occurring in the surveillance scene based on the metadata.The metadata associated with the predicted aggressive behavior is labeled as such and control commands are issued accordingly.
Owner:IQSIGHT BV

Method and system for detecting whale overlapped sound events based on adaptive multi-scale synthetic attention

The invention discloses a whale overlapped sound event detection method and system based on adaptive multi-scale synthetic attention, and belongs to the technical field of ocean engineering and ocean signals, and the method comprises the steps: collecting historical whale sound signals, and carrying out the preprocessing and ACT transformation of the historical whale sound signals, and obtaining a transformed data set; constructing a time-frequency perception cross-offset feature extraction network, and training the time-frequency perception cross-offset feature extraction network by using the transformed data to obtain an adaptive multi-scale synthesis attention network; constructing an adaptive window size prediction network, and predicting window parameters required for constructing a multi-scale attention fusion network according to a feature map output by the adaptive multi-scale synthesis attention network; constructing a multi-scale attention fusion network based on the window parameters, and carrying out iteration to obtain a whale overlapped sound event detection model; and acquiring a real-time whale sound signal, and detecting the real-time whale sound signal by using the detection model to obtain a detection result.
Owner:ANHUI UNIV

Sound event detection method and device and computer readable storage medium

The invention relates to the technical field of sound detection, and discloses a sound event detection method and device and a computer readable storage medium. The method comprises the steps that spectrum features of audio to be detected are acquired, and the dimensions of the spectrum features comprise the number of frequency points and the number of continuous frames; inputting the spectrum feature into an audio feature extraction model to obtain a general audio feature which is a feature vector used for indicating the audio information; inputting the general audio features into a multi-sound event classification model to obtain a first prediction probability of each sound event in preset multi-class sound events; and judging the existence condition of each sound event in the to-be-detected audio according to the first prediction probability. According to the invention, the number of types of sound event detection is effectively expanded while the bearable computing power requirement of the end side equipment is maintained.
Owner:SHENZHEN BAICHUAN SECURITY TECH CO LTD

A smart audio non-invasive monitoring method and system for a home environment

The application discloses a smart audio non-inductive monitoring method and system for a home environment, and belongs to the technical field of smart home and health monitoring, which comprises the following steps: collecting audio data in a room and extracting features to obtain an input feature tensor; based on the input feature tensor, sound event detection and sound source positioning are performed through a SELD model based on a deep neural network to obtain local observation data, which is converted to a global coordinate system; the position of a target object in the global coordinate system is continuously tracked and state estimation is performed based on a particle filter to obtain a global motion trajectory; based on the local observation data and the global motion trajectory, combined with trigger information of infrared and radar sensors, decision-level fusion is performed by using Dempster-Shafer evidence theory to output a final activity state and an environmental abnormal event, thereby solving the problems that in the current monitoring scheme, normal life of a user is disturbed, privacy and self-esteem are highly invasive, and monitoring data is not comprehensive.
Owner:NORTH CHINA UNIVERSITY OF TECHNOLOGY

User profile generation methods, devices, electronic devices, and computer-readable media

This application discloses a user profile generation method, apparatus, electronic device, and computer-readable medium. An embodiment of the method includes: segmenting a user's call audio into an audio segment without human voice and an audio segment containing human voice; extracting a first audio feature from the audio segment without human voice, inputting the first audio feature into a pre-trained first ambient sound event detection model to obtain a first detection result; extracting a second audio feature from the audio segment containing human voice, inputting the second audio feature into a pre-trained second ambient sound event detection model to obtain a second detection result; and generating a user profile based on the first and second detection results. This implementation enriches the methods for generating user profiles. User profiles generated in this way can provide users with services related to their environment, thereby improving service quality.
Owner:CHINA TELECOM CORP LTD