Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

84 results about "Sound source separation" patented technology

Sound acquisition and processing system based on cooperation of multiple microphone arrays

The invention discloses a sound acquisition and processing system based on cooperation of multiple microphone arrays. The system comprises a sound acquisition module, a multi-channel signal preprocessing module, a sound signal feature extraction module, an abnormal sound recognition module, a sound source positioning module and an alarm module. A multi-channel mixed data signal is collected through a circularly-arranged multi-microphone array formed by a plurality of microphones, after echo cancellation, wave beam domain noise reduction and multi-sound-source separation, a single-sound-source feature vector is extracted, according to the single-sound-source feature vector, abnormal sound including explosion, screaming or glass breakage is recognized through a BiLSTM and an attention mechanism model, and the abnormal sound is recognized through an attention mechanism model. And the GCC-PHAT and MDS-MUSIC algorithms are combined to position abnormal sound production, and alarm information is generated. According to the invention, accurate identification, positioning and alarm of the abnormal sound can be realized, and the real-time performance, the accuracy and the multi-target processing capability of abnormal sound monitoring in a complex environment can be obviously improved.
Owner:HANGZHOU DIANZI UNIV

Intelligent early warning method and system for preventing external damage of power transmission line

The invention relates to the technical field of intelligent operation and maintenance of a power system, in particular to an intelligent early warning method and system for preventing external damage of a power transmission line. The method comprises the following steps: acquiring acoustic data of a power transmission line through an acoustic sensor; performing voiceprint feature sparse reconstruction according to the transmission line acoustic data to obtain voiceprint feature data; performing sparse causal graph matching on the voiceprint feature data to obtain sound source behavior causal graph data; performing double-domain attention cross recognition according to the sound source behavior causal atlas data to obtain sound source separation recognition data; and performing scene map generation on the sound source separation identification data to obtain sound source scene map data. According to the invention, by constructing the nested causal chain graph and fusing the propagation weight mechanism, the development evolution and multi-stage trigger relationship of the external damage event of the power transmission line can be accurately captured. Compared with a traditional single-point detection method, the method has higher chain reasoning ability and early risk identification ability, and the response foresight and causal interpretability of an early warning system are effectively improved.
Owner:HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER

Audio processing method and device, computer equipment, storage medium and program product

The invention relates to an audio processing method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: acquiring an original audio which is acquired by a microphone array and comprises a multi-channel audio; according to the spectrum feature of each channel audio included in the original audio and the physical spacing of each microphone in the microphone array, determining direction features corresponding to the original audio in a plurality of preset virtual sound source directions; inputting the frequency spectrum features of the channel audio and the direction features corresponding to the directions of the pseudo sound sources into a pre-trained sound source separation model; the pre-trained sound source separation model is used for carrying out filtering operation on the spectrum features of the channel audios according to the direction features corresponding to the virtual sound source directions to obtain filtered spectrum features of the channel audios, and the filtered spectrum features of the channel audios are utilized to generate and output a target spectrum. By adopting the method, the accuracy of sound enhancement and separation in a complex acoustic scene can be improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Multi-voice separation method based on lightweight dual-path Transform network

The invention discloses a multi-voice separation method based on a lightweight dual-path Transform network, and the method comprises the steps: collecting audio multi-voice data, and carrying out the preprocessing of the data, and forming a data set; the method comprises the following steps: constructing a dual-path Transform network model DPTNet, and introducing a recurrent neural network to optimize the dual-path Transform network model DPTNet; and training the dual-path network model DPTNet, and performing engineering deployment based on the trained model. The method is beneficial to obtaining higher-quality audio fingerprint recognition capability, sound source separation capability and voice enhancement function, can be used for tracking and positioning the position of a sound source, helps positioning and tracking related applications, can be expanded to the medical field, can be used for heart sound segmentation, namely, recognition of specific signals of the heart, and can be applied to the field of medical science. The method helps to diagnose cardiovascular and other medical problems, and has technical innovation and practical application value.
Owner:NANTONG UNIV

Sound source separation method and device

The embodiment of the invention provides a sound source separation method and device, and relates to the technical field of data processing. The method comprises the following steps: converting a to-be-separated audio signal from a time domain signal into a time-frequency domain signal; frequency band segmentation is carried out on the time-frequency domain signal, so that the time-frequency domain signal is segmented into a plurality of sub-band signals, and frequency bands of the plurality of sub-band signals are not overlapped; respectively acquiring spectrum characteristics of the plurality of sub-band signals; acquiring a frequency spectrum mask of at least one sound source of the audio signal to be separated according to the frequency spectrum characteristics of the plurality of sub-band signals; and acquiring an audio signal of the at least one sound source according to the spectrum mask of the at least one sound source and the time-frequency domain signal. The embodiment of the invention is used for improving the robustness of a sound source separation algorithm.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Digital conference voice processing method, system and device and storage medium

The invention relates to a digital conference voice processing method, system and device and a storage medium, and the method comprises the following steps: carrying out the pickup of a conference voice, obtaining a mixed voice signal, carrying out the framing sampling, and forming a voice sampling sequence; carrying out sound source direction estimation based on the sequence to obtain multi-sound-source position information, and carrying out beam forming and spatial filtering on the voice sampling sequence according to the multi-sound-source position information to obtain a sound source separation signal; performing voice segment segmentation on the signal to obtain a voice segment sequence, extracting voiceprint features of each segment, and generating a speaker feature mark; and performing time sequence recombination on the voice fragment sequence by using the mark, constructing a speaking time sequence table, selectively outputting the voice fragment sequence according to the table, and finally generating clear and ordered conference voice. The technical problems that due to the fact that a traditional voice processing method lacks effective space-voiceprint joint constraint, voice separation is not thorough, identities of speakers are confused, and the speaking time sequence is disordered are solved.
Owner:SHENZHEN YUXUN IOT CO LTD

Generative audio anonymization reconstruction method and device based on sound source separation and semantic preservation, equipment and program product

The invention discloses a generative audio anonymization reconstruction method, device, equipment and program product based on sound source separation and semantic preservation, and relates to the technical field of voice privacy protection and audio signal processing. The method comprises the following steps: acquiring an original audio, and carrying out sound source separation on the original audio to obtain at least one speaker sound track and an environment background sound track; performing authentication on each speaker sound track, and determining an authentication result of each speaker sound track; wherein the authentication result of the speaker sound track comprises an authorized person sound track and an unauthorized person sound track; carrying out anonymization processing on the unauthorized human voice track to obtain an anonymized human voice track; and carrying out re-synthesis on the anonymized human voice track and the environment background audio track to obtain an anonymized scene audio track. According to the technical scheme provided by the embodiment of the invention, an audio stream breakage phenomenon can be avoided, the scene continuity of the audio is improved, and the intelligibility and the overall quality of the audio are further improved.
Owner:SHENZHEN JIAYZ PHOTO IND LTD

A recording processing method and related apparatus

The application provides a recording processing method and related devices. The method can include: an electronic device can perform sound source positioning based on the sound collected by a microphone, obtain the position of a target sound source and the number of sound sources in the recording environment, and then perform sound source separation on the sound collected by the microphone according to the position of the target sound source and the number of sound sources in the recording environment to obtain the sound corresponding to the target sound source, i.e., a target audio signal. The electronic device can also determine the signal-to-noise ratio and display the current sound pickup quality to the user. This method can monitor and display the sound pickup quality to the user in real time, so that the user can adjust in time when the sound pickup quality is poor, thereby obtaining high-quality audio and improving the user experience.
Owner:BEIJING HONOR DEVICE CO LTD

System and method for automating design of sound source separation deep learning model

Disclosed are a system and method for automating the design of a sound source separation deep learning model. A method of automating a design of a sound source separation deep learning model, which is performed by a design automation system, may include automatically searching for a combination of hyper parameters of a separation model constructed in a sound source separation deep learning model by using a neural architecture search (NAS) algorithm and reconstructing the sound source separation deep learning model based on the retrieved combination of the hyper parameters of the separation model.
Owner:INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY

Sound file generation device, sound file reproduction device, sound file generation method, and sound file reproduction method

A sound data acquisition unit 110 acquires recording data obtained by recording environmental sound. A sound source separation unit 112 separates the recording data into environmental sound data and sound effect data. A first generation unit 116 generates an environmental sound file including the separated environmental sound data. A second generation unit 118 generates a sound effect file including the separated sound effect data.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Speech enhancement network training, speech enhancement method, apparatus and electronic device

This disclosure relates to a speech enhancement network training, speech enhancement method, apparatus, and electronic device. The method includes acquiring sample speech information, a clean speech signal, an original impulse response signal, a first device frequency response, and target sound source distribution data corresponding to target sound source characteristics; performing sound source separation processing on the original impulse response signal to obtain an original direct source response signal and an original reflected source response signal; performing sound source characteristic alignment processing based on the target sound source distribution data, the original direct source response signal, and the original reflected source response signal to obtain a target impulse response signal; generating target enhanced speech information corresponding to the sample speech information based on the target impulse response signal, the clean speech signal, and the first device frequency response; and training a preset neural network for speech enhancement based on the sample speech information and the target enhanced speech information to obtain a target speech enhancement network corresponding to the target sound source characteristics. Utilizing embodiments of this disclosure can improve speech enhancement effects.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

A lithium ion battery thermal runaway acoustic early warning method based on feature reconstruction, medium and system

This invention provides a method, medium, and system for acoustic early warning of thermal runaway in lithium-ion batteries based on feature reconstruction, belonging to the field of lithium-ion battery technology. The invention constructs a positive sample set by collecting safety valve opening sounds through multi-condition thermal runaway experiments, and expands the samples using data augmentation. Multi-resolution Mel spectra are extracted from the audio signals. These Mel spectra are then input into a complex-domain phase-aware separation model for joint estimation of complex-domain amplitude masking and phase residuals. Physical prior corrections are applied to the reconstruction results using a sound source separation algorithm based on wave equation time-frequency inverse scattering and a low-rank sparse time-frequency matrix decomposition algorithm based on random matrix theory. The three corrected signals are weighted and fused to obtain a corrected Mel spectra, which are finally input into a temporal convolutional network to classify and identify the safety valve opening sounds and output a thermal runaway early warning signal. This invention solves the technical problem of insufficient accuracy in thermal runaway early warning caused by acoustic feature reconstruction distortion in complex noise environments.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Audio processing method, training method of sound source separation model, and electronic device

The application provides an audio processing method, a training method of a sound source separation model and an electronic device. The audio processing method comprises: fully extracting audio features of to-be-processed audio data by using a two-dimensional convolution-based feature extraction network, a frequency band interaction layer and a frequency point interaction layer of a sound source separation model, and then performing sound source separation based on the fully extracted audio features. The scheme adds the frequency band interaction layer and the frequency point interaction layer on the basis of the two-dimensional convolution-based feature extraction network, extracts frequency band interaction features between multiple sub-frequency bands and global frequency point interaction features in each sub-frequency band, so that the advantages of small calculation amount of the two-dimensional convolution-based feature extraction network are retained, and the defects of insufficient extraction capability of the two-dimensional convolution-based feature extraction network for global interaction features are made up, so that the entire sound source separation process can achieve the effect of relatively high separation precision and relatively short processing delay.
Owner:HONOR DEVICE CO LTD

Audio control method and system, storage medium, electronic equipment and vehicle

The invention discloses an audio control method and system, a storage medium, electronic equipment and a vehicle. The audio control method comprises the step of controlling an audio playing effect of at least one sound source in fused audio according to a command action of a user. According to the invention, the audio data of the plurality of sound sources can be separated from the fused audio, and the audio data of the plurality of sound sources can be played, so that the playing of the fused audio is realized. Through sound source separation, the position of each sound source and the audio playing effect can be flexibly adjusted, so that a user can feel sound from different directions and different sound sources, the sense of space and immersion of audio playing are enhanced, and the immersive listening experience is realized. Besides, according to the embodiment of the invention, the command action of the user is identified, the audio playing effect of at least one sound source is controlled according to the command action, and the command process of the user is fused during audio playing, so that the user can feel that the user becomes a music commander, the interactivity and entertainment are enhanced, and the personalized sound listening experience is realized.
Owner:BYD CO LTD

Voice control screen display method and device, computer equipment and storage medium

The embodiment of the invention provides a voice control screen display method and device, computer equipment and a storage medium, and is used for the technical field of display control. The voice control screen display method comprises the steps of performing audio separation and voiceprint comparison on mixed audio data to determine a target sounding direction, performing sound source tracking collection according to the target sounding direction, and filtering collected audio information to obtain an audio data stream, semantic recognition and sentiment analysis are carried out on the audio data stream to obtain a semantic command set and a demonstration sentiment label, and a screen display adjustment instruction is generated according to the semantic command set and the demonstration sentiment label through an instruction format. According to the technical scheme, through sound source separation and voiceprint comparison, sound source tracking, semantic recognition, sentiment analysis and instruction mapping, screen display can be controlled in real time only by an authorized user in a complex noisy environment, so that the false triggering rate is remarkably reduced, and the command recognition accuracy and the consistency of system interaction experience are improved.
Owner:SHENZHEN DOCTORS OF INTELLIGENCE & TECH CO LTD

Sound source separation system

A sound source separation system is configured to perform blind source separation of individual sound source signals by an independent component analysis (ICA) method from a plurality of mixed signals in which two or more sound source signals are mixed. The sound source separation system includes an n number of microphones, where n≥2; a virtual microphone signal generator configured to generate, from output signals of the n number of the microphones, an m number of the virtual microphone signals that are output signals of the m number of virtual unidirectional microphones having directivities in different directions, where m>n; and an ICA processor configured to separate an L number of the sound source signals that are signals of the L number of different sound sources, by the ICA method, from the m number of the virtual microphone signals that are the plurality of the mixed signals, where L>n and L≤m.
Owner:ALPS ALPINE CO LTD

Data processing method and device, equipment, medium and product

The invention provides a data processing method and device, equipment, a medium and a product, and the method comprises the steps: obtaining multi-modal data collected in a business scene, and the multi-modal data comprises audio data and visual data; the audio data comprises multi-sound-source audio signals obtained by carrying out audio signal collection on N objects in the service scene, the visual data comprises visual signals of the N objects synchronously collected in the process of collecting the multi-sound-source audio signals, and N is an integer greater than 1; based on the audio data and the visual data, sound source separation is carried out on the multi-sound-source audio signals in a multi-modal fusion processing mode, an audio separation result is obtained, and the audio separation result comprises the audio signal of each object; and performing optimization processing on the audio signal of each object in the audio separation result to obtain an optimized audio signal of each object. The audio signals and the visual signals are fused, sound source separation is carried out on the multi-sound-source audio signals, and the accuracy of audio separation is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Rotating sound source separation method based on fast Bayesian inference

The invention discloses a rotating sound source separation method based on fast Bayesian inference, and the method comprises the steps: dividing a plane where a sound source is located into uniform or non-uniform equivalent source grid points, enabling each grid point to represent an equivalent sound source so as to describe the sound field energy contribution of a rotating sound source and a static sound source, and dividing scanning grid points which coincide with the equivalent source grid points; for mixed sound signals collected by a microphone array, scanning grid point output is calculated by adopting a traditional beam forming algorithm and a modal composition beam forming method considering the Doppler effect of a rotating sound source, key features of two kinds of beam forming output are extracted based on a convolution kernel, and preliminary acceleration is achieved by calculating an approximate transfer matrix through convolution.
Owner:ZHEJIANG SHANGFENG SPECIAL BLOWER IND CO LTD

Refinement step for beamforming for acoustic source separation

Aspects of the subject technology relate to systems, methods, and computer readable media for estimating acoustic spectra. Acoustic data can be received at a hydrophone array from a first acoustic source and a second acoustic source in a downhole environment. An initial noise spatial correlation matrix estimation can be generated based on the acoustic data. The initial noise spatial correlation matrix estimation can be applied to a beamformer to generate a first source spectra estimation for the first acoustic source and the second acoustic source. A revised noise spatial correlation matrix estimation can be generated based on the first source spectra estimation. The revised noise spatial correlation matrix estimation can be applied to the beamformer to generate a second source spectra estimation for the first acoustic source and the second acoustic source in the downhole environment based on the first source spectra estimation.
Owner:HALLIBURTON ENERGY SERVICES INC

Audio processing method, device and system based on multi-modal noise reduction and sound source separation

The embodiment of the invention discloses an audio processing method, device and system based on multi-mode noise reduction and sound source separation. The method comprises the following steps: acquiring an original audio signal acquired by an array consisting of at least two microphones; performing beam forming processing on the original audio signal to obtain an initial audio signal; dividing the residual noise in the initial audio signal into steady-state noise and unsteady-state noise, and respectively suppressing the noise by adopting a corresponding noise reduction strategy to obtain a denoised first audio signal; performing sound source separation on the first audio signal to obtain multiple paths of separated second audio signals of different sound sources; automatically adjusting a compression ratio and a compression range according to audio characteristics of the second audio signal to compress the second audio signal to obtain compressed audio data; and transmitting the audio data to the pre-binding device based on a preset transmission protocol. According to the embodiment of the invention, the technical problem of how to improve the audio quality and meet the low-delay transmission of the real-time transcription requirement is solved.
Owner:SHENZHEN JIAYZ PHOTO IND LTD

Sound source separation system

To provide a "sound source separation system" for separating sound source signals of sound sources more than the number of microphones by an ICA method.SOLUTION: Five beam formers 21 of a virtual microphone sound generation part 2 generate five virtual microphone signals which are outputs of five virtual unidirectional microphones having directivities in different directions from outputs of four omnidirectional microphones 11 of a microphone set 1. The ICA processing section 3 receives the five virtual microphone signals input from the virtual microphone sound generation section 2 as inputs from the five different microphones 11, and separates and outputs five sound source signals which are signals of five different sound sources by the ICA method.SELECTED DRAWING: Figure 1
Owner:ALPS ALPINE CO LTD

Zero sample singing sound conversion method and system

The invention provides a zero sample singing sound conversion method and system, and the method comprises the following steps: carrying out the fundamental frequency disturbance and wet sound simulation of an input audio, and generating enhanced training data; obtaining speaker embedding by using a singing enhanced tone and style extractor; and based on enhanced training data and speaker embedding, harmonic waves and noise excitation are generated in combination with a neural source filter, and a target singing waveform is synthesized. The problems of environmental noise, sound source separation residue and unnatural singing style expression in a real scene are solved.
Owner:GIANT MOBILE TECH CO LTD

Audio playback control method and control apparatus

ActiveCN116208899BElectric megaphonesSound source separationEnvironmental noise
The application discloses an audio playing control method, comprising: acquiring first audio data; acquiring an environmental noise decibel value; performing sound source separation on the first audio data to obtain a plurality of second audio data with different frequency structures; adjusting the volume of the plurality of second audio data according to the environmental noise decibel value; performing audio mixing on the second audio data after volume adjustment; and adjusting the playing volume of the second audio data after audio mixing according to the environmental noise decibel value. The audio playing control method provided by the application can realize sound source separation of playing audio data, and adjust the volume of, for example, target sound source data and environmental sound source data according to the environmental noise condition, thereby improving the automatic adjustment capability of the system. The volume of sound of different sound sources is adjusted, so that sound containing substantial content such as a human voice sound source is highlighted in a noisy environment, and all kinds of sound sources are balanced in a quiet environment, giving people a sense of stability and improving the auditory experience.
Owner:SHENZHEN HONGHE INNOVATION INFORMATION TECH CO LTD

Information processing device and information processing method

An information processing device, wherein the information processing device includes circuitry configured to: obtain video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and wherein the circuitry is further configured to add, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.
Owner:SONY GROUP CORP +1

Bluetooth earphone supporting AI voice intelligence

The invention discloses a Bluetooth earphone supporting AI voice intelligence. The Bluetooth earphone comprises a sound source separation and noise reduction module, an edge model compression module, a multi-mode emotion interaction module, a heat dissipation control module and a predictive maintenance module. Belongs to the technical field of Bluetooth earphones, and particularly relates to a Bluetooth earphone supporting AI voice intelligence. According to the Bluetooth earphone supporting AI voice intelligence, a bone conduction signal and an air conduction signal are fused through a three-dimensional sound field mapping and dynamic mask generation technology, and pure voice can be extracted from a complex sound field; the translation model is compressed based on progressive neural architecture search and mixed precision quantification, and low-delay edge calculation of dialects and foreign languages is achieved; a space-time diagram convolutional network and near-end strategy optimization are combined, translation data, tactile operation frequency data and eyeball trajectory data are coordinated, and a heat dissipation strategy is dynamically regulated and controlled, so that equipment stably runs in the environment.
Owner:SHENZHEN AOWAN TECH CO LTD

Speech recognition method based on multi-sound-source separation and related equipment

The invention is suitable for the technical field of artificial intelligence, and provides a voice recognition method based on multi-sound-source separation, and the method comprises the steps: obtaining voice data in a multi-sound-source environment; performing time-frequency feature extraction processing on the voice data to obtain time-frequency spectrum features; performing feature focusing processing on the time-frequency spectrum feature to obtain a target sound source focusing feature; based on the target sound source focusing feature, performing multi-sound source separation on the voice data, and determining target voice data; and performing voice recognition and correction processing on the target voice data, and outputting a voice recognition result corresponding to the target voice data. The method solves the problem that the existing voice recognition method cannot effectively distinguish the space and spectrum characteristics of different sound sources under the condition of multi-sound-source interference, so that the voice recognition accuracy is low.
Owner:SHENZHEN LUKA DR TECHNOLOGY CO LTD

Audio processing method, apparatus, device, and storage medium

The application discloses an audio processing method, device and equipment and a storage medium, relates to the technical field of audio processing, and discloses an audio processing method, which comprises the following steps: inputting a stereo audio signal into a preset sound source separation model to obtain a plurality of independent audio objects; obtaining original spatial information of each audio object from the stereo audio signal, and obtaining music structure information of the stereo audio signal; generating target spatial parameters of the audio objects in a three-dimensional immersive sound field based on the original spatial information and the music structure information, wherein the target spatial parameters comprise position coordinates and / or movement trajectories; and rendering each audio object based on the target spatial parameters to obtain a multi-channel immersive audio signal. The application can improve the generation quality of the multi-channel immersive audio signal.
Owner:WEIFANG GOERDYNA TECH CO LTD

A Smart Early Warning Method and System for Protecting Transmission Lines from External Damage

The present application relates to the technical field of power system intelligent operation and maintenance, and particularly relates to a power transmission line external damage prevention intelligent early warning method and system. The method comprises the following steps: collecting power transmission line acoustic data through an acoustic sensor; performing acoustic fingerprint feature sparse reconstruction according to the power transmission line acoustic data to obtain acoustic fingerprint feature data; performing sparse causal graph matching on the acoustic fingerprint feature data to obtain acoustic source behavior causal graph atlas data; performing double-domain attention cross identification according to the acoustic source behavior causal graph atlas data to obtain acoustic source separation identification data; and generating a scene atlas from the acoustic source separation identification data to obtain acoustic source scene atlas data. By constructing a nested causal chain graph and fusing a propagation weight mechanism, the present application can accurately capture the development and evolution of power transmission line external damage events and multi-level triggering relationships. Compared with traditional single-point detection methods, the present application has stronger chain reasoning ability and early risk identification ability, effectively improving the response foresight and causal explainability of the early warning system.
Owner:HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER

Industrial noise source positioning and separating method, device and equipment and storage medium

The invention relates to an industrial noise source positioning and separating method and device, equipment and a storage medium, and the method comprises the steps: collecting a multi-channel audio signal, and carrying out the preprocessing and local feature extraction of the multi-channel audio signal, and obtaining a local feature vector packet; carrying out sound source localization based on the local feature vector packet by adopting a hybrid localization strategy to generate space position information of a sound source, and calculating localization confidence; sound source separation is carried out on the multi-channel audio signals, and separated audio signals and sound source separation quality confidence coefficients of all independent sound sources are obtained; collaborative management is carried out on the output of each agent through the arbitration agent, and confidence fusion and conflict judgment are carried out; and when the confidence coefficient is lower than a first preset threshold value, triggering an artificial verification process or scheduling a mobile detection unit to carry out auxiliary measurement so as to determine target sound source positioning and separation data, and storing the target sound source positioning and separation data in a vector database. The method has the advantages of high-precision positioning and reliable separation of the industrial noise source, and can adapt to the change of a complex acoustic environment.
Owner:E SURFING IOT CO LTD

Sound source separation device

A sound source separation device according to an embodiment of the present disclosure includes a plurality of microphones, a matrix unit, and an output unit. The plurality of microphones may receive a plurality of microphone input signals transmitted from a plurality of sound sources. The matrix unit generates an objective function according to an estimated source vector and an estimated noise vector estimated based on the plurality of microphone input signals, and replace a first term and a second term included in the objective function using a log-likelihood function to estimate a demixing matrix. The output unit provides output vectors calculated based on the microphone input signals and the demixing matrix.
Owner:SOGANG UNIV RES & BUSINESS DEV FOUND