Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1042 results about "Microphone array" patented technology

A microphone array is any number of microphones operating in tandem. Typically, an array is made up of omnidirectional microphones, directional microphones, or a mix of omnidirectional and directional microphones distributed about the perimeter of a space, linked to a computer that records and interprets the results into a coherent form. Arrays may also be formed using numbers of very closely spaced microphones. Given a fixed physical relationship in space between the different individual microphone transducer array elements, simultaneous DSP (digital signal processor) processing of the signals from each of the individual microphone array elements can create one or more "virtual" microphones. Different algorithms permit the creation of virtual microphones with extremely complex virtual polar patterns and even the possibility to steer the individual lobes of the virtual microphones patterns so as to home-in-on, or to reject, particular sources of sound. The application of these algorithms can produce varying levels of accuracy when calculating source level and location, and as such, care should be taken when deciding how the individual lobes of the virtual microphones are derived.

Microphone array sound source localization method and system based on cross-correlation-beam forming closed-loop optimization

The invention relates to a microphone array sound source positioning method and system based on cross-correlation-beam forming closed-loop optimization, and belongs to the technical field of sound source positioning. The method comprises the following steps: collecting multichannel sound signals through a microphone array and preprocessing the multichannel sound signals to extract time-frequency features and suppress noise interference; time delay information among the microphones is estimated by adopting a generalized cross-correlation phase transformation algorithm, and an optimization strategy is introduced to improve estimation stability and anti-interference performance; enhancing the target sound source signal in combination with a minimum variance undistorted response beam forming algorithm and an adaptive Kalman filtering mechanism; constructing a closed-loop feedback optimization mechanism based on the beam output signal to realize feedback adjustment; and adopting a hybrid network architecture, taking the beam output signal amplitude spectrum as input, and outputting the frequency spectrum or mask of the obtained target sound source signal. The method has the advantages of high calculation efficiency, high positioning precision and strong anti-interference capability, and is suitable for real-time acoustic signal processing in a complex environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Sound acquisition and processing system based on cooperation of multiple microphone arrays

The invention discloses a sound acquisition and processing system based on cooperation of multiple microphone arrays. The system comprises a sound acquisition module, a multi-channel signal preprocessing module, a sound signal feature extraction module, an abnormal sound recognition module, a sound source positioning module and an alarm module. A multi-channel mixed data signal is collected through a circularly-arranged multi-microphone array formed by a plurality of microphones, after echo cancellation, wave beam domain noise reduction and multi-sound-source separation, a single-sound-source feature vector is extracted, according to the single-sound-source feature vector, abnormal sound including explosion, screaming or glass breakage is recognized through a BiLSTM and an attention mechanism model, and the abnormal sound is recognized through an attention mechanism model. And the GCC-PHAT and MDS-MUSIC algorithms are combined to position abnormal sound production, and alarm information is generated. According to the invention, accurate identification, positioning and alarm of the abnormal sound can be realized, and the real-time performance, the accuracy and the multi-target processing capability of abnormal sound monitoring in a complex environment can be obviously improved.
Owner:HANGZHOU DIANZI UNIV

Voice recognition processing method, system and equipment based on conference scene and medium

The invention relates to a voice recognition processing method, system and device based on a conference scene and a medium, and belongs to the technical field of voice processing. The voice recognition processing method comprises the following steps: acquiring an original conference audio stream collected by a microphone array; performing signal preprocessing on the original conference audio stream collected by the main channel, and outputting a pure voice signal; generating a sound source orientation thermodynamic diagram based on the original conference audio stream; extracting multi-dimensional voiceprint feature vectors from the pure voice signals, performing dynamic grouping, outputting a voice fragment set marked with voiceprint IDs, and generating an initial transcription text; dynamically correcting the initial transliteration text, and outputting a transliteration text stream with an industry term tag; and performing periodic memory enhancement processing on the transliteration text stream, outputting and analyzing a long text, and generating structured conference summary data. According to the invention, the automation level and accuracy of conference voice processing can be improved.
Owner:CHINA TRANSPORT INFORMATION TECH GRP CO LTD

AI scene conversion method based on visitor behavior

The invention discloses an AI scene conversion method and system based on visitor behaviors. According to the method, visitor behavior data are collected in real time through a multi-mode sensor network, and a UWB positioning base station, a millimeter wave radar module, a high-definition camera, a microphone array and a pressure sensor matrix are included. The system adopts an edge computing architecture, is equipped with an NVIDIAJetsonAGXXavier processor and a 512-core VoltaGPU (Graphics Processing Unit), and operates three core AI models to realize intelligent analysis. The LSTM residence time prediction model adopts a three-layer bidirectional architecture, each layer is configured with 256 hidden units, and visitor behavior sequences within 30 minutes are analyzed to predict the residence time. The Transform interaction intention recognition model comprises six layers of encoders and eight multi-head attention modules, and visual, voice and behavior three-mode characteristics are fused to judge visitor interaction requirements. According to the technical scheme, self-adaptive adjustment of the exhibition hall environment is achieved, the visitor satisfaction degree is improved by 28%, the operation cost is reduced by 15%, and important technical support is provided for intelligent transformation of modern exhibition places.
Owner:元羽兽数字科技(上海)有限公司

Gas leakage identification method and system based on sound positioning, medium and equipment

The invention relates to the field of gas leakage identification, and discloses a gas leakage identification method and system based on sound localization, a medium and equipment, and the method comprises the steps: collecting a leakage sound wave signal through a microphone array, and extracting acoustic features in a complex background; simulating a diffusion process of gas in a turbulence environment after gas leakage through an established leakage gas convection-diffusion model, and combining gas concentration distribution in the simulated diffusion process with acoustic characteristics to predict sound source localization at a leakage position; modeling a sound source localization search process as POMDP, and iteratively updating a confidence state of a sound source position through a particle filter; spatial distribution features are extracted from the confidence state through DBSCAN clustering to serve as input of the LSTM-DQN network, spatial and temporal features are fused through the constructed LSTM-DQN network, cross-scene generalization is achieved through transfer learning, and sound source coordinates are dynamically optimized and output through the Q network. According to the invention, the positioning accuracy and real-time performance in a complex environment are improved.
Owner:SUZHOU SHENGTENG ROBOT CO LTD +1

Simultaneous interpretation data processing method and system based on POE microphone array

The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation data processing method and system based on a POE microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and meeting place environment noise spectrum features through a distributed microphone array powered by the Ethernet; after time domain framing is carried out on the audio stream, adaptive filtering is carried out by using a dynamic noise reduction weight coefficient to obtain a primary pure voice segment; dividing the multi-language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-language speech endpoint detection model, and matching a corresponding acoustic model to generate a phoneme-level time alignment sequence; comparing and outputting a term replacement instruction stream in real time in combination with a simultaneous transfer term library, and generating an intermediate semantic representation vector after fusion; and the low-delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual meeting place scene.
Owner:SUZHOU FUCHUAN TECH

Sound source positioning and detecting method and device

The invention belongs to the technical field of sound source processing, and provides a sound source positioning and detecting method and device. The method comprises the following steps: acquiring delay estimation of microphones I and II and delay estimation of microphones I and III based on a three-path linear uniform microphone array; based on the delay estimation sum, respectively carrying out positioning calculation under near-field and far-field conditions; according to the sound source distance under the near-field condition, comparing the sound source distance with a distance judgment threshold value, and determining a far-field / near-field output sound source position; and carrying out feature extraction on the signals of any microphone array, and carrying out event classification based on a pre-trained convolutional neural network. According to the method, far / near field model selection is carried out according to the distance judgment threshold, large deviation generated by a single model in a critical region is avoided, continuous and stable positioning from short distance to long distance is achieved, time classification can be achieved while position calculation is carried out, and integrated output is achieved.
Owner:YANGZHOU YUAN ELECTRONICS TECH CO LTD

Audio data processing method and device and electronic equipment

The invention discloses an audio data processing method and device and electronic equipment, and relates to the field of audio processing. The method comprises the following steps: acquiring original multichannel audio data through a microphone array; generating a main beam output signal pointing to the target sound source direction and at least one sidelobe suppression output signal; extracting voiceprint feature vectors; comparing the voiceprint feature vector of the current frame with the center vector of the existing voiceprint class cluster to obtain a voiceprint clustering result; if it is determined that the first voiceprint class cluster and the second voiceprint class cluster exist, extracting a first audio frame belonging to the first voiceprint class cluster and a second audio frame belonging to the second voiceprint class cluster, and performing audio reconstruction to obtain a first voice stream and a second voice stream; and carrying out residual noise elimination on the sidelobe suppression output signal to obtain a first isolated audio stream and a second isolated audio stream. By implementing the technical scheme provided by the invention, the audio processing quality is improved.
Owner:SHENZHEN SOUNDFIT TECH CO LTD

Power transmission line external damage risk identification method fusing image and sound features

The invention provides a power transmission line external damage risk identification method fusing image and sound features, and belongs to the technical field of power supply. Image frames collected by a camera are processed based on a target detection algorithm, when external damage machinery is identified, a target bounding box is output, and the timestamp of the current image frame is recorded; the microphone array is triggered to carry out audio acquisition, and an acoustic reference signal pointing to the external breaking mechanical direction is generated; visual features and Mel frequency cepstrum coefficient MFCC features are extracted, and acousto-optic time delay is calculated; and determining the vertical distance from the external breaking machine to the power transmission line based on the sound velocity, the acousto-optic time delay and the installation parameters of the camera, and carrying out early warning according to the vertical distance and the voltage grade of the power transmission line. Therefore, through a collaborative fusion mechanism of visual sense and acoustic perception, the problems of high false alarm rate, poor positioning precision and large response delay in external damage monitoring of the power transmission line in the prior art are effectively solved, and the monitoring accuracy and practicability are improved.
Owner:HENAN TIEYUE DIGITAL TECHNOLOGY CO LTD

AI multi-mode voice interaction method based on vehicle-mounted intelligent terminal and electronic equipment

The invention provides an AI multi-mode voice interaction method based on a vehicle-mounted intelligent terminal and electronic equipment, and the method comprises the steps: synchronously collecting voice information, facial micro-expressions and hand touch actions of a user through a microphone array, an in-vehicle camera and a steering wheel touch sensor on the vehicle-mounted intelligent terminal, and forming multi-mode interaction data; fusing the multi-modal interaction data, and judging a current driving scene through a pre-trained driving scene recognition model; dynamically adjusting voice interaction parameters according to the driving scene judgment result; recognizing a voice instruction of the user based on the adjusted voice interaction parameter; and in combination with instruction preferences in historical interaction data of the user, intention completion is performed on the recognized voice instruction, and a personalized execution instruction is generated. According to the invention, the recognition accuracy and interaction efficiency of the voice instruction are improved.
Owner:SHENZHEN BEIBO INTELLIGENT TECH

Voice data recognition method and system based on AI voice algorithm

The invention discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem that the voice data recognition capability is low. The method comprises the following steps: S1, multi-mode cooperative triggering collection: synchronously collecting lip electromyographic signals and voiceprint features through a multi-mode sensor, an activation instruction is generated through feature fusion, and voice acquisition starting is triggered; s2, AI adaptive noise reduction processing: carrying out noise separation on the original audio signal by adopting a generative adversarial network, separating environmental noise features to generate a dynamic noise reduction mask, and keeping the integrity of human voice features; s3, beam dynamic optimization adjustment: analyzing real-time audio quality based on a reinforcement learning algorithm, dynamically adjusting beam pointing and gain parameters of a microphone array, and focusing a target sound source; and S4, semantic association cache enhancement: carrying out real-time semantic analysis on the collected voice data. According to the invention, the voice data recognition capability of an AI voice algorithm is greatly improved.
Owner:HUAQIAO UNIVERSITY

Industrial equipment operation state monitoring and abnormity early warning method and system based on acoustic characteristics

The invention discloses an industrial equipment operation state monitoring and abnormity early warning method and system based on acoustic characteristics. The system comprises a multi-channel acoustic data acquisition module, an acoustic data preprocessing and enhancing module, a robustness acoustic feature extraction module, a deep learning abnormal mode recognition module, an abnormal early warning and auxiliary positioning module, an intelligent diagnosis and early warning enhancing module and a data storage and management module. The method comprises the following steps: collecting multichannel acoustic data, and extracting time domain, frequency domain and time-frequency domain robustness features after noise reduction, screening and framing preprocessing; identifying equipment abnormity through a deep learning model; a microphone array is combined to analyze the sound source direction and position abnormal equipment, and early warning information is generated; and integrating SCADA / DCS process data, and generating graded early warning and operation and maintenance suggestions. The method can adapt to a complex noise environment, accurately recognize early abnormity, provide intelligent diagnosis, support predictive maintenance and reduce fault risks.
Owner:YANGTZE ECOLOGY & ENVIRONMENT CO LTD

Speech enhancement and high-precision recognition method and system in complex environment

PendingCN121641016ASpeech recognitionSpectral density estimationNerve network
The invention provides a voice enhancement and high-precision recognition method and system in a complex environment, and relates to the technical field of voice processing, and the method comprises the steps: collecting a time domain signal in an off-road parking sentry box environment for preprocessing, detecting a mute segment signal in a standard time domain signal for noise power spectral density estimation, and obtaining a noise power spectral density value; a reverberation parameter is obtained by combining voice onset information and noise spatial correlation estimation, prediction is performed by using a deep neural network model, voice masking is applied to microphone array signals to perform enhancement processing, adaptive feature extraction is performed on time domain enhanced voice signals, and a voice signal is obtained. And performing high-precision recognition on the voice adaptive feature sequence based on an acoustic model and a language model, and outputting a target recognition text. The technical problems of poor voice signal quality and low recognition accuracy in a complex noise environment in the prior art are solved. The technical effects of improving the voice signal quality and the recognition accuracy and realizing clear, accurate and real-time voice interaction are achieved.
Owner:INTELLIGENT INTER CONNECTION TECH CO LTD

Bridge expansion joint acoustic monitoring method based on beam forming microphone array

The invention discloses a bridge expansion joint acoustic monitoring method based on a beam forming microphone array, and the method comprises the steps: synchronously collecting acoustic signals through multiple channels, and achieving the high-precision signal alignment through clock offset compensation; the noise analysis module is started by combining a pre-triggering signal of a vehicle passing through the detection unit, the optimal microphone distance is dynamically determined according to sound field coherence and a signal-to-noise ratio maximization model, and a stepping motor drives a movable microphone to achieve real-time adjustment. On the basis, a spatial filtering weight is generated by adopting a self-adaptive beam forming algorithm, and directional enhancement and interference suppression of the expansion joint sound source are realized. Through collaborative optimization of a hardware structure and a signal processing method, the sensitivity and accuracy of acoustic monitoring of the bridge expansion joint are effectively improved, the abnormal sound source recognition capability can be remarkably enhanced, the environmental noise interference is reduced, and the method is suitable for intelligent monitoring and diagnosis of the long-term service state of the bridge expansion joint.
Owner:SOUTHEAST UNIV

Wake-Word Processing in an Electronic Device

Wake-word processing by a wearable electronic device could be carried out when the device is worn by a user and is in a device sleep state, the device including a linear microphone array having at least two microphones vertically spaced from each other, and the device also including a processor. And the example method could involve (i) the at least two microphones of the linear microphone array receiving an audio waveform representing a wake-word utterance, (ii) the processor making a determination, based at least on an angle of arrival of the audio waveform at the at least two microphones of the linear microphone array and / or an energy level of the audio waveform received at the at least two microphones of the linear array, of whether to accept the wake-word utterance or rather to reject the wake-word utterance, and (iii) the processor controlling operation of the device based on the determination.
Owner:STRYKER CORP

Multi-modal interaction control method and system of multifunctional teaching assisting robot and robot

The invention relates to the technical field of intelligent education and Internet of Things fusion, in particular to a multi-modal interaction control method and system of a multifunctional teaching-assistant robot and the robot. According to the system, a tablet AI processor runs an Android system as a core, a display module, a man-machine interaction module, an audio input module, an audio output module, an image acquisition module, a temperature and humidity sensor module, a WIFI or Bluetooth module, an Ethernet module and an Internet of Things module are connected, and the Internet of Things module supports RS485 wired access and Zigbee wireless access at the same time; the tablet AI processor executes voice interaction, video call, face recognition, environment monitoring, network interaction and equipment linkage, and provides a face recognition alignment and depth feature matching algorithm and an annular microphone array beam forming and sound source direction estimation algorithm, so that multi-modal interaction and multi-equipment linkage in a teaching scene are realized. The convenience and the safety are improved; and the equipment access cost is reduced.
Owner:SHENZHEN YUXIN DIGITAL TECH CO LTD

Method for wall direction estimation

A system that performs wall direction estimation to determine a position of an acoustically reflective surface relative to a device. For example, the device may detect a direct sound received from an active sound source and a first reflection reflected from a nearby wall. Based on a time delay, the device can estimate a direction and / or distance to the wall. For example, the device determines a time differential of arrival (TDOA) between the direct sound and the first reflection for each microphone in a microphone array. The device combines TDOA values of the entire microphone array into a distance vector, which the device can compare to reference distance data to determine the best wall direction estimate. For example, reference distance values are fixed based on a geometry of the microphone array and the device can identify an azimuth value that is most similar to the distance vector.
Owner:AMAZON TECH INC

Conference sound amplification system based on AI intelligent algorithm and 360-degree omnidirectional noise reduction

The invention discloses a conference sound reinforcement system based on an AI intelligent algorithm and 360-degree omni-directional noise reduction, and relates to the technical field of audio signal processing, the conference sound reinforcement system comprises a conference management center, the conference management center is in communication connection with the following modules: a multi-sound-source sensing module used for capturing 360-degree omni-directional sound field information in a conference environment and constructing a sound source space distribution model; according to the invention, the omnidirectional microphone array unit covers all directions of a conference space, synchronously collects audio data streams, eliminates the limitation of a conventional unidirectional microphone, combines the sound source positioning space mapping unit, constructs a sound source space distribution model based on the time difference of arrival and the phase difference through a deep learning model, and improves the sound source positioning accuracy. The azimuth angle, pitch angle and distance parameters of the sound source are accurately analyzed, the position of the sound source is mapped to a virtual space coordinate system, a dynamically updated 3D sound source distribution diagram is generated, the position change and intensity distribution of the sound source are reflected in real time, and the positioning precision in a complex acoustic environment is remarkably improved.
Owner:JUSHENG (YANGJIANG) TECHNOLOGY CO LTD

Hearing aid with own-voice mitigation

A system for hearing assistance includes an array of microphones, which are configured be mounted in proximity to a head of a user of the system and to output electrical signals in response to first acoustic waves that are incident on the microphones. A speaker is configured for mounting in proximity to an ear of the user and configured to output second acoustic waves in response to a drive signal applied to the speaker. Processing circuitry is configured to generate the drive signal by amplifying and filtering the electrical signals using a beamforming filter that emphasizes first sounds that impinge on the array of microphones within a selected angular range while suppressing second sounds that are spoken by the user.
Owner:NUANCE HEARING LTD

Array decoupling sound source localization method and device based on deep learning, and readable medium

The invention discloses a formation decoupling sound source localization method and device based on deep learning and a readable medium, and the method comprises the steps: constructing and training a sound source localization model, and obtaining a trained sound source localization model; acquiring a first sound source signal received by a first microphone and a second sound source signal received by a second microphone in the microphone array; calculating generalized cross-correlation frequency domain representation between the first sound source signal and the second sound source signal based on the first sound source signal and the second sound source signal; obtaining a frequency domain feature based on the guide vector between the first microphone and the second microphone and the generalized cross-correlation frequency domain representation; determining an input feature based on the frequency domain feature, and inputting the input feature into a trained sound source localization model to obtain a candidate sound source angle and a confidence coefficient corresponding to the candidate sound source angle; and post-processing the candidate sound source angle to obtain a sound source positioning result. The invention solves the problems that the existing sound source localization method based on deep learning is large in calculated parameter quantity and cannot be applied to embedded equipment and the like.
Owner:YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD

Fault diagnosis method, device and equipment of transformer and storage medium

The invention relates to the technical field of electric power operation and maintenance, in particular to a transformer fault diagnosis method, device and equipment and a storage medium, and the method comprises the steps: collecting a transformer acoustic monitoring signal based on an acoustic microphone array, carrying out blind source separation and decoupling, and constructing an acoustic signal feature set; transformer operation condition parameters are extracted, mechanical state degradation evaluation is carried out according to the acoustic signal feature set, and a mechanical state health degree index is generated; and calculating an acoustic signal receiving timestamp of the acoustic microphone array, performing sound source distribution mapping, and constructing a sound source spatial distribution diagram. According to the method, blind source separation decoupling, mechanical state degradation evaluation, sound source distribution mapping and fault situation diagnosis are performed by acquiring the acoustic signals and the operation condition parameters, so that tiny mechanical faults in the transformer can be accurately identified, intelligent maintenance decision is realized, and the method has the advantages that the tiny mechanical faults in the transformer can be detected earlier; the real-time performance and reliability of fault diagnosis are improved, and the accident risk caused by diagnosis lag is reduced.
Owner:STATE GRID HENAN ELECTRIC POWER ELECTRIC POWER SCI RES INST +2

Voiceprint monitoring method and system integrating bird recognition and noise traceability

The invention discloses a voiceprint monitoring method and system integrating bird recognition and noise traceability, and the method comprises the following steps: a position optimization deployment stage: collecting environment parameters of a destination, and carrying out the evaluation of a plurality of candidate positions through a scoring model based on the environment parameters, selecting the candidate position with the highest score as the final deployment position of the microphone array; in the voiceprint data acquisition stage, at the final deployment position, a solar-powered microphone array is used for acquiring mixed sound signals in the environment; a voiceprint analysis processing stage: carrying out parallel processing on the mixed sound signal, the parallel processing process comprising: a bird recognition processing sub-stage: extracting voiceprint features of bird buzzing from the mixed sound signal, and inputting the voiceprint features to a deep learning classification model integrated with a feature enhancement module, and outputting an identification result containing the bird species and the calibrated confidence. According to the invention, high-precision bird monitoring can be realized, and meanwhile, precise noise control can be synchronously carried out based on bird monitoring equipment.
Owner:SANYA BOFAN ECOLOGICAL ENVIRONMENT TECHNOLOGY CO LTD

Electric welding equipment fault diagnosis method and system based on voiceprint features

The invention relates to an electric welding equipment fault diagnosis method and system based on voiceprint features. The method comprises the following steps: acquiring multi-channel original sound signals received by a microphone array deployed on electric welding equipment; carrying out collaborative filtering processing on the multichannel original sound signals by adopting a weight vector adaptive adjustment collaborative filtering method so as to obtain filtered multichannel sound signals; extracting feature vectors corresponding to different time scales and different frequency sub-bands in the filtered multichannel sound signals, and performing feature fusion processing to obtain multi-scale fusion features; performing feature dimension reduction processing on the multi-scale fusion features based on an attention mechanism to obtain global voiceprint features; performing feature extraction on the multichannel original sound signals based on an extended feature extraction model obtained by training of a BYOL model training framework to obtain extended voiceprint features; and obtaining equipment fault information based on the global voiceprint features and the extended voiceprint features. The anti-interference capability of voiceprint features can be improved, and a foundation is laid for fault diagnosis of electric welding equipment.
Owner:HANGZHOU PINGZHI TECH CO LTD +1

Emotional diary intelligent glasses system driven by multi-mode intelligent body

The invention discloses an emotional diary intelligent glasses system driven by a multi-modal agent. The emotional diary intelligent glasses system comprises a multi-modal perception layer, a unified vector space processing layer and an agent collaborative decision-making layer. The multi-modal sensing layer constructs a full-dimensional sensing network by integrating a camera, a microphone array and an eye movement tracker, and collects image, sound and eye movement information; the unified vector space processing layer realizes cross-modal information unified representation by using a multi-modal large model; the agent collaborative decision-making layer comprises an emotion perception agent, a scene understanding agent, a diary generation agent and a personalized adaptive agent. According to the method, deep understanding of the emotional state and the environment scene of the user is realized through a multi-modal fusion perception and agent cooperation mechanism, the personalized emotional diary is generated in combination with chain thinking and reflection, and the content insight and the user experience are effectively improved.
Owner:LINKER

Three-dimensional acoustic visualization rendering method and system based on sound field modeling

PendingCN121544772A3D-image rendering3D modellingVoxelVisual expression
The invention discloses a three-dimensional acoustic visualization rendering method and system based on sound field modeling. The method comprises the following steps: arranging a microphone array to collect a multi-channel time domain audio signal, and carrying out short-time Fourier transform to obtain frequency domain complex sound pressure data; performing inversion based on a beam forming algorithm to obtain complex sound pressure of each voxel point in a three-dimensional space, and constructing a four-dimensional acoustic voxel matrix; further utilizing a volume ray projection algorithm to map the volume element data into color and transparency, and realizing three-dimensional rendering and visual expression of the sound field; the system supports frequency screening, time evolution playback, transparency adjustment and three-dimensional visual angle control, and the visual analysis capability of the sound field is improved; and meanwhile, a structured data output interface is provided and can be linked with an AI algorithm, so that intelligent functions such as acoustic anomaly detection and tuning optimization are realized. The method has the capabilities of high-precision modeling, deep rendering expression and intelligent expansion, and is suitable for various acoustic application scenes such as intelligent sound boxes, vehicle-mounted sound boxes, conference systems and the like.
Owner:NANJING JIQIDAO INTELLIGENT TECH CO LTD

Extracting audio signal from audio mixed signal using neural network

The present disclosure provides an audio system, method, and system for facilitating machine operation. The machine includes an actuator that assists the tool in performing the task. In an example, an audio system is configured to receive an audio mixed signal of a signal generated by an audio source that includes at least one of a tool or an actuator that is performing a task. An audio source forming an audio mixed signal is identified by a relative position to each microphone of a microphone array that measures the audio mixed signal. The audio system is configured to extract an audio signal from an audio mixed signal generated by the identified audio source based on a correlation of spectral features in a multi-channel spectrogram of the audio mixed signal and directional information indicative of a relative position of the identified audio source. The audio system outputs the extracted audio signal to facilitate operation of the machine.
Owner:MITSUBISHI ELECTRIC CORP

Doll intelligent speech recognition method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence and voice recognition, and discloses a doll intelligent voice recognition method and system based on artificial intelligence, and the method comprises the steps: collecting audio through a multi-microphone array; locally executing speech enhancement, acoustic feature extraction and lightweight wake-up word detection; local speech recognition is started, semantic confidence is evaluated, and desensitized data are uploaded to a cloud only when the confidence is insufficient; and the cloud end performs secondary recognition of context perception by using the large model in combination with the language habits of the children, and returns an optimization result. The system comprises an audio acquisition module, a local processing module, a cloud collaboration module and a state management module. According to the invention, through an end-cloud collaborative architecture, the identification robustness and the interaction intelligence level in a complex environment are significantly improved while privacy and low delay are guaranteed.
Owner:ZHUHAI ZHIHUI HUACHUANG TECHNOLOGY CO LTD

Noise reduction in audio mixing systems including a beamformer

This disclosure provides methods, devices, and systems for audio signal mixing. The present implementations more specifically relate to mixing audio signals from a microphone array by performing fixed beamforming to generate beams, reducing noise on the beams, and mixing the beams to generate a final audio signal for playback. In some aspects, an audio mixing system includes a fixed beamformer to generate beams from audio signals from a microphone array and noise reduction units (NRUs) to reduce a noise component of each audio beam. The system also includes logic to calculate a signal characteristic of each reduced noise audio beam to determine, based on the signal characteristics, the reduced noise audio beams that include a speech component. The logic also generates a gain for each audio beam based on the selection, with the gains used in beam mixing. In some aspects, the NRU includes a neural network noise reduction unit.
Owner:SYNAPTICS INC