Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1512 results about "Microphone array" patented technology

A microphone array is any number of microphones operating in tandem. Typically, an array is made up of omnidirectional microphones, directional microphones, or a mix of omnidirectional and directional microphones distributed about the perimeter of a space, linked to a computer that records and interprets the results into a coherent form. Arrays may also be formed using numbers of very closely spaced microphones. Given a fixed physical relationship in space between the different individual microphone transducer array elements, simultaneous DSP (digital signal processor) processing of the signals from each of the individual microphone array elements can create one or more "virtual" microphones. Different algorithms permit the creation of virtual microphones with extremely complex virtual polar patterns and even the possibility to steer the individual lobes of the virtual microphones patterns so as to home-in-on, or to reject, particular sources of sound. The application of these algorithms can produce varying levels of accuracy when calculating source level and location, and as such, care should be taken when deciding how the individual lobes of the virtual microphones are derived.

Multi-modal shared teleoperation system and method for three-arm space robot

Disclosed in the present invention are a multi-modal shared teleoperation system and method for a three-arm space robot. The system at least comprises a master-side teleoperation system, a communication module and a slave-side robot system, wherein the master-side teleoperation system at least comprises two force feedback hand controllers, a microphone array and upper computer software, and the slave-side robot system comprises two operating arms each equipped with a gripper at the end, an observation arm having a binocular camera mounted at the end, a vision unit, a force sensor and lower computer software. The method comprises: an operator controlling two operating arms of an extravehicular robot to execute a task, and controlling an observation arm to acquire a better local field of view. In the method, a multi-modal teleoperation method comprising pose control, voice control and force control is fused with autonomous control of a robot by means of a shared control algorithm, and thus, human-robot collaborative control over the position, orientation and contact force of a robotic arm can be realized on the basis of the requirements of the operator, and the robot autonomously executes other relatively simple tasks, thereby reducing the operation burden of operators, and improving the control efficiency.
Owner:SOUTHEAST UNIV

Multimodal shared telerobotic system and method for three-arm space robot

A multimodal shared telerobotic system and method for a three-arm space robot, the system at least includes a local-site system, a communication module, and a remote-site system, where the local-site system includes two force-feedback haptic devices for left and right hands, a microphone array, and upper computer software; the remote-site system includes two robotic arms provided with end-effectors, an observation arm with a stereo camera installed at an end thereof, two force sensors, a vision unit and lower computer software; an operator can control the two robotic arms of the robot outside a cabin for performing operations, and control the observation arm to obtain a better local view; and a multimodal telerobotic control method of pose control, voice control, and force control is integrated with the robot's autonomous control through a shared control algorithm.
Owner:SOUTHEAST UNIV

Robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation

The invention discloses a robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation. The method comprises the following steps: S1, dynamically fusing multi-modal emotions; the method comprises the following steps: S1, synchronously acquiring voice, visual and text signals through a multi-source heterogeneous sensor, capturing a user voice stream by a high-fidelity microphone array, and extracting acoustic characteristics such as intonation and speed, S2, performing cross-modal reasoning; s3, synchronously generating contents; step S4: style migration; step S5, anthropomorphic voice and expression generation; according to the method, man-machine interaction emotion is analyzed and generated by utilizing a large language model and multi-modal information fusion, the singleness of interaction emotion and the deficiency of emotional sharing ability are avoided, a strong emotion interaction characteristic is achieved, the image of the robot is obtained through a generative technology and can be migrated to any image, the limitation that a specific image is independently made is broken through, and the interaction effect of the robot is improved. The advantage that one robot can be suitable for different scenes is achieved.
Owner:JIANGSU YUNMU ZHIZAO TECH CO LTD

Motor fault detection method and system based on voiceprint recognition

The invention discloses a motor fault detection method and system based on voiceprint recognition. According to the method, an annular microphone array is adopted to collect motor sound signals in a non-contact mode, a three-channel time-frequency data set is constructed through empirical mode decomposition (EMD) and a Mel-frequency cepstral coefficient (MFCC), fault diagnosis is carried out in combination with a CNN + ResNet network, and dynamic time warping (DTW) and CNN fusion matching is supported. The system comprises a preprocessing module, a fault template library and a matching algorithm, integrates wavelet denoising and multi-beam acquisition technologies, covers a frequency band of 50Hz-20kHz, can display a fault type and trend analysis in real time, and triggers secondary verification when the confidence coefficient is insufficient. According to the scheme, the anti-interference capability is improved through array signal processing, model parameters are optimized in combination with transfer learning, non-contact detection is achieved, the real-time performance and accuracy of fault diagnosis are remarkably improved, and the method is suitable for industrial motor health monitoring.
Owner:GUANGZHOU DAYIN ZHIYUAN DIGITAL TECH CO LTD

Leakage detection and partial discharge digital imaging detection method and system based on acousto-optic fusion

The invention relates to the technical field of nondestructive testing, in particular to a leak detection and partial discharge digital imaging detection method and system based on acousto-optic fusion, and the method comprises the following steps: based on channel microphone array sound wave data in a partial discharge signal suspicious region, extracting a sound wave abnormal section, positioning a sound source, matching image edge features, and synchronously marking; and analyzing the phase change of the multi-frequency signal to judge a sound source concentration area, tracking the moving trend of a disturbance point, and outputting an acousto-optic fusion positioning trend track. According to the method, the partial discharge feature recognition sensitivity is improved through high-frequency peak paragraph screening and dominant frequency recognition, the abnormal region positioning precision is enhanced in combination with image edge extraction and sound source space matching, and acousto-optic synchronous positioning and trend trajectory display are achieved through multi-frequency signal phase analysis and image frame disturbance tracking. Through fusion of frequency domain feature extraction, image recognition, dynamic comparison and other actions, the spatial precision of abnormal source recognition and the multi-source fusion analysis efficiency are improved, and the partial discharge traceability and dynamic monitoring capability are enhanced.
Owner:李美娟 +1

Rescue method and system of rescue robot for exploration

The invention discloses a rescue method and system of a rescue robot for exploration, and relates to the technical field of underground space rescue, and the method comprises the following steps: obtaining multi-path acoustic echo data of a karst cave and motion track data of the robot, and constructing a three-dimensional point cloud model according to the multi-path acoustic echo data and the motion track data of the robot; and identifying unmatched data in the multi-path acoustic echo data and the robot motion trail data, and taking an area where the unmatched data is located as an abnormal area. According to the method, the spatial resolution of signal acquisition is enhanced through the multi-microphone array, robust acoustic fingerprints are extracted in combination with a noise reduction algorithm and short-time Fourier transform, high-confidence human body sound source signals are screened out by using a feature template matching mechanism, accurate extraction and recognition of human body acoustic features in a karst cave complex noise environment are realized, and the accuracy of human body acoustic feature recognition is improved. The problem of misjudgment caused by confusion of sound source features and environmental noise in a traditional method is effectively solved.
Owner:NORTH CHINA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Multi-language cross-culture communication auxiliary method and system based on large model

The invention provides a multi-language cross-culture communication assisting method and system based on a large model. The method comprises the following steps: receiving a source language audio stream during a call, calling a multi-language sound frequency harmonic modulation feature library to extract fundamental frequency harmonic intensity distribution and tone turning features, and generating a cultural acoustic fingerprint vector; based on the vector, controlling a microphone array phase difference, directionally enhancing a fundamental frequency harmonic component of a speaker and suppressing noise, and outputting a high signal-to-noise ratio spectrogram; analyzing the pronunciation rhythm and tone turning characteristics of the spectrogram, capturing the pitch jump and duration of the syllable boundary, and generating an acoustic culture label; associating the spectrogram with a target semantic library, matching harmonic distribution and a cultural context rule based on a large model, and outputting a cultural interpretation prompt containing an ambiguity resolution suggestion; and generating a calibration result according to the acoustic tag and the semantic prompt, and overlapping the dynamic floating subtitles to the face area of the speaker in the video conference picture. According to the invention, cultural tone ambiguity in multi-language communication is eliminated.
Owner:LUSTER LIGHTWAVE CO LTD

Microphone array sound source localization method and system based on cross-correlation-beam forming closed-loop optimization

The invention relates to a microphone array sound source positioning method and system based on cross-correlation-beam forming closed-loop optimization, and belongs to the technical field of sound source positioning. The method comprises the following steps: collecting multichannel sound signals through a microphone array and preprocessing the multichannel sound signals to extract time-frequency features and suppress noise interference; time delay information among the microphones is estimated by adopting a generalized cross-correlation phase transformation algorithm, and an optimization strategy is introduced to improve estimation stability and anti-interference performance; enhancing the target sound source signal in combination with a minimum variance undistorted response beam forming algorithm and an adaptive Kalman filtering mechanism; constructing a closed-loop feedback optimization mechanism based on the beam output signal to realize feedback adjustment; and adopting a hybrid network architecture, taking the beam output signal amplitude spectrum as input, and outputting the frequency spectrum or mask of the obtained target sound source signal. The method has the advantages of high calculation efficiency, high positioning precision and strong anti-interference capability, and is suitable for real-time acoustic signal processing in a complex environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Multi-element microphone array sound source localization method based on multistage signal preprocessing and subspace spectrum optimization

The invention relates to a multi-element microphone array sound source localization method based on multistage signal preprocessing and subspace spectrum optimization, and belongs to the technical field of acoustic detection. Aiming at the problems of poor noise immunity, weak multi-sound-source resolution capability and low calculation efficiency of the existing sound source positioning technology, a triple signal preprocessing and subspace collaborative optimization scheme is provided; firstly, incoherent noise is suppressed through phase coherent filtering, a signal is reconstructed through principal component analysis, and phase deviation is calibrated through fundamental frequency; then constructing a guiding matrix and decomposing a noise subspace, and extracting a coarse positioning result; and finally, high-precision angle optimization is realized based on a chaos initialization differential evolution algorithm, and the efficiency is improved by combining a dynamic search range and an early stop mechanism. According to the method, the anti-interference capability in a low signal-to-noise ratio environment is remarkably enhanced, the problems of missing detection and false detection during dense distribution of multiple sound sources are effectively solved, meanwhile, the positioning precision and the real-time performance are considered, and the method is suitable for acoustic fault detection of complex scenes such as power transmission line inspection.
Owner:CHONGQING UNIV

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

New energy aluminum alloy precision casting multi-mode quality monitoring device and method

The invention relates to the technical field of new energy aluminum alloy precision casting multi-modal quality monitoring, in particular to a new energy aluminum alloy precision casting multi-modal quality monitoring device and method. According to the technical scheme, the new energy aluminum alloy precision casting multi-mode quality monitoring device comprises a quality monitoring device body, an infrared thermal imager probe, a germanium single crystal protection lens, an integrated spectrometer probe debugging port, an acoustic detection microphone array and a honeycomb acoustic shielding cover. An infrared thermal imager probe is arranged on one side of the quality monitoring device main body, and a germanium single crystal protection lens is arranged on the outer side of the infrared thermal imager probe; synchronous acquisition and fusion analysis of multi-physical field data are realized by integrating six modes of spectrum, infrared thermal imaging, vibration, machine vision, acoustics and pressure deformation, and meanwhile, when the shrinkage risk is predicted by the infrared thermal imaging, a die-casting process compensation mechanism is automatically triggered to kill defects in the germination stage, so that the defect rate is greatly reduced.
Owner:GUANGYUAN YINGHE AUTO PARTS MANUFACTURING CO LTD

Sound acquisition and processing system based on cooperation of multiple microphone arrays

The invention discloses a sound acquisition and processing system based on cooperation of multiple microphone arrays. The system comprises a sound acquisition module, a multi-channel signal preprocessing module, a sound signal feature extraction module, an abnormal sound recognition module, a sound source positioning module and an alarm module. A multi-channel mixed data signal is collected through a circularly-arranged multi-microphone array formed by a plurality of microphones, after echo cancellation, wave beam domain noise reduction and multi-sound-source separation, a single-sound-source feature vector is extracted, according to the single-sound-source feature vector, abnormal sound including explosion, screaming or glass breakage is recognized through a BiLSTM and an attention mechanism model, and the abnormal sound is recognized through an attention mechanism model. And the GCC-PHAT and MDS-MUSIC algorithms are combined to position abnormal sound production, and alarm information is generated. According to the invention, accurate identification, positioning and alarm of the abnormal sound can be realized, and the real-time performance, the accuracy and the multi-target processing capability of abnormal sound monitoring in a complex environment can be obviously improved.
Owner:HANGZHOU DIANZI UNIV

Casting industry abnormal sound detection and grading response method based on voiceprint recognition

The invention provides a casting industry abnormal sound detection and grading response method based on voiceprint recognition, and belongs to the technical field of casting industry detection. A high-temperature-resistant microphone array is arranged at an easy-to-leak part of cast aluminum equipment, and three-stage filtering noise reduction and amplitude normalization preprocessing are adopted, so that the problem of poor signal quality caused by noise interference in a complex environment is effectively solved; the characteristics of the molten aluminum leakage sound in different frequency bands and different stages can be captured through variable window long-short time Fourier transform and extended Mel frequency cepstrum coefficient in combination with extraction of an energy change rate and a frequency spectrum gravity center; a Transform-CNN hybrid deep learning model based on an attention mechanism is constructed, and feature screening is optimized through principal component analysis and recursive feature elimination, so that the recognition and generalization ability of the model to the abnormal sound in the casting industry is significantly improved; and meanwhile, graded response measures are made based on the detection result, so that the abnormal conditions of molten aluminum leakage with different severity degrees are processed.
Owner:SHENZHEN POLYTECHNIC

Voice recognition processing method, system and equipment based on conference scene and medium

The invention relates to a voice recognition processing method, system and device based on a conference scene and a medium, and belongs to the technical field of voice processing. The voice recognition processing method comprises the following steps: acquiring an original conference audio stream collected by a microphone array; performing signal preprocessing on the original conference audio stream collected by the main channel, and outputting a pure voice signal; generating a sound source orientation thermodynamic diagram based on the original conference audio stream; extracting multi-dimensional voiceprint feature vectors from the pure voice signals, performing dynamic grouping, outputting a voice fragment set marked with voiceprint IDs, and generating an initial transcription text; dynamically correcting the initial transliteration text, and outputting a transliteration text stream with an industry term tag; and performing periodic memory enhancement processing on the transliteration text stream, outputting and analyzing a long text, and generating structured conference summary data. According to the invention, the automation level and accuracy of conference voice processing can be improved.
Owner:CHINA TRANSPORT INFORMATION TECH GRP CO LTD

Intelligent conference memo generation method based on robot

The invention provides an intelligent conference memo generation method based on a robot, and the method comprises the steps: collecting a multi-channel voice signal through a microphone array, and enhancing the voice of a target speaker through a beam forming technology; background noise is separated by adopting a self-adaptive filtering algorithm, and the voice signal quality is improved; speech features are extracted, language model parameters are adjusted, and a transliteration text is generated; analyzing the text structure, and extracting conference themes, participants and decision contents to form structured information; calculating semantic similarity among the knowledge graph nodes, and if the semantic similarity is higher than a preset threshold value, associating historical records to generate extension information; constructing a structured memorandum based on the extended information, organizing a conference theme, participants, decision contents and associated historical records, and generating an initial memorandum; and monitoring a memorandum editing operation, updating a knowledge graph node relationship, and generating a final memorandum document. According to the method, the accuracy, integrity and availability of conference records are effectively improved, and the conference efficiency is remarkably improved.
Owner:HUNAN HEXIN ANHUA BLOCKCHAIN TECH CO LTD

Intelligent voice recognition and analysis system based on universal smart phone chip

The invention discloses an intelligent voice recognition and analysis system based on a universal smart phone chip, and the system comprises a data collection module which synchronously captures a voice signal and motion sensor data through a built-in microphone array and a motion sensor, and generates an original multi-mode data package with a time sequence stamp; and the noise reduction processing module is used for receiving the original multi-mode data packet, executing environmental noise spectrum analysis, generating an anti-phase sound wave, performing signal enhancement and outputting a pure voice stream. According to the method, through multi-modal noise separation, nonlinear signal enhancement and hierarchical privacy protection, the contradiction between speech recognition precision and privacy security in a complex environment is solved, and meanwhile, by means of dynamic resource scheduling and a lightweight model, the finite computing power of a mobile phone chip is utilized to the maximum extent in the aspect of speech recognition.
Owner:BEIJING ZHIMAI TECHNOLOGY CO LTD

Partial discharge positioning method, partial discharge positioning equipment, partial discharge positioning device and readable storage medium

The invention relates to a partial discharge positioning method, equipment and device and a readable storage medium. The method comprises the following steps: collecting a visible light image and an infrared image of a to-be-detected area containing a partial discharge phenomenon, fusing the visible light image and the infrared image to generate a double-spectrum positioning map, collecting a sound wave signal of the to-be-detected area through a spiral microphone array, carrying out fusion processing on the double-spectrum positioning map and a sound source thermodynamic diagram of the sound wave signal, and obtaining a sound source thermodynamic diagram of the sound wave signal. Generating a comprehensive discharge positioning map, adopting a clustering algorithm to identify a plurality of features of the visible light image, the infrared image and the sound wave signal, determining the discharge position and the discharge type of the partial discharge phenomenon in the comprehensive discharge positioning map based on the plurality of features, and outputting the comprehensive discharge positioning map. Marking the discharge position and the discharge type of the partial discharge phenomenon in the comprehensive discharge positioning map; the optical, thermal and acoustic multi-mode information of the discharge is integrated, the positioning precision of the partial discharge is greatly improved, the fault condition can be conveniently and quickly judged, and the quick inspection requirement is met.
Owner:SHUOHUANG RAILWAY DEV +1

Single-microphone multi-array pickup method and system based on sound attenuation simulation

The invention provides a single-microphone multi-array pickup method and system based on sound attenuation simulation, and the method comprises the steps: obtaining an original audio signal recorded by a single microphone and sound source direction information, and carrying out the sound attenuation simulation based on the sound source direction information and a preset virtual microphone array orientation parameter, calculating the included angle between the sound source direction and each virtual microphone; generating an audio intensity attenuation coefficient of each virtual microphone according to the included angle; and acting the audio intensity attenuation coefficient on the original audio signal to generate multi-channel audio data simulating the multi-microphone array. By adopting the method, the response difference of different microphone positions to the sound source can be simulated, the pickup result of the multi-microphone array is obtained, and high hardware cost and complex installation flow caused by actual deployment of a complex microphone array are avoided; and low-cost and high-diversity training data sources are provided for sound source localization, noise suppression and other models based on deep learning.
Owner:BEIJING YUANZHI DIGITAL INFORMATION TECHNOLOGY CO LTD

Vocal cord problem identification feedback system for ophthalmology and otorhinolaryngology department

PendingCN120531329APhysical therapies and activitiesBronchoscopesDiseaseEarly Cancer Detection
The invention discloses a vocal cord problem recognition and feedback system for the ophthalmology and otorhinolaryngology department. The vocal cord problem recognition and feedback system comprises a sound collection module, an image collection module, a biological feedback module, a data processing module, an AI diagnosis module and a rehabilitation guidance module. The sound acquisition module comprises a microphone array and a self-adaptive noise reduction unit; the image acquisition module is provided with an endoscope camera and an image enhancement processor; the biological feedback module integrates a laryngeal myoelectricity sensor and a three-dimensional motion simulator; the data processing module executes multi-modal feature extraction and fusion; through mutual cooperation of the sound acquisition module, the image acquisition module, the biological feedback module and the data processing module, data can be accurately acquired, the early canceration detection rate is improved and the misdiagnosis rate is reduced through multi-modal fusion, the data acquired by the multi-modal structure is analyzed through the data processing module, the model is combined with a weekly updated disease map, and the early canceration detection rate is improved. And the recurrence prediction accuracy is improved.
Owner:SHANGHAI XINERYUE TEACHING MOULD CO LTD

AI scene conversion method based on visitor behavior

The invention discloses an AI scene conversion method and system based on visitor behaviors. According to the method, visitor behavior data are collected in real time through a multi-mode sensor network, and a UWB positioning base station, a millimeter wave radar module, a high-definition camera, a microphone array and a pressure sensor matrix are included. The system adopts an edge computing architecture, is equipped with an NVIDIAJetsonAGXXavier processor and a 512-core VoltaGPU (Graphics Processing Unit), and operates three core AI models to realize intelligent analysis. The LSTM residence time prediction model adopts a three-layer bidirectional architecture, each layer is configured with 256 hidden units, and visitor behavior sequences within 30 minutes are analyzed to predict the residence time. The Transform interaction intention recognition model comprises six layers of encoders and eight multi-head attention modules, and visual, voice and behavior three-mode characteristics are fused to judge visitor interaction requirements. According to the technical scheme, self-adaptive adjustment of the exhibition hall environment is achieved, the visitor satisfaction degree is improved by 28%, the operation cost is reduced by 15%, and important technical support is provided for intelligent transformation of modern exhibition places.
Owner:元羽兽数字科技(上海)有限公司

Multi-scene illumination management system integrating voice control and remote monitoring

The invention provides a multi-scene illumination management system integrating voice control and remote monitoring, and relates to the technical field of illumination management, and the system comprises the steps: receiving a scene wake-up instruction collected by a microphone array and a scene switching instruction sent by a remote monitoring terminal; adding priority labels and timestamps to the scene wakeup instruction and the scene switching instruction, and when it is detected that a timestamp difference value between the scene wakeup instruction and the scene switching instruction is smaller than a conflict time threshold value, determining an instruction arbitration coefficient through preemptive priority weights of the scene wakeup instruction and the scene switching instruction, determining an effective instruction according to the instruction arbitration coefficient and all the priority labels; and when the instruction priority of the effective instruction is greater than the instruction priority of the currently executed illumination scene instruction, generating a coverage execution signal, otherwise, generating an instruction discarding signal, further outputting an illumination control signal, and adjusting light parameters of the illumination equipment through the illumination control signal. According to the invention, multi-control-source conflicts in a lighting management system can be eliminated.
Owner:SHENZHEN BIAOMEI LIGHTING DESIGN ENG CO LTD

Intelligent interaction control method for AI glasses

The invention discloses an intelligent interaction control method for AI glasses, and relates to the technical field of AI glasses, and the method comprises the following steps: synchronously collecting user input signals and environmental parameters through a multi-source sensor group built in the glasses, the sensor group at least comprising an IMU, a binocular camera, a microphone array and a physiological sensor; carrying out real-time fusion processing on the multi-modal input data by adopting a lightweight neural network to generate an interaction intention feature vector; based on real-time detection results of environmental noise intensity and illumination conditions, dynamically distributing weight coefficients of all interaction modes; generating a hierarchical response instruction according to the weight coefficient and a confidence threshold; and continuously optimizing the personalized interaction strategy of the user through a federal learning framework. According to the AI glasses intelligent interaction control method, through multi-modal weighted fusion, edge AI acceleration and federated learning optimization, the interaction success rate in a complex environment is improved to 93.6%; the end-to-end delay is controlled within 28ms; and a user-defined interaction strategy is supported.
Owner:EMDOORVR TECH CO LTD

Speech recognition method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence, and particularly provides an artificial intelligence-based speech recognition method, which comprises the following steps of: acquiring multiple paths of speech signals through a microphone array; performing noise reduction processing on each path of voice signal, removing background noise, enhancing the voice signal by adopting an adaptive beam forming algorithm, and extracting a target voice signal; performing short-time Fourier transform on the preprocessed voice signals, extracting voice spectrum features, and extracting high-level semantic features of the voice through a deep learning model; inputting the extracted speech features into a speech recognition model based on an attention mechanism, generating text transcription of target speech, dynamically adjusting model parameters through an adaptive learning module, and optimizing a recognition result; according to the method, the recognition result is corrected according to the real-time feedback of the user, the corrected data is used for online updating of the model, the robustness and adaptability of the system are improved, and the method has the effects that training and evaluation are conducted through high-quality data, and delay is reduced in real-time processing.
Owner:BEIJING HURRICANE SOFTWARE CO LTD

Electric pump control method and system based on instruction perception

The invention discloses an electric pump control method and system based on instruction perception, and relates to the technical field of electric pump control, comprising: starting a voice recognition module, a wireless communication module and a sensor acquisition module, and establishing a bidirectional communication link with a remote controller; an AI-driven noise reduction algorithm is adopted to pre-process the environment audio collected by the microphone array, and an improved hidden Markov model is utilized to perform keyword recognition on the pre-processed audio to obtain a voice instruction signal; the remote controller sends a control instruction to a receiving end, analyzes the instruction content, records the current communication channel quality, dynamically adjusts the transmitting power, the frequency hopping strategy or the coding mode according to the channel quality, and ensures stable communication between the remote controller and the electric pump; and the main control unit simultaneously receives a voice command signal and a remote controller command signal, sets a command priority rule in combination with data acquired by the sensor, and executes corresponding operation according to the priority.
Owner:REDDY CO LTD

Sound source separation method and system for microphone array to pick up voice signals

The invention discloses a sound source separation method and system for a microphone array to pick up voice signals. The method comprises the following steps: picking up voice signals sent by a target reactor sound source and an interference sound source in a space by using the microphone array; performing segmentation and Fourier transform on the voice signal to obtain a time-frequency domain signal; calculating the time difference of arrival and the phase difference between each pair of microphones, and obtaining the direction estimation result of each sound source in the space by combining the geometric arrangement information of the microphone array; generating a time-frequency mask, applying the time-frequency mask to the time-frequency domain signal, and separating the time-frequency domain signal of each sound source; performing inverse short-time Fourier transform and fragment splicing to obtain a corresponding continuous time domain voice signal; and outputting to different audio output channels. According to the invention, the specific sound source signal generated by the reactor can be effectively separated, the operation state of the reactor can be accurately analyzed, the abnormal condition can be timely found, and the reliability and precision of monitoring the reactor can be further improved.
Owner:WUXI POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD

Iron tower voiceprint intelligent detection method and system

The invention provides an iron tower voiceprint intelligent detection method and system, and belongs to the voiceprint detection technology. The method comprises the following steps: firstly, outputting a sweep frequency signal in a specific frequency range, collecting iron tower vibration response data, extracting candidate frequency through spectral analysis, and dynamically adjusting excitation frequency by using a gradient descent algorithm; hardware filtering is carried out on collected sound signals, frequency band signals related to excitation frequency are reserved, and space beam forming is carried out through a microphone array to enhance iron tower voiceprint signals. Separating iron tower vibration components from the mixed signals by utilizing independent component analysis, performing time-frequency analysis, extracting an energy ratio of a specified frequency band by adopting wavelet packet transformation, and performing deep learning processing in combination with a one-dimensional convolutional neural network to generate an energy ratio and zero-crossing rate feature vector; and finally, inputting the feature vectors into a support vector data description model, setting an initial threshold value and dynamically adjusting the initial threshold value, thereby realizing graded judgment of bolt looseness and improving detection efficiency and stability.
Owner:DEZHOU POWER SUPPLY COMPANY OF STATE GRID SHANDONG ELECTRIC POWER

Gas leakage identification method and system based on sound positioning, medium and equipment

The invention relates to the field of gas leakage identification, and discloses a gas leakage identification method and system based on sound localization, a medium and equipment, and the method comprises the steps: collecting a leakage sound wave signal through a microphone array, and extracting acoustic features in a complex background; simulating a diffusion process of gas in a turbulence environment after gas leakage through an established leakage gas convection-diffusion model, and combining gas concentration distribution in the simulated diffusion process with acoustic characteristics to predict sound source localization at a leakage position; modeling a sound source localization search process as POMDP, and iteratively updating a confidence state of a sound source position through a particle filter; spatial distribution features are extracted from the confidence state through DBSCAN clustering to serve as input of the LSTM-DQN network, spatial and temporal features are fused through the constructed LSTM-DQN network, cross-scene generalization is achieved through transfer learning, and sound source coordinates are dynamically optimized and output through the Q network. According to the invention, the positioning accuracy and real-time performance in a complex environment are improved.
Owner:SUZHOU SHENGTENG ROBOT CO LTD +1

Simultaneous interpretation data processing method and system based on POE microphone array

The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation data processing method and system based on a POE microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and meeting place environment noise spectrum features through a distributed microphone array powered by the Ethernet; after time domain framing is carried out on the audio stream, adaptive filtering is carried out by using a dynamic noise reduction weight coefficient to obtain a primary pure voice segment; dividing the multi-language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-language speech endpoint detection model, and matching a corresponding acoustic model to generate a phoneme-level time alignment sequence; comparing and outputting a term replacement instruction stream in real time in combination with a simultaneous transfer term library, and generating an intermediate semantic representation vector after fusion; and the low-delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual meeting place scene.
Owner:SUZHOU FUCHUAN TECH

Aerial engine compressor rotation stall voiceprint monitoring method, system, medium and equipment

The invention discloses a voiceprint monitoring method, system, medium and equipment for rotation stall of an aero-engine compressor, and the method comprises the steps: determining the layout of a microphone array through the model parameters of the compressor of the aero-engine, and carrying out the synchronous collection of sound pressure time domain signals of multiple channels of the compressor of the aero-engine through the microphone array; arranging the multi-channel sound pressure time-domain signals in sequence to construct a time-domain signal matrix; performing spectral analysis on the sound pressure time-domain signals by adopting a fast Fourier transform algorithm to obtain a spectrogram of the sound pressure time-domain signals of each channel, monitoring whether an abnormal single-tone frequency peak value lower than one time of rotation frequency appears in the spectrogram or not, obtaining an acoustic mode spectrogram of the abnormal single-tone frequency peak value by adopting a single-frequency acoustic mode decomposition method, and obtaining an acoustic mode spectrum of the abnormal single-tone frequency peak value; and whether the rotating stall characteristic of the gas compressor occurs in the acoustic mode spectrogram is monitored to judge whether the rotating stall characteristic of the gas compressor occurs or not.
Owner:XI AN JIAOTONG UNIV +1