Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6044 results about "Audio signal" patented technology

An audio signal is a representation of sound, typically using a level of electrical voltage for analog signals, and a series of binary numbers for digital signals. Audio signals have frequencies in the audio frequency range of roughly 20 to 20,000 Hz, which corresponds to the lower and upper limits of human hearing. Audio signals may be synthesized directly, or may originate at a transducer such as a microphone, musical instrument pickup, phonograph cartridge, or tape head. Loudspeakers or headphones convert an electrical audio signal back into sound.

Headset antenna and connector for the same

ActiveUS20090033574A1Increase the equivalent impedanceImprove rendering capabilitiesAntenna supports/mountingsAntenna adaptation in movable bodiesHeadphonesAudio frequency
A headset antenna and a connector for the same are provided. The headset antenna includes an audio signal line, an antenna and a high impedance element in specified application frequency ranges. The audio signal line is adapted for transmitting an audio signal and the antenna is adapted for receiving an RF signal. The high impedance element is disposed on a transmission path of the audio signal and generates a high impedance at a specified frequency band of the RF signal, so that the audio signal line is equivalent to an open circuit and the antenna obtains a better receiving capability.
Owner:HTC CORP

Compensation for nonuniform delayed group communications

A method for synchronizing audio reproduction in collocated end devices is presented. Each of the devices auto-correlates using noise to determine a threshold prior to the antenna receiving an audio signal. When the devices receive a common audio signal, they provide audio outputs. Each device cross-correlates its audio output with the audio outputs of the other devices. The timing of the audio output of each device is then adjusted such that the audio outputs of all of the devices align temporally with the lagging or leading device.
Owner:MOTOROLA SOLUTIONS INC

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Wind turbine generator voiceprint fault recognition method

The invention provides a wind turbine generator voiceprint fault recognition method, and relates to the technical field of wind turbine generator state monitoring and fault diagnosis, and the method comprises the steps: carrying out the noise reduction of an original audio signal through variational mode decomposition, screening a target mode of which the frequency, energy and kurtosis accord with features, and reconstructing the signal; extracting a Mel frequency cepstrum coefficient and a sensing noise robust coefficient, and generating multi-dimensional voiceprint data in combination with statistical characteristics such as a frequency spectrum gravity center, a spectrum entropy, energy, kurtosis and a zero-crossing rate; constructing a support set based on the prototype network, realizing small sample fault classification by calculating the Euclidean distance between the feature vector and the prototype vector, and outputting a preliminary result; judging whether the voiceprint is abnormal according to a preset threshold value, if so, storing the voiceprint into a dynamic abnormal voiceprint knowledge base; frequently occurring abnormal samples are manually labeled and added into a support set, the prototype network is retrained to update the model, and continuous optimization of the fault recognition capability is achieved.
Owner:CGN (SHANXI) NEW ENERGY INVESTMENT CO LTD

Unsupervised wind power equipment blade fault detection method based on phase perception parallel attention mechanism

The invention relates to a wind power equipment blade fault detection technology, discloses an unsupervised wind power equipment blade fault detection method based on a phase perception parallel attention mechanism, and solves the problems that an existing wind power equipment blade fault detection method is high in dependence on labeled data, insufficient in generalization ability under strong noise and variable working conditions and high in fault detection efficiency. And a weak transient fault signal and a dynamic change characteristic are difficult to capture robustly. According to the scheme of the invention, the method comprises the steps: collecting a blade operation audio signal, and extracting a dual-channel time-frequency feature containing an amplitude spectrum and a phase spectrum through improved short-time Fourier transform; a deep adversarial auto-encoder is constructed by using an encoder containing a phase perception parallel attention module, a decoder and an auxiliary encoder, and normal working condition feature distribution is learned by reconstructing an error loss, potential representation consistency loss, adversarial loss and phase consistency loss optimization model during off-line training; in the reasoning stage, the fault is judged based on the feature distance score and the reconstruction error score.
Owner:CHINA HYDROELECTRIC ENGINEERING CONSULTING GROUP CHENGDU RESEARCH HYDROELECTRIC INVESTIGATION DESIGN AND INSTITUTE

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

Security camera abnormal behavior identification method and system based on multi-modal fusion

The invention relates to a security camera abnormal behavior identification method and system based on multi-modal fusion. The method comprises the steps of obtaining multi-modal data such as visible light image data, infrared image data and audio signal data of a security camera in a target monitoring period; performing feature extraction on the data of different modes to obtain respective feature vector sets; the features of all the modes are fused, and in the multi-mode feature fusion process, a weighted fusion algorithm for dynamically adjusting the fusion weight according to the confidence coefficient weight of feature vectors of all the modes is adopted; inputting the fusion feature vector set into an abnormal behavior classification model for classification and early warning; according to the scheme, the accuracy and effectiveness of abnormal behavior recognition in a complex environment can be improved, and then the overall efficiency of security monitoring is improved.
Owner:SHENZHEN KEAN DIGITAL CO LTD

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Intelligent predictive maintenance system for audio equipment fault

The invention discloses an intelligent predictive maintenance system for an audio equipment fault, and the system comprises a data collection module which collects the internal sensor data during the operation of audio equipment, and outputs an audio signal feature parameter, an external environment parameter, and a historical operation log; the feature preprocessing module is used for carrying out standardization and noise reduction processing on the multi-source data based on a multi-modal feature fusion algorithm; the health state evaluation module outputs a health state evaluation result of the audio equipment in real time through a hybrid analysis model combining a convolutional neural network, a long and short-term memory network and an attention mechanism; the fault risk prediction module is used for performing real-time fault risk prediction according to the evaluation result and generating predictive maintenance decision parameters; and the maintenance decision and early warning module is used for outputting fault early warning information according to the prediction parameters and automatically generating maintenance operation suggestions when the early warning level reaches a preset condition. According to the invention, the operation reliability of the audio equipment can be effectively improved, and intelligent prediction and advanced maintenance of faults are realized.
Owner:SHENZHEN JIEYU INFORMATION TECH CO LTD

Physiological monitoring soundbar

A soundbar for medical monitoring which may comprise a speaker, a sensor, and a hardware processor. The speaker can be configured to emit audio signals. The sensor can be configured to obtain sensor data relating to a physiology of a subject. The sensor can include a camera and the sensor data can include image data. The hardware processor can be configured to access the sensor data and determine a health status of the subject based on at least the sensor data.
Owner:MASIMO CORP

Noise reduction method combining different noise reduction algorithms of motorcycle riding earphone

The invention relates to the technical field of earphone noise reduction processing, and discloses a motorcycle riding earphone noise reduction method combining different noise reduction algorithms, comprising the following steps: acquiring an original audio signal, and distributing the original audio signal to each processing module; a preprocessing signal is generated, voice activity information is determined, and environmental noise energy information is calculated; according to the environmental noise energy or the auxiliary information, adaptively adjusting a high-low noise energy threshold value; the original signals are input into the traditional and AI noise reduction module in parallel to generate two paths of noise reduction signals; based on the environment noise, the voice information and the threshold value, performing weighted fusion to generate an output signal; and carrying out equalization and compression processing on the output signal, and driving the loudspeaker to output. According to the invention, comprehensive judgment on environmental noise energy information, voice activity information and riding speed is introduced, an output signal of a traditional noise reduction module is set to be zero in a high-speed voice-free environment, and AI noise reduction is combined, so that tone quality distortion of a traditional noise reduction algorithm in a noise environment is effectively avoided.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD +1

Method, apparatus and system for neural network hearing aid

The disclosure generally relates to a method, system and apparatus to improve a user's understanding of speech in real-time conversations by processing the audio through a neural network contained in a hearing device. The hearing device may be a headphone or hearing aid. In one embodiment, the disclosure relates to an apparatus to enhance incoming audio signal. The apparatus includes a controller to receive an incoming signal and provide a controller output signal; a neural network engine (NNE) circuitry in communication with the controller, the NNE circuitry activatable by the controller, the NNE circuitry configured to generate an NNE output signal from the controller output signal; and a digital signal processing (DSP) circuitry to receive one or more of controller output signal or the NNE circuitry output signal to thereby generate a processed signal; wherein the controller determines a processing path of the controller output signal through one of the DSP or the NNE circuitries as a function of one or more of predefined parameters, incoming signal characteristics and NNE circuitry feedback.
Owner:FORTELL RESEARCH INC

Method, apparatus and system for neural network hearing aid

The disclosure generally relates to a method, system and apparatus to improve a user's understanding of speech in real-time conversations by processing the audio through a neural network contained in a hearing device. The hearing device may be a headphone or hearing aid. In one embodiment, the disclosure relates to an apparatus to enhance incoming audio signal. The apparatus includes a controller to receive an incoming signal and provide a controller output signal; a neural network engine (NNE) circuitry in communication with the controller, the NNE circuitry activatable by the controller, the NNE circuitry configured to generate an NNE output signal from the controller output signal; and a digital signal processing (DSP) circuitry to receive one or more of controller output signal or the NNE circuitry output signal to thereby generate a processed signal; wherein the controller determines a processing path of the controller output signal through one of the DSP or the NNE circuitries as a function of one or more of predefined parameters, incoming signal characteristics and NNE circuitry feedback.
Owner:FORTELL RESEARCH INC

Bird identification method and device based on sound-image multi-modal fusion

The invention discloses a bird identification method based on sound-image multi-modal fusion. The bird identification method comprises the following steps: S1, carrying out standardized frame-level preprocessing on bird audio signals; s2, acoustic features are extracted and enhanced, and an acoustic high-level feature vector which highlights birdsong discrimination information and suppresses environmental noise is obtained; s3, visual image standardization preprocessing; s4, performing visual feature extraction and multi-scale fusion to obtain a visual high-level feature vector which enhances correspondence to the bird key form area and inhibits background interference; s5, performing dynamic weighted fusion on the decision-making layer to obtain a bird existence probability; and S6, comparing the bird existence probability with a preset threshold value of the corresponding bird, and judging whether the bird exists or not and the type of the existing bird. Through cross-modal feature enhancement and adaptive fusion, the precision, robustness and real-time performance of bird recognition in a complex orchard environment are significantly improved, and a core technical support is provided for green intelligent bird repelling.
Owner:NANJING FORESTRY UNIV

Method, apparatus and system for neural network enabled hearing aid

The disclosure generally relates to a method, system and apparatus for processing audio through a neural network contained in a hearing device. In one embodiment, the disclosure relates to an apparatus to enhance incoming audio signal. The apparatus includes a controller to receive an incoming signal and provide a controller output signal; neural network engine (NNE) circuitry in communication with the controller, the NNE circuitry activatable by the controller, the NNE circuitry configured to generate an NNE output signal from the controller output signal; and digital signal processing (DSP) circuitry to receive one or more of controller output signal or the NNE circuitry output signal to thereby generate a processed signal; wherein the controller determines a processing path of the controller output signal through one of the DSP or the NNE circuitries as a function of one or more of predefined parameters, incoming signal characteristics and NNE circuitry feedback.
Owner:FORTELL RESEARCH INC

Automation for inserting a reference accessing a screen-shared file in meeting summaries and transcripts

The disclosed techniques provide a system for automatically inserting reference to a file in meeting transcripts or meeting summaries. In general, the disclosed techniques manage and enrich meeting transcripts, meeting summaries, meeting recordings, and screen-shared files during an online meeting. During an online meeting, when a presenter screenshares an application file, such as Word doc, PowerPoint, Excel, etc., a system creates meeting transcripts or a summary for the real-time discussion based on audio signals and / or chat messages relating to the shared contents of the application file. The system also determines the location of the application file and inserts a reference to the application file in the transcript or summary.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Grid telephone traffic quality inspection intelligent analysis system and method based on large language model

The invention relates to the technical field of intelligent telephone traffic quality inspection, and discloses a grid telephone traffic quality inspection intelligent analysis system and method based on a large language model, and the system comprises the steps: collecting and obtaining a multi-role call audio signal in real time, and preliminarily carrying out the speaker separation and role marking through a voice recognition module and a voiceprint recognition module; forming a preliminary role recognition result; detecting a suspected role identity error region in combination with multi-dimensional features, and when a detection result meets a preset condition, triggering a dynamic correction mechanism, generating auxiliary judgment information in combination with identity declaration keywords, business term matching and dialogue context logic inference, adopting a multi-dimensional weight decision strategy, and re-correcting a role identity tag, so as to obtain a role identity error region. Updating a role recognition result; and setting an accurate evaluation module of a dynamic correction result, feeding back and adjusting a multi-feature weight and a trigger threshold in real time, and forming an iterative optimization mechanism of dynamic correction. The method has the advantage of improving the dynamic correction capability.
Owner:XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER

Method, apparatus and system for neural network hearing aid

The disclosure generally relates to a method, system and apparatus to improve a user's understanding of speech in real-time conversations by processing the audio through a neural network contained in a hearing device. The hearing device may be a headphone or hearing aid. In one embodiment, the disclosure relates to an apparatus to enhance incoming audio signal. The apparatus includes a controller to receive an incoming signal and provide a controller output signal; a neural network engine (NNE) circuitry in communication with the controller, the NNE circuitry activatable by the controller, the NNE circuitry configured to generate an NNE output signal from the controller output signal; and a digital signal processing (DSP) circuitry to receive one or more of controller output signal or the NNE circuitry output signal to thereby generate a processed signal; wherein the controller determines a processing path of the controller output signal through one of the DSP or the NNE circuitries as a function of one or more of predefined parameters, incoming signal characteristics and NNE circuitry feedback.
Owner:FORTELL RESEARCH INC

Device working state recognition method based on voiceprint recognition model

The invention discloses an equipment working state recognition method based on a voiceprint recognition model, and relates to the technical field of industrial equipment operation state recognition. The equipment working state recognition method based on the voiceprint recognition model comprises the following steps: collecting operation audio waveform data of target equipment, extracting acoustic representation data containing parameters such as short-time energy, a frequency spectrum centroid, a spectrum flux, MFCC and a zero crossing rate, inputting the acoustic representation data into a pre-trained voiceprint recognition model to extract voiceprint feature representation vectors, and carrying out voiceprint feature representation on the target equipment; according to the method, the audio signal is divided into the frames, the acoustic features such as short-time energy, spectrum centroid, spectrum flux, MFCC and zero crossing rate are extracted, the inter-frame evolution relation is modeled in combination with the bidirectional neural network, and the attention mechanism is introduced to highlight the key frame segment, so that the real-time performance of the audio signal is improved, and the real-time performance of the audio signal is improved. The recognition capability of working conditions such as fuzzy state boundary or unobvious transition is effectively enhanced, and the time sequence analysis and state judgment precision is improved.
Owner:FUJIAN RUIXIN TECH CO LTD

Adaptive noise reduction method and system for multi-mode audio SoC main control chip

The invention relates to the field of adaptive noise reduction, in particular to an adaptive noise reduction method and system for a multi-mode audio SoC main control chip. The method comprises the following steps: acquiring an original audio input signal according to an SoC main control chip, performing time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and constructing a full-frequency domain noise interference topology table; carrying out multi-mode interference factor separation on the full-frequency-domain noise interference topology table, carrying out multi-noise environment modeling, and constructing a real-time noise scene model; time sequence noise slope fluctuation modeling is carried out on the real-time noise scene model, dynamic noise reduction response learning optimization is carried out, and a dynamic noise reduction optimization strategy is constructed; and performing neural coding gain and audio boundary line reconstruction according to the original audio input signal to obtain a coding gain key audio signal and a weak audio optimization signal. According to the invention, by flexibly adjusting the noise reduction intensity of the audio signal, the real-time scene noise reduction performance is optimized.
Owner:HANK ELECTRONICS

Respiration abnormity identification method, system and equipment and medium

The invention relates to the technical field of intelligent medical monitoring, and discloses a breathing abnormity identification method, system and device and a medium, and the method comprises the steps: obtaining a multi-modal original data set; extracting a spectrum feature of the audio signal, a periodic feature of the physiological motion signal and a change feature of the environmental monitoring data from the multi-modal original data set, and constructing a multi-dimensional feature matrix based on the extracted features; performing cross-modal fusion processing on the multi-dimensional feature matrix by adopting a multi-channel convolutional network to generate a joint feature vector; and in combination with a pre-constructed knowledge graph, identifying an abnormal feature mode from the joint feature vector through a density clustering algorithm, screening an abnormal mode in combination with a medical diagnosis rule, and outputting a breathing abnormality identification result. According to the method, the defects that a single signal is sensitive to noise and a non-stationary signal is insufficient in analysis capability are overcome; meanwhile, a medical knowledge graph and a density clustering algorithm are combined, non-pathological breathing modes are screened out, and the accuracy and clinical credibility of complex breathing anomaly recognition are remarkably improved.
Owner:GENERAL HOSPITAL OF NUCLEAR IND

Processing parametrically coded audio

A method comprising receiving a first input bit stream for a first parametrically coded input audio signal, the first input bit stream including data representing a first input core audio signal and a first set including at least one spatial parameter relating to the first parametrically coded input audio signal. A first covariance matrix of the first parametrically coded audio signal is determined based on the spatial parameter(s) of the first set. A modified set including at least one spatial parameter is determined based on the determined first covariance matrix, wherein the modified set is different from the first set. An output core audio signal is determined, which is based on, or constituted by, the first input core audio signal. An output bit stream for a parametrically coded output audio signal is generated, the output bit stream including data representing the output core audio signal and the modified set.
Owner:DOLBY LABORATORIES LICENSING CORP +1

Earthquake early warning equipment and method based on campus broadcast

The invention discloses an earthquake early warning device and method based on campus broadcast, and the method comprises the steps: carrying out the building response coupling analysis through receiving an early warning signal of an earthquake monitoring center, generating an earthquake magnitude-building coupling feature, and mapping the earthquake magnitude-building coupling feature into an enhanced early warning source adaptive to an acoustic environment; acoustic cavity detection is carried out based on threatened area identification, and differential playing topology is constructed; extracting a high-risk gathering area through personnel density scanning, generating personalized evacuation guiding content and determining a playing priority; high-precision time sequence control is realized by adopting a synchronous reference anchor point technology, and a multi-mode playing domain is constructed through carrier modulation and frequency division multiplexing; and finally, an acoustic navigation field is utilized to guide accurate scheduling and playing of audio signals, differential playing and accurate coverage of early warning information can be realized according to the anti-seismic characteristics and acoustic environment characteristics of different buildings, and an intelligent acoustic solution is provided for campus earthquake early warning.
Owner:FUZHOU BENYANG INFORMATION TECH CO LTD

Atmosphere lamp control method, electronic device and program product

The invention discloses an atmosphere lamp control method, electronic equipment and a program product. The method comprises the following steps: acquiring PCM data of an audio signal in real time in the music playing process of a vehicle; performing frequency domain analysis on the PCM data corresponding to each audio frame, and extracting frequency information and loudness information; the method comprises the following steps: calculating frequency information and loudness information of a preset number of continuous audio frames, and respectively calculating a frequency change rate and a loudness change rate within a preset time; respectively comparing the frequency change rate and the loudness change rate with corresponding change rate thresholds; and if at least one of the frequency change rate and the loudness change rate exceeds the corresponding change rate threshold value, a corresponding light updating instruction is sent to the atmosphere lamp control module. The music rhythm function of the automotive interior atmosphere lamp can be enhanced, and the cooperative interaction ability of music and the atmosphere lamp is improved.
Owner:ANHUI KAIYANG TECHNOLOGY CO LTD +1

Human body gesture generation method and related equipment

The invention discloses a human body gesture generation method and related equipment, and relates to the technical field of computer vision, and the method comprises the steps: obtaining multi-modal input information, the multi-modal input information comprises a voice audio signal, text transcription information and a reference gesture sequence, the text transcription information comprises a semantic annotation, and the reference gesture sequence comprises a reference gesture sequence; the reference gesture sequence is a basic gesture template in the target scene; generating a multi-modal feature based on the multi-modal input information; performing time alignment processing operation on the multi-modal features to obtain alignment condition features; performing space-time decoupling modeling operation on the alignment condition features to obtain optimized gesture potential features; inputting the optimized gesture potential features into a diffusion model for iterative denoising to obtain denoised gesture potential features; and generating a target human body gesture sequence based on the de-noised gesture potential features.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Vehicle NVH comfort assessment method and device based on sound and vibration fusion and electronic equipment

The invention provides a vehicle NVH comfort assessment method and device based on sound and vibration fusion and electronic equipment. The method comprises the following steps: performing in-vehicle interference noise identification processing on window-level data in a candidate window set to obtain a candidate effective window set; calculating a sound and vibration consistency index between the in-vehicle audio signal and the in-vehicle vertical acceleration signal for the window-level data in the candidate effective window set, and screening an external noise dominant window to obtain an effective window set; respectively calculating an acoustic index and a vibration index based on the effective window set, and performing calibration compensation on the acoustic index and the vibration index based on the sensor calibration information; and performing fusion calculation based on the data volume information of the effective window set, the working condition coverage information, the uncertainty information corresponding to the sensor calibration information and the proportion information of the effective window set in the candidate window set to obtain a comprehensive confidence index. According to the method, the evaluation accuracy and stability can be improved, the cross-equipment comparability is enhanced, and the result credibility is improved.
Owner:CAR CONTROL (BEIJING) TECH CO LTD

Audio upmixing method and audio apparatus

An audio upmixing method and an audio apparatus are disclosed. The method comprises: performing feature extraction on a stereophonic audio signal to obtain a stereophonic audio feature; and extracting channel audio signals and right channel audio signals from the stereophonic audio feature based on audio output channels. The audio output channels are independent of each other, and each audio output channel is configured to output the corresponding left channel audio signal or the corresponding right channel audio signal. The method further comprises fusing the left channel audio signal and the right channel audio signal corresponding to a target channel to obtain first audio signals of a plurality of target channels. Each target channel corresponds to two audio output channels. The method further comprises outputting an audio upmixing signal of a target format based on the first audio signals. The target format corresponds to the plurality of target channels.
Owner:ANKER INNOVATIONS TECH CO LTD

Voice data recognition method and system based on AI voice algorithm

The invention discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem that the voice data recognition capability is low. The method comprises the following steps: S1, multi-mode cooperative triggering collection: synchronously collecting lip electromyographic signals and voiceprint features through a multi-mode sensor, an activation instruction is generated through feature fusion, and voice acquisition starting is triggered; s2, AI adaptive noise reduction processing: carrying out noise separation on the original audio signal by adopting a generative adversarial network, separating environmental noise features to generate a dynamic noise reduction mask, and keeping the integrity of human voice features; s3, beam dynamic optimization adjustment: analyzing real-time audio quality based on a reinforcement learning algorithm, dynamically adjusting beam pointing and gain parameters of a microphone array, and focusing a target sound source; and S4, semantic association cache enhancement: carrying out real-time semantic analysis on the collected voice data. According to the invention, the voice data recognition capability of an AI voice algorithm is greatly improved.
Owner:HUAQIAO UNIVERSITY

Real-time noise reduction method and system supporting Bluetooth audio interaction

The invention discloses a real-time noise reduction method and system supporting Bluetooth audio interaction, and the method comprises the steps: collecting an environment audio signal and a Bluetooth interaction audio signal, and carrying out the timestamp alignment and spatial calibration of the environment audio signal and the Bluetooth interaction audio signal; an improved multivariate variational mode decomposition algorithm is adopted to decompose and extract noise features, the noise features are combined with user historical noise data, a dynamic noise model is constructed through a long-short-term memory network, parameters are updated in real time, Bluetooth interaction audio content features and a user behavior data recognition scene are analyzed, and a noise reduction strategy is determined. And according to the dynamic noise model and the scene result, reinforcing learning to optimize filter parameters, generating a directional anti-noise signal to carry out noise reduction on the Bluetooth interaction audio signal, collecting the noise reduction intensity manually adjusted by a user, calculating voice definition and total harmonic distortion, and carrying out noise reduction on the Bluetooth interaction audio signal. And inputting a feedback and evaluation result into the modeling and noise reduction steps, carrying out dynamic range adjustment, sound channel balance and power amplification on the audio after noise reduction, and outputting the audio through a loudspeaker.
Owner:SHENZHEN GAOWEI COMM TECH CO LTD

Audio signal generation model and training method using generative adversarial network

A generative adversarial network-based audio signal generation model for generating a high quality audio signal may comprise: a generator generating an audio signal with an external input; a harmonic-percussive separation model separating the generated audio signal into a harmonic component signal and a percussive component signal; and at least one discriminator evaluating whether each of the harmonic component signal and the percussive component signal is real or fake.
Owner:ELECTRONICS & TELECOMM RES INST +1