Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

73 results about "Voice activity" patented technology

Voice activity detection is an essential component of many audio systems, such as automatic speech recognition and speaker recognition.

Noise reduction method combining different noise reduction algorithms of motorcycle riding earphone

The invention relates to the technical field of earphone noise reduction processing, and discloses a motorcycle riding earphone noise reduction method combining different noise reduction algorithms, comprising the following steps: acquiring an original audio signal, and distributing the original audio signal to each processing module; a preprocessing signal is generated, voice activity information is determined, and environmental noise energy information is calculated; according to the environmental noise energy or the auxiliary information, adaptively adjusting a high-low noise energy threshold value; the original signals are input into the traditional and AI noise reduction module in parallel to generate two paths of noise reduction signals; based on the environment noise, the voice information and the threshold value, performing weighted fusion to generate an output signal; and carrying out equalization and compression processing on the output signal, and driving the loudspeaker to output. According to the invention, comprehensive judgment on environmental noise energy information, voice activity information and riding speed is introduced, an output signal of a traditional noise reduction module is set to be zero in a high-speed voice-free environment, and AI noise reduction is combined, so that tone quality distortion of a traditional noise reduction algorithm in a noise environment is effectively avoided.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD +1

Real-time duplex translation method based on multi-channel parallel processing and corresponding product

The invention relates to the field of real-time translation, and provides a real-time duplex translation method based on multichannel parallel processing and a corresponding product, and the method comprises the steps: collecting multipath voice signals of at least two user groups in real time through a group of audio collection modules; dynamically adjusting beam forming parameters of each audio acquisition module in one group of audio acquisition modules based on a sound source positioning result, and feeding back the beam forming parameters to the corresponding audio acquisition modules; monitoring the voice activity of each audio acquisition module corresponding to each audio channel; when it is monitored that the voice activity of any audio channel reaches a preset condition, automatically activating the translation processing flow of the audio channel and keeping the monitoring state of the other audio channels; a parallel processing mechanism is adopted for voice signals of the activated audio channels, and meanwhile real-time translation of the currently activated audio channels and voice activity monitoring of the other audio channels are executed; and transmitting a translation result of the current speaking user to other users participating in dialogue in the user group to realize synchronous coordination of multichannel data.
Owner:MEIG SMART TECH CO LTD +1

Bluetooth mesh-based riding helmet earphone multi-terminal talkback synchronization method and system

The invention relates to the technical field of wireless communication, in particular to a Bluetooth mesh-based riding helmet earphone multi-terminal talkback synchronization method and system, and the method comprises the steps: constructing a self-organizing network through a plurality of terminal nodes, carrying out the error detection of the self-organizing network, obtaining an error value, carrying out the time compensation of the terminal nodes, and obtaining a calibration terminal node, carrying out voice activity monitoring on the plurality of calibration terminal nodes to obtain a plurality of marked terminal nodes, carrying out time slot allocation on the plurality of ordered voice requests to obtain a plurality of voice frames with timestamps, carrying out multi-hop forwarding on the plurality of voice frames with timestamps to obtain a plurality of verification voice frames, carrying out timestamp sorting to obtain a plurality of sorted voice frames, and sending the sorted voice frames to a server; and performing synchronous error correction to obtain a plurality of synchronous voice frames, decoding and playing the plurality of synchronous voice frames to obtain a synchronous voice stream, and realizing multi-terminal synchronous talkback. According to the invention, the problems of difficult network establishment, asynchronous voice and speaking right conflict in the application of the riding helmet Bluetooth earphone can be solved.
Owner:SHENZHEN WEIMAITONG ELECTRONIC TECH CO LTD

Intelligent sentence segmentation active speech detection method and device based on multi-state temporal modeling

This application discloses an intelligent method and apparatus for detecting active speech with sentence segmentation based on multi-state temporal modeling. The method includes: receiving audio signals from at least one channel; extracting acoustic feature sequences from the audio signals using a target speech recognition model corresponding to the number of channels; determining the probability distribution of each speech frame corresponding to the acoustic feature sequences belonging to different speech activity states, obtaining a state sequence corresponding to each channel, wherein the speech activity state includes at least one of the following: initial silence state, speech state, intra-turn pause silence state, and inter-turn sentence segmentation silence state; and determining the time of sentence segmentation in the audio signal based on the state sequence. This application solves the technical problem of erroneous sentence segmentation in speech activity detection based on a fixed silence threshold in related technologies.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Voice noise reduction method and device for multi-person scene, electronic equipment and storage medium

The invention discloses a voice noise reduction method and device for a multi-person scene, electronic equipment and a storage medium, and relates to the technical field of voice signal processing. The method comprises the following steps: acquiring an audio signal and a video image of a target space scene; determining audio perception information matched with the audio signal and a visual voice activity detection result corresponding to the video image; fusing the audio perception information and the visual voice activity detection result, and determining a target voice sounding object when the voice activity fusion result indicates that the voice activity exists; and positioning a human face corresponding to the target voice sounding object to obtain position change information, updating a pickup direction, controlling beam forming processing, and outputting a denoised target voice signal. According to the scheme provided by the invention, the voice of the current effective speaker can be accurately separated and enhanced from the aliasing audio signals, meanwhile, the interference of other speakers and environmental noise is inhibited, and help is provided for realizing high-quality voice interaction and recognition.
Owner:ANHUI JISHEN YINSAI TECHNOLOGY CO LTD

Multi-mode voice interaction method and device, intelligent equipment and readable storage medium

The invention provides a multi-mode voice interaction method and device and intelligent equipment, is suitable for the technical field of intelligent voice interaction, is applied to the intelligent equipment, and comprises the following steps: in response to a detected voice activity, extracting first voice data in the voice activity, and obtaining video data shot synchronously with the first voice data, a plurality of different users are shot in the video data. And screening out a target user sending the first voice data from the video data. And performing semantic integrity analysis on the text content corresponding to the first voice data. And when the semantic integrity analysis result is that the semantics is incomplete, continuously acquiring second voice data of the target user for a plurality of times, and generating a corresponding target statement with complete semantics after the first voice data and the second voice data are combined. Generating reply data according to the target statement, and outputting the reply data through voice. According to the embodiment of the invention, accurate, coherent and real-time voice interaction with the user needing interaction can be realized.
Owner:浙江人形机器人创新中心有限公司

Full-duplex intelligent voice interaction system and method based on voice activity detection and intention recognition

The invention discloses a full-duplex intelligent voice interaction system based on voice activity detection and intention recognition, and the system comprises a voice recognition module which is used for converting a voice stream into a text; the voice activity detection module is used for detecting voice activity; the intention recognition module is used for judging whether the user voice has a clear intention or not; the interruption control module is used for controlling pause and recovery of voice broadcast according to the voice activity detection result and the intention recognition result; the speech synthesis module is used for converting the text into a speech stream; the buffer area management module is used for managing a voice stream buffer area; and the resource management module is used for managing different tasks through the thread pool and processing voice recognition, voice activity detection and voice synthesis tasks in parallel. The invention further provides a full-duplex intelligent voice interaction method based on voice activity detection and intention recognition. Full-duplex voice interaction is realized; a mode of combining voice activity detection and intention recognition is adopted; and a high-concurrency scene is supported.
Owner:BEIJING HOLLYCRM TECH

An electronic device and method for audio processing

PCT designated stageWO2026116709A1MicrophonesLoudspeakersVoice activitySpeech sound
A method for audio processing performed by an electronic device is provided. The method includes detecting at least one of a first voice activity near a first device and a second voice activity near a second device using a voice recognition module associated with the first device, comparing the first voice activity and the second voice activity to determine whether the first voice activity and the second voice activity exceed a predetermined threshold, and outputting the at least one of the first voice activity or the second voice activity through the second device in response to determining that the first voice activity and the second voice activity exceed the predetermined threshold.
Owner:SAMSUNG ELECTRONICS CO LTD

Real-time duplex translation method and corresponding product based on multi-channel parallel processing

The application relates to the field of real-time translation, and provides a real-time duplex translation method based on multi-channel parallel processing and a corresponding product.The method comprises the following steps: collecting multi-channel voice signals of at least two user groups in real time through a group of audio acquisition modules respectively; dynamically adjusting beam forming parameters of each audio acquisition module in the group of audio acquisition modules based on a sound source positioning result and feeding back to the corresponding audio acquisition module; monitoring voice activity of each audio channel corresponding to each audio acquisition module; when the voice activity of any audio channel reaches a predetermined condition, automatically activating a translation processing procedure of the audio channel and keeping the remaining audio channels in a monitoring state; using a parallel processing mechanism for the voice signals of the activated audio channel, simultaneously performing real-time translation of the currently activated audio channel and voice activity monitoring of the remaining audio channels; and transmitting the translation result of the current speaker to other users participating in the conversation in the user group, so that multi-channel data is synchronously coordinated.
Owner:MEIG SMART TECH CO LTD +1

A voice and visual interaction control method for safe driving

This invention discloses a safe driving voice and visual interaction control method. The method includes: synchronously collecting and preprocessing the driver's visual data and voice interaction data; extracting key visual features based on the preprocessed visual data and calculating visual state feature values; triggering standardized voice interaction based on the visual state feature values, calculating a voice activity score in conjunction with the preprocessed voice interaction data, and determining the driver's voice response delay level based on the voice activity score; matching based on a predefined 3D virtual guide action sequence according to the voice response delay level, triggering a linkage response after matching to form a non-intrusive driving reminder with light, sound, and shape linkage; and completing closed-loop feedback control based on the execution state of the 3D virtual guide action sequence and the non-intrusive driving reminder with light, sound, and shape linkage. The method provided by this invention can reduce the monitoring misjudgment rate and achieve non-intrusive safety reminders.
Owner:SHANGHAI CHANGXING SOFTWARE CO LTD

A voice data processing method and device, electronic equipment and storage medium

The present disclosure discloses a speech data processing method and device, electronic equipment and storage medium. The method comprises: obtaining target speech data to be processed; detecting the target speech data to obtain target phonemes corresponding to each audio frame; determining the target type of the audio frame corresponding to the target phoneme based on the phoneme type corresponding to the target phoneme, wherein the target type includes a silence type or a non-silence type; determining the voice activity boundary in the target speech data based on the target type corresponding to the audio frame, and taking the voice activity boundary as the detection result of the target speech data. The method provided by the present disclosure can more accurately detect the type of audio frame by identifying the phonemes of each audio frame in the speech data and determining the silence type audio frame and the non-silence type audio frame using the phonemes, and can more accurately locate the voice activity boundary in the speech data compared with the existing technology using the two-classification method.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Audio system and method for voice activity detection

Audio systems, methods, and processor instructions are provided herein that detect voice activity of a user and provide an output voice signal. The systems, methods, and instructions receive a plurality of microphone signals and combine the plurality of microphone signals according to a first combination and a second combination. The first combination produces a primary signal having an enhanced response in a direction of a mouth of the user, and the second combination produces a reference signal having a reduced response in the direction of the mouth of the user. The primary signal and the reference signal are added and subtracted to produce a sum signal and a difference signal, respectively. The sum signal is compared to the difference signal and an output voice signal is provided based on the comparison.
Owner:BOSE CORP

Electronic device and method for audio processing

A method for audio processing performed by an electronic device is provided. The method includes detecting at least one of a first voice activity near a first device and a second voice activity near a second device using a voice recognition module associated with the first device, comparing the first voice activity and the second voice activity to determine whether the first voice activity and the second voice activity exceed a predetermined threshold, and outputting the at least one of the first voice activity or the second voice activity through the second device in response to determining that the first voice activity and the second voice activity exceed the predetermined threshold.
Owner:SAMSUNG ELECTRONICS CO LTD

Intelligent outbound robot dialogue regulation and control system fused with multi-mode emotion recognition

The invention discloses an intelligent outbound robot dialogue regulation and control system fused with multi-modal emotion recognition, and relates to the technical field of artificial intelligence, the intelligent outbound robot dialogue regulation and control system is composed of an audio sensing and structuring module, a multi-modal emotion feature deconstruction module and a dynamic emotion interaction regulation and control module; the method comprises the following steps: firstly, recording target human voice in real time by using recording equipment, judging a voice activity frame through short-time energy and a zero-crossing rate, and accurately segmenting the voice activity frame into a voice segment sequence; secondly, extracting speech speed and loudness features from each segment of audio, and constructing an emotion processing feature sequence in combination with normalization processing; and finally, based on comparison of the emotional feature sequence and a preset threshold table, the emotional state of the target person is judged in real time, dialogue emotional regulation and control parameters of the outbound robot are automatically adjusted according to a dynamic regulation and control strategy, and intelligent and humanized outbound dialogue regulation and control are achieved.
Owner:BEIJING HAOFENG CHUANGYUAN TECH CO LTD

Motorcycle riding earphone different noise reduction algorithm combination noise reduction method

The application relates to the technical field of earphone noise reduction processing, and discloses a motorcycle riding earphone different noise reduction algorithm combination noise reduction method, which comprises the following steps: obtaining an original audio signal and distributing the original audio signal to each processing module; generating a pretreatment signal, determining speech activity information, and calculating environmental noise energy information; adaptively adjusting high and low noise energy thresholds according to the environmental noise energy or auxiliary information; inputting the original signal into traditional and AI noise reduction modules in parallel to generate two-way noise reduction signals; generating an output signal through weighted fusion based on the environmental noise, speech information and thresholds; and performing equalization and compression processing on the output signal and driving a loudspeaker to output. The application introduces comprehensive judgment on environmental noise energy information, speech activity information and riding speed, the output signal of the traditional noise reduction module is set to zero in a high-speed non-speech environment, and AI noise reduction is combined, so that the sound quality distortion of the traditional noise reduction algorithm in a noise environment is effectively avoided.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD +1

End-to-end speech diarization via iterative speaker embedding

A method includes receiving an input audio signal corresponding to utterances spoken by multiple speakers. The method also includes encoding the input audio signal into a sequence of T temporal embeddings. During each of a plurality of iterations each corresponding to a respective speaker of the multiple speakers, the method includes selecting a respective speaker embedding for the respective speaker by determining a probability that the corresponding temporal embedding includes a presence of voice activity by a single new speaker for which a speaker embedding was not previously selected during a previous iteration and selecting the respective speaker embedding for the respective speaker as the temporal embedding. The method also includes, at each time step, predicting a respective voice activity indicator for each respective speaker of the multiple speakers based on the respective speaker embeddings selected during the plurality of iterations and the temporal embedding.
Owner:GOOGLE LLC

Audio processing method, training method of audio processing model and electronic equipment

The invention discloses an audio processing method, an audio processing model training method and an electronic device, and relates to the technical field of audio processing, the method is applied to a first electronic device, and the method comprises the following steps: the first electronic device performs feature extraction on an input audio, and obtains a voiceprint feature sequence corresponding to the input audio; and the first electronic equipment performs voiceprint extraction on the voiceprint feature sequence, and identifies the voiceprint of the speaker corresponding to the voiceprint feature sequence. And the first electronic equipment takes the voiceprint of the speaker as auxiliary query, performs voiceprint clustering on the voiceprint feature sequence, and obtains voice activity information of the speaker corresponding to the voiceprint feature sequence. In the application, the first electronic device can update the voiceprint library by using the voiceprint information output in real time, and does not need to actively register the voiceprint information of a speaker. Moreover, the voiceprint of the speaker is used as auxiliary query and is applied to the voiceprint clustering part, so that the whole audio processing performance and the speaker recognition accuracy can be improved.
Owner:HONOR DEVICE CO LTD +1

Voice activity detection apparatus, electronic device, and voice activity detection method

The present disclosure relates to a voice activity detection apparatus, an electronic device and a voice activity detection method. The voice activity detection apparatus comprises a gain and filter module configured to band-pass filter a signal to be detected so that the signal to be detected is within a pre-set passband, wherein a passband gain of the gain and filter module is adjustable; a comparison module communicatively connected with the gain and filter module, and the comparison module is configured to compare an intensity of the signal to be detected from the gain and filter module with a pre-set intensity threshold to generate a comparison signal for detecting voice activity; and a gain adjustment module communicatively connected with the gain and filter module, and the gain adjustment module is configured to generate a first gain adjustment signal according to the signal to be detected from the gain and filter module, wherein the passband gain of the gain and filter module is configured to be adjusted according to the first gain adjustment signal.
Owner:SHANGHAI PANSILICON SEMICONDUCTOR TECHNOLOGY CO LTD

Methods and devices for encoding and / or decoding spatial background noise within a multi-channel input signal

The present document describes a method (600) for encoding a multi-channel input signal (101) which comprises N different channels. The method (600) comprises, for a current frame of a sequence of frames, determining (601) whether the current frame is an active frame or an inactive frame using a signal and / or a voice activity detector, and determining (602) a downmix signal (103) based on the multi-channel input signal (101), wherein the downmix signal (103) comprises N channels or less. In addition, the method (600) comprises determining (603) upmixing metadata (105) comprising a set of parameters for generating, based on the downmix signal (103), a reconstructed multi-channel signal (111) comprising N channels, wherein the upmixing metadata (105) is determined in dependance of whether the current frame is an active frame or an inactive frame. The method (600) further comprises encoding (604) the upmixing metadata (105) into a bitstream.
Owner:DOLBY LABORATORIES LICENSING CORP

A low-delay streaming voice interaction system with break handling functionality

The application provides a low-delay streaming voice interaction system with a breaking processing function, and relates to the technical field of artificial intelligence.The application carries out necessary preprocessing and acoustic feature extraction through a real-time acoustic processing module, and uses robustness enhancement technology to resist distortion and complex environmental noise introduced by an interaction channel.A streaming acoustic decoding module carries out acoustic modeling, language model application and decoding in real time and in parallel, and outputs an ultralow-delay text transcription result stream.The real-time acoustic processing module is responsible for high-precision and ultralow-delay detection of user voice activity, especially user voice activity during AI playback of voice, to determine the real-time voice activity state of the user.The system uses an efficient and low-delay bidirectional streaming network transmission mode between modules and between the system and a communication platform, so that audio streams, acoustic feature streams, text streams and control signals can be transmitted and processed in real time with extremely low end-to-end delay.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Speech recognition method and apparatus, storage medium, and electronic device

A speech recognition method, device, storage medium and electronic equipment. The method comprises: obtaining speech information to be recognized, the speech information to be recognized comprising speech information of at least two channels; performing speech activity segment detection on the speech information to be recognized to obtain a plurality of speech activity segments corresponding to each channel; determining noise information of each channel based on cross information of the speech activity segments corresponding to different channels in the time dimension; and performing speech recognition on the speech information of the corresponding channel according to the noise information of each channel to obtain a speech recognition result corresponding to the speech information to be recognized. The method can greatly improve the accuracy of speech recognition.
Owner:IFLYTEK CO LTD

Robust voice activity detector system for use with an earphone

An electronic device or method for adjusting a gain on a voice operated control system can include one or more processors and a memory having computer instructions. The instructions, when executed by the one or more processors causes the one or more processors to perform the operations of receiving a first microphone signal, receiving a second microphone signal, updating a slow time weighted ratio of the filtered first and second signals, and updating a fast time weighted ratio of the filtered first and second signals. The one or more processors can further perform the operations of calculating an absolute difference between the fast time weighted ratio and the slow time weighted ratio, comparing the absolute difference with a threshold, and increasing the gain when the absolute difference is greater than the threshold. Other embodiments are disclosed.
Owner:ST PORTFOLIO HLDG LLC

Method for determining an activity of an intrinsic voice of a user of a hearing device, hearing device, and hearing device system

A method for detecting activity of the own voice of a wearer of a hearing device by way of a signal processing apparatus of the hearing device. A first input signal is generated by a first input transducer, and a second input signal is generated by a second input transducer. The two input signals are supplied to a detection unit of the signal processing apparatus, which has a neural network and an input stage, which is connected in front of the neural network. Information signals are generated by the input stage on the basis of the two input signals and the information signals are evaluated by the neural network. A detection result is output by the detection unit based on the evaluation of the information signals by the neural network.
Owner:SIVANTOS PTE LTD

Automated multi-speaker and multi-lingual speech analysis

PCT designated stageWO2026142921A1Semantic vectorSystems analysis
Exemplary system and methods use a combination of application modules and neural network architecture for multi-speaker and multi-language speech analysis. The exemplary system can receive a natural language input, which it decomposes into plural segments. A sub-group of the plural segments are accumulated in a buffer where each segment representing a period during which voice activity is detected. The sub-groups are analyzed for voice activity of multiple speakers and one or more text segments are generated based on the speakers. A semantic vector for each text segment is generated and stored in vector memory. Relevant data associated with each semantic vector is retrieved from the vector memory based on a similarity measure; and a response including specified information extracted from the one or more text segments is generated based on at least the relevant data.
Owner:ERESTECH

Exercise teaching method and device based on split-screen display, electronic equipment and product

The invention relates to an exercise teaching method based on split-screen display. The exercise teaching method comprises the following steps: extracting a standard action vector sequence and a voice activity intermittent interval in a teaching video; in response to the received split-screen instruction, the teaching video and the user video collected in real time are displayed on a display interface of the terminal equipment in a split-screen mode at the same time, and the user video comprises a user image and skeleton key points overlapped on the user image; performing normalization processing on the skeleton key point data, and extracting an action vector sequence according to body groups; aligning the action vector sequence of each body group of the user with a standard action vector sequence, and calculating an action score of each body group; and in the voice activity intermittent interval, prompting information related to the user action is played based on the action score so as to realize real-time voice feedback to the user. According to the method, the teaching video and the real-time picture of the user are displayed on the same interface in a split-screen manner, so that the intuition and immersion of practice are remarkably improved.
Owner:BEIJING XIAOTANG TECH CO LTD

Speech enhancement method and device based on double microphone array, equipment and medium

PendingCN122392555ABeam directionNoise
The specification provides a speech enhancement method based on a dual microphone array, the method comprising: acquiring surrounding acoustic signals collected by two microphones in a microphone array respectively. Based on the acoustic signals collected by the two microphones respectively, coherence information between the two microphones is determined. Based on the coherence information, a speech activity state is judged, and in the case that the speech activity state is a speech missing state, a noise covariance matrix of the spatial filter is updated; and a gain function value of the post-filter is determined based on the coherence information. The two collected acoustic signals are subjected to beamforming processing by the spatial filter to obtain acoustic signals of at least one beam direction. The acoustic signals of any beam direction are subjected to filtering processing by the post-filter to obtain an estimation result of a speech component contained in the acoustic signals of the any beam direction.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Context-aware hardware-based voice activity detection

Certain aspects of the present disclosure provide a method for performing voice activity detection, comprising: receiving audio data from an audio source of an electronic device; generating a plurality of model input features based on the received audio data using a hardware-based feature generator; providing the plurality of model input features to a hardware-based voice activity detection model; receiving an output value from the hardware-based voice activity detection model; and determining a presence of voice activity in the audio data based on the output value.
Owner:QUALCOMM INC

Voice fraud analysis method, device and equipment and storage medium

The application relates to a voice fraud analysis method, which comprises the following steps: acquiring a voice signal, collecting voice data of different sources through two independently configured microphone arrays, and recording the voice data into two independent and isolated audio channels; detecting the voice data in the two audio channels through a voice activity detection technology, identifying and marking a voice activity time period in each audio channel, performing a fragmentation operation on the voice data based on the voice activity time period, and extracting a plurality of effective voice segments; generating a dialogue text according to the extracted effective voice segments, and screening a text part to be analyzed in the generated text; selecting a pre-trained fraud analysis model according to the data characteristics of the text to be analyzed, and determining a fraud probability of the voice data through the analysis model. The application can realize accurate processing and analysis of voice signals, significantly improve the quality of voice data and the accuracy of fraud detection, and is suitable for complex voice interaction scenes.
Owner:PING AN TECH (SHENZHEN) CO LTD