Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

101 results about "Voice activity" patented technology

Voice activity detection is an essential component of many audio systems, such as automatic speech recognition and speaker recognition.

Low-delay streaming voice interaction system with interruption processing function

The invention provides a low-delay streaming voice interaction system with an interruption processing function, and relates to the technical field of artificial intelligence. Necessary preprocessing and acoustic feature extraction are performed through a real-time acoustic processing module, and distortion and complex environmental noise introduced by an interaction channel are resisted through a robustness enhancement technology; the streaming acoustic decoding module performs acoustic modeling, language model application and decoding in real time in parallel, and outputs an ultra-low-delay text transfer result stream; the real-time acoustic processing module is combined with a signal processing technology and is responsible for detecting user voice activity in a high-precision and ultra-low-delay manner, and particularly judging the real-time voice activity state of a user through the user voice activity in an AI voice playing period; an efficient and low-delay bidirectional streaming network transmission mode is adopted between modules of the system and between the modules and a communication platform, and it is ensured that audio streams, acoustic feature streams, text streams and control signals can be transmitted and processed in real time with extremely low end-to-end delay.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Noise reduction method combining different noise reduction algorithms of motorcycle riding earphone

The invention relates to the technical field of earphone noise reduction processing, and discloses a motorcycle riding earphone noise reduction method combining different noise reduction algorithms, comprising the following steps: acquiring an original audio signal, and distributing the original audio signal to each processing module; a preprocessing signal is generated, voice activity information is determined, and environmental noise energy information is calculated; according to the environmental noise energy or the auxiliary information, adaptively adjusting a high-low noise energy threshold value; the original signals are input into the traditional and AI noise reduction module in parallel to generate two paths of noise reduction signals; based on the environment noise, the voice information and the threshold value, performing weighted fusion to generate an output signal; and carrying out equalization and compression processing on the output signal, and driving the loudspeaker to output. According to the invention, comprehensive judgment on environmental noise energy information, voice activity information and riding speed is introduced, an output signal of a traditional noise reduction module is set to be zero in a high-speed voice-free environment, and AI noise reduction is combined, so that tone quality distortion of a traditional noise reduction algorithm in a noise environment is effectively avoided.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD +1

Real-time duplex translation method based on multi-channel parallel processing and corresponding product

The invention relates to the field of real-time translation, and provides a real-time duplex translation method based on multichannel parallel processing and a corresponding product, and the method comprises the steps: collecting multipath voice signals of at least two user groups in real time through a group of audio collection modules; dynamically adjusting beam forming parameters of each audio acquisition module in one group of audio acquisition modules based on a sound source positioning result, and feeding back the beam forming parameters to the corresponding audio acquisition modules; monitoring the voice activity of each audio acquisition module corresponding to each audio channel; when it is monitored that the voice activity of any audio channel reaches a preset condition, automatically activating the translation processing flow of the audio channel and keeping the monitoring state of the other audio channels; a parallel processing mechanism is adopted for voice signals of the activated audio channels, and meanwhile real-time translation of the currently activated audio channels and voice activity monitoring of the other audio channels are executed; and transmitting a translation result of the current speaking user to other users participating in dialogue in the user group to realize synchronous coordination of multichannel data.
Owner:MEIG SMART TECH CO LTD +1

Bluetooth mesh-based riding helmet earphone multi-terminal talkback synchronization method and system

The invention relates to the technical field of wireless communication, in particular to a Bluetooth mesh-based riding helmet earphone multi-terminal talkback synchronization method and system, and the method comprises the steps: constructing a self-organizing network through a plurality of terminal nodes, carrying out the error detection of the self-organizing network, obtaining an error value, carrying out the time compensation of the terminal nodes, and obtaining a calibration terminal node, carrying out voice activity monitoring on the plurality of calibration terminal nodes to obtain a plurality of marked terminal nodes, carrying out time slot allocation on the plurality of ordered voice requests to obtain a plurality of voice frames with timestamps, carrying out multi-hop forwarding on the plurality of voice frames with timestamps to obtain a plurality of verification voice frames, carrying out timestamp sorting to obtain a plurality of sorted voice frames, and sending the sorted voice frames to a server; and performing synchronous error correction to obtain a plurality of synchronous voice frames, decoding and playing the plurality of synchronous voice frames to obtain a synchronous voice stream, and realizing multi-terminal synchronous talkback. According to the invention, the problems of difficult network establishment, asynchronous voice and speaking right conflict in the application of the riding helmet Bluetooth earphone can be solved.
Owner:SHENZHEN WEIMAITONG ELECTRONIC TECH CO LTD

Care information generation system

To provide a care information generation system which recognizes required voice from voice of a care site and can generate care information.SOLUTION: A care information generation system 1000 includes a generative AI engine G and at least one of terminal devices 10A and 10B for generating care information from voice of a care site. The terminal devices A and B include microphones 11A and 11B for converting voice of the care site into a voice signal, recognize generation of voice activity data based on the voice signal and transmit the voice activity data to the generative AI engine. The voice activity data received from the terminal devices is inputted to the generative AI engine G and the generative AI engine G generates care information based on the voice activity data.SELECTED DRAWING: Figure 1
Owner:MAI SYSTEM PLANNING CO LTD

Hearing Device-Based Systems and Methods for Monitoring a Listening State of a User

An illustrative hearing system may be configured to receive, from an input transducer included in a hearing device configured to be worn by a user, audio data representative of one or more audio signals presented to the user and acquire motion data representative of head movements of the user while the user wears the hearing device and / or own-voice data representative of an own-voice activity of the user. The hearing system may be further configured to determine, based the motion data and / or own-voice data, a listening state of the user with respect to the one or more audio signals and perform, based on the listening state, an operation associated with the hearing device.
Owner:SONOVA AG

Intelligent sentence segmentation active speech detection method and device based on multi-state temporal modeling

This application discloses an intelligent method and apparatus for detecting active speech with sentence segmentation based on multi-state temporal modeling. The method includes: receiving audio signals from at least one channel; extracting acoustic feature sequences from the audio signals using a target speech recognition model corresponding to the number of channels; determining the probability distribution of each speech frame corresponding to the acoustic feature sequences belonging to different speech activity states, obtaining a state sequence corresponding to each channel, wherein the speech activity state includes at least one of the following: initial silence state, speech state, intra-turn pause silence state, and inter-turn sentence segmentation silence state; and determining the time of sentence segmentation in the audio signal based on the state sequence. This application solves the technical problem of erroneous sentence segmentation in speech activity detection based on a fixed silence threshold in related technologies.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Voice noise reduction method and device for multi-person scene, electronic equipment and storage medium

The invention discloses a voice noise reduction method and device for a multi-person scene, electronic equipment and a storage medium, and relates to the technical field of voice signal processing. The method comprises the following steps: acquiring an audio signal and a video image of a target space scene; determining audio perception information matched with the audio signal and a visual voice activity detection result corresponding to the video image; fusing the audio perception information and the visual voice activity detection result, and determining a target voice sounding object when the voice activity fusion result indicates that the voice activity exists; and positioning a human face corresponding to the target voice sounding object to obtain position change information, updating a pickup direction, controlling beam forming processing, and outputting a denoised target voice signal. According to the scheme provided by the invention, the voice of the current effective speaker can be accurately separated and enhanced from the aliasing audio signals, meanwhile, the interference of other speakers and environmental noise is inhibited, and help is provided for realizing high-quality voice interaction and recognition.
Owner:ANHUI JISHEN YINSAI TECHNOLOGY CO LTD

Multi-mode voice conversion method and system based on user behaviors

The invention relates to the technical field of speech recognition and synthesis, and discloses a multi-mode speech conversion method and system based on user behaviors, and the method comprises the steps: carrying out the vibration separation of a bone conduction signal, obtaining a separation vibration signal, mapping the face muscle micro-current into a muscle deformation gradient, and analyzing the intention intensity of a user through implicit behavior data; mapping the muscle deformation gradient into a voice fundamental frequency of the user, linearly converting the separation vibration signal into a voice formant of the user, and performing noise injection on a non-voice active section based on intention intensity to obtain noise injection voice; constructing user behavior voice corresponding to the voice fundamental frequency, the voice formant and the noise injection voice through an orthogonal projection layer in a preset orthogonal regularization vocoder; and synthesizing the speech feature vector and the user behavior speech into bone conduction propagation speech through a bone conduction synthesis layer in an orthogonal regularization vocoder. According to the invention, the accuracy of the multi-mode voice conversion technology can be improved.
Owner:SHENGZHEN BEIHAI RALL TRANSIT CENTURY TECHNOLOGY CO LTD

Multi-mode voice interaction method and device, intelligent equipment and readable storage medium

The invention provides a multi-mode voice interaction method and device and intelligent equipment, is suitable for the technical field of intelligent voice interaction, is applied to the intelligent equipment, and comprises the following steps: in response to a detected voice activity, extracting first voice data in the voice activity, and obtaining video data shot synchronously with the first voice data, a plurality of different users are shot in the video data. And screening out a target user sending the first voice data from the video data. And performing semantic integrity analysis on the text content corresponding to the first voice data. And when the semantic integrity analysis result is that the semantics is incomplete, continuously acquiring second voice data of the target user for a plurality of times, and generating a corresponding target statement with complete semantics after the first voice data and the second voice data are combined. Generating reply data according to the target statement, and outputting the reply data through voice. According to the embodiment of the invention, accurate, coherent and real-time voice interaction with the user needing interaction can be realized.
Owner:浙江人形机器人创新中心有限公司

Full-duplex intelligent voice interaction system and method based on voice activity detection and intention recognition

The invention discloses a full-duplex intelligent voice interaction system based on voice activity detection and intention recognition, and the system comprises a voice recognition module which is used for converting a voice stream into a text; the voice activity detection module is used for detecting voice activity; the intention recognition module is used for judging whether the user voice has a clear intention or not; the interruption control module is used for controlling pause and recovery of voice broadcast according to the voice activity detection result and the intention recognition result; the speech synthesis module is used for converting the text into a speech stream; the buffer area management module is used for managing a voice stream buffer area; and the resource management module is used for managing different tasks through the thread pool and processing voice recognition, voice activity detection and voice synthesis tasks in parallel. The invention further provides a full-duplex intelligent voice interaction method based on voice activity detection and intention recognition. Full-duplex voice interaction is realized; a mode of combining voice activity detection and intention recognition is adopted; and a high-concurrency scene is supported.
Owner:BEIJING HOLLYCRM TECH

Voice activity detection method and related equipment

The invention provides a voice activity detection method and related equipment. The method comprises the following steps: acquiring an audio signal to be processed, and determining a voice activity sequence of the audio signal based on a diffusion model; wherein the processing of the audio signal based on the diffusion model comprises the following steps: obtaining a time frame signal in a sequence form according to audio features, calculating a gradient of the time frame signal based on a neural network model, predicting a sampling signal according to the gradient, and obtaining a voice activity sequence for indicating whether a voice signal exists in each time frame based on the sampling signal. Compared with a related detection method, the method focuses on probability distribution of voice signals, is not affected by the signal-to-noise ratio, and can work normally under the extremely low signal-to-noise ratio; noise signals are not concerned, influence of noise types is avoided, and generalization performance is better. Therefore, an accurate voice activity detection result can be stably obtained.
Owner:HONOR DEVICE CO LTD

An electronic device and method for audio processing

PCT designated stageWO2026116709A1MicrophonesLoudspeakersVoice activitySpeech sound
A method for audio processing performed by an electronic device is provided. The method includes detecting at least one of a first voice activity near a first device and a second voice activity near a second device using a voice recognition module associated with the first device, comparing the first voice activity and the second voice activity to determine whether the first voice activity and the second voice activity exceed a predetermined threshold, and outputting the at least one of the first voice activity or the second voice activity through the second device in response to determining that the first voice activity and the second voice activity exceed the predetermined threshold.
Owner:SAMSUNG ELECTRONICS CO LTD

Real-time duplex translation method and corresponding product based on multi-channel parallel processing

The application relates to the field of real-time translation, and provides a real-time duplex translation method based on multi-channel parallel processing and a corresponding product.The method comprises the following steps: collecting multi-channel voice signals of at least two user groups in real time through a group of audio acquisition modules respectively; dynamically adjusting beam forming parameters of each audio acquisition module in the group of audio acquisition modules based on a sound source positioning result and feeding back to the corresponding audio acquisition module; monitoring voice activity of each audio channel corresponding to each audio acquisition module; when the voice activity of any audio channel reaches a predetermined condition, automatically activating a translation processing procedure of the audio channel and keeping the remaining audio channels in a monitoring state; using a parallel processing mechanism for the voice signals of the activated audio channel, simultaneously performing real-time translation of the currently activated audio channel and voice activity monitoring of the remaining audio channels; and transmitting the translation result of the current speaker to other users participating in the conversation in the user group, so that multi-channel data is synchronously coordinated.
Owner:MEIG SMART TECH CO LTD +1

A voice and visual interaction control method for safe driving

This invention discloses a safe driving voice and visual interaction control method. The method includes: synchronously collecting and preprocessing the driver's visual data and voice interaction data; extracting key visual features based on the preprocessed visual data and calculating visual state feature values; triggering standardized voice interaction based on the visual state feature values, calculating a voice activity score in conjunction with the preprocessed voice interaction data, and determining the driver's voice response delay level based on the voice activity score; matching based on a predefined 3D virtual guide action sequence according to the voice response delay level, triggering a linkage response after matching to form a non-intrusive driving reminder with light, sound, and shape linkage; and completing closed-loop feedback control based on the execution state of the 3D virtual guide action sequence and the non-intrusive driving reminder with light, sound, and shape linkage. The method provided by this invention can reduce the monitoring misjudgment rate and achieve non-intrusive safety reminders.
Owner:SHANGHAI CHANGXING SOFTWARE CO LTD

A voice data processing method and device, electronic equipment and storage medium

The present disclosure discloses a speech data processing method and device, electronic equipment and storage medium. The method comprises: obtaining target speech data to be processed; detecting the target speech data to obtain target phonemes corresponding to each audio frame; determining the target type of the audio frame corresponding to the target phoneme based on the phoneme type corresponding to the target phoneme, wherein the target type includes a silence type or a non-silence type; determining the voice activity boundary in the target speech data based on the target type corresponding to the audio frame, and taking the voice activity boundary as the detection result of the target speech data. The method provided by the present disclosure can more accurately detect the type of audio frame by identifying the phonemes of each audio frame in the speech data and determining the silence type audio frame and the non-silence type audio frame using the phonemes, and can more accurately locate the voice activity boundary in the speech data compared with the existing technology using the two-classification method.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Audio system and method for voice activity detection

Audio systems, methods, and processor instructions are provided herein that detect voice activity of a user and provide an output voice signal. The systems, methods, and instructions receive a plurality of microphone signals and combine the plurality of microphone signals according to a first combination and a second combination. The first combination produces a primary signal having an enhanced response in a direction of a mouth of the user, and the second combination produces a reference signal having a reduced response in the direction of the mouth of the user. The primary signal and the reference signal are added and subtracted to produce a sum signal and a difference signal, respectively. The sum signal is compared to the difference signal and an output voice signal is provided based on the comparison.
Owner:BOSE CORP

Electronic device and method for audio processing

A method for audio processing performed by an electronic device is provided. The method includes detecting at least one of a first voice activity near a first device and a second voice activity near a second device using a voice recognition module associated with the first device, comparing the first voice activity and the second voice activity to determine whether the first voice activity and the second voice activity exceed a predetermined threshold, and outputting the at least one of the first voice activity or the second voice activity through the second device in response to determining that the first voice activity and the second voice activity exceed the predetermined threshold.
Owner:SAMSUNG ELECTRONICS CO LTD

Intelligent outbound robot dialogue regulation and control system fused with multi-mode emotion recognition

The invention discloses an intelligent outbound robot dialogue regulation and control system fused with multi-modal emotion recognition, and relates to the technical field of artificial intelligence, the intelligent outbound robot dialogue regulation and control system is composed of an audio sensing and structuring module, a multi-modal emotion feature deconstruction module and a dynamic emotion interaction regulation and control module; the method comprises the following steps: firstly, recording target human voice in real time by using recording equipment, judging a voice activity frame through short-time energy and a zero-crossing rate, and accurately segmenting the voice activity frame into a voice segment sequence; secondly, extracting speech speed and loudness features from each segment of audio, and constructing an emotion processing feature sequence in combination with normalization processing; and finally, based on comparison of the emotional feature sequence and a preset threshold table, the emotional state of the target person is judged in real time, dialogue emotional regulation and control parameters of the outbound robot are automatically adjusted according to a dynamic regulation and control strategy, and intelligent and humanized outbound dialogue regulation and control are achieved.
Owner:BEIJING HAOFENG CHUANGYUAN TECH CO LTD

Motorcycle riding earphone different noise reduction algorithm combination noise reduction method

The application relates to the technical field of earphone noise reduction processing, and discloses a motorcycle riding earphone different noise reduction algorithm combination noise reduction method, which comprises the following steps: obtaining an original audio signal and distributing the original audio signal to each processing module; generating a pretreatment signal, determining speech activity information, and calculating environmental noise energy information; adaptively adjusting high and low noise energy thresholds according to the environmental noise energy or auxiliary information; inputting the original signal into traditional and AI noise reduction modules in parallel to generate two-way noise reduction signals; generating an output signal through weighted fusion based on the environmental noise, speech information and thresholds; and performing equalization and compression processing on the output signal and driving a loudspeaker to output. The application introduces comprehensive judgment on environmental noise energy information, speech activity information and riding speed, the output signal of the traditional noise reduction module is set to zero in a high-speed non-speech environment, and AI noise reduction is combined, so that the sound quality distortion of the traditional noise reduction algorithm in a noise environment is effectively avoided.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD +1

Audio control for extended-reality shared space

Methods, systems, computer-readable media, and apparatuses for audio signal processing are presented. Some configurations include determining that first audio activity in at least one microphone signal is voice activity; determining whether the voice activity is voice activity of a participant in an application session active on a device; based at least on a result of the determining whether the voice activity is voice activity of a participant in the application session, generating an antinoise signal to cancel the first audio activity; and by a loudspeaker, producing an acoustic signal that is based on the antinoise signal. Applications relating to shared virtual spaces are described.
Owner:QUALCOMM INC

End-to-end speech diarization via iterative speaker embedding

A method includes receiving an input audio signal corresponding to utterances spoken by multiple speakers. The method also includes encoding the input audio signal into a sequence of T temporal embeddings. During each of a plurality of iterations each corresponding to a respective speaker of the multiple speakers, the method includes selecting a respective speaker embedding for the respective speaker by determining a probability that the corresponding temporal embedding includes a presence of voice activity by a single new speaker for which a speaker embedding was not previously selected during a previous iteration and selecting the respective speaker embedding for the respective speaker as the temporal embedding. The method also includes, at each time step, predicting a respective voice activity indicator for each respective speaker of the multiple speakers based on the respective speaker embeddings selected during the plurality of iterations and the temporal embedding.
Owner:GOOGLE LLC

Audio processing method, training method of audio processing model and electronic equipment

The invention discloses an audio processing method, an audio processing model training method and an electronic device, and relates to the technical field of audio processing, the method is applied to a first electronic device, and the method comprises the following steps: the first electronic device performs feature extraction on an input audio, and obtains a voiceprint feature sequence corresponding to the input audio; and the first electronic equipment performs voiceprint extraction on the voiceprint feature sequence, and identifies the voiceprint of the speaker corresponding to the voiceprint feature sequence. And the first electronic equipment takes the voiceprint of the speaker as auxiliary query, performs voiceprint clustering on the voiceprint feature sequence, and obtains voice activity information of the speaker corresponding to the voiceprint feature sequence. In the application, the first electronic device can update the voiceprint library by using the voiceprint information output in real time, and does not need to actively register the voiceprint information of a speaker. Moreover, the voiceprint of the speaker is used as auxiliary query and is applied to the voiceprint clustering part, so that the whole audio processing performance and the speaker recognition accuracy can be improved.
Owner:HONOR DEVICE CO LTD +1

Voice activity detection apparatus, electronic device, and voice activity detection method

The present disclosure relates to a voice activity detection apparatus, an electronic device and a voice activity detection method. The voice activity detection apparatus comprises a gain and filter module configured to band-pass filter a signal to be detected so that the signal to be detected is within a pre-set passband, wherein a passband gain of the gain and filter module is adjustable; a comparison module communicatively connected with the gain and filter module, and the comparison module is configured to compare an intensity of the signal to be detected from the gain and filter module with a pre-set intensity threshold to generate a comparison signal for detecting voice activity; and a gain adjustment module communicatively connected with the gain and filter module, and the gain adjustment module is configured to generate a first gain adjustment signal according to the signal to be detected from the gain and filter module, wherein the passband gain of the gain and filter module is configured to be adjusted according to the first gain adjustment signal.
Owner:SHANGHAI PANSILICON SEMICONDUCTOR TECHNOLOGY CO LTD

Methods and devices for encoding and / or decoding spatial background noise within a multi-channel input signal

The present document describes a method (600) for encoding a multi-channel input signal (101) which comprises N different channels. The method (600) comprises, for a current frame of a sequence of frames, determining (601) whether the current frame is an active frame or an inactive frame using a signal and / or a voice activity detector, and determining (602) a downmix signal (103) based on the multi-channel input signal (101), wherein the downmix signal (103) comprises N channels or less. In addition, the method (600) comprises determining (603) upmixing metadata (105) comprising a set of parameters for generating, based on the downmix signal (103), a reconstructed multi-channel signal (111) comprising N channels, wherein the upmixing metadata (105) is determined in dependance of whether the current frame is an active frame or an inactive frame. The method (600) further comprises encoding (604) the upmixing metadata (105) into a bitstream.
Owner:DOLBY LABORATORIES LICENSING CORP

A low-delay streaming voice interaction system with break handling functionality

The application provides a low-delay streaming voice interaction system with a breaking processing function, and relates to the technical field of artificial intelligence.The application carries out necessary preprocessing and acoustic feature extraction through a real-time acoustic processing module, and uses robustness enhancement technology to resist distortion and complex environmental noise introduced by an interaction channel.A streaming acoustic decoding module carries out acoustic modeling, language model application and decoding in real time and in parallel, and outputs an ultralow-delay text transcription result stream.The real-time acoustic processing module is responsible for high-precision and ultralow-delay detection of user voice activity, especially user voice activity during AI playback of voice, to determine the real-time voice activity state of the user.The system uses an efficient and low-delay bidirectional streaming network transmission mode between modules and between the system and a communication platform, so that audio streams, acoustic feature streams, text streams and control signals can be transmitted and processed in real time with extremely low end-to-end delay.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Directional activity mask detector for a vehicle

PendingUS20250259638A1Speech analysisTarget signalRelative transfer function
A method for a directional activity mask detector for a vehicle includes generating a blocking matrix based on pre-recorded signals from a target zone, receiving, at a voice activity detector, audio frames from a microphone array, and applying the blocking matrix to one or more zones within a vehicle. The method also includes detecting signals from unblocked zones of the vehicle, determining an activity of a target signal based on the detected signals from the unblocked zones, and estimating, by a beamformer, a relative transfer function (RTF) vector based on the received audio frames and the determined activity of the target signal.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Intelligent sentence segmentation activity voice detection method and device based on multi-state time sequence modeling

The invention discloses an intelligent sentence segmentation activity voice detection method and device based on multi-state time sequence modeling. The method comprises the following steps: receiving an audio signal of at least one channel; extracting an acoustic feature sequence of the audio signal by adopting a target speech recognition model corresponding to the channel number; probability distribution that each voice frame corresponding to the acoustic feature sequence belongs to different voice activity states is determined, a state sequence corresponding to each channel is obtained, and the voice activity states comprise at least one of the following states: an initial mute state, a voice state, an in-speech-round pause mute state and a speech round discontinuous sentence mute state; and determining the sentence segmentation time in the audio signal according to the state sequence. According to the method and the device, the technical problem that wrong sentence segmentation exists in voice activity detection based on a fixed silence threshold in related technologies is solved.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Adaptive comfort noise parameter determination

PendingUS20250299683A1Speech analysisComfort noiseNoise
A method for generating a comfort noise (CN) parameter is provided. The method includes receiving an audio input; detecting, with a Voice Activity Detector (VAD), a current inactive segment in the audio input; as a result of detecting, with the VAD, the current inactive segment in the audio input, calculating a CN parameter CNused; and providing the CN parameter CNused to a decoder. The CN parameter CNused is calculated based at least in part on the current inactive segment and a previous inactive segment.
Owner:TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)