Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Stuttering" patented technology

Stuttering is a speech disorder characterized by repetition of sounds, syllables, or words; prolongation of sounds; and interruptions in speech known as blocks.

Ego-dysphoric voice transformation for stuttering reduction

This disclosure relates to an apparatus, system, computer program, and computer implementation of a speech conversion method for reducing stuttering. The speech conversion is preferably performed in a mobile electronics user device, a speech processing device incorporated into a wearable hearing system, or a server. The apparatus, system, computer program, and method are preferably configured to perform the following steps: receive input speech information from a speech sensor device (106), including at least one speech utterance in the user's natural voice; perform speech conversion (118) to generate output speech information with an ego-dysphoric target speech, such that the at least one speech utterance is converted as if the same utterance were produced by different speakers; and prompt the user to play back the speech-converted output speech information, particularly binaural playback, at least in near real-time, as feedback to the user's speech.
Owner:ベルケ·ベンノ

Self-incongruous speech conversion for reducing stuttering

The present disclosure relates to devices, systems, computer programs, and computer-implemented voice conversion methods for reducing stutters. The voice conversion is preferably performed in a mobile electronic user device, an audio processing device integrated into the wearable hearing system, or a server. The devices, systems, computer programs and methods are preferably configured for performing the steps of: receiving input audio information from an audio sensor device (106), the input audio information comprising at least one spoken utterance of natural speech of a user; performing a speech conversion (118) to generate output audio information with a self-discordant target speech as at least one spoken utterance is converted as if the same speech content was produced by a different speaker; reproduction, in particular binaural reproduction, of the voice-converted output audio information is presented to the user at least approximately in real time as feedback to the user's speaking.
Owner:本诺·贝尔克

Stuttering speech recognition method and device, computer device and storage medium

The application relates to the technical field of artificial intelligence, and provides a stuttering speech recognition method and device, computer equipment and a storage medium, the method comprising the following steps: acquiring initial text information for recognizing target stuttering speech; inputting the initial text information into a target stuttering predictor for prediction, generating target text information according to a prediction result of the target stuttering predictor; and inputting the target text information into a preset speech recognition model, recognizing the target stuttering speech corresponding to the target text information in the preset speech recognition model, so that the accuracy of stuttering speech recognition is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Traffic regulation method and device of audio and video stream, electronic equipment and storage medium

PendingCN122293892ABandwidth capData loss
This application relates to a method, apparatus, electronic device, and storage medium for traffic control of audio and video streams. The method includes: acquiring an audio and video stream to be sent and parsing the audio and video stream to obtain the frame type and single-frame data size in the audio and video stream; determining the key frame data size based on the frame type and single-frame data size; acquiring the bandwidth threshold of the link used to send the audio and video stream, wherein the bandwidth threshold is used to characterize the upper limit of bandwidth that can be allocated to the link; determining whether the triggering conditions of the traffic control strategy are met based on the key frame data size and the bandwidth threshold; and if the triggering conditions of the traffic control strategy are met, splitting each key frame in the audio and video stream into multiple data sub-packets for transmission. This can achieve traffic control of the audio and video stream, avoiding the peak traffic of the audio and video stream from exceeding the bandwidth threshold of the link, which could lead to key frame data loss and problems such as video distortion, stuttering, and audio disconnection.
Owner:BEIJING FEIXUN DIGITAL TECH CO LTD

Priority control method for garbage collection and related equipment

The invention provides a priority control method for garbage collection and related equipment. According to the method, for the important stage where a main thread is easily blocked in garbage collection, the operation priority of the important stage is improved under the condition that the operation priority adjustment permission for adjusting the important stage is set, safe and orderly utilization of time slices of a central processing unit can be guaranteed, meanwhile, operation of the important stage can be ended as soon as possible, and the efficiency of garbage collection is improved. Therefore, the time for blocking other threads by the GC thread can be shortened, the phenomenon of picture or sound lagging caused by blocking of other threads is reduced, and the user experience is improved. In some embodiments, when the important stage starts running or is about to start running, the priority improvement time period is determined according to the historical running time (time consumption of a running state) of the important stage, and the priority of the important stage is improved according to the priority improvement time period, so that the running priority of the important stage can be more accurately regulated and controlled.
Owner:HONOR DEVICE CO LTD

Emergency rescue video call system and method for people trapped in a malfunctioning elevator

PendingCN122317227ANoise (video)Packet loss
This invention relates to the field of edge-cloud collaboration technology, specifically to an emergency rescue video call system and method for people trapped in malfunctioning elevators. The system includes: a state-aware coding module that calculates the signal-to-noise ratio slope to compress macroblocks and construct a video stream; a keyframe loss assessment module that calculates packet loss density to obtain truncation loss rate; an asymmetric error correction scheduling module that stretches the step size to construct a hybrid stream; a delay mapping retransmission module that monitors delay parameters to obtain detection results; and a cross-modal time-domain synchronization module that stretches audio and shifts and calibrates time markers. In this invention, spatial redundancy compression is performed based on dynamic image and signal-to-noise ratio patterns to adapt to physical channel attenuation. The load missing density is analyzed and the step distance is dynamically stretched to perform asymmetric error correction scheduling. Combined with the delay detection results, periodic peak search and non-modulation extension are performed on the waveform to drive the retransmission packet marker shift calibration. In a weak network transmission environment, the audio-visual data is kept aligned in the time domain and the screen stuttering and disconnection faults are eliminated.
Owner:GORE ELEVATOR TIANJIN

Data processing method, apparatus and system

PCT designated stageWO2026045580A1TransmissionSpeech synthesisSpeech soundStuttering
The present application provides a data processing method, apparatus and system. The method comprises: receiving a first code stream and a second code stream, the first code stream comprising data obtained by encoding text corresponding to first audio data, and the second code stream comprising data obtained by encoding the first audio data; when reconstructed audio data is incomplete, performing text-to-speech processing on reconstructed text data to obtain second audio data; and fusing the second audio data and the reconstructed audio data to obtain third audio data, and outputting the third audio data, the reconstructed text data being obtained by decoding the first code stream, and the reconstructed audio data being generated on the basis of the second code stream. In this way, even if a network state becomes poor, a receiving end can still recover complete audio data, thereby reducing problems such as stuttering, dropped audio frames, and even silence, and improving the fluency of audio in a real-time communication scenario.
Owner:HUAWEI TECH CO LTD

A Bluetooth-based wireless headphone audio transmission method

This invention discloses a Bluetooth-based wireless earphone audio transmission method. It acquires environmental audio matrix data through a sensor array in the wireless earphone; extracts noise spectrum features from the environmental audio matrix data using an improved CNN convolutional neural network; and eliminates motion friction noise in the matrix data using an adaptive filtering algorithm to obtain dual-channel noise-reduced data. The dual-channel noise-reduced data is input into a fusion model to identify the user's current scene. Combining Bluetooth RSS I signals and historical channel data, it predicts the available bandwidth within the next 5ms. Based on the scene identification result and available bandwidth, it switches the codec. Based on the codec, it segments the audio data stream into high-priority and low-priority data packets, which are transmitted through the Bluetooth main channel and auxiliary channel, respectively. This improves the audio data transmission success rate, effectively reduces audio stuttering and distortion, and enhances the user's listening experience and satisfaction.
Owner:SHENZHEN LANQI CHUANGFA TECH CO LTD

Lightweight double-diaphragm wireless interconnection sound box

The utility model discloses a lightweight double-diaphragm wireless interconnection sound box, which can ensure that audio streams can be efficiently and stably transmitted in a complex environment through a wireless communication unit, delay and lagging are reduced, user experience is improved, an audio processing unit accurately restores high and low frequency components of audio signals, tone quality is optimized, and sound quality is improved. And the main control unit generates a double-diaphragm driving control signal according to the decoded audio signal, realizes dynamic frequency division driving and enhances the tone quality layering sense by accurately controlling the vibration of the high-frequency and low-frequency diaphragms, and is also responsible for dynamically adjusting the output voltage of the power supply unit, optimizing the energy consumption and prolonging the endurance time according to the load demand. The double-diaphragm driving unit adopts a high-frequency and low-frequency diaphragm separation design, interference is avoided, high-fidelity audio restoration is realized, and the power supply unit outputs stable voltage to a double-diaphragm load according to an instruction of the main control unit, so that stable performance of the sound box is ensured.
Owner:GUANGZHOU YISON ELECTRON TECH CO LTD

Speech extraction method and device, electronic equipment and storage medium

The present application relates to the technical field of artificial intelligence, and provides a voice extraction method and device, electronic equipment and storage medium, the method comprising: obtaining a mixed voice segment at a current time and identity representation information of a target speaker; obtaining a historical voice segment of the target speaker at at least one historical time; fusing mixed voice features extracted from the mixed voice segment, historical context features extracted from the historical voice segment, and the identity representation information to obtain target fusion features; and extracting a target voice segment of the target speaker at the current time from the mixed voice segment according to the target fusion features. The present application introduces the historical voice segment containing rich phonemes, prosody and other sound states as a dynamic context reference, breaking the limitations of traditional stateless models, not only giving the model a memory ability for voice content, significantly improving the continuity of the output, thereby effectively avoiding the generation of stuttering, jumps and artifacts at the block boundary.
Owner:IFLYTEK CO LTD

StutterNet stutter detection algorithm based on improvement

The invention discloses an improved StutterNet-based stutter detection algorithm, relates to the technical field of stutter detection, and aims to solve the problems that a fixed-length context in a basic StutterNet framework may not be an optimal choice for detecting all types of stutters, and a larger context can improve the performance of extension and unsmoothness of repeated types, so that the stutter detection efficiency is improved. According to the technical scheme, the method is characterized by comprising the following steps of S1, improving the class imbalance problem; s2, network architecture optimization: introducing a channel attention mechanism into a related layer of the StuterNet; s3, collecting data; s4, data enhancement and model training; and S5, experimental verification. The effects that precious experience is provided for stutter detection technology development, model parameters can be deeply optimized according to follow-up research, strategy fusion is perfected and improved, and the stutter detection performance is continuously improved are achieved.
Owner:NANJING UNIV OF SCI & TECH

Children stutter speech recognition method and system based on improved attention mechanism

The invention discloses a child stutter speech recognition method and system based on an improved attention mechanism, and belongs to the technical field of child stutter speech recognition. The method specifically comprises the following steps: collecting a stutter voice signal of a child, and obtaining an Fbank feature through feature extraction; inputting the Fbank features into a shared encoder composed of two CNN layers and a hierarchical BLSTM layer, and extracting and fusing local features and context time sequence features to obtain high-order features; based on the high-order features, collaborative optimization is carried out on a mixed attention mechanism decoder and a CTC decoder through a joint loss function, and a joint loss value is calculated; and constructing a speech recognition model based on the joint loss value, and obtaining a final recognition result, thereby improving the recognition accuracy of the child stuttering speech signal, and improving the alignment capability of the input child speech signal and the output recognition text.
Owner:NANJING UNIV OF SCI & TECH

Ego dystonic voice conversion for reducing stuttering

The present disclosure relates to devices, systems, computer programs as well as a computer-implemented voice conversion method for reducing stuttering. The voice conversion preferably takes place in a mobile electronic user device, an audio processing device integrated into a wearable hearing system or a server. The devices, systems, computer programs and methods are preferably configured for carrying out the following steps: receiving input audio information from an audio sensor device (106), which comprises at least one verbal utterance in a natural voice of a user; carrying out a voice conversion (118) for generating output audio information in an ego-dystonic target voice, in that the at least one verbal utterance is converted as if the same speech content was produced by a different speaker; prompting a reproduction, in particular a binaural reproduction, of the voice-converted output audio information to the user at least approximately in real time as feedback to the speaking of the user.
Owner:BELKE BENNO

Audio stuttering analysis method, audio terminal, electronic device, and storage medium

The application provides an audio stall analysis method, an audio terminal, an electronic device and a storage medium, and relates to the technical field of computers. The audio stall analysis method comprises the following steps: first, acquiring log information related to an audio data stream; then, determining a stall time and a record time corresponding to each error log in the log information; and finally, obtaining an audio stall analysis result based on the stall time and the record time corresponding to each error log. By determining the stall time and the record time corresponding to each error log in the log information related to the audio data stream, the audio stall analysis result can be obtained. Since the log information related to the audio data volume is not limited to the data sending end, the present scheme can detect the data receiving end. Meanwhile, by determining the stall time corresponding to different error logs, the stall condition caused by each type of error log can be determined, so that the stall condition of audio playing can be comprehensively detected.
Owner:HENGXUAN TECH (BEIJING) CO LTD

A low-latency, high-reliability converged communication audio and video dynamic adaptation encoding and decoding transmission system

PendingCN122317276AControl cellData acquisition
This invention relates to the field of communication technology, specifically to a low-latency, high-reliability converged communication audio and video dynamic adaptation encoding and decoding transmission system. It includes: a link data acquisition and storage unit; a dynamic adaptation control unit employing a lightweight causal timing prediction engine with an improved temporal convolution + sparse attention architecture; an encoding and decoding unit; a transmission adaptation unit; and a reliability verification and feedback unit. This invention uses a lightweight causal timing prediction module to perform feedforward prediction on continuous link state sequences, and uses the prediction results as feedforward compensation signals input to the adaptation parameter decision module, replacing the traditional passive backward adjustment logic. This allows for early adaptation to time-varying link fluctuations. Simultaneously, combined with a confidence decay and mode switching mechanism, it automatically reverts to a reactive adjustment mode when the prediction confidence falls below a set threshold, mitigating audio and video stuttering and excessive latency issues caused by sudden changes in link state.
Owner:BEIJING SANYONGHUATONG TECH CO LTD

Speech therapy system and method therefor

PCT designated stageWO2026035406A1Stammering correctionSpeech recognitionEngineeringAcoustics
A speech therapy system and method therefor are disclosed. The system includes graduated speaking exercise modules and a computer system including a processor and a memory. The modules are arranged sequentially and are collectively configured to provide graduated speaking exercises, or GSEs, of increasing conversational realism for a stuttering user. The processor executes the app and the modules, and each of the modules create an associated GSE that defines a different state of the app. When the app is in a current state defined by a current GSE, the app obtains or determines a fluency metric from user speech or from a user fluency self-rating. When the metric meets an upper fluency threshold of the current GSE, the app transitions to a next app state defined by a next GSE, and the app can conclude that the user is fluent if the upper threshold is met for a final GSE.
Owner:FLUENCYAI LLC

Vocal practice apparatus, in particular for the stuttering treatment

A vocal practice apparatus, in particular for the treatment of stuttering; wherein the apparatus comprises: a microphone configured to generate a first audio signal from vocal sounds emitted by a user while performing the exercise; a camera configured to film the user while performing the exercise and to generate a video signal; a first monitor configured to be observed by the user while performing the exercise; a vibrating device configured to transmit a vibratory impulse to the user while performing the exercise; a first headset configured to be worn by the user while performing the exercise; a control unit connected to the microphone, the camera, the first
Owner:VIVAVOCE SRL

Speech Therapy System and Method Therefor

PendingUS20260045177A1Electrical appliancesTeaching apparatusSpeech therapy treatmentAcoustics
A speech therapy system and method therefor are disclosed. The system includes graduated speaking exercise modules and a computer system including a processor and a memory. The modules are arranged sequentially and are collectively configured to provide graduated speaking exercises, or GSEs, of increasing conversational realism for a stuttering user. The processor executes the app and the modules, and each of the modules create an associated GSE that defines a different state of the app. When the app is in a current state defined by a current GSE, the app obtains or determines a fluency metric from user speech or from a user fluency self-rating. When the metric meets an upper fluency threshold of the current GSE, the app transitions to a next app state defined by a next GSE, and the app can conclude that the user is fluent if the upper threshold is met for a final GSE.
Owner:FLUENCY AI LLC

Intelligent methods and systems for preventing and optimizing audio and video playback stuttering.

This application relates to the field of audio and video processing technology, and provides an intelligent method and system for preventing and optimizing audio and video playback stuttering. In this application, firstly, a playback status data set is collected in real time, including bitstream parameters, hardware resource usage data, and network transmission status parameters. Then, a playback load correlation graph is constructed based on this data set, describing the mutual influence relationships between parameters. Next, a dynamic resource allocation strategy is generated based on the playback load correlation graph, adjusting the hardware resource allocation ratio and network transmission priority. Finally, the strategy is executed, and data changes are continuously monitored, dynamically adjusting the strategy to maintain playback smoothness. Thus, through multi-parameter correlation modeling and dynamic optimization, precise stuttering prevention is achieved, improving the audio and video playback experience.
Owner:SHENZHEN ZIDOO TECH CO LTD