Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

521 results about "Delay" patented technology

Delay is an audio effect and an effects unit which records an input signal to an audio storage medium, and then plays it back after a period of time. The delayed signal may either be played back multiple times, or played back into the recording again, to create the sound of a repeating, decaying echo.

Low-delay audio input switching method and system, storage medium and equipment

The invention relates to the technical field of audio control, and discloses a low-delay audio input switching method and system, a storage medium and equipment, and the method comprises the steps: carrying out the parallel pre-initialization of a plurality of pieces of audio input equipment, and building an independent parallel audio data cache for each piece of equipment; monitoring the state of each audio input device and the audio stream quality in real time, and judging whether switching is triggered or not based on a multi-factor decision model; after the switching decision is triggered, seamless audio data stream switching is executed, and format unification processing and cross fade-in and fade-out transition are included; a unified equipment operation interface is provided through the hardware abstraction layer, and system resources are optimized and managed; a delay sensing closed-loop control mechanism is constructed, processing delay of each link is monitored in real time, a caching strategy, a processing algorithm and resource allocation parameters are dynamically adjusted, self-adaptive balance of low delay and high tone quality is achieved, and through the method, quick and smooth switching of audio input equipment is achieved, delay is remarkably reduced, and real-time audio experience is improved.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Audio and video low-delay return method and system in extreme environment

The invention relates to the technical field of audio and video emergency transmission, and discloses an audio and video low-delay return method and system in an extreme environment. The method comprises the following steps: acquiring original multi-modal data of audio and video acquisition equipment in a target area, and analyzing a data state of the original multi-modal data; and meanwhile, available network transmission links are monitored, and the quality is evaluated. Self-adaptive coding parameters are generated in combination with the link quality and the data state, dynamic coding is executed on original data, and a coding stream suitable for redundant transmission is formed. And carrying out cooperative distribution transmission based on the network link set and the coded stream to generate return data. According to the generation process, a device and energy management instruction is formed and issued to the acquisition and relay device. According to the method, the network state and the content characteristics are optimized in a dynamic coding link in a collaborative manner, and the audio and video quality, the real-time performance and the system energy efficiency which are transmitted back in an extreme environment are improved through a feedback closed loop for transmitting a result to equipment management.
Owner:XIAN YUNKAI INFORMATION TECHNOLOGY CO LTD

Sound source positioning and detecting method and device

The invention belongs to the technical field of sound source processing, and provides a sound source positioning and detecting method and device. The method comprises the following steps: acquiring delay estimation of microphones I and II and delay estimation of microphones I and III based on a three-path linear uniform microphone array; based on the delay estimation sum, respectively carrying out positioning calculation under near-field and far-field conditions; according to the sound source distance under the near-field condition, comparing the sound source distance with a distance judgment threshold value, and determining a far-field / near-field output sound source position; and carrying out feature extraction on the signals of any microphone array, and carrying out event classification based on a pre-trained convolutional neural network. According to the method, far / near field model selection is carried out according to the distance judgment threshold, large deviation generated by a single model in a critical region is avoided, continuous and stable positioning from short distance to long distance is achieved, time classification can be achieved while position calculation is carried out, and integrated output is achieved.
Owner:YANGZHOU YUAN ELECTRONICS TECH CO LTD

Distributed audio group mute and recovery method and system, storage medium and equipment

The invention relates to the technical field of audio control, and discloses a distributed audio group mute and recovery method and system, a storage medium and a device, and the method comprises the steps: determining a master device in a distributed audio network, and enabling the master device to find and register a plurality of slave devices to establish an audio group; receiving a group mute or recovery instruction sent by a user through the master device, and performing validity verification on the instruction; the master device allocates a uniform timestamp to the verified instruction, and sends a synchronous control command to all slave devices in the group based on the timestamp; the slave equipment executes mute or recovery operation at the moment corresponding to the timestamp, synchronous response of the equipment in the group is realized, all the equipment is ensured to accurately and synchronously execute the mute / recovery operation through unified management of the master equipment and the adoption of global timestamp synchronization and adaptive delay prediction technologies, and an intelligent state monitoring and exception recovery mechanism is provided, so that the service life of the slave equipment is prolonged. Finally, efficient and reliable multi-room audio group control with excellent user experience is realized.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Music source separation method and wearable device

The invention discloses a music source separation method and wearable equipment, and relates to the technical field of signal processing, music source separation is performed by adopting a neural network model of a UNet architecture, the neural network model at least comprises an encoder and a decoder, and the method comprises the following steps: obtaining an original signal of a to-be-separated music source, and converting the original signal to a frequency domain to obtain a frequency domain signal; performing feature coding processing on the frequency domain signal through an encoder to obtain coded feature data, the feature coding processing including feature extraction, and the encoder performing feature extraction processing at least by using a TFC-TDF block of causal convolution; and performing feature decoding processing on the coded feature data through a decoder, and outputting to obtain a music source separation result. According to the method and the device, low-delay music source separation is realized in the end side equipment with limited resources.
Owner:GOERTEK INC

Automatic delay calibration method for smart speaker system, device, and storage medium

PCT designated stageWO2026066023A1Signal processingTransducer circuitsSignal onTime alignment
Disclosed in the present invention are an automatic delay calibration method for a smart speaker system, a device, and a storage medium. The method comprises: initializing a smart speaker system and performing environmental testing to preliminarily configure system parameters; exciting a target sound channel and playing a target test signal; locating a first test signal on the basis of a voice activity detection (VAD) algorithm, determining the start and end of the first test signal, and identifying a time difference of the first test signal to obtain a delay of main channels; locating a time of arrival of a second test signal on the basis of a fast peak search algorithm, and identifying a delay of a subwoofer; and performing delay calibration on the subwoofer and the main channels, and dynamically adjusting delay calibration parameters of the channels on the basis of a delay calibration amount, thereby achieving accurate time alignment between the subwoofer and the main channels. The present invention solves the problem of low-frequency signal delays that are difficult to process in conventional technologies, and particularly makes a breakthrough progress in the coordination of a subwoofer and other sound channels.
Owner:LINKPLAY TECHNOLOGY INC NANJING

High-resolution low-delay audio and video transmission method and system

The invention discloses a high-resolution and low-delay audio and video transmission method and system, which realize high-quality and low-delay remote audio and video transmission by dynamically negotiating transmission resolution and self-adaptive compression coding. The method comprises the following steps: constructing a resolution negotiation matrix through extended display identification data based on receiving end equipment, and determining an optimal transmission resolution parameter; the compression ratio is dynamically adjusted in combination with network bandwidth fluctuation, a compressed data stream is generated by adopting inter-frame prediction coding, and the anti-interference capability is enhanced through forward error correction and time division multiplexing packaging; and finally, the differential signal pair is transmitted to a receiving end to be decoded and restored into a standard audio and video signal. According to the method, the transmission delay is remarkably reduced while the high-resolution image quality is ensured, the method effectively adapts to a complex network environment, the transmission distance and stability are considered, and the method is suitable for application scenes such as remote conferences and real-time monitoring which have high real-time requirements.
Owner:SHENZHEN DE SHENG DA ELECTRONIC SCI & TECH CO LTD

Multi-channel audio calibration method, system and device and storage medium

The invention discloses a multi-channel audio calibration method, system and device and a storage medium, and relates to the technical field of audio, and the method comprises the steps: collecting and separating a main sound channel signal and a bass sound channel signal in a multi-channel audio signal; detecting time delay between a main sound channel and a bass sound channel based on a phase correlation algorithm, and performing adaptive delay compensation; the state change of the audio signal is monitored in real time, and before or when a signal switching event is detected, gradient gain envelope is applied to the bass sound channel signal, so that the signal amplitude gradually changes to a target value according to a preset nonlinear curve; a sonic boom event is analyzed and identified based on multi-dimensional characteristics, a self-adaptive amplitude limiter with a prospective processing capability is started for suppression, and harmonic distortion compensation and phase correction are carried out on a processed signal, so that the problems of signal asynchronization caused by low-frequency output delay of an audio system in the prior art, and the system reliability is improved are solved. Therefore, the low-frequency sound quality and the system stability are improved.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

Computer-implemented sound calibration method, system and storage medium

A computer-implemented sound calibration method, a system and a storage medium are provided. A test audio is played by a subwoofer and a main speaker, transmitted audio is collected by an audio collection device at a primary listening position, a sound generation delay and a volume parameter difference are analyzed, and the subwoofer is calibrated based on of a frequency response curve of a scenario, so as to solve the problem in the related art where the user experience is adversely affected due to low sound calibration accuracy between the subwoofer and the main speaker.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Sound source localization method based on multi-frequency separation and Newton optimization deconvolution

The invention discloses a sound source localization method based on multi-frequency separation and Newton optimization deconvolution, relates to the technical field of array acoustic signal processing, and is used for solving the problem that weak sound sources and multiple sound sources are difficult to identify. According to the method, dominant frequency is extracted through multichannel frequency domain analysis, a cross-spectrum matrix is constructed in combination with a near-field propagation model and a guide vector, delay summation beam forming is executed to obtain sound source preliminary distribution, then the distribution is regarded as a convolution result, a maximum likelihood model is introduced, and a two-stage deconvolution strategy of coarse estimation and Newton method fine optimization is adopted to obtain a high-resolution sound source. According to the multi-sound-source positioning method, subgrid-level analysis of sound source positions and amplitudes is achieved, finally, all frequency results are fused, continuous sound source images are smoothly output through a two-dimensional Gaussian kernel, the resolution and real-time performance of multi-sound-source positioning are remarkably improved, and the multi-sound-source positioning method is suitable for high-precision acoustic imaging in a complex sound field.
Owner:STATE GRID JIANGXI ELECTRIC POWER CO LTD

Single-microphone acoustic echo and noise suppression

This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques for separating microphone signals into speech, echo, and noise signals. In some aspects, a speech enhancement system may include a delay estimator and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a microphone signal via a microphone and a far-end audio signal for output via a speaker and estimates a reference audio signal based on a delay between the microphone signal and the far-end audio signal. In some aspects, the AEN decoupling filter may determine a speech mask, an echo mask, and a noise mask based on the microphone signal and the reference audio signal and may suppress an echo component and a noise component of the microphone signal based on the determined set of masks.
Owner:SYNAPTICS INC

Earphone and sound box multichannel audio synchronous transmission method and system

The invention relates to the technical field of telecommunication, and discloses an earphone and sound box multichannel audio synchronous transmission method and system, and the method comprises the steps: firstly obtaining an initial audio stream of a target device (earphone and sound box), recognizing format features, analyzing a frequency spectrum component, generating a left / right sound channel signal, then detecting the transmission priority of the sound channel signal, analyzing a channel distribution demand, and carrying out the transmission of the sound channel signal; the query equipment responds to the time delay and calculates a synchronization reference value, then determines a signal competition hotspot based on the value, analyzes the interference intensity and the frequency band occupancy rate, calculates a clock offset index, then determines a collaborative playing path according to the index, queries a phase tolerance threshold, formulates a time sequence calibration process, and finally corrects a sound channel signal by using the process. Obtaining, recording and optimizing clock reference parameters, and generating a synchronous transmission scheme. The overall sound effect stability of cooperative playing of the earphone and the sound box can be improved.
Owner:SHENZHEN XUSHENG TECH CO LTD

Video conference echo suppression method based on low-delay adaptive learning model

The invention relates to a video conference echo suppression method based on a low-delay adaptive learning model. The method comprises the following steps: acquiring acoustic scene noise features, and matching a lightweight LSTM neural network model by using the acoustic scene noise features; the method comprises the following steps: collecting real-time sound signals of a video conference, preprocessing and extracting multi-modal features of the real-time sound signals; performing adaptive filtering by using an NLMS algorithm to estimate an echo path between a near-end microphone signal and a far-end reference signal of the real-time sound signal, performing convolution on the echo path and the far-end reference signal to generate an echo estimation signal, and subtracting the echo estimation signal from the near-end microphone signal to obtain a preliminary residual signal; learning a gain by using the lightweight LSTM neural network model, and further suppressing the residual signal by using the gain to obtain a time domain enhanced signal; comfortable noise is introduced, and incremental learning is adaptively triggered. And efficient linear and nonlinear echo cancellation is supported.
Owner:CHINA LIFE INSURANCE CO LTD

Low-delay lossless digital audio transmission method and system and storage medium

The invention discloses a low-delay lossless digital audio transmission method, a low-delay lossless digital audio transmission system and a storage medium, which relate to the technical field of digital audio transmission and are used for monitoring network environment data in real time and dynamically adjusting audio transmission parameters according to changes of network bandwidth Bw, network delay Dr, packet loss rate Pd and network jitter Wd. Comprising an audio transmission rate Rsv, a data packet size Spd and a redundancy coding ratio Rbf. The adaptive adjustment can ensure that the data transmission quantity is reduced and the packet loss probability is reduced when the network condition is poor, and meanwhile, the transmission efficiency is improved and the audio quality is optimized when the bandwidth condition is good. Compared with a traditional fixed transmission parameter method, lossless audio transmission can be achieved under different network conditions, and meanwhile unnecessary delay is reduced. According to the invention, intelligent compression of the audio data is realized, the data volume is reduced under the condition that the bandwidth is limited, the network burden is reduced, and the stability and integrity of audio transmission in a low-bandwidth environment are ensured.
Owner:JIAXING WANSHENG ELECTRONICS TECH CO LTD

Simultaneous interpretation method and system based on large model and electronic equipment

The invention discloses a simultaneous interpretation method and system based on a large model and electronic equipment, and the method comprises the steps: extracting bilingual parallel corpora related to terms from professional resources based on a standardized professional dictionary, obtaining qualified corpora through data enhancement processing and manual screening, and constructing a multi-level corpus according to the levels of words, sentences and paragraphs; the method comprises the following steps: receiving an input audio stream in real time, extracting acoustic features through preprocessing, inputting a pre-established large-scale speech recognition model, and carrying out incremental decoding on the acoustic features in a sliding window mode; and calling a sentence boundary prediction network to judge a pause point, and outputting a text stream with a timestamp. Performing fine tuning on the large-scale speech recognition model by using a multi-level corpus, translating a text stream based on the fine-tuned large-scale speech recognition model, and constraining term translation according to a standardized professional dictionary; and synchronously displaying the audio output in the translation result and the subtitles. According to the scheme, the terminology recognition and translation accuracy is improved, and simultaneous interpretation delay is reduced.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Sound sound field orientation system for immersive audio experience

The invention, which relates to the technical field of audio experience, discloses an immersive audio experience-oriented sound sound field orientation system comprising a position sensing module, a sound field modeling module, a characteristic sound wave generation module, an array control module and a space feedback adjustment module. The position and orientation information of a target user in a three-dimensional space is obtained in real time through a non-contact sensor, a real-time acoustic propagation model containing reflection surface distribution and material characteristics is constructed, a multi-channel acoustic signal is generated based on the model, a loudspeaker array is driven to produce sound through a beam control parameter set, and the sound transmission effect is achieved. According to the invention, the spatial focusing of the sound beam in the target direction is realized, the sound production delay and phase difference of each loudspeaker unit are dynamically adjusted based on the direction deviation information of the received sound pressure, and a closed-loop control path is formed. The method is suitable for various scenes such as virtual reality, intelligent cinemas and voice interaction.
Owner:FOSHAN HIPP TECH CO LTD

An audio and video signal processing system and a video conference terminal device using the same

The application provides an audio and video signal processing system and a video conference terminal device using the same, and relates to the technical field of Internet.The system comprises an audio and video processing module, a signal preprocessing module, a signal delay measurement module, a signal synchronization processing module and a master control module, the audio and video processing module is used for separating audio signals and video signals in input audio and video data signals, the signal preprocessing module is used for preprocessing the audio signals, the signal delay measurement module is used for measuring relative delay amounts of the audio signals and the video signals, the signal synchronization processing module is used for performing delay compensation according to the relative delay amounts, realizing the synchronization of audio and video display, and the video conference terminal device further comprises a microphone, a camera, a display screen and a sound box.The system and the video conference terminal device can test the relative delay time of audio and video signals, perform delay compensation, ensure the synchronization of output audio and video pictures, and improve the video conference quality and user experience.
Owner:BEIJING ZOBO ELECTRONIC TECH CO LTD

WebRTC speech enhancement system and method based on multi-modal large model

The invention discloses a WebRTC voice enhancement system and method based on a multi-modal large model, and belongs to the field of artificial intelligence and real-time communication cross technology, and the system comprises an audio and video collection module which is used for synchronously collecting original voice signals and corresponding video image data of a user side through a WebRTC protocol stack, high-precision alignment is realized through timestamp marking and a cache mechanism; the multi-modal feature extraction module is used for extracting audio features from the voice signals, extracting visual lip movement features from the video images and generating text semantic features through a voice recognition engine; a noise matching and updating module; a multi-modal semantic perception enhancement module; an audio reconstruction module; and a WebRTC integration module. According to the invention, high-precision noise suppression and millisecond-level delay are realized while the voice semantic integrity is guaranteed, and the requirements of industrial inspection, vehicle-mounted communication and other scenes on high fidelity, low delay and strong robustness are met.
Owner:浪潮智慧城市科技有限公司

Priority-based traffic broadcast switching method and device, equipment and storage medium

According to the priority-based traffic broadcast switching method, device and equipment and the storage medium provided by the invention, by receiving and playing multiple paths of audio input signals and generating the emergency broadcast trigger signal when the rail transit operation state is abnormal, unified scheduling and control of multi-source audios are realized. Furthermore, according to the method, the emergency broadcast trigger signal is compared with the currently played audio signal through a priority judgment mechanism, and when the priority of the emergency broadcast trigger signal is higher, the low-priority audio is automatically interrupted and the emergency broadcast signal is preferentially played, so that the emergency information is ensured to be transmitted to the passenger at the first time, and the passenger safety is ensured. And delay or omission caused by other audio interferences can be avoided. The method provided by the invention not only improves the reliability and timeliness of emergency broadcast, but also realizes the dynamic management of the broadcast task through a priority mechanism, and gives consideration to the flexibility and practicability of the system while guaranteeing the safety.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Information processing device, information processing method, information processing system, and program

For example, reverberation processing in consideration of sound quality can be performed.Provided is an information processing device including a training processing unit that generates training data by convolving a measurement signal representing acoustic characteristics collected by a same sound collection unit as a sound collection unit used for collecting an observation signal with a reference signal having sound quality and reverberation characteristics different from the observation signal, generates teaching data by adapting, to the reference signal, an average level and a delay value of a convolution signal generated by convolving a direct sound component of the measurement signal with the reference signal, and trains a learning model for performing reverberation processing of the observation signal collected by the sound collection unit by using the training data and the teaching data as input data.
Owner:SONY GROUP CORP

Video delay measurement device and program

To provide a program production system that measures a delay time of a video image with respect to audio and auxiliary data by a simple method so as to realize synchronization among the video image, the audio and the auxiliary data generated at the same time.SOLUTION: The video time difference calculating unit 11 - 1 of the video delay measuring apparatus 2 extracts the time stamp V1 from the video RTP packet VTS1 and generates the video RTP packet V2 including the timing information. The video RTP packet V2 is transmitted to the video delay time difference measuring path. A video time difference calculation section 11-1 extracts a time stamp V3 from a video RTP packet VTS2 to obtain a video time difference VT. The audio time difference calculator 11-2 obtains the audio time difference ST, and the ancillary data time difference calculator 11 - 3 obtains the ancillary data time difference AT. The delay time calculating unit 13 subtracts the audio time difference ST from the video time difference VT to obtain the audio delay time SD, and subtracts the ancillary data time difference AT from the video time difference VT to obtain the ancillary data delay time AD.SELECTED DRAWING: Figure 5
Owner:NIPPON HOSO KYOKAI

Audio signal squeal suppression method and device, storage medium and electronic equipment

The invention discloses an audio signal howling suppression method and device, a storage medium and electronic equipment, and the method comprises the steps that a control terminal receives a current data frame sent by intelligent equipment, and the current data frame comprises a current audio signal collected by the intelligent equipment, the sending delay of the intelligent equipment and the latest network propagation delay; the control terminal determines end-to-end time delay based on the sending time delay, the network propagation time delay, the playing time delay and the receiving time delay of the control terminal, and determines final time delay based on the end-to-end time delay and a time delay estimation value calculated by the control terminal; and the control terminal performs howling suppression processing on the current audio signal based on the final delay. By applying the howling suppression method and device, the howling suppression effect can be guaranteed, and meanwhile the processing calculation amount is reduced.
Owner:HANGZHOU EZVIZ SOFTWARE CO LTD

End-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method

The invention discloses an end-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method, and relates to the technical field of voice interaction control, and the method comprises the steps: an end side collects microphone audio, obtains loudspeaker reference audio, judges a broadcast state, and switches a voice activity detection strategy; end-side framing detection is carried out, voice segment events are output through a three-state state machine, characteristics such as duration, relative volume and broadcast overlapping degree are extracted, and candidate interruption events are generated in combination with an automatic voice recognition intermediate text; the end side coarse screening reports a gray area event to the cloud side semantic research and judgment, and pre-suppression, buffer foresight contraction and block buffer are executed during waiting; and the cloud side issues an instruction containing a period of validity, a deduplication key and an interruption type, and the execution side stops generating, synthesizing and playing and empties buffer so as to play a completion receipt, submit a dialogue history and roll back unplayed content, thereby reducing false triggering and missed triggering, shortening interruption delay, reducing residual broadcast and inhibiting context drift.
Owner:SUZHOU MENGWU INTELLIGENT TECHNOLOGY CO LTD

Audio and video processing system and method supporting AI intelligent repair technology

The invention discloses an audio and video processing system and method supporting an AI intelligent repair technology, and relates to the technical field of audio and video data processing, and the method comprises the following steps: carrying out the scene recognition based on a playing source type, a user interaction form, an audio and video code rate, an audio sampling rate and network transmission time delay, and generating live broadcast, film watching and singing scene labels; determining a time delay upper limit, a resolution target, a detail reduction level, an audio fidelity index and a beat alignment threshold according to the scene label, and forming a performance constraint vector; according to the invention, a closed-loop mechanism of scene identification, performance constraint setting, repair link arrangement, cloud and terminal task allocation, real-time adaptive adjustment and parameter iterative optimization is constructed for different requirements of live broadcast, film watching, singing and other scenes, so that low delay, high image quality, high tone quality and computing power adaptation are realized; and the stability and the continuous optimization capability of the system are improved through versioning and staged release.
Owner:SICHUAN YINCHUANG WEIYE TECH CO LTD

Speech recognition and speech synthesis optimization method and system based on large model

The invention provides a voice recognition and voice synthesis optimization method and system based on a large model, and the method comprises the steps: extracting lip motion features, lip feature timestamps and audio feature timestamps through obtaining a real-time voice input signal and an image frame sequence of a user face region, so as to generate an initial space-time offset sequence; generating a dynamic offset compensation parameter sequence based on the initial space-time offset sequence in combination with a reference alignment template in the voice and vision synchronization data set; and performing joint processing on the dynamic offset compensation parameter sequence, the lip motion characteristics and the real-time voice input signal by using a large model to reconstruct a target voice segment, generating a corrected phoneme sequence in combination with the lip motion characteristics, retrieving a mouth shape parameter group corresponding to the corrected phoneme sequence from a phoneme mouth shape mapping rule base, and performing mouth shape correction on the mouth shape parameter group. To generate a voice waveform in phase synchronization with the lip motion; according to the invention, the naturalness and immersion of man-machine interaction and the robustness in a voice missing or delay scene are improved.
Owner:LUSTER LIGHTWAVE CO LTD

Display device and sound synchronization method

Embodiments of the invention provide a display device and a sound synchronization method. The method comprises the steps of obtaining an original audio delay value of a loudspeaker; the original audio delay value is the processing duration required by the loudspeaker from receiving the audio data to playing the audio data; calculating a link time delay value for transmitting the audio data to the sound box; the link time delay value is a one-way time delay value corresponding to audio data transmitted to a data link where the sound box is located; performing difference calculation on the link delay value and the original audio delay value to obtain an additional delay compensation value; and setting the additional delay compensation value to the loudspeaker so as to synchronize the output time of the loudspeaker and the output time of the sound box for outputting the audio data. According to the method, the additional delay compensation value is set to the loudspeaker, so that the loudspeaker actively lags behind the corresponding time when playing the audio data so as to keep synchronous with the output of the loudspeaker box, namely, the audio synchronization between the loudspeaker box and the television loudspeaker is realized, and the problem that the sound of the loudspeaker and the sound of the loudspeaker box is not synchronous in a sound playing scene is solved.
Owner:VIDAA (NETHERLANDS) INT HLDG LTD

Video call real-time audio and video synchronous monitoring system

The invention relates to the technical field of audio and video synchronization, in particular to a video call real-time audio and video synchronous monitoring system which comprises a frame time extraction module, a main time sequence locking module, a delay structure analysis module, a synchronous node generation module and a synchronous control instruction module. According to the invention, through real-time extraction and analysis of timestamp difference calculation of audio and video frames and establishment of a transmission delay interval comparison sequence, the time difference in audio and video transmission can be effectively identified, the monitoring accuracy of an audio and video asynchronization phenomenon is enhanced, and through real-time adjustment of timestamp offset, the synchronization accuracy of the audio and video in a playing process is ensured; in a dynamic network environment, the offset rate is monitored and adjusted in real time, the continuity and fluency of interactive experience are ensured, the synchronization state of the audio and video frames can be controlled more accurately by analyzing the offset difference value centralized distribution section and marking the synchronization fluctuation section, and the situation that the large-scale user experience is reduced due to synchronization error accumulation is avoided.
Owner:DONGGUAN CLOUD FOX INTELLIGENCE LTD

Multimedia playback synchronization

A device includes one or more processors configured to obtain output delay data based on a transmission from an audio output device. The output delay data indicates a playback delay associated with audio output by the audio output device. The one or more processors are further configured to determine, based on the output delay data and host delay data, a synchronization delay to coordinate audio output by the audio output device and video output at a display. The one or more processors are further configured to initiate sending of audio data to the audio output device and to send video data to the display. The video data is delayed, based on the synchronization delay, relative to the sending of the audio data.
Owner:QUALCOMM INC

Multi-event cooperative control method and device for audio equipment, equipment and storage medium

PendingCN121255399AProgram initiation/switchingComputer hardwareEvent cycle
The invention relates to the technical field of embedded equipment control, and discloses a multi-event cooperative control method and device for audio equipment, equipment and a storage medium, and the method comprises the steps: constructing a queue structure body of multi-event circulation according to hardware parameters of the audio equipment, and achieving the efficient cooperative processing of events through a classification rule and a time scheduling method. The method has the advantages of effectively managing event queues, reducing high-priority event delay, avoiding equipment downtime and improving user experience and equipment performance.
Owner:XIAMEN LEYUNRUI TECHNOLOGY CO LTD