Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

46 results about "Time alignment" patented technology

Loudspeaker time-alignment usually simply referred to as "time-alignment or Time-Align" is a term applied to loudspeaker systems which use multiple drivers (like woofer, mid-range and tweeter) to cover the entire audio range.

Automatic delay calibration method for smart speaker system, device, and storage medium

PCT designated stageWO2026066023A1Signal processingTransducer circuitsSignal onTime alignment
Disclosed in the present invention are an automatic delay calibration method for a smart speaker system, a device, and a storage medium. The method comprises: initializing a smart speaker system and performing environmental testing to preliminarily configure system parameters; exciting a target sound channel and playing a target test signal; locating a first test signal on the basis of a voice activity detection (VAD) algorithm, determining the start and end of the first test signal, and identifying a time difference of the first test signal to obtain a delay of main channels; locating a time of arrival of a second test signal on the basis of a fast peak search algorithm, and identifying a delay of a subwoofer; and performing delay calibration on the subwoofer and the main channels, and dynamically adjusting delay calibration parameters of the channels on the basis of a delay calibration amount, thereby achieving accurate time alignment between the subwoofer and the main channels. The present invention solves the problem of low-frequency signal delays that are difficult to process in conventional technologies, and particularly makes a breakthrough progress in the coordination of a subwoofer and other sound channels.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Fan main shaft abnormity identification method based on sound and vibration signal conjoint analysis

The invention relates to the field of fan spindle state monitoring, and aims at synchronously acquiring sound and vibration signals through multiple channels, realizing nanosecond time alignment by adopting a precision time protocol, and performing denoising and normalization preprocessing on the signals to improve data integration and signal fidelity. Furthermore, short-time Fourier transform and continuous wavelet transform are combined to extract multi-scale time-frequency features, and a high-dimensional combined feature vector is generated in combination with cross-correlation analysis. And mapping the feature vectors to a low-dimensional manifold space through a local linear embedding algorithm, and constructing a dynamic mode reference template. Indexes such as curvature, track length and direction entropy are monitored in real time, whether the spindle has an abnormal evolution trend or not is judged through a self-adaptive curvature threshold, and abnormity judgment is achieved in combination with track backtracking verification. According to the scheme, the abnormal starting boundary of the spindle state can be caught in a refined mode, and the early warning and operation and maintenance response capacity of fan operation is effectively improved.
Owner:GUANGDONG ZHONGHUI ZHIWEI ENERGY MANAGEMENT CO LTD

Aerial work platform intelligent supervision method and system based on data analysis

The invention discloses an aerial work platform intelligent supervision method and system based on data analysis. The method comprises the following steps: S1, collecting operation state data of an aerial work platform and behavior video stream data of an operator; s2, inputting a dynamic time warping network to carry out cross-modal time alignment, and outputting a key state feature sequence; s3, inputting the behavior video stream into a double-path Transform model, and generating an operation behavior representation sequence; s4, jointly modeling key state features and behavior characterization, and constructing a risk scoring function; s5, the risk scoring result is compared with a grade threshold value, an early warning mechanism is triggered, and a control instruction is generated; s6, executing platform control actions, such as speed limiting, descent, locking, pause and recording feedback; and S7, updating the model by taking the key state characteristics, the behavior characterization, the scoring result and the feedback as samples. According to the invention, intelligent identification and early warning of high-altitude operation states and behaviors are realized, and operation safety and supervision efficiency are improved.
Owner:JIANGYIN HUACHENG SPECIAL MASCH ENG CO LTD

Audio quality automatic scoring method and system combining MFCC and time domain statistical characteristics

The invention discloses an automatic audio quality scoring method and system combining MFCC and time domain statistical characteristics, and the method comprises the steps: carrying out the segmentation according to seconds on the basis of reference and to-be-detected signal resampling and time alignment, and further extracting cepstrum and time domain statistical characteristics according to 100 ms sub-segments; the method comprises the following steps: forming 170-dimensional second-level features by performing hierarchical aggregation of mean value + maximum value on sub-segment features, splicing the 170-dimensional second-level features with corresponding reference segments to form 340-dimensional joint features, and inputting the 340-dimensional joint features into a pre-trained support vector machine model to output second-by-second discrete scores from 0 to 5, thereby realizing high-time-resolution automatic quality evaluation and anomaly positioning under the conditions of light calculation power and small samples.
Owner:深圳联康测控有限公司

Acoustic echo cancellation method and processing terminal

The invention discloses an acoustic echo cancellation method and a processing terminal, and the method comprises the following steps: 1, carrying out the time alignment of an input far-end signal and a near-end signal, and obtaining a near-end input signal; 2, framing the near-end input signal according to the time domain signal, and performing Fourier transform on each frame to obtain the near-end input signal of the frequency domain corresponding to each frame; 3, superposing the near-end input signal and the far-end signal in the frequency domain on the frequency dimension to obtain a combined signal, extracting a real part of the combined signal, filtering, and taking a logarithm to obtain a filtered logarithm filtering signal; and step 4, inputting the logarithmic filtering signal and the combined signal into a trained deep neural network, outputting a preliminary audio signal in a frequency domain by the deep neural network, and performing inverse Fourier transform on the preliminary audio signal to obtain a final audio signal in a time domain. According to the method, the echo component is effectively suppressed, and the key characteristics of the near-end voice signal are reserved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Tennis service recognition method and device based on time interval

The invention discloses a tennis serving recognition method and device based on a time interval, and the method comprises the steps: collecting video frames in parallel through a plurality of cameras, and carrying out the time alignment of the video frames, thereby obtaining a plurality of synchronous frames; recognizing two-dimensional pixel coordinates of the tennis ball in the multiple paths of synchronous frames, and constructing three-dimensional coordinates of the tennis ball in the space through a ray intersection method; analyzing the continuous three-dimensional coordinates of the tennis ball in the sliding window in the space, and judging the movement trend of the tennis ball on the y axis; if the motion trend on the y axis is continuous rising or falling, calculating the frame interval between the current frame and the last effective motion; finally, judging whether the frame interval is greater than an effective motion interval threshold value or not; if yes, new serving is judged, and the coordinates of the serving starting point are recorded. According to the method, the y-axis change trend of the tennis movement is analyzed through the multi-frame sliding window, the serving movement characteristics are accurately recognized in combination with the frame interval threshold value, serving and common hitting can be accurately distinguished according to the serving movement characteristics, and the misjudgment rate is effectively reduced.
Owner:BEIJING GIVERNY SPORTS TECHNOLOGY CO LTD

Multi-modal data real-time analysis and feedback method and system

The invention provides a multi-modal data real-time analysis and feedback method and system, and the method comprises the steps: collecting a video frame and an audio frame, reading a count value of the same monotonic timer when the collection of the video frame and the audio frame is completed, and generating an audio and video sequence; calculating the behavior popularity of each region in each time slice based on the sequence, and generating a region set for multi-modal event analysis in combination with a preset threshold and a quantity upper limit; scheduling the corresponding video sub-blocks to a visual computing power unit for target positioning and action classification, and executing voice activity detection and keyword category judgment on the time-aligned audio clips at the same time; visual and audio results are fused, and a structured event sequence is generated according to time slice and region compression; and constructing a classroom teaching chain through the sequence, identifying a key event, and finally generating a teaching event description. According to the invention, under the condition of domestic chip combination, cost, time delay, multi-modal consistency and data security controllability are considered, and real-time perception and feedback of teaching behaviors are effectively supported.
Owner:GUANGZHOU KINDLINK INTELLIGENT TECHNOLOGY CO LTD

A distributed audio and video synchronization method based on multi-field coupling

This application discloses a distributed audio and video synchronization method based on multi-field coupling, relating to the field of distributed audio and video synchronization technology. It obtains the time progress field, rhythm field, and semantic field through multi-source node state data broadcast by neighboring nodes, and obtains the time progress error, rhythm error, and semantic error based on these fields. The optimal playback speed is determined based on the time progress error, rhythm error, and semantic error, and the synchronization mode is determined based on the semantic error. Finally, an audio and video synchronization control strategy is generated based on the optimal playback speed and synchronization mode. This strategy is executed, and the process enters the next synchronization cycle until the convergence condition is met, completing synchronization. Through multi-field coupling, synchronization from time alignment to perceptual consistency is achieved, balancing playback continuity, rhythm consistency, and semantic consistency, thus improving the audio and video synchronization effect and user experience in a distributed environment.
Owner:SICHUAN HUSHAN ELECTRIC APPLIANCE

Multi-sensor monitoring system and method based on time alignment

The invention discloses a multi-sensor monitoring system and method based on time alignment, and belongs to the technical field of multi-sensor monitoring, and the system comprises a sensor collection module which is used for obtaining an electrocardiosignal, a photoelectric volume pulse wave signal, a heart sound signal and a respiratory impedance signal containing a main clock receiving timestamp; the time alignment module is used for performing acquisition time alignment on the electrocardiosignal, the pulse oximeter signal, the heart sound signal and the respiratory impedance signal containing the main clock receiving timestamp to form the electrocardiosignal, the pulse oximeter signal, the heart sound signal and the respiratory impedance signal which are in time alignment; the multi-modal fusion analysis module is used for extracting local morphological features in the electrocardiosignal, the pulse oximeter signal, the heart sound signal and the respiratory impedance signal after time alignment to form a feature map, and extracting dynamic weights between different signals and between different time points / features in the same signal; according to the invention, the interpretability and accuracy of judgment are improved.
Owner:SHANGHAI JIAOTONG UNIV

Audio automatic detection system and method based on multi-dimensional calibration

The invention discloses an audio automatic detection system and method based on multi-dimensional calibration, and the method comprises the steps: arranging recording equipment through a standardized recording position rule, carrying out the calibration of space parameters through a range finder, an angle ruler and other tools, and guaranteeing the consistency of collected signals; a multi-channel acquisition architecture is constructed to support synchronous recording and detection of multiple devices; by comparing multi-dimensional acoustic feature extraction of standard audio and recorded audio and adopting a two-stage time alignment mechanism, similarity is accurately calculated to identify abnormities such as sound break and sound break, and a detection result is formed. And finally, analyzing and generating a report based on a detection result, and performing visual display and automatic pushing. The accuracy, efficiency and reliability of audio detection are remarkably improved, and the method is particularly suitable for a batch automatic detection scene before consumer electronics leave a factory.
Owner:TPV DISPLAY TECH (XIAMEN) CO LTD

AI director assistant method and system capable of implementing intelligent interaction effect

The invention belongs to the technical field of director, and particularly relates to an AI director assistant method and system capable of implementing an intelligent interaction effect. The method comprises the following steps: synchronously accessing multiple paths of audio and video signals and carrying out space-time alignment to generate a synchronized audio and video stream set and space mapping parameters; and carrying out hierarchical analysis and semantic feature extraction on the audio-visual content to form an individual behavior portrait. And identifying a social function role and predicting a behavior intention based on the individual behavior portrait. A scene significance scoring system is constructed according to social function roles and behavior intentions, and a key scene with a propagation value is automatically identified. And finally, candidate lens schemes are generated and sorted based on the key scene, and lens switching and dynamic composition adjustment are executed. According to the method, automatic closed loop from signal perception to behavior understanding to picture output is realized, the event response time is remarkably shortened, the accuracy of wonderful instant capture and the continuity of picture presentation are improved, and the problems of response delay and one-sided judgment existing in manual director are effectively solved.
Owner:杭州羿贝科技有限公司

Audio data reading alignment method and device for multi-sound card, equipment and medium

The application provides a multi-sound card audio data reading alignment method, device, equipment and medium, the method comprises the steps of: an application program is started, and at least two threads are created; a sampling attribute configured by the application program using ALSA is acquired, and the audio data reading time consumption is determined according to the sampling attribute; the thread sleep duration is determined based on the audio data reading time consumption, and the at least two threads are kept in a sleep state and continuously sleep for the thread sleep duration; the at least two threads are started at the same time point according to a multi-thread synchronization starting mechanism; the audio data of the corresponding sound card is read by the at least two threads through ALSA-API; and the audio data is stored in the user buffer area corresponding to the at least two threads in the application program. The application can ensure that multiple sound cards can simultaneously transmit audio data to the application program when the application program reads audio data in a multi-thread environment, so that the audio data read by the application program is time-aligned.
Owner:BEI DOU ZHI LIAN KE JI YOU XIAN GONG SI

An audio quality automatic scoring method and system combining MFCC and time domain statistical features

The application discloses an audio quality automatic scoring method and system combining MFCC and time domain statistical features, comprising the following steps: on the basis of resampling and time alignment of a reference signal and a to-be-detected signal, extracting cepstrum and time domain statistical features according to second segmentation and further 100 ms subsegment extraction, forming 170-dimensional second-level features through mean value+maximum value hierarchical aggregation of subsegment features, splicing the corresponding reference segment into 340-dimensional joint features, inputting the pre-trained support vector machine model to output a second-by-second discrete score of 0 to 5, and realizing high-time-resolution automatic quality evaluation and abnormal positioning under the conditions of light computing power and small samples.
Owner:深圳联康测控有限公司

Speech recognition with accurate time alignment of speech units

PendingUS20260120690A1Speech recognitionTimestampTime alignment
Disclosed are apparatuses, systems, and techniques that use one or more artificial intelligence models for time-aligned automatic speech recognition (ASR) of speech. The techniques include processing, an ASR model, one or more audio frames representative of a speech to generate, for a transcription unit (TU) of the speech a first set of likelihood values and a second set of likelihood values. An individual likelihood value of the first set characterizes a probability that the TU corresponds to a vocabulary token. An individual likelihood value of the second set characterizes a probability that the TU corresponds to a timestamp token. The techniques further include generating, using the first set of likelihood values and the second set of likelihood values, a timed transcription of the speech.
Owner:NVIDIA CORP

Psychological condition recognition method based on multi-modal data

The application discloses a psychological condition recognition method based on multi-modal data, comprising the following steps: S1, using a linear space-time detector to improve the quality of video and audio features; S2, using a cross-modal time aligner to ensure time synchronization between modes; S3, using Transformers to detect key segments, and then obtaining the final prediction result through a linear layer; and inputting the prediction result. The problems that the existing deep learning method has low recognition accuracy for psychological conditions are solved.
Owner:YUNNAN UNIV

A hearing aid usage behavior analysis system based on multi-dimensional time sequence features

ActiveCN121935643BChannel dataHearing aid
The application relates to the technical field of user behavior analysis, and discloses a hearing aid use behavior analysis system based on multi-dimensional time sequence characteristics, which comprises a time alignment module, parameter channel data, environment channel data and behavior channel data are acquired, the data are sorted according to server time alignment to form multi-channel time sequence data, asynchronous acquisition is continuously represented on the same time axis, a feature statistical module is used for calculating short-term features of statistical quantities in a preset time window, long-term features of wearing time and average hearing aid power are statistically calculated according to natural days, the features are combined to form a multi-dimensional time sequence feature sequence, scene and operation changes are stably described, a channel modeling module is driven by the sequence to form a multi-channel time sequence model, user behavior feature representation and hearing loss trend prediction results are output, the short-term fluctuation and the long-term evolution are connected, a behavior clustering module is used for clustering to form a user behavior portrait according to the representation, and individualized test and fitting suggestions are generated in combination with the hearing loss trend prediction results and the parameter channel data.
Owner:HANGZHOU HUIER HEARING INSTR & TECH CO LTD

Key travel control method and system of mechanical keyboard and electronic equipment

The invention relates to a key travel control method and system for a mechanical keyboard and electronic equipment, and the method comprises the steps: carrying out the time alignment and feature extraction of collected pressure data, audio data, vibration data, posture data and application data, and generating a multi-dimensional feature vector; inputting the multi-dimensional feature vector into a self-attention converter model, and outputting a scene classification identifier and triggering probability prediction; searching in a hand feeling mapping library according to the scene classification identifier to obtain a target key travel parameter and a target damping parameter, and generating a target control instruction according to the trigger probability prediction, the target key travel parameter and the target damping parameter; and the magnetorheological fluid damper and the linear brake respond to the target control instruction, the magnetorheological fluid damper is controlled to reach target damping, and the linear brake is controlled to stretch out and draw back to the target position. Therefore, the problems that a mechanical keyboard is fixed and single in hand feeling, cannot adapt to dynamic changeable scenes, cannot meet personalized requirements of users and the like are solved, the keyboard input efficiency and the operation accuracy are improved, and the use experience is improved.
Owner:SHENZHEN HENGCHANGTONG ELECTRONICS CO LTD

Decoupling type recording and broadcasting method and system based on key frame index and difference additional recording

The invention discloses a decoupling type recording and broadcasting method and system based on a key frame index and difference additional recording, and the method comprises the steps: firstly continuously collecting multiple paths of audio and video streams, segmenting the audio and video streams into video fragment files, meanwhile, carrying out the parallel real-time detection of a board key event, and generating a key frame index file in which an event type and a millisecond timestamp are recorded; after class, according to the effective starting and ending time of the actual course and the index file, precisely extracting videos of corresponding time periods from the stored fragments; if fragment missing is found, fragments of other machine positions in the same time period are searched based on indexes, patch video fragments are generated through frame-level difference encoding, finally, the patch video fragments and normal video fragments are spliced to form a complete classroom record, and through a decoupling architecture of full-amount recording, event indexing, post-event accurate extraction and additional difference recording, the complete classroom record is obtained. According to the invention, the problem of missed recording and wrong recording caused by equipment delay, class schedule change or recording interruption in the prior art is solved, and the integrity of recorded content, the time alignment precision and the system fault-tolerant capability are remarkably improved.
Owner:CHENGDU SOBEY DIGITAL TECH CO LTD

Unmanned station equipment fault pre-diagnosis method and system based on voiceprint recognition

PendingCN121331161ASpeech analysisTime alignmentSelf adaptive
The embodiment of the invention relates to the technical field of unmanned station equipment fault diagnosis, and provides an unmanned station equipment fault pre-diagnosis method and system based on voiceprint recognition. According to the embodiment of the invention, accurate time alignment is carried out on the original audio stream and the telemetering stream, a parallel network is adopted to extract acoustic and working condition features, deep interaction and time dimension aggregation of the working condition-acoustic features are realized through a Transformer type fusion encoder taking the working condition as a condition, and finally the fault probability is output by a classification head. And the self-adaptive representation and discrimination of the voiceprint under the dynamic working condition are realized. By means of the mode, false alarms caused by normal working condition changes can be remarkably reduced, the recognition rate and diagnosis precision of early-stage tiny fault acoustic symptoms are improved, and meanwhile real-time performance and robustness are considered.
Owner:BEIJING HUANENG XINRUI CONTROL TECH

Centralized management and control software platform based on multi-source heterogeneous data fusion

The invention discloses a centralized management and control software platform based on multi-source heterogeneous data fusion, and relates to the technical field of data management and control, first, real-time interconnection of video and audio collectors is ensured by using a heartbeat monitoring mechanism, and position coordinates of the collectors are determined through feedback information; then, a video-audio interconnection mapping network is constructed, and a foundation is laid for subsequent data processing; in a data review stage, accurately positioning abnormal videos and audio segments by analyzing a video frame pixel difference rate and an audio decibel value waveform curve; then, performing time alignment and matching on the abnormal segments to generate a high-value data set; and finally, surveying and mapping a collection area of the collector by means of a plane-coordinate system, and locking target position coordinates of the audio and video source in real time by combining information such as a sound source direction angle and the like, thereby assisting efficient acquisition of key information.
Owner:NANJING BEIYE ELECTROMECHANICAL EQUIP CO LTD

Methods, systems, apparatus, and articles of manufacture to perform time alignment for watermarks

Methods, apparatus, systems, and articles of manufacture are disclosed to perform time alignment for watermarks. An example apparatus adjusts a power value of an element of a template based on respective average magnitudes and respective tonality ratios corresponding to a plurality of frequency representations of a media signal, the media signal to be encoded with at least one watermark, the element corresponding to one of the plurality of frequency representations. Additionally, the example apparatus computes an alignment of the template to the media signal based on respective power values of elements of the template, the template corresponding to a type of the at least one watermark. The example apparatus also encodes the media signal with the at least one watermark according to the alignment.
Owner:THE NIELSEN CO (US) LLC

A multi-task speech enhancement method and device

ActiveCN120766701BImprove noise characteristicsimprove accuracySpeech analysisFeature extractionNoise
The embodiment of the application provides a multi-task speech enhancement method and device, which utilizes a speech enhancement model to perform feature extraction processing on two signals of a microphone signal and a reference signal, utilizes a dynamic time delay alignment module to perform adaptive time alignment on the two signals, utilizes an adaptive gating module to fuse the features extracted from the two signal branches, extracts multi-scale time features and frequency domain features from the time and frequency dimensions, and improves the accuracy of noise feature extraction, echo feature extraction and speech feature extraction. The application can jointly perform noise suppression and echo cancellation tasks, realize collaborative optimization of multi-task speech enhancement, improve speech call quality, reduce required computing resources, can realize lightweight deployment, and can be suitable for resource-limited application scenarios.
Owner:BEIJING UNIV OF POSTS & TELECOMM +1

Methods and apparatuses for SRS configuration with validity area

Various aspects of the present disclosure relate to methods and apparatuses for sounding reference signal (SRS) configuration with validity area. According to an embodiment of the present disclosure, a user equipment (UE) can include: at least one memory; and at least one processor coupled with the at least one memory and configured to cause the UE to: receive a first configuration including an area-specific time alignment timer associated with an SRS validity area; and transmit an indication to acquire, activate, or deactivate a second configuration associated with the SRS validity area when the UE is in a non-connected state, wherein the second configuration includes an SRS configuration, or a time alignment configuration, or a timing advance (TA) command.
Owner:LENOVO (BEIJING) LTD

Signal sequence-based time alignment of media content

A system includes a computing platform having a hardware processor, a memory storing software code and one or more signal emission device(s) controlled by the computing platform. The software code is executed to determine, using at least one predetermined signal frequency, a calibration signal sequence including an alignment initiation signal and a unique sequence of synchronization signals identified with a unique time interval, emit during the unique time interval, the alignment initiation signal and emit during the unique time interval after emitting the alignment initiation signal, the unique sequence of synchronization signals. The software code is further executed to receive first and second media content produced by first and second recording devices, the first and second media content produced while the first and second recording devices are situated so as to detect the calibration signal sequence, and time align, using the calibration signal sequence, the first and second media content.
Owner:DISNEY ENTERPRISES INC

A method and system for echo cancellation and noise reduction in steel plant workshop environments.

This invention belongs to the field of industrial noise control and relates to a method and system for echo cancellation and noise reduction in steel plant workshop environments. The method includes the following steps: (1) real-time acquisition of sound signals from on-site monitoring points in the steel plant workshop environment during operation; the on-site monitoring points include various noise sources and various target sources; (2) time alignment of the original sound signal and the workshop echo signal from the same on-site monitoring point to obtain an aligned signal; (3) filtering the aligned signal obtained in step (2) to suppress superimposed noise and echo signals in the channel; (4) training a neural network mask, using the mask frequency domain to process the sound signal to suppress residual noise signals, and outputting a clean sound signal. The method provided by this invention effectively solves the limitations of traditional sound noise reduction technology in complex industrial environments.
Owner:UNIV OF SCI & TECH BEIJING

Hearing aid use behavior analysis system based on multi-dimensional time sequence characteristics

The invention relates to the technical field of user behavior analysis, and discloses a hearing aid use behavior analysis system based on multi-dimensional time sequence characteristics, which comprises a time alignment module for acquiring parameter channel data, environment channel data and behavior channel data, and sequencing the data into multi-channel time sequence data according to server time alignment; the characteristic statistics module calculates short-term characteristics of statistical magnitude in a preset time window and counts long-term characteristics of wearing duration and average hearing aid power according to natural days, the short-term characteristics are combined into a multi-dimensional time sequence characteristic sequence, scene and operation changes are stably described, and the real-time performance of the hearing aid is improved. The channel modeling module drives a multi-channel time sequence model through the sequence, user behavior feature representation and hearing loss trend prediction results are output, the penetration from short-term fluctuation to long-term evolution is achieved, and the behavior clustering module forms a user behavior portrait according to representation clustering. And personalized test matching suggestions are generated in combination with the hearing loss trend prediction result and the parameter channel data.
Owner:HANGZHOU HUIER HEARING INSTR & TECH CO LTD

Conference pickup method, terminal, system and computer storage medium

The embodiment of the invention provides a conference pickup method, a terminal, a system and a computer storage medium, and the method comprises the steps: a main conference terminal carries out the collection of audio data in a conference process, and obtains main audio data; receiving slave audio data collected for the conference and sent by a slave conference terminal corresponding to the master conference terminal; obtaining time delay information between the master conference terminal and the slave conference terminal, and performing time alignment processing on the master audio data and the slave audio data based on the time delay information to obtain aligned master audio data and aligned slave audio data; and selecting target audio data from the aligned main audio data and the aligned slave audio data to obtain a conference pickup result based on the target audio data. According to the embodiment of the invention, the method can effectively improve the definition of conference pickup, provides high-quality audio experience, and enables participants to be more focused on the conference content.
Owner:DINGTALK (CHINA) INFORMATION TECH CO LTD

Time alignment of qmf-based processing data

To provide time alignment of encoded data of an audio encoder with associated metadata, such as spectral band replication (SBR) metadata.SOLUTION: An audio decoder (100, 300) configured to determine reconstructed frames of an audio signal (237) from access units (110) of a received data stream is described. The access unit (110) includes waveform data (111) and metadata (112), where the waveform data (111) and the metadata (112) are associated with a same reconstructed frame of the audio signal (127). An audio decoder (100, 300) comprises a waveform processing path (101, 102, 103, 104, 105) configured to generate a plurality of waveform subband signals (123) from waveform data (111), and a metadata processing path (108, 109) configured to generate decoded metadata (128) from metadata (111).SELECTED DRAWING: Figure 1
Owner:DOLBY INTERNATIONAL AB