Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

28 results about "Reference tone" patented technology

A reference tone is a pure tone corresponding to a known frequency, and produced at a stable sound pressure level (volume), usually by specialized equipment.

Sound effect evaluation method, device and equipment and storage medium

The application discloses an audio effect evaluation method and device, equipment and a storage medium. The application inputs a preset reference sound source into an audio device to be tested for playing, wherein the preset reference sound source is a sound source determined by statistical analysis of sound sources, in combination with a parameter corrected according to a preset evaluation result and a peak factor corresponding to a music playing platform; collects audio to be tested output by the audio device to be tested when playing the preset reference sound source; determines a relative frequency response and an estimated loudness value according to a spectrum corresponding to the audio to be tested and a spectrum corresponding to the preset reference sound source; and generates an audio effect evaluation result of the audio device to be tested according to the relative frequency response and the estimated loudness value. The application evaluates the audio effect of the audio device to be tested by using the reference sound source constructed in advance, solves the problem of traditional sweep frequency signal tuning deviation, greatly reduces the number of repeated trial and error, and significantly improves the efficiency, accuracy and reliability of audio product tuning work.
Owner:GEER TECH CO LTD

A speech deepfake attribution method and system under few-shot data

PendingCN122347960AReference sampleMedicine
The application provides a voice deep forgery attribution method and system under few sample data, and relates to the technical field of voice forgery attribution. The method comprises: obtaining a to-be-tested audio and a reference sample set; performing feature processing on the to-be-tested audio and the reference audio to obtain an embedding vector of the to-be-tested audio and an embedding vector of the reference audio; performing feature aggregation on the embedding vector of the reference audio to obtain a multi-prototype representation of the reference audio; and generating an attribution result of the to-be-tested audio according to the embedding vector of the to-be-tested audio and the multi-prototype representation of all reference audios, wherein the attribution result is a forgery category of the to-be-tested audio. The application overcomes the defect that a traditional single-prototype representation is difficult to cover a complex distribution, can effectively adapt to a few sample data scene, improves the attribution robustness in an unknown forgery scene, and gets rid of excessive dependence on a known category statistical hypothesis.
Owner:HEFEI UNIV OF TECH +1

A method and system for evaluating spatial resolution of a binaural rendering system

The application discloses a binaural rendering system spatial resolution evaluation method and system. The method is: generating test audio stimulus signals, including reference audio stimulus signals and comparison audio stimulus signals; using a binaural rendering system to be evaluated to process the test audio stimulus signals to obtain corresponding binaural audio; constructing a comparison stimulus azimuth angle corresponding to each reference azimuth angle; selecting corresponding binaural audio as the reference stimulus signal and the comparison stimulus signal from each binaural audio according to each reference azimuth angle and the comparison stimulus azimuth angle; playing the reference stimulus signal and the comparison stimulus signal respectively; confirming whether there is a difference according to a perceived position change; adjusting the angle difference between the reference azimuth angle corresponding to the reference stimulus signal and the comparison azimuth angle corresponding to the comparison stimulus signal until a preset termination condition is met, and determining the spatial resolution evaluation result of the binaural rendering system to be evaluated according to the minimum audible angle threshold of each reference azimuth angle.
Owner:PEKING UNIV

Electronic device and operating method to compensate for in-phase / quadrature imbalance

An electronic device includes a feedback oscillator configured to output a first oscillation signal and a second oscillation signal, the second oscillation signal having a defined phase difference from the first oscillation signal, the feedback oscillator including a phase shifter configured to receive the first oscillation signal and output the second oscillation signal, an up-conversion mixer configured to output a first loopback signal obtained by mixing the first oscillation signal with a reference tone signal, and output a second loopback signal obtained by mixing the second oscillation signal with the reference tone signal, and a receiver configured to generate a first reference IQ signal from the first loopback signal, generate a second reference IQ signal from the second loopback signal, and compare an actual phase difference between the first reference IQ signal and the second reference IQ signal with the defined phase difference.
Owner:SAMSUNG ELECTRONICS CO LTD

Song discrimination method, computer device, and storage medium

The application relates to a song identification method, a computer device and a storage medium, relates to the technical field of audio processing, and can improve the accuracy of an identification result of a synthesized song voice. The method comprises the following steps: acquiring a feature parameter corresponding to a human voice signal in a to-be-identified song; inputting the feature parameter corresponding to the human voice signal into a trained song identification model, outputting first identification information of the to-be-identified song by the song identification model according to the feature parameter corresponding to the human voice signal; comparing a timbre corresponding to the human voice signal in the to-be-identified song with each reference timbre, and determining second identification information of the to-be-identified song according to a comparison result; the reference timbre comprises a timbre corresponding to a synthesized song voice; and determining an identification result of the to-be-identified song according to the first identification information and the second identification information.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Information processing device, information processing method, and information processing program

PCT designated stageWO2026141037A1Information processingSound sources
An information processing device according to the present disclosure comprises: an analysis unit that acquires a reference sound source and analyzes the acquired reference sound source; an extraction unit that extracts, on the basis of a result of the analysis by the analysis unit, style information similar to the reference sound source from style information created in advance and used as training data when executing machine learning of a composition model; and a display unit that displays the style information extracted by the extraction unit.
Owner:SONY GROUP CORP

Voice interaction linkage control sound system

This invention discloses a voice-interactive linkage control audio system, relating to the field of voice interaction and audio-visual equipment linkage control technology. In a living room linkage playback scenario, the control host of this invention simultaneously uses playback reference audio, ambient audio, and device linkage status to determine the current interaction state. In the interstitial interaction state, it concatenates the pre-buffered audio with the subsequently acquired audio to form a command segment to be parsed. Then, it combines the media session state and device linkage state to generate a command execution graph containing immediate control items, sound effect control items, delayed execution items, and execution constraints. The graph is then validated based on the device linkage state, available media events, current sound effect mode, and execution constraints. This reduces misidentification in interstitial scenarios, avoids mutual rewriting of multiple actions, and minimizes the mismatch between delayed execution and the current playback state.
Owner:GUANGZHOU CHUANGKAI ELECTRONIC TECH CO LTD

Method of determining a perceptual impact of reverberation on the perceptual quality of a signal, and computer program product

The invention relates to a method of determining a perceptual impact of an amount of echo or an amount of reverberation in a degraded audio signal on a perceptual quality of the degraded audio signal, wherein the degraded audio signal is received from an audio transmission system, the degraded audio signal being obtained by a transmission of a reference audio signal by the audio transmission system to provide the degraded audio signal. The method comprises performing a windowing operation on the degraded audio signal and the reference audio signal by multiplying the degraded audio signal and the reference audio signal with a window function to produce degraded digital audio samples and reference digital audio samples. A local estimate of the amount of echo or the amount of reverberation is determined based on the degraded digital audio samples and the reference digital audio samples.
Owner:NEDERLANDSE ORG VOOR TOEGEPAST NATUURWETENSCHAPPELIJK ONDERZOEK TNO

Information interaction method and apparatus, and device and storage medium

The embodiments of the present disclosure relate to an information interaction method and apparatus, and a device and a storage medium. The method provided herein comprises: presenting an interaction interface with a virtual object; and in response to a first message received in the interaction interface satisfying a preset condition, providing music content, wherein a song corresponding to the music content is determined on the basis of the first message, and at least one audio attribute of the music content is generated on the basis of reference audio content associated with the virtual object.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Spatial region based audio separation

A method includes receiving target audio data captured by a first audio input device, the target audio data comprising a target audio signal and a first version of an interfering audio signal, and receiving reference audio data captured by a second audio input device different from the first audio input device, the reference audio data comprising a second version of the interfering audio signal. The method also includes processing, using a trained neural network, the target audio data and the reference audio data to generate enhanced audio data, the neural network attenuating the interfering audio signal in the enhanced audio data.
Owner:GOOGLE LLC

A speech echo cancellation system

This invention provides a speech echo cancellation system, relating to the field of audio signal processing technology. The system includes acquiring target audio data and reference audio data and importing them into an audio storage buffer area; reading the target audio data based on an FIR echo cancellation filter and performing convolutional noise cancellation processing using the audio filtering coefficients of the FIR echo cancellation filter to generate audio echo cancellation data; if the audio echo cancellation data meets preset filtering conditions, calculating the cancellation error between the audio echo cancellation data and the reference audio data; if the cancellation error meets preset update conditions, using an adaptive algorithm to calculate the audio filtering coefficient increment based on the cancellation error and updating the audio filtering coefficients; and performing echo component cancellation processing on the acquired current audio data using the FIR echo cancellation filter, outputting the current echo cancellation data to the playback end, thereby enabling accurate identification and cancellation of echo components in the audio signal.
Owner:SOUTH CHINA UNIV OF TECH

A multi-channel audio upmix method, apparatus, device and storage medium

PendingCN122177137ASpeech analysisStereophonic systemsNoiseDiffusion network
This application discloses a multi-channel audio upmixing method, apparatus, device, and storage medium, relating to the field of signal processing technology. The method includes: converting multi-track reference audio and stereo audio into Mel spectrograms; compressing the Mel spectrograms to a low-dimensional continuous latent space to obtain latent vectors; determining audio style features and audio content features based on stereo audio and multi-track reference audio using an encoder; performing noise prediction based on a diffusion network, the noisy latent vectors, the current time step, the audio style features, and the audio content features; comparing the prediction results with the actual noise; training the diffusion network based on the comparison results; using the trained diffusion network to output a target latent vector based on the target multi-track reference audio and target stereo audio; decoding the target latent vector into a multi-track Mel spectrogram; and converting the multi-track Mel spectrogram into a multi-channel audio waveform to complete the multi-channel audio upmixing. This achieves high-quality upmixing adaptation for different multi-channel systems.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Speech synthesis methods and apparatus, computer equipment, storage media

This invention relates to the field of artificial intelligence technology and discloses a speech synthesis method and apparatus, computer device, and storage medium, comprising: serializing reference text to obtain a text token sequence; generating an emphasis mask token sequence, and updating the token of each target specified character in the emphasis mask token sequence from 0 to 1 in response to detecting the presence of a target specified character in the reference text; performing emotion intensity recognition on reference audio to obtain an emotion intensity token; concatenating a pre-acquired emotion category token, emotion intensity token, emphasis mask token sequence, and text token sequence to obtain a target fusion sequence; performing attention processing on the target fusion sequence using a large language model to obtain a target speech token sequence; and generating speech from the target speech token sequence to obtain the target speech. This invention can be applied to speech synthesis scenarios in fintech and healthcare, solving the technical problem of poor emotional expression in synthesized speech.
Owner:PING AN TECH (SHENZHEN) CO LTD

Arias generation method based on neural-symbolic dual driving and programmed logic mapping

The application discloses a singing aria generation method based on neural-symbolic dual driving and programmed logic mapping, selects classic singing sections of various schools of opera to establish a reference audio library, extracts acoustic characteristics of each singing section, statistically analyzes style constraint features in the acoustic characteristics, generates a school and an acoustic fingerprint confidence interval mapping table, analyzes a natural language description requirement input by a user, generates a structured prompt word containing a style constraint feature confidence interval, retrieves a best matching singing section from the reference audio library, and finally generates a required singing aria according to the structured prompt word and acoustic characteristics of the best matching singing section.
Owner:ZHEJIANG UNIV OF TECH

In-vehicle apparatus

An in-vehicle apparatus according to an embodiment of the disclosure includes: a storage circuit configured to store pieces of reference audio data indicating sounds that are likely to be a hindrance to a driving operation of a driver; and a processing circuit configured to compare audio data included in a content with each of the pieces of reference audio data and perform a determination as to whether a sound indicated by the audio data is likely to be the hindrance to the driving operation, and configured to reduce, based on a result of the determination, a signal component that is included in the audio data and that is likely to be the hindrance to the driving operation.
Owner:SUBARU CORP

Audio processing methods, apparatus, electronic devices and storage media

This application discloses an audio processing method, apparatus, electronic device, and storage medium, comprising: performing frequency domain transformation on a first audio, a second audio, and a reference audio to obtain a first transformation result corresponding to the first audio, a second transformation result corresponding to the second audio, and a third transformation result corresponding to the reference audio; the first audio is acquired by a first microphone, the second audio is acquired by a second microphone, and the reference audio is a single-channel audio; based on the first transformation result, a first amplitude and a first phase are obtained; based on the second transformation result, a second amplitude and a second phase are obtained; based on the first phase and the second phase, a phase difference is obtained; the phase difference, the first amplitude, the second amplitude, and the third transformation result are combined to obtain a combined result; the combined result is input into a processing model to obtain audio separation results for multiple sound zones, wherein the multiple sound zones are the sound zones in a vehicle. This application embodiment achieves audio separation using only two microphones, saving costs.
Owner:TENCENT TECH WUHAN

Audio synthesis method and model training method, apparatus, device, medium and product

PendingCN122392481AAudio synthesisSynthesis methods
An audio synthesis method, model training method, apparatus, device, medium, and product are disclosed. The method includes: converting reference audio into a first audio with a specified timbre, wherein the timbre of the reference audio differs from the specified timbre, and the specified timbre is selected from a variety of candidate timbres; obtaining first text, extracting semantic features from the first text, and converting the first text into a first phoneme, wherein the first text describes the content of the audio to be synthesized; extracting first timbre features and first expression features from the first audio, wherein the first expression features include at least one of a first prosodic feature or a first emotional feature; concatenating the first text semantic features, the first phoneme, and the first expression features to predict the audio content, thereby obtaining a first audio content representation; and fusing the first audio content representation, the first phoneme, and the first timbre features to obtain a target audio with the specified timbre. This method can improve the controllability of timbre and the accuracy of audio synthesis.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Music generation method, apparatus, and electronic device

This invention provides a music generation method, apparatus, and electronic device, relating to the field of music technology. The method includes: acquiring music generation instructions and lyrics text, the music generation instructions including music composition constraint information; generating composition parameters based on the music composition constraint information; generating a target music style skeleton based on the music composition constraint information and composition parameters, the target music style skeleton including track information of multiple instruments; generating a mixing reference audio based on the track information of multiple instruments; and generating target music based on the lyrics text, composition parameters, and mixing reference audio. This invention can generate mixing reference audio based on music generation instructions, and ultimately generate target music by fusing multi-source information such as lyrics text, composition parameters, and mixing reference audio. This results in target music with natural and smooth vocals and accompaniment, a unified style, and a clear musical structure, thereby improving the accuracy of music generation.
Owner:HANGZHOU LIANYING NETWORK TECHNOLOGY CO LTD

Audio data processing method and apparatus, device, and medium

PendingCN122369482AFrequency spectrumNoise
The embodiment of the application provides an audio data processing method, device and equipment and medium, the method comprises: the target audio data and reference audio data are obtained Fourier transform to obtain relevant time-frequency spectrum data, to carry out feature extraction to obtain relevant amplitude data and phase data, then, after speech optimization processing is carried out to the target original input feature data constructed, the target optimization feature data after reference time-frequency spectrum data as noise data is separated and removed from target time-frequency spectrum data can be obtained, to carry out filtering processing to the fusion time-frequency spectrum data obtained based on the multi-path audio data to which the target audio data belongs, then, Fourier inverse transform is carried out to the filter time-frequency spectrum data corresponding to the target time-frequency spectrum data obtained, that is, the audio optimization data of the target audio data can be obtained and is used to carry out speech control. By adopting the embodiment of the application, noise in the target audio data can be effectively removed, so that the accuracy of speech control is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Methods and apparatus for affiliate interrupt detection

Methods, apparatus, systems and articles of manufacture are disclosed for affiliate interrupt detection. An example method disclosed herein includes determining whether a first time period of a first audio signal corresponds to a first affiliate interrupt period based on whether (1) a first type of watermark is detected in the first time period of the first audio signal, and (2) a second type of watermark is detected in the first audio signal outside the first time period but not in the first time period of the first audio signal, and determining whether the first time period of the first audio signal corresponds to the first affiliate interrupt period when watermarks are not detected in the first time period of the first audio signal based on comparison of first signatures with second signatures representing a corresponding first time period of a reference audio signal.
Owner:THE NIELSEN CO (US) LLC

Digital content management in virtual environments

Implementations described herein relate to methods, systems, and computer-readable media for digital content management. In some implementations, the method includes receiving, by a processor, an audio file, determining whether there is a match of a segment of the audio file with one or more reference audio files, if it is determined that there is no match, classifying the audio file as an authentic audio file, and if it is determined that there is the match, identifying, one or more designated audio segments that are semantically similar to the segment, and replacing the segment of the audio file with a particular designated audio segment of the one or more designated audio segments.
Owner:ROBLOX CORP

Video classification method and apparatus

Provided is a method for classifying videos. The method comprises: acquiring a target audio and a corresponding target video comprising human body actions; determining, based on a human body action matching degree of each target frame in the target video relative to a corresponding reference frame in a reference video, a total human body action matching degree score of the target video relative to the reference video; determining, based on an audio matching degree of each target audio segment in the target audio relative to a corresponding reference audio segment in a reference audio of the reference video, a total audio matching degree score of the target audio relative to the reference audio; and determining, based on the total human body action matching degree score and the total audio matching degree, a comprehensive classification result.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Intelligent dubbing generation method and device fusing voiceprint features and semantic understanding

The application discloses a method and device for generating intelligent dubbing by fusing voiceprint features and semantic understanding, and relates to the technical field of artificial intelligence and speech synthesis. The application firstly acquires semantic features of a target text and voiceprint features of a reference audio; secondly, a feature alignment module based on a multi-head self-attention mechanism is constructed to map the voiceprint features to a semantic hidden space and generate semantic vectors with speaker identity attributes; subsequently, a sentiment reasoning engine is introduced to dynamically predict sentiment labels and prosody parameters according to semantic contexts; finally, an improved neural vocoder is used to generate high-fidelity audio waveforms. The application solves the problem of disconnection between tone color and sentiment and harsh prosody in existing speech synthesis technology, realizes deep fusion of tone color cloning and text sentiment expression, and significantly improves the realism, expressiveness and degree of personalization of the generated dubbing.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

AI noise reduction speech enhancement method for multi-mode wireless earphone

The application discloses an AI noise reduction voice enhancement method of a multi-mode wireless earphone, and relates to the field of audio signal processing, which comprises the following steps: synchronously collecting feedforward, feedback microphone signals and reference audio signals, and obtaining a working mode instruction, and obtaining an acoustic feature vector after processing; performing wind noise analysis based on the vector to generate a wind noise scene description vector, simultaneously generating a sound quality deviation degree vector based on the feedback and reference signals, and analyzing a multi-objective optimization weight vector based on the working mode; calculating a mode fidelity and comfort degree conflict factor based on the wind noise scene description vector and the multi-objective optimization weight vector; splicing all the vectors and the conflict factor into an enhancement system state vector, inputting the enhancement system state vector into a second neural network model which is pre-trained, and outputting a cooperative control vector which is used for coordinating downstream processing; and performing adaptive noise reduction, voice enhancement or environmental sound mixing processing and dynamic equalization compensation based on the cooperative control vector, and synthesizing a final output signal.
Owner:SHENZHEN BOLUKE ELECTRONIC TECH CO LTD

Audio denoising method for a bluetooth communication earphone

The application discloses an audio denoising method of a Bluetooth communication earphone and belongs to the technical field of audio processing. The method comprises the following steps: acquiring a mixed audio stream containing human voice and background sound and a reference audio stream mainly containing background sound, extracting spatial sound correlation coefficients, real-time sound energy deviation degrees and sound frequency distribution concentration degrees respectively; generating a frequency point update confidence value for controlling an adjustment process; synchronously calculating an overall working state index of a current processing time period; dynamically adjusting a historical data dependence weight in a sound intelligibility estimation process; and combining temporarily stored audio phase data to perform reverse conversion to obtain a time sequence sound output signal. The application solves the problem of false update or false freezing of an adaptive filter when a sound field changes due to rigid gating of a speech activity detector by performing joint analysis of spatial correlation degrees and energy characteristics of the mixed audio stream and the reference audio stream in a frequency point by frequency point manner.
Owner:SHENZHEN HENGCLOUDS ECOLOGY TECH CO LTD

Voice processing collaboration method based on dual-cuda stream

This invention relates to a collaborative speech processing method based on dual CUDA streams, belonging to the field of speech processing technology. The method includes: creating a first CUDA computation stream and a second CUDA computation stream on a GPU; the first CUDA computation stream is used for acoustic echo cancellation and speech activity detection, and the second CUDA computation stream is used for speaker separation; receiving near-end microphone audio data packets and far-end reference audio data packets through a WebSocket access layer; a signal alignment module performing timestamp verification, queue buffering, and frame-level pairing output on the near-end microphone audio data packets and far-end reference audio data packets to form aligned frame data; performing acoustic echo cancellation frame-by-frame on the aligned frame data in the first CUDA computation stream and generating echo-cancelled audio, while simultaneously generating speech activity detection results; this invention achieves hardware-level parallel execution of the dual computation streams and on-demand dynamic allocation of computing resources, significantly reducing end-to-end processing latency and improving the efficiency of graphics processor resource utilization.
Owner:SHENZHEN XINZHI FUTURE TECHNOLOGY CO LTD

Unstructured business data interaction method and system based on large language model

PendingCN122392537AEngineeringBusiness data
The present application relates to the technical field of data processing, in particular to an unstructured business data interaction method and system based on a large language model, which solves the technical problem of poor interaction effect caused by recognition errors in the prior art. The method comprises: obtaining a to-be-processed audio signal carrying user speech and at least one historical reference audio signal; performing audio feature analysis on the audio signals of each time period in the to-be-processed audio signal based on the at least one historical reference audio signal, and determining the speech retention degree corresponding to the audio signals of each time period in the to-be-processed audio signal; the speech retention degree is used to represent the attention weight of the audio signals of the corresponding time period in the filtering process; filtering the to-be-processed audio signal according to the speech retention degree corresponding to the audio signals of each time period to obtain a target speech signal; converting the target speech signal into text information, inputting the text information into a large language model, and outputting a feedback result for the text information.
Owner:BEIJING DEXUN AVIATION SERVICE CO LTD