Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

195 results about "Reference tone" patented technology

A reference tone is a pure tone corresponding to a known frequency, and produced at a stable sound pressure level (volume), usually by specialized equipment.

Single-microphone acoustic echo and noise suppression

This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques for separating microphone signals into speech, echo, and noise signals. In some aspects, a speech enhancement system may include a delay estimator and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a microphone signal via a microphone and a far-end audio signal for output via a speaker and estimates a reference audio signal based on a delay between the microphone signal and the far-end audio signal. In some aspects, the AEN decoupling filter may determine a speech mask, an echo mask, and a noise mask based on the microphone signal and the reference audio signal and may suppress an echo component and a noise component of the microphone signal based on the determined set of masks.
Owner:SYNAPTICS INC

Audio generation method and apparatus based on large language model, electronic device, and storage medium

A method of audio generation based on a large language model is disclosed, which involves the fields of artificial intelligence such as large language models, natural language processing, deep learning, and audio generation. The method of audio generation based on a large language model comprises: acquiring a text to be processed; parsing the text to be processed using the large language model to obtain role information and emotional information corresponding to the text to be processed; obtaining a target reference text and a target reference audio according to the role information and the emotional information; and generating a target audio corresponding to the text to be processed according to the text to be processed, the target reference text, and the target reference audio.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Parallel tacotron non-autoregressive and controllable TTS

A method for training a non-autoregressive TTS model includes receiving training data that includes a reference audio signal and a corresponding input text sequence. The method also includes encoding the reference audio signal into a variational embedding that disentangles the style / prosody information from the reference audio signal and encoding the input text sequence into an encoded text sequence. The method also includes predicting a phoneme duration for each phoneme in the input text sequence and determining a phoneme duration loss based on the predicted phoneme durations and a reference phoneme duration. The method also includes generating one or more predicted mel-frequency spectrogram sequences for the input text sequence and determining a final spectrogram loss based on the predicted mel-frequency spectrogram sequences and a reference mel-frequency spectrogram sequence. The method also includes training the TTS model based on the final spectrogram loss and the corresponding phoneme duration loss.
Owner:GOOGLE LLC

Method and device for generating digital human video

The invention relates to the technical field of digital people, and provides a method and equipment for generating a digital people video, which can be applied to a digital marketing scene. According to the method, the generation of the digital human video is divided into an audio generation process and a video generation process, when the audio is generated, a target text, an emotion parameter and a reference audio are taken as input, and the target audio of the target text is broadcasted by the sound of a reference person in the reference video and a set emotion, so that sound cloning is realized, the degree of violation of the generated digital human video is reduced, and the user experience is improved. And the text features are extracted from two dimensions of characters and pronunciation, so that the pronunciation accuracy during sound cloning is ensured. When the video is generated, the cloned target audio and the cloned reference video are used as input, and the fidelity of the digital human video is improved through expression and mouth shape matching, so that a reference person can quickly and efficiently generate vivid digital human videos with different broadcast contents only by shooting one reference video without shooting one by one, thereby reducing the manufacturing cost of the digital human video, and improving the user experience. And the video production efficiency is improved.
Owner:JUHAOKAN TECH CO LTD

Audio processing method and device

The invention provides an audio processing method and device, and the method comprises the steps: collecting to-be-detected audio data related to a target audio based on to-be-detected audio equipment, and determining reference audio data related to the target audio; determining a to-be-tested audio index corresponding to the to-be-tested audio data in an audio consistency dimension, determining a reference audio index corresponding to the reference audio data, and generating an index difference score based on the to-be-tested audio index and the reference audio index; and under the condition that the to-be-tested audio data is determined to be abnormal based on the index difference score, generating difference processing prompt information corresponding to the to-be-tested audio data based on a difference mode rule and the index difference score. The audio processing method can be widely applied to the field of digital music in the digital creative industry.
Owner:GUANGZHOU SEASUN ENTERTAINMENT NETWORK TECHCO

ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion

PendingCN121354597ASpeech synthesisFeature extractionSpeaking style
The invention provides an ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion, and the method comprises the steps: obtaining a reference audio of a target user, and carrying out the decoupling extraction of a volume feature vector, a pitch feature vector, and a speaking style feature vector from the reference audio through a multi-dimensional pronunciation feature extraction network; a text to be verified is converted into a phoneme sequence, the phoneme sequence is input into a pronunciation feature prediction network for predicting a frame-level dynamic pronunciation feature sequence based on the phoneme sequence, and three pronunciation feature vectors are injected into the network through an adaptive instance normalization mechanism to dynamically modulate a prediction process. Generating a frame-level dynamic pronunciation feature sequence containing time sequence rhythm change; inputting a VAE-GAN synthesis trunk, re-injecting the three pronunciation feature vectors through an adaptive instance normalization mechanism, and generating a test voice sample consistent with the voiceprint characteristics of the target user; and initiating an identity verification query for the ASV system, and calculating safety indexes of different user groups according to a verification result so as to evaluate the risk of the ASV system.
Owner:FUJIAN NORMAL UNIV

A method, apparatus, device, medium and program product for generating an audio file

Embodiments of the present disclosure provide a method, device, medium and program product for generating an audio file. The method comprises: displaying an editing control on a playing page; in response to an interaction operation on the editing control, displaying an audio editing page, the audio editing page comprising audio information of a preset audio file and a generation control; in response to a selection operation on the audio information, determining a reference audio segment in the preset audio file; and in response to an interaction operation on the generation control, generating a target audio segment according to the reference audio segment, and determining a target audio file according to the target audio segment. The technical solution of the embodiments of the present disclosure generates a target audio segment of similar style according to a reference audio segment selected by a user, which can meet the user's expectation of music generation, reduce the difficulty of music creation, and improve the user experience.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Speech synthesis methods, devices, computer equipment and storage media

This application provides a speech synthesis method, apparatus, computer device, and storage medium. The method relates to speech synthesis technology and is applied in the financial field. It includes: acquiring initial text to be synthesized and reference audio; inputting the initial text into a preset phoneme encoder, outputting multiple phonemes corresponding to the initial text; inputting the reference audio into a preset prosodic encoder, outputting multiple prosodices corresponding to the reference audio; aligning each phoneme and multiple prosodices based on a self-attention mechanism, acquiring embedding information corresponding to each phoneme; the embedding information includes at least one prosodic and a weight corresponding to each prosodic; generating a Mel spectrum based on each phoneme and its corresponding embedding information; generating synthesized audio based on the Mel spectrum, thus completing the prosodic alignment of the initial text and the reference audio. After synthesis, the prosodices will appear in accurate positions in the audio, improving the naturalness, credibility, persuasiveness, and appeal of the speech, and adapting to different application scenarios and user preferences.
Owner:PING AN TECH (SHENZHEN) CO LTD

Story audio timbre processing method and related device

The invention discloses a story audio timbre processing method and a related device, and relates to the technical field of audio processing, and the method comprises the steps: extracting a story voice from a to-be-processed story audio before carrying out the timbre processing of the to-be-processed story audio through employing a reference audio, and then carrying out the timbre conversion processing of the story voice based on the reference audio, thereby achieving the timbre processing of the to-be-processed story audio. And finally, determining a final story audio based on the target story voice. In the whole processing process, voice recognition is not needed, the influence of voice recognition accuracy on the tone processing effect is avoided, the to-be-processed audio is subjected to story voice extraction and story voice processing, the influence of story background voice on the tone processing effect is avoided, and therefore the voice processing efficiency is improved. According to the scheme, the tone processing effect of the story audio can be improved.
Owner:HEFEI IFLYTEK TOYCLOUD TECH

Sound effect evaluation method, device and equipment and storage medium

The application discloses an audio effect evaluation method and device, equipment and a storage medium. The application inputs a preset reference sound source into an audio device to be tested for playing, wherein the preset reference sound source is a sound source determined by statistical analysis of sound sources, in combination with a parameter corrected according to a preset evaluation result and a peak factor corresponding to a music playing platform; collects audio to be tested output by the audio device to be tested when playing the preset reference sound source; determines a relative frequency response and an estimated loudness value according to a spectrum corresponding to the audio to be tested and a spectrum corresponding to the preset reference sound source; and generates an audio effect evaluation result of the audio device to be tested according to the relative frequency response and the estimated loudness value. The application evaluates the audio effect of the audio device to be tested by using the reference sound source constructed in advance, solves the problem of traditional sweep frequency signal tuning deviation, greatly reduces the number of repeated trial and error, and significantly improves the efficiency, accuracy and reliability of audio product tuning work.
Owner:GEER TECH CO LTD

Artificial intelligence systems and methods for detecting musical infringement in symbolic music

The present disclosure relates to a system and method for detecting musical infringement in symbolic music. The system receives a first musical composition and, when provided as audio, converts it into a symbolic format using a transcription neural network trained to extract a main melody. It then generates k-mer sequences comprising consecutive notes, indexes these sequences in a data structure configured for dynamic conditioning, and compares them to reference musical compositions. Upon estimating a similarity measure that exceeds a threshold, the system performs a refined local sequence alignment adapted for music, accounting for key shifts, rests, and melodic or rhythmic variations. Based on this refined alignment, the system determines whether the first musical composition includes a musical fragment that infringes upon or regurgitates a portion of at least one reference composition.
Owner:SOUND PATROL INC

Music detection method, apparatus, device, and medium

The application relates to the technical field of data processing, and particularly provides a music detection method and device, equipment and a medium. The music detection method comprises the following steps: obtaining target music description information of target music to be detected; obtaining reference music description information of reference music corresponding to the target music; determining a score similarity between the reference music and the target music according to first score information in the reference music description information and second score information in the target music description information; and obtaining an originality detection result of the target music according to the score similarity; and the originality detection result indicates whether the target music has originality. In this way, the cost of music originality detection can be reduced, and the accuracy of music originality detection can be improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Portrait dialogue video generation method, multi-person dialogue video generation method, product, equipment and storage medium

The invention discloses a portrait dialogue video generation method, a multi-person dialogue video generation method, a product, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the portrait dialogue video generation method comprises the steps: extracting a face parameter of a first dialogue object from a reference face image; determining a primary audio of the first dialogue object and a secondary audio of the second dialogue object based on the reference audio; fusing the primary audio of the first dialogue object and the secondary audio of the second dialogue object to obtain a fused audio feature; constructing a three-dimensional portrait geometric sequence of the first dialogue object according to the fused audio features and the face parameters of the first dialogue object; and generating a portrait dialogue video of the first dialogue object based on the reference audio, the reference face image, the fused audio features and the three-dimensional portrait geometric sequence. The method aims at supporting dynamic switching between the speaking state and the listening state of the dialogue object in the portrait dialogue video, the portrait expression modeling precision of the dialogue object is improved, and the interaction fluency of multi-person dialogue is improved.
Owner:GUANGDONG-HONG KONG-MACAO GREATER BAY AREA DIGITAL ECONOMY RESEARCH INSTITUTE (INTERNATIONAL ADVANCED TECHNOLOGY APPLICATION PROMOTION CENTER (SHENZHEN) +1

Multi-channel speech compression system and method

A method, computer program product, and computing system for encoding audio encounter information of a reference audio acquisition device of a plurality of audio acquisition devices of an audio recording system, thus defining encoded reference audio encounter information. Location information may be estimated, via a machine vision system, for an acoustic source within an acoustic environment. One or more acoustic relative transfer functions may be selected from a plurality of acoustic relative transfer functions for the plurality of audio acquisition devices of the audio recording system based upon, at least in part, the location information. The encoded reference audio encounter information and a representation of the selected one or more acoustic relative transfer function may be transmitted.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Audio processing method and device

The embodiment of the invention provides an audio processing method and device, computer equipment, a computer readable storage medium and a computer program product, and belongs to the field of data processing. The audio processing method comprises the following steps: acquiring a pitch feature and a bottleneck feature of an initial audio; generating respective Gaussian distributions of the plurality of frequency bands according to the pitch features, the bottleneck features and the reference timbre, and generating a plurality of segmented spectrums based on the Gaussian distributions of the corresponding frequency bands through a plurality of segmented decoders; synthesizing the plurality of segmented frequency spectrums to generate a synthesized frequency spectrum; a target frequency spectrum is generated through a target decoder according to the synthetic frequency spectrum and the pitch characteristics, and the smoothness of the target frequency spectrum is higher than that of the synthetic frequency spectrum; and taking the target frequency spectrum as model input, and generating a target audio through a pre-trained audio generation model. According to the technical scheme provided by the embodiment of the invention, fine processing can be carried out according to different characteristics of audios of different frequency bands, so that the audio processing quality is improved.
Owner:SHANGHAI HODE INFORMATION TECH CO LTD

Sound replication method and related apparatus

The application provides a sound replication method and related device, and relates to the technical field of sound processing. The reference audio is subjected to audio verification to obtain first audio, the first audio is subjected to a speech enhancement operation to achieve the purposes of noise reduction, dereverberation and improvement of the signal-to-noise ratio of the audio, thereby obtaining second audio, the second audio is subjected to a speech activity detection and segment division operation to obtain a candidate speech segment, a target speech segment meeting the sound replication requirement is selected from the candidate speech segment, the optimal segment with clear timbre, high signal-to-noise ratio and stable pronunciation is selected from the candidate speech segment for timbre embedding extraction and speech generation, the probability of timbre deviation of the generated speech is reduced, the accuracy of the speech generation is improved, and the user experience is improved.
Owner:BEIJING SOHU NEW MEDIA INFORMATION TECH

High-performance zero-sample text-to-speech conversion method and system based on vLLM acceleration

The invention discloses a high-performance zero-sample text-to-speech method and system based on vLLM acceleration, and belongs to the technical field of intelligent speech, and the system comprises a dynamic streaming sentence segmentation module which carries out the intelligent segmentation of an input text, and obtains a to-be-converted text; the multi-reference audio fusion module is used for receiving a plurality of reference audios of a speaker and combining weight fusion to obtain a final audio coding feature; the session management and control module is used for judging whether the to-be-converted text is a new session request according to the session ID of the to-be-converted text, and if the to-be-converted text is the new session request, distributing a speaker ID and an associated audio signaling feature to the to-be-converted text; if not, searching a speaker ID (Identity) and an audio conditioning feature; and the text-to-voice module is used for performing voice synthesis on the text to be converted by adopting the IndexTTS model accelerated by the vLLM. According to the method, the reasoning performance can be remarkably improved, the audio quality and the response time delay stability in a multi-round dialogue scene are ensured, and the natural continuity of long text synthesis is ensured.
Owner:JIANGSU HAOBAI INFORMATION SERVICE CO LTD

A speech deepfake attribution method and system under few-shot data

PendingCN122347960AReference sampleMedicine
The application provides a voice deep forgery attribution method and system under few sample data, and relates to the technical field of voice forgery attribution. The method comprises: obtaining a to-be-tested audio and a reference sample set; performing feature processing on the to-be-tested audio and the reference audio to obtain an embedding vector of the to-be-tested audio and an embedding vector of the reference audio; performing feature aggregation on the embedding vector of the reference audio to obtain a multi-prototype representation of the reference audio; and generating an attribution result of the to-be-tested audio according to the embedding vector of the to-be-tested audio and the multi-prototype representation of all reference audios, wherein the attribution result is a forgery category of the to-be-tested audio. The application overcomes the defect that a traditional single-prototype representation is difficult to cover a complex distribution, can effectively adapt to a few sample data scene, improves the attribution robustness in an unknown forgery scene, and gets rid of excessive dependence on a known category statistical hypothesis.
Owner:HEFEI UNIV OF TECH +1

Audio noise reduction result detection method and device, electronic equipment and storage medium

The invention relates to the technical field of audio signal processing, and provides an audio noise reduction result detection method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a first time-frequency feature data pair corresponding to homologous double-channel audio data; performing time sequence alignment on the first time-frequency characteristic data pair to obtain a second time-frequency characteristic data pair after time sequence alignment; generating audio input data based on the second time-frequency characteristic data pair; inputting the audio input data into a deep learning model, outputting a first quality parameter through a regression head of the deep learning model, and outputting text diagnosis information through a generation head of the deep learning model; and based on the first quality parameter and the text diagnosis information, determining a detection result of noise reduction of the audio data. According to the embodiment of the invention, the method achieves the detection of the audio noise reduction result under the condition of no clean reference audio, and improves the detection universality of the audio noise reduction result and the interpretability of the detection result.
Owner:ZHUHAI MOJIE TECH CO LTD

Sound duplicating method, device and equipment and storage medium

The embodiment of the invention provides a sound copying method and device, equipment and a storage medium, which can be applied to scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving, audio and video, in the method, feature extraction is performed on reference audio in advance, and reference timbre features and reference rhythm features are obtained. And obtaining and storing a sound feature file of the reference audio based on the reference timbre feature and the reference rhythm feature, so that when the reference audio is used as input for multiple times of audio synthesis, only the sound feature file of the reference audio needs to be read in each time of audio synthesis, and based on the first text feature of the text to be synthesized and the sound feature file, the sound feature file of the reference audio is read. According to the method and the device, the synthetic audio corresponding to the to-be-synthesized text is generated without repeatedly reading the reference audio and repeatedly calculating the reference audio, so that the time consumption and resource consumption of sound copying are effectively reduced, and the waiting time of synthesizing the audio by using a sound copying model is also effectively reduced, thereby improving the use experience and enhancing the controllability of the system.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method, apparatus and device for generating dynamic image based on audio, and storage medium

Embodiments of the present application provide a method and device for generating a dynamic image based on audio, an apparatus, and a storage medium, relating to the field of natural human-computer interaction. The method comprises: first obtaining a reference image and a reference audio input by a user; then, based on the reference image and a trained generation network model, determining a target head action feature and a target expression coefficient feature, and adjusting the trained generation network model based on the target head action feature and the target expression coefficient feature to obtain a target generation network model; finally, based on the reference audio, the reference image, and the target generation network model, processing a to-be-processed image to obtain a target dynamic image; wherein the to-be-processed image is the same as an image object in the reference image; in this way, a corresponding digital person can be obtained based on a single picture of a target person; in this way, video acquisition work and data cleaning work are not required, the production cost of the digital person can be reduced, and the production cycle of the digital person is shortened.
Owner:JIAXING SILICON INTELLIGENT TECHNOLOGY CO LTD

Voice cloning method and system

The invention provides a voice cloning system and method, and the system comprises a data preprocessing module, a feature extraction module, a model training module, and a reasoning generation module, and the data preprocessing module is used for carrying out the human voice accompaniment separation, audio cutting, and automatic marking of an original audio; the feature extraction module is used for extracting a self-encoding feature and a Mel spectrum feature from the audio and pairing the self-encoding feature and the Mel spectrum feature with the text length; the model training module is used for performing fine tuning training on the basis of a small amount of reference audio data by combining a GPT-like model and a VITS and Valle model architecture; and the reasoning generation module is used for generating target audio according to the reference audio and the text provided by the user. According to the method and the device, high-quality voice cloning and text-to-voice conversion can be realized only by a small amount of data sets, so that the user experience is improved, and the development cost is also reduced.
Owner:SHANGHAI MANJU NETWORK TECHNOLOGY CO LTD

Audio signal self-adaptive compensation adjusting system and method applied to earphone

The invention relates to the technical field of audio signal processing, and discloses an audio signal adaptive compensation adjustment system and method applied to an earphone, and the system comprises a multi-mode sensing module which is configured to be used for collecting an actual acoustic signal, a physiological vibration signal and a reference audio signal; a theoretical compensation calculation module configured to calculate a theoretical compensation gain based on the reference audio and the actual acoustic signal; the leakage source identification module is configured to be used for extracting acoustic and vibration characteristics and executing cross-modal fusion analysis so as to output a leakage source classification state; and the dynamic compensation state machine and application module is configured for selecting different smoothing parameters and smoothing the theoretical gain according to the leakage source state so as to calculate the final compensation gain and applying the final compensation gain to the reference audio signal. According to the invention, through cross-modal fusion analysis of acoustic and physiological vibration, accurate distinguishing and dynamic smooth compensation of leakage sources are realized, and hearing comfort is improved.
Owner:BESING TECH SHENZHEN CO LTD

Tone conversion method and device, storage medium and program product

The invention discloses a timbre conversion method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining original voice recognition features of original voice and reference timbre recognition features of reference voice, carrying out the timbre removal of the original voice recognition features, obtaining target original voice recognition features, and carrying out the timbre conversion of the target original voice recognition features; and generating a target voice at least based on the target original voice recognition feature and the reference tone recognition feature, the tone of the target voice being the same as the tone of the reference voice, and the voice content of the target voice being the same as the voice content of the original voice. According to the method and the device, the timbre is removed from the voice recognition features of the original voice, so that the problem of timbre leakage is relieved.
Owner:IFLYTEK CO LTD

Large language model processing method, audio processing method and related device

The embodiment of the invention provides a large language model processing method, an audio processing method and a related device, which are used for reducing the number of models on the premise of generating the same features. The method provided by the embodiment of the invention comprises the following steps: acquiring a sample pitch sequence, a sample lyric sequence and music score data of a sample audio; acquiring a first reference feature of the sample reference audio; performing mask processing on the sample pitch sequence and the sample lyric sequence to obtain a processed mask pitch sequence and a processed mask lyric sequence; processing the music score data, the first reference feature, the mask pitch sequence and the mask lyric sequence through a large language model to obtain a predicted pitch sequence and a predicted lyric sequence; determining training loss based on the predicted pitch sequence and the predicted lyric feature sequence; and optimizing the large language model based on the training loss until the large language model converges to obtain a target large language model.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Terminal equipment and external audio equipment echo cancellation method

Some embodiments of the invention show a terminal device and an external audio device echo cancellation method, and the method comprises the steps: adding a first mark signal to a reference audio signal, and transmitting the reference audio signal to an external audio device, so that the external audio device plays a recombined audio signal. Receiving a collected environment audio signal; separating a second mark sequence and a main audio signal from the environment audio signal; calculating time delay, frequency offset and a sound wave arrival angle based on the second mark sequence; compensating the reference audio signal according to the time delay and the frequency offset; cancelling the compensated reference audio signal from the main audio signal; and suppressing residual echo according to the sound wave arrival angle to obtain user audio. According to the embodiment of the invention, the mark signal with the mark sequence is implanted at the sending end, the mark sequence is separated at the receiving end, and the time delay, the frequency offset and the sound wave arrival angle are calculated through the mark sequence, so that time delay and frequency offset compensation is provided for echo cancellation, residual echo is suppressed based on the sound wave arrival angle, and the echo cancellation effect is improved.
Owner:JUHAOKAN TECH CO LTD

Passenger recognition device mounted in vehicle, and method by which passenger recognition device performs calibration

A passenger recognition device for recognizing a passenger in a vehicle, and a method by which the passenger recognition device performs calibration are provided. The passenger recognition device can perform a calibration operation in which an available frequency region is determined by acquiring frequency response characteristics related to ultrasonic signals in an inaudible frequency band for each of a plurality of speakers, volume values output by each of the plurality of speakers are adjusted for each frequency within the available frequency region, and thus offset data for adjusting the output volume of each of the plurality of speakers is acquired so that ultrasonic signals having signal level values greater than or equal to a reference volume is received by a microphone, and offset data about the available frequency region and the output volume for each frequency are stored for each of the plurality of speakers.
Owner:SAMSUNG ELECTRONICS CO LTD

Method for learning an audio quality metric combining labeled and unlabeled data

Described is a method of training a neural-network-based system for determining an indication of an audio quality of an audio input. The method includes obtaining, as input, at least one training set comprising audio samples. The audio samples include audio samples of a first type and audio samples of a second type, wherein each of the first type of audio samples is labelled with information indicative of a respective predetermined audio quality metric, and wherein each of the second type of audio samples is labelled with information indicative of a respective audio quality metric relative to that of a reference audio sample. The method further includes: inputting the training set to the neural-network-based system; and iteratively training the system to predict the respective label information of the audio samples in the training set.
Owner:DOLBY INTERNATIONAL AB

Pronunciation feedback generator

A device includes a memory configured to store input audio that corresponds to speech representing a target sentence spoken by a user. The device also includes one or more processors configured to detect a prosody component of the speech. The one or more processors are also configured to detect a phonetic component of the speech. The one or more processors are configured to perform a prosody comparison of a reference prosody component and the detected prosody component. The one or more processors are configured to perform a phonetics comparison of a reference phonetic component and the detected phonetic component. Each of the reference prosody component and the reference phonetic component is based on the target sentence with speech characteristics of the user and having a target pronunciation. The one or more processors are configured to generate an output based on the prosody comparison and the phonetics comparison.
Owner:QUALCOMM INC