Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

301 results about "Reference tone" patented technology

A reference tone is a pure tone corresponding to a known frequency, and produced at a stable sound pressure level (volume), usually by specialized equipment.

Voice enhancement method and device based on noise perception, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice enhancement method, device and equipment based on noise perception and a medium. Environment feature information is extracted and input into an audio enhancement model to generate an enhanced audio signal; obtaining a reference audio sample, extracting a personalized feature vector, and carrying out personalized processing on the enhanced audio signal; and collecting playing feedback data, determining a playing time domain adjustment parameter and a playing frequency domain adjustment parameter, adjusting the personalized enhanced audio signal, and generating an optimized audio signal. According to the method, dynamic adjustment is realized in combination with the feedback parameters in the playing process by fusing the environmental perception information and the personalized speaker characteristics, clear and natural optimized audio output with personalized styles can be generated in a complex environment, and the voice interaction quality and adaptability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice cloning method and device based on emotion enhancement and related medium

The invention discloses a speech cloning method and device based on emotion enhancement, and a related medium. The method comprises the following steps: respectively obtaining a reference audio and a prediction text corresponding to the reference audio; respectively preprocessing the reference audio and the prediction text to generate standardized audio data for feature extraction; inputting the standardized audio data into an emotion enhancement module so as to generate acoustic features matched with a target voice style through multi-round feature fusion and an autoregression generation mechanism; and inputting the acoustic features into a speech synthesis module for decoding processing, and outputting predicted speech. According to the invention, by introducing the emotion enhancement module, accurate modeling of the emotion style of the speaker in the acoustic feature generation process is realized, so that the emotion simulation degree and the personality restoration capability of the synthesized speech are significantly improved.
Owner:AFIRSTSOFT CO LTD

Fault server positioning method and device, storage medium and electronic equipment

The invention discloses a fault server positioning method and device, a storage medium and electronic equipment, and relates to the technical field of computers, and the method comprises the steps: detecting the operation state of a server in a server room, responding to a target server in a fault operation state, and controlling the target server to output an audio signal meeting a target audio parameter. The characteristic that an audio collector can collect audio signals is utilized, and reference audio signals output by all servers and collected by the audio collector are received. The audio signal meeting the target audio parameter is screened out from the reference audio signal, so that the audio collector which collects the audio signal output by the target server is determined, and the target server is positioned through the position of the audio collector. The technical problem that the positioning efficiency of the fault server is low is solved, and the technical effect of improving the positioning efficiency of the fault server is achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Neural network generation of speech

The invention relates to neural network speech generation. Apparatus, systems, and techniques for generating speech audio are disclosed. In at least one embodiment, a processor generates a first speech audio based at least in part on a second speech audio and a reference audio using one or more neural networks.
Owner:NVIDIA CORP

End side voice model deployment method and device, equipment and storage medium

The invention relates to the technical field of end-side model deployment, in particular to an end-side voice model deployment method and device, equipment and a storage medium. Comprising the steps of recording reference audio at an end side; extracting a reference semantic token and an embedded vector of the reference audio; obtaining a training text set, integrating and inputting the training text code, the embedded vector and the reference semantic token corresponding to each training text in the training text set into a preset language model, and outputting a comprehensive token sequence corresponding to the training text; performing Mel spectrum conversion on each comprehensive token sequence to obtain a first spectrum representation corresponding to the comprehensive token sequence; generating an audio signal corresponding to the training text according to each first spectrum representation, and integrating to obtain a training data set; inputting the training data set into a to-be-trained model, and training to obtain a lightweight voice model which only retains timbre modeling parameters; and updating the lightweight voice model to the end side to complete end side deployment. According to the invention, the voice synthesis model deployment of the end-side equipment can be realized.
Owner:SHENZHEN RAISOUND TECH

Single-microphone acoustic echo and noise suppression

This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques for separating microphone signals into speech, echo, and noise signals. In some aspects, a speech enhancement system may include a delay estimator and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a microphone signal via a microphone and a far-end audio signal for output via a speaker and estimates a reference audio signal based on a delay between the microphone signal and the far-end audio signal. In some aspects, the AEN decoupling filter may determine a speech mask, an echo mask, and a noise mask based on the microphone signal and the reference audio signal and may suppress an echo component and a noise component of the microphone signal based on the determined set of masks.
Owner:SYNAPTICS INC

Speech synthesis device and training method thereof, electronic equipment and storage medium

The invention relates to a voice synthesis device and a training method thereof, electronic equipment and a storage medium, and the method comprises a voiceprint coding module which receives an input reference audio and extracts voiceprint features of the reference audio; the audio code generation module is used for determining a target expert network from a plurality of expert networks to execute a voice synthesis operation according to the input text to be synthesized and the voiceprint features so as to obtain a generated audio code sequence; wherein the plurality of expert networks correspond to different application scenes; and the audio decoding module is used for converting the audio coding sequence output by the audio coding generation module into an audio signal. According to the embodiment of the invention, through the dynamically selected target expert network, the voice synthesis of the application scene suitable for the input to-be-synthesized text and the voiceprint feature can be executed, and the scene adaptability of the synthesized voice is remarkably improved.
Owner:MOORE THREADS TECH CO LTD

Audio generation method and apparatus based on large language model, electronic device, and storage medium

A method of audio generation based on a large language model is disclosed, which involves the fields of artificial intelligence such as large language models, natural language processing, deep learning, and audio generation. The method of audio generation based on a large language model comprises: acquiring a text to be processed; parsing the text to be processed using the large language model to obtain role information and emotional information corresponding to the text to be processed; obtaining a target reference text and a target reference audio according to the role information and the emotional information; and generating a target audio corresponding to the text to be processed according to the text to be processed, the target reference text, and the target reference audio.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Parallel tacotron non-autoregressive and controllable TTS

A method for training a non-autoregressive TTS model includes receiving training data that includes a reference audio signal and a corresponding input text sequence. The method also includes encoding the reference audio signal into a variational embedding that disentangles the style / prosody information from the reference audio signal and encoding the input text sequence into an encoded text sequence. The method also includes predicting a phoneme duration for each phoneme in the input text sequence and determining a phoneme duration loss based on the predicted phoneme durations and a reference phoneme duration. The method also includes generating one or more predicted mel-frequency spectrogram sequences for the input text sequence and determining a final spectrogram loss based on the predicted mel-frequency spectrogram sequences and a reference mel-frequency spectrogram sequence. The method also includes training the TTS model based on the final spectrogram loss and the corresponding phoneme duration loss.
Owner:GOOGLE LLC

Method for realizing personalized singing synthesis model training through multi-mode voice driving

The invention relates to the technical field of voice signal processing, and discloses a method for realizing personalized singing synthesis model training through multi-modal voice driving, and the method comprises the following steps: obtaining multi-modal input data, including text, reference audio, speaker and emotional features; processing the data to obtain coding features; the inter-modal redundant information is evaluated and suppressed through redundant perception coding; compressing coding features by using an information bottleneck model, and retaining effective information; fusing the compressed features to generate personalized singing features; the input decoding module generates Mel spectrum features; and converting the Mel spectrum into an audio waveform through a vocoder, and outputting the audio waveform. And converting the Mel spectrum into an audio waveform through a vocoder, and outputting the audio waveform. According to the method, multi-modal data can be effectively fused, the individuation and emotion expression ability of the singing sound is improved, redundant information interference is reduced, the quality and efficiency of singing sound synthesis are improved, and finally, high-fidelity individualized singing sound is output.
Owner:SHENZHEN ZHONGLU CULTURE COMMUNICATION CO LTD

Systems and methods to automate trust delivery

The present disclosure relates to systems, methods, and products for using machine-learning networks to generate trustworthy audio and face mesh. A system, serving as a digital avatar, generates a trust audio and trust face mesh corresponding to an input text. A method includes generating a set of trust embedding vectors based on a reference audio; generate a text embedding vector based on the input text; generate a conditioned vector based on the set of trust embedding vectors and the text embedding vector; synthesize an audio representation based on the conditioned vector; generate the trust audio based on the synthesized audio representation; obtain a speech feature representation based on the trust audio; obtain an abstract feature vector based on the speech feature representation; and generate positions of vertices based on the abstract feature vector, the positions of vertices being used for generating the trust face mesh.
Owner:ACCENTURE GLOBAL SOLUTIONS LTD

Method and device for generating digital human video

The invention relates to the technical field of digital people, and provides a method and equipment for generating a digital people video, which can be applied to a digital marketing scene. According to the method, the generation of the digital human video is divided into an audio generation process and a video generation process, when the audio is generated, a target text, an emotion parameter and a reference audio are taken as input, and the target audio of the target text is broadcasted by the sound of a reference person in the reference video and a set emotion, so that sound cloning is realized, the degree of violation of the generated digital human video is reduced, and the user experience is improved. And the text features are extracted from two dimensions of characters and pronunciation, so that the pronunciation accuracy during sound cloning is ensured. When the video is generated, the cloned target audio and the cloned reference video are used as input, and the fidelity of the digital human video is improved through expression and mouth shape matching, so that a reference person can quickly and efficiently generate vivid digital human videos with different broadcast contents only by shooting one reference video without shooting one by one, thereby reducing the manufacturing cost of the digital human video, and improving the user experience. And the video production efficiency is improved.
Owner:JUHAOKAN TECH CO LTD

Generating gesture reenactment video from video motion graphs using machine learning

Embodiments are disclosed for generating a gesture reenactment video sequence corresponding to a target audio sequence using a trained network based on a video motion graph generated from a reference speech video. In particular, in one or more embodiments, the disclosed systems and methods comprise receiving a first input including a reference speech video and generating a video motion graph representing the reference speech video, where each node is associated with a frame of the reference video sequence and reference audio features of the reference audio sequence. The disclosed systems and methods further comprise receiving a second input including a target audio sequence, generating target audio features, identifying a node path through the video motion graph based on the target audio features and the reference audio features, and generating an output media sequence based on the identified node path through the video motion graph paired with the target audio sequence.
Owner:ADOBE INC

Audio processing method and device

The invention provides an audio processing method and device, and the method comprises the steps: collecting to-be-detected audio data related to a target audio based on to-be-detected audio equipment, and determining reference audio data related to the target audio; determining a to-be-tested audio index corresponding to the to-be-tested audio data in an audio consistency dimension, determining a reference audio index corresponding to the reference audio data, and generating an index difference score based on the to-be-tested audio index and the reference audio index; and under the condition that the to-be-tested audio data is determined to be abnormal based on the index difference score, generating difference processing prompt information corresponding to the to-be-tested audio data based on a difference mode rule and the index difference score. The audio processing method can be widely applied to the field of digital music in the digital creative industry.
Owner:GUANGZHOU SEASUN ENTERTAINMENT NETWORK TECHCO

Human voice quality optimization method and device, equipment and medium

The invention provides a voice quality optimization method and device, equipment and a medium, and relates to the technical field of audio data processing. In the application, firstly, Fourier transformation is performed on original audio data to form original audio frequency spectrum data; secondly, determining a target harmonic frequency from the original audio frequency spectrum data, and extracting harmonic audio data from the original audio data based on the target harmonic frequency; then, performing joint semantic coding on the reference audio data and the original audio data to form joint audio coding features; thirdly, performing semantic coding on the harmonic audio data to form harmonic audio coding features; further, optimizing the joint audio coding feature based on the harmonic audio coding feature to form an optimized audio coding feature; and finally, performing semantic decoding on the optimized audio coding features to form optimized audio data. On the basis of the content, the problem that in the prior art, the voice quality optimization effect is poor can be solved.
Owner:CHENGDU XIAOCHANG TECH CO LTD

Editing method and device, equipment and storage medium

According to the embodiment of the invention, a method and device for editing, equipment and a storage medium are provided. The method comprises the following steps: in response to an operation of adding a target timbre on a timbre configuration interface, recording a reference audio for reading a reference text with the target timbre; in response to an operation of triggering audio generation on the editing interface, obtaining an input text for audio generation; in response to a selection operation on the target tone, generating a target audio based on the reference audio and the input text, the target audio including a voice for reading the input text with the target tone; and generating an audio editing result or a video editing result based on the target audio. In this manner, at least a portion of the target audio may be generated using timbres in the speech data. This may facilitate a user to use a desired tone in media content generation, thereby advantageously improving the efficiency of editing.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion

The invention provides an ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion, and the method comprises the steps: obtaining a reference audio of a target user, and carrying out the decoupling extraction of a volume feature vector, a pitch feature vector, and a speaking style feature vector from the reference audio through a multi-dimensional pronunciation feature extraction network; a text to be verified is converted into a phoneme sequence, the phoneme sequence is input into a pronunciation feature prediction network for predicting a frame-level dynamic pronunciation feature sequence based on the phoneme sequence, and three pronunciation feature vectors are injected into the network through an adaptive instance normalization mechanism to dynamically modulate a prediction process. Generating a frame-level dynamic pronunciation feature sequence containing time sequence rhythm change; inputting a VAE-GAN synthesis trunk, re-injecting the three pronunciation feature vectors through an adaptive instance normalization mechanism, and generating a test voice sample consistent with the voiceprint characteristics of the target user; and initiating an identity verification query for the ASV system, and calculating safety indexes of different user groups according to a verification result so as to evaluate the risk of the ASV system.
Owner:FUJIAN NORMAL UNIV

A method, apparatus, device, medium and program product for generating an audio file

Embodiments of the present disclosure provide a method, device, medium and program product for generating an audio file. The method comprises: displaying an editing control on a playing page; in response to an interaction operation on the editing control, displaying an audio editing page, the audio editing page comprising audio information of a preset audio file and a generation control; in response to a selection operation on the audio information, determining a reference audio segment in the preset audio file; and in response to an interaction operation on the generation control, generating a target audio segment according to the reference audio segment, and determining a target audio file according to the target audio segment. The technical solution of the embodiments of the present disclosure generates a target audio segment of similar style according to a reference audio segment selected by a user, which can meet the user's expectation of music generation, reduce the difficulty of music creation, and improve the user experience.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Audio generation method and device

The invention discloses an audio generation method and device, and the method comprises the steps: obtaining an audio generation model, converting a reference audio of a target object into a reference sound feature vector, and converting a to-be-processed text into a to-be-processed text feature vector, and the reference audio comprises reference sound information needed when the target audio is synthesized; through the audio generation model, at least based on a to-be-processed text feature vector and a reference sound feature vector, performing model reasoning to obtain audio lexical elements matched with a to-be-processed text; and through the audio generation model, decoding the audio lexical elements matched with the to-be-processed text to obtain a target audio simulating the timbre of the target object. The natural degree of the synthesized audio can be improved.
Owner:SHANGHAI XIYU JIZHI TECH CO LTD

Speech synthesis methods, devices, computer equipment and storage media

This application provides a speech synthesis method, apparatus, computer device, and storage medium. The method relates to speech synthesis technology and is applied in the financial field. It includes: acquiring initial text to be synthesized and reference audio; inputting the initial text into a preset phoneme encoder, outputting multiple phonemes corresponding to the initial text; inputting the reference audio into a preset prosodic encoder, outputting multiple prosodices corresponding to the reference audio; aligning each phoneme and multiple prosodices based on a self-attention mechanism, acquiring embedding information corresponding to each phoneme; the embedding information includes at least one prosodic and a weight corresponding to each prosodic; generating a Mel spectrum based on each phoneme and its corresponding embedding information; generating synthesized audio based on the Mel spectrum, thus completing the prosodic alignment of the initial text and the reference audio. After synthesis, the prosodices will appear in accurate positions in the audio, improving the naturalness, credibility, persuasiveness, and appeal of the speech, and adapting to different application scenarios and user preferences.
Owner:PING AN TECH (SHENZHEN) CO LTD

Story audio timbre processing method and related device

The invention discloses a story audio timbre processing method and a related device, and relates to the technical field of audio processing, and the method comprises the steps: extracting a story voice from a to-be-processed story audio before carrying out the timbre processing of the to-be-processed story audio through employing a reference audio, and then carrying out the timbre conversion processing of the story voice based on the reference audio, thereby achieving the timbre processing of the to-be-processed story audio. And finally, determining a final story audio based on the target story voice. In the whole processing process, voice recognition is not needed, the influence of voice recognition accuracy on the tone processing effect is avoided, the to-be-processed audio is subjected to story voice extraction and story voice processing, the influence of story background voice on the tone processing effect is avoided, and therefore the voice processing efficiency is improved. According to the scheme, the tone processing effect of the story audio can be improved.
Owner:HEFEI IFLYTEK TOYCLOUD TECH

Sound effect evaluation method, device and equipment and storage medium

The application discloses an audio effect evaluation method and device, equipment and a storage medium. The application inputs a preset reference sound source into an audio device to be tested for playing, wherein the preset reference sound source is a sound source determined by statistical analysis of sound sources, in combination with a parameter corrected according to a preset evaluation result and a peak factor corresponding to a music playing platform; collects audio to be tested output by the audio device to be tested when playing the preset reference sound source; determines a relative frequency response and an estimated loudness value according to a spectrum corresponding to the audio to be tested and a spectrum corresponding to the preset reference sound source; and generates an audio effect evaluation result of the audio device to be tested according to the relative frequency response and the estimated loudness value. The application evaluates the audio effect of the audio device to be tested by using the reference sound source constructed in advance, solves the problem of traditional sweep frequency signal tuning deviation, greatly reduces the number of repeated trial and error, and significantly improves the efficiency, accuracy and reliability of audio product tuning work.
Owner:GEER TECH CO LTD

Artificial intelligence systems and methods for detecting musical infringement in symbolic music

The present disclosure relates to a system and method for detecting musical infringement in symbolic music. The system receives a first musical composition and, when provided as audio, converts it into a symbolic format using a transcription neural network trained to extract a main melody. It then generates k-mer sequences comprising consecutive notes, indexes these sequences in a data structure configured for dynamic conditioning, and compares them to reference musical compositions. Upon estimating a similarity measure that exceeds a threshold, the system performs a refined local sequence alignment adapted for music, accounting for key shifts, rests, and melodic or rhythmic variations. Based on this refined alignment, the system determines whether the first musical composition includes a musical fragment that infringes upon or regurgitates a portion of at least one reference composition.
Owner:SOUND PATROL INC

Music detection method, apparatus, device, and medium

The application relates to the technical field of data processing, and particularly provides a music detection method and device, equipment and a medium. The music detection method comprises the following steps: obtaining target music description information of target music to be detected; obtaining reference music description information of reference music corresponding to the target music; determining a score similarity between the reference music and the target music according to first score information in the reference music description information and second score information in the target music description information; and obtaining an originality detection result of the target music according to the score similarity; and the originality detection result indicates whether the target music has originality. In this way, the cost of music originality detection can be reduced, and the accuracy of music originality detection can be improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Audio data processing method and device, electronic equipment and storage medium

Embodiments of the present application provide an audio data processing method and device, electronic equipment and storage medium, the method comprising: acquiring first audio data collected by a microphone, the microphone having a delay time; the first audio data including echo data of background audio produced after the background audio played inside a terminal is diffused through a loudspeaker; acquiring second audio data; the second audio data being original data of the background audio played inside the terminal device; acquiring the size of audio data collected by the microphone within the delay time; aligning the second audio data with the first audio data according to the size of the audio data; and performing echo cancellation processing on the first audio data with the aligned second audio data as reference audio data. The microphone data and the reference audio data can be aligned to the same time point for echo cancellation processing, improving the effect of echo cancellation.
Owner:THUNDERSOFT (NANJING) CO LTD

Portrait dialogue video generation method, multi-person dialogue video generation method, product, equipment and storage medium

The invention discloses a portrait dialogue video generation method, a multi-person dialogue video generation method, a product, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the portrait dialogue video generation method comprises the steps: extracting a face parameter of a first dialogue object from a reference face image; determining a primary audio of the first dialogue object and a secondary audio of the second dialogue object based on the reference audio; fusing the primary audio of the first dialogue object and the secondary audio of the second dialogue object to obtain a fused audio feature; constructing a three-dimensional portrait geometric sequence of the first dialogue object according to the fused audio features and the face parameters of the first dialogue object; and generating a portrait dialogue video of the first dialogue object based on the reference audio, the reference face image, the fused audio features and the three-dimensional portrait geometric sequence. The method aims at supporting dynamic switching between the speaking state and the listening state of the dialogue object in the portrait dialogue video, the portrait expression modeling precision of the dialogue object is improved, and the interaction fluency of multi-person dialogue is improved.
Owner:GUANGDONG-HONG KONG-MACAO GREATER BAY AREA DIGITAL ECONOMY RESEARCH INSTITUTE (INTERNATIONAL ADVANCED TECHNOLOGY APPLICATION PROMOTION CENTER (SHENZHEN) +1

Sleep-aid audio generation method, device, equipment, and storage medium

The present application provides a method, device, equipment, and storage medium for generating sleep-aiding audio. The method includes: collecting the brain wave sequence and respiratory signal sequence of the sleep-aiding subject; after time-aligning the brain wave sequence and the respiratory signal sequence, sampling at a preset sampling frequency, and connecting the brain wave and respiratory signal at the sampling time point as a state signal each time; determining the sleep state feature vector and the current sleep state based on the time series signal composed of the state signal; determining the reference audio that matches the current sleep state; and generating sleep-aiding audio that assists the sleep-aiding subject to enter the next sleep state based on the reference audio and the sleep state feature vector through a diffusion model. The method of the present application generates sleep-aiding audio based on the brain wave sequence and respiratory signal sequence of the sleep-aiding subject, and can adaptively generate sleep-aiding audio based on the sleep state of the sleep-aiding subject, thereby achieving a better sleep-aiding effect.
Owner:ZHIXIANG FUTURE (HEFEI) TECHNOLOGY CO LTD

Multi-channel speech compression system and method

A method, computer program product, and computing system for encoding audio encounter information of a reference audio acquisition device of a plurality of audio acquisition devices of an audio recording system, thus defining encoded reference audio encounter information. Location information may be estimated, via a machine vision system, for an acoustic source within an acoustic environment. One or more acoustic relative transfer functions may be selected from a plurality of acoustic relative transfer functions for the plurality of audio acquisition devices of the audio recording system based upon, at least in part, the location information. The encoded reference audio encounter information and a representation of the selected one or more acoustic relative transfer function may be transmitted.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Vlog generation methods and related devices

This application discloses a Vlog generation method and related apparatus. The method includes: acquiring video footage, the duration of the target Vlog, reference transition points, video emotional information, and a reference audio signal; determining the edited video footage, the transition points of the target Vlog, and reference background music information based on the acquired information; performing track splitting and transcription processing on the reference audio signal to obtain the MIDI score and phrase / section segmentation points of the reference audio signal; processing the reference audio signal according to the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog; and obtaining the target Vlog based on the edited video footage and the background music of the target Vlog. Using the method of this application, Vlogs that meet the personalized needs of users can be obtained.
Owner:HUAWEI TECH CO LTD