Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

93 results about "Audio synthesis" patented technology

Audio synthesis method, audio synthesis model training method, apparatus, electronic device, computer-readable storage medium, and computer program product

An audio synthesis method, an audio synthesis model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which relate to artificial intelligence technology. The method includes: invoking an audio synthesis model based on language information and preset style information of a target text to perform following processing, the audio synthesis model including a prior encoder and a waveform decoder: generating audio features corresponding to the target text based on the language information and the preset style information by using the prior encoder; performing normalizing flow processing on the audio features by using the prior encoder, to obtain a hidden variable of the target text; and performing waveform decoding on the hidden variable of the target text by using the waveform decoder, to obtain a synthetic waveform conforming to an audio style described in the preset style information and corresponding to the target text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio synthesis method and device, medium and equipment

The embodiment of the invention provides an audio synthesis method and device, a medium and equipment, and relates to the technical field of speech synthesis. The method comprises the steps of obtaining a target text to be subjected to audio synthesis; splitting the target text into a plurality of text units with complete semantics to obtain a task sequence; for the ith text unit in the task sequence, scheduling influence factors are obtained before the audio synthesis task is executed, the scheduling influence factors comprise application layer semantic information and / or equipment real-time performance indexes, and i is a positive integer; determining a target synthetic link from a cloud synthetic link and an end-side synthetic link based on the scheduling influence factors; and completing audio synthesis of the ith text unit through the target synthesis link. According to the scheme provided by the embodiment of the invention, in the audio synthesis process of the target text, adaptive switching between the cloud audio synthesis link and the end-side audio synthesis link can be realized in a low-delay manner.
Owner:IFLYTEK CO LTD

Interactive digital human generation method and system based on picture and audio synthesis

The invention discloses an interactive digital human generation method based on picture and audio synthesis, and belongs to the crossing field of artificial intelligence and computer graphics. According to the method, a full-link process of'feature extraction-model construction-emotion driving-real-time interaction-video output 'can be automatically completed only by uploading a character picture and a section of audio by a user: key point and semantic feature extraction is performed on the picture to obtain face / posture information; performing voice recognition, semantic analysis and emotion recognition on the audio to obtain a semantic tag and an emotion parameter; generating a personalized three-dimensional digital human based on the information, and driving the personalized three-dimensional digital human to generate facial expressions and limb actions which are synchronous with emotions; the user intention is analyzed in real time through natural language understanding and computer vision, and multi-modal interaction is achieved. According to the method, digital human creation can be completed without professional modeling and motion capture equipment, the generation cost is reduced, and the method can be widely applied to virtual anchors, online education, intelligent customer service and movie and television entertainment scenes.
Owner:NEW ONE (BEIJING) TECH CO LTD

Cross-culture melody feature extraction and musical instrument matching system based on deep learning

The invention discloses a cross-culture melody feature extraction and musical instrument matching system based on deep learning, and belongs to the technical field of music information processing and artificial intelligence. The system comprises an audio acquisition module, an audio preprocessing module, a melody feature extraction module, a cultural context understanding module, an emotion semantic analysis module, a musical instrument timbre database, a musical instrument matching recommendation module, a distributor scheme generation module, a user interaction module and an audio synthesis module. Multi-dimensional melody features are extracted through a multi-scale convolutional neural network and a bidirectional long-short term memory network, cultural semantic understanding is realized through reasoning on a music knowledge graph by using a graph neural network, and musical instrument matching is performed by using a multi-objective optimization algorithm in combination with three dimensions of timbre integrating degree, cultural consistency and emotional expressive power. According to the system, automation of the whole process from humming melody to musical instrument configuration is achieved, the efficiency of the instrument is improved by 400%, the accuracy rate reaches 92%, and an intelligent tool is provided for cross-culture music creation and national music modernization reorganization.
Owner:SHENZHEN UNIV

Hearing aid health analysis method based on physiological signal monitoring function

The application discloses a hearing aid health analysis method based on physiological signal monitoring function, comprising the following steps: collecting physiological signal data of a user in real time through a multi-modal biological sensor array integrated in a hearing aid body; the hearing aid body comprises a digital signal processor and a wireless communication module; the physiological signal data at least comprises an electroencephalogram signal, a galvanic skin response signal and a temporal artery pulse signal; the physiological signal data is subjected to motion artifact elimination processing through an adaptive filtering algorithm; a preprocessed physiological signal is subjected to feature extraction through an embedded health analysis engine, and a composite health feature vector comprising a sympathetic nerve activity index, a cognitive load metric and a cardiovascular function index is generated; a health risk assessment report is dynamically generated based on the association between the composite health feature vector and a preset hearing scene; and a graded health alarm signal is output through a bone conduction vibration unit and an audio synthesis module of the hearing aid body.
Owner:SHENZHEN XINZHENGYU TECH

A method, system, and computer device for generating text-based videos.

This invention proposes a method, system, and computer device for generating text-based videos, relating to the technical field of video processing. The method includes: acquiring user input parameters; generating anchor point prompts and anchor point character diagrams based on the user input parameters; evaluating the anchor point character diagrams until they pass evaluation; extracting information from the anchor point prompts to obtain a storyboard; evaluating the storyboards until they pass evaluation; generating audio and storyboard video based on the storyboards; evaluating the storyboard video until it passes evaluation; post-processing the storyboard video to obtain a post-processed video; and combining the post-processed video with the audio to generate a target video. This invention effectively provides a closed-loop evaluation node for text-based video generation, improving the quality of generated text-based videos.
Owner:GUANGDONG HENGQIN SHUSHUSHUO STORY INFORMATION TECH CO LTD

Vocoder training method, audio synthesis method, medium, apparatus, and computing device

Embodiments of the present disclosure provide a vocoder training method. The vocoder training method comprises: obtaining a first fundamental frequency sequence of audio in an audio corpus; performing fundamental frequency perturbation processing on the first fundamental frequency sequence to obtain a second fundamental frequency sequence; performing mapping processing on the second fundamental frequency sequence to obtain a target tensor; and training the target tensor, an acoustic feature sequence corresponding to the audio, and the audio to obtain a vocoder used for audio synthesis. The method of the present disclosure improves the robustness of the vocoder to fundamental frequency prediction errors in actual application by introducing fundamental frequency perturbation, thereby significantly improving the accuracy and quality of audio synthesis and providing a better experience for users. In addition, embodiments of the present disclosure provide an audio synthesis method, a medium, a device and a computing device.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

A smart management system and method based on video privacy information identification

The application discloses a kind of wisdom management system and method based on video privacy information identification, it is related to information processing identification technical field, including: time stamp matching test is carried out to audio of different length, matching test model is established, audio in court video is extracted, audio is divided and optimization division result is obtained, each segment audio after division is converted into text separately, private content in text is identified, the location of private content in audio is found by time stamp matching and is carried out mute, complete audio is synthesized after mute processing, and the division process of subsequent audio to be muted is planned, after frame extraction is carried out to court video image, private information in each frame image is identified and is carried out coding, audio-video is synthesized after coding video and mute processing audio, the probability that mute is not complete even part of private information in audio is not muted due to that time stamp matching deviation is larger is reduced, the reliability of court audio private information processing is guaranteed.
Owner:JIANGSU XINSHIYUN TECH CO LTD

Sound scene enhancement device based on environmental perception

The utility model discloses a sound scene enhancing device based on environmental perception, and relates to the technical field of acoustics. The system comprises a microphone array, a signal conditioning circuit, an analog-to-digital converter, a DSP processor, a sound scene database memory, a master control MCU, an FPGA beam controller, an audio synthesis chip, a power amplifier, a loudspeaker array and a noise classification coprocessor, the microphone array is electrically connected with the signal conditioning circuit, and the signal conditioning circuit is electrically connected with the analog-to-digital converter. The analog-to-digital converter is electrically connected with the DSP processor, the analog-to-digital converter is electrically connected with the master control MCU, the DSP processor is electrically connected with the sound scene database memory, the DSP processor is electrically connected with the noise classification coprocessor, the noise classification coprocessor is electrically connected with the master control MCU, the master control MCU is electrically connected with the sound scene database memory, and the master control MCU is electrically connected with the FPGA beam controller. The sound scene can be automatically adjusted according to environment changes.
Owner:JIANGSU ACOUSTIC IND TECH INNOVATION CENT

Audio synthesis method, apparatus, device, computer-readable medium, and program product

Embodiments of the present disclosure disclose an audio synthesis method, device, equipment, computer readable medium and program product. A specific implementation of the method comprises: obtaining target audio corresponding to a first language and accent correction information corresponding to an audio synthesis scene, wherein the audio synthesis scene is a scene for synthesizing audio corresponding to a second language based on the target audio; extracting audio feature information corresponding to the target audio; and generating the synthesized audio corresponding to the second language by using a pre-trained speech synthesis large model according to the audio feature information and the accent correction information. The implementation is related to artificial intelligence, and by using the accent correction information, in the process of synthesizing the audio corresponding to the second language based on the target audio corresponding to the first language, the disturbance of the accent noise can be avoided, so that the synthesized audio without the accent noise can be accurately and efficiently generated.
Owner:BEIJING YIJING INFORMATION TECHNOLOGY CO LTD

Model training methods, devices, electronic equipment, computer-readable storage media, and computer program products

This application provides a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product. The method includes: extracting audio feature sequences from raw audio data; determining a first phoneme feature sequence based on the raw audio data and text data; aligning the first phoneme feature sequence based on the audio feature sequence to obtain a second phoneme feature sequence; configuring an omission flag in the second phoneme feature sequence based on text data to obtain a third phoneme feature sequence; structurally associating the audio feature sequence and the third phoneme feature sequence to obtain structured data; storing the structured data in a target storage file; and upon receiving a training instruction, reading the structured data from the target storage file and updating the model parameters of the audio synthesis model based on the data in the structured data that does not have an omission flag configured, thereby obtaining the trained audio synthesis model. This application can improve the training efficiency and accuracy of audio synthesis models.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio processing method, system and equipment based on multi-order frequency adjustment and medium

The invention relates to the technical field of audio processing based on multi-order frequency adjustment, and discloses an audio processing method, system and device based on multi-order frequency adjustment and a medium, and the method comprises the steps: obtaining an initial audio; frequency division is carried out on the initial audio frequency to obtain initial frequency band data, and the initial frequency band data comprise low-frequency data, intermediate-frequency data and high-frequency data; performing denoising processing on the initial frequency band data to obtain target frequency band data; performing audio synthesis on the target frequency band data to obtain a to-be-adjusted audio; and based on target cavity characteristic data of the sound production equipment, performing frequency adjustment on the to-be-adjusted audio by adopting at least two EQ adjustment modules for multi-order frequency adjustment to obtain a target audio suitable for the sound production equipment. Therefore, the man-made effect of the frequency band edge is effectively avoided, the frequency curve of the finally output audio is more continuous and smooth, and the audio better meets the hearing requirement of a user.
Owner:深圳市鸿宇光电有限公司

Audio synthesis method, audio synthesis system, computer device and storage medium

The application relates to an audio synthesis method, an audio synthesis system, a computer device and a computer readable storage medium. The method comprises the following steps: acquiring a current buffer data amount of an audio buffer; the current buffer data amount is used for indicating a playing time length of accompaniment audio data in the audio buffer; based on the difference between the current buffer data amount and a reference buffer data amount, the data amount of the accompaniment audio data buffered in the audio buffer is adjusted to obtain an adjusted audio buffer; the accompaniment audio data is read out from the adjusted audio buffer, and the read accompaniment audio data is played; and when a user sings along with the played accompaniment audio data, the dry audio data of the user is recorded; the read accompaniment audio data and the recorded dry audio data are subjected to audio synthesis processing to obtain synthesized audio data. By adopting the method, the read accompaniment audio data and the recorded dry audio data can be ensured to correspond to each other, and the audio quality of the synthesized audio is improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

An audio cloning method based on cloud computing technology and related equipment

PendingCN122313989AAudio synthesisData center
This application provides an audio cloning method and related equipment based on cloud computing technology to ensure the security of cloned voice and prevent user privacy leaks. The method is applied to a cloud platform, which manages the infrastructure running cloud computing services. The infrastructure includes at least one data center, and each data center includes multiple servers. The method includes: acquiring original audio from a first tenant; encrypting the original audio to obtain encrypted audio; and storing the encrypted audio in the infrastructure. Acquiring a first text from the first tenant. Decrypting the encrypted audio to obtain the original audio. Synthesizing the original audio and the first text to obtain cloned audio, where the sound features of the cloned audio are identical to those of the original audio, and the content of the cloned audio is the content of the first text.
Owner:HUAWEI TECH CO LTD

An adaptive sleep environment system

PendingCN122297868AAudio synthesisNoise
This invention provides an adaptive sleep environment system, belonging to the field of smart home technology. It includes: a sound acquisition module for acquiring the wake-up voice of a monitored subject and the ambient noise within the sleep space of the monitored subject; a vital sign monitoring module for monitoring the sleep vital sign parameters of the monitored subject; a directional audio playback module for playing sleep-aiding audio; and a system host that, in response to the wake-up voice, generates reverse audio based on the ambient noise and synthesizes the reverse audio into the sleep-aiding audio. The directional audio playback module is also used to directionally play the synthesized reverse audio towards the direction of the monitored subject and adjust the volume of the sleep-aiding audio based on the sleep state determined by the sleep vital sign parameters. Beneficial effects: By directionally playing sleep-aiding audio, it avoids interference with other sleepers sharing the same space, thus improving their sleep environment.
Owner:SHANGHAI SIXTH PEOPLES HOSPITAL

Dubbing method and device, electronic equipment and storage medium

The embodiment of the invention provides a dubbing method and device, electronic equipment and a storage medium, and relates to the technical field of data processing. The specific implementation scheme is as follows: obtaining a target line text and an original line audio; performing acoustic feature extraction on the original line audio to obtain at least one acoustic feature; based on the at least one acoustic feature, generating natural language description information about the acoustic feature of the original line audio; and based on the natural language description information, performing audio synthesis processing on the target line text to obtain a dubbing audio which has acoustic characteristics of the original line audio and belongs to the target language. Visibly, according to the scheme of the application, the dubbing efficiency for the target line text can be improved.
Owner:CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD

Multi-target 3D audio rapid synthesis method based on matrix sparse decomposition

The invention discloses a multi-target 3D audio rapid synthesis method based on matrix sparse decomposition. The multi-target 3D audio rapid synthesis method comprises the following steps: S1, performing PCA decomposition on an HRIR database for 3D audio synthesis to obtain a principal component score matrix and a principal component vector matrix; s2, performing sparse decomposition on the PCA principal component vector matrix obtained in the step S1 to obtain a kernel matrix and a sparse basis matrix; and S3, during multi-target 3D audio synthesis, selecting a principal component score corresponding to each target, sequentially multiplying sound source data to be processed by the principal component score matrix, the sparse basis matrix and the kernel matrix, and taking a trace to obtain processed 3D audio data. The method has the beneficial effects that the HRIR database is compressed under the condition that the reconstruction error is very small, and the calculation efficiency of multi-target 3D audio synthesis is greatly improved.
Owner:SHANGHAI AVIATION ELECTRIC

Self-adaptive audio coding automatic driving state auditory prompting method and self-adaptive audio coding automatic driving state auditory prompting system

The invention provides a self-adaptive audio coding automatic driving state auditory prompting method and system, and is applied to the technical field of intelligent driving. Internal operation state data and environment target data of an automatic driving system are collected in the vehicle operation process; based on the current decision-making result, decision-making related targets which directly influence the decision-making are screened out; global audio characteristic parameters are generated by performing time sequence modeling on the internal operation state and are used for expressing the overall operation situation of the automatic driving system; meanwhile, importance evaluation is conducted on decision-related targets, corresponding local audio attribute parameters are generated and used for highlighting auditory significance of the key targets, the global audio feature parameters serve as auditory backgrounds to be continuously output, the local audio attribute parameters serve as auditory foregrounds to be dynamically overlapped, and auditory prompts are formed through audio synthesis and playing. In this way, continuous and explainable auditory expression of the running state and decision causal information of the automatic driving system is achieved.
Owner:JILIN UNIVERSITY

Audio dialogue method and device, equipment and storage medium

The embodiment of the invention provides an audio dialogue method and device, equipment, a storage medium and a program product. The method comprises the following steps: encoding a first audio stream collected from an environment into an audio feature sequence by using a streaming audio encoder; generating a sequence of text units as a response to the first audio stream based on the system cue word and the audio feature sequence using a trained machine learning model; and in response to the sequence of text units satisfying the audio synthesis condition, using the streaming audio synthesizer to generate a second audio stream from the sequence of text units for playing. In this way, the audio conversation can be realized in a full duplex and streaming manner, and discrete audio coding does not need to be introduced in the audio generation process. This may improve the performance and efficiency of the audio dialog.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +2

Context aware sound field mapping system and method for enhancing display interaction

The invention provides a context-aware sound field mapping system and method for enhancing display interaction, and the system comprises a context analysis module which is used for analyzing the voice content to be broadcasted by AI, and extracting key semantics, emotion, intention and hidden scenes; the acoustic environment sensing module is used for analyzing acoustic characteristics of a real environment where a user is located through a microphone array; the background sound generation and selection module is used for generating or retrieving a most matched background sound material from a local or cloud sound database based on the analyzed context; and the audio synthesis and spatialization rendering module is used for intelligently mixing the dry sound signals of the AI with background sound, and applying a spatial audio technology, so that the finally output sound has a sense of direction, a sense of distance and a sense of environmental fusion. The method is used for dynamically generating and rendering the spatialized background sound matched with the dialogue context for artificial intelligence voice interaction, so that immersion of digital entity existence in a physical environment is created.
Owner:BEIJING LANGZHIWAN INTELLIGENT TECHNOLOGY CO LTD

A voice signal reconstruction method, device, equipment and storage medium thereof

The embodiment of the application belongs to the technical field of speech processing, is applied to a speech signal reconstruction scene, and relates to a speech signal reconstruction method, device and equipment and a storage medium thereof. All subband features contained in a mel spectrogram are identified. Independent encoders are used to respectively perform downsampling processing on different subband features, so as to obtain a low-dimensional feature vector of each subband feature. The low-dimensional feature vector corresponding to each subband feature is quantized to obtain a discrete numerical result. The discrete numerical results corresponding to all subband features are input into a decoder for upsampling recovery processing, so as to obtain a reconstructed speech waveform corresponding to an original speech signal. The downsampling processing is first performed on each subband feature, and then the upsampling recovery is performed in combination with all downsampling quantization results, so that the audio information in each subband feature is fully utilized, the speech high-frequency part is reconstructed in a more detailed manner, and the audio synthesis quality of the intelligent speech customer service in the financial field is ensured.
Owner:PING AN TECH (SHENZHEN) CO LTD

Lightweight binaural audio synthesis method based on implicit neural network

The invention discloses a lightweight binaural audio synthesis method based on an implicit neural network, which efficiently generates high-fidelity binaural audio from monaural audio and sound source pose information through a two-stage framework. An initial binaural signal containing a main time clue is roughly generated by using a neural time domain distortion module, a frequency spectrum of the signal is finely corrected by using an implicit binaural correction module, and a frequency spectrum clue is introduced. The implicit correction module models a complex frequency spectrum correction process into a continuous mapping function which is represented by a small multilayer perceptron and is from space-time-frequency coordinates to complex correction values. According to the method, the parameter quantity and the calculation requirement of the model are reduced, so that the model can be efficiently deployed on virtual reality, augmented reality and other edge devices with limited calculation resources, the problem of efficiency and quality balancing in real-time high-fidelity synthesis is solved, the method is suitable for virtual reality (VR), augmented reality (AR) and other resource-limited scenes needing real-time 3D audio rendering, and the real-time 3D audio rendering efficiency is improved. Such as mobile games, online conferences and smart headphones.
Owner:EAST CHINA NORMAL UNIV

Model training method and device, electronic equipment, computer readable storage medium and computer program product

The invention provides a model training method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: extracting an audio feature sequence from original audio data, and determining a first phoneme feature sequence based on the original audio data and text data; aligning the first phoneme feature sequence based on the audio feature sequence to obtain a second phoneme feature sequence; configuring an ignoring identifier in the second phoneme feature sequence based on the text data to obtain a third phoneme feature sequence; performing structured association on the audio feature sequence and the third phoneme feature sequence to obtain structured data, and storing the structured data to a target storage file; and when a training instruction is received, reading the structured data from the target storage file, and updating the model parameters of the audio synthesis model based on the data not configured with the neglect identifier in the structured data to obtain a trained audio synthesis model. The training efficiency and accuracy of the audio synthesis model can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

System and method for monitoring and producing audio feedback based on respiratory movements

PCT designated stageWO2026060246A1Broadcast transmission systemsInertial sensorsAudio synthesismuscle spasm
A system and method for monitoring respiratory movements and generating real-time audio feedback reflects a user's breathing pattern. The system includes a respiratory sensor, a signal processor, an audio synthesis unit, a wireless transmitter, and an audio output device. Respiratory signals are captured, filtered, and analyzed to determine inhalation and exhalation phases. The audio synthesis unit produces corresponding sound signals that are dynamically modulated based on the user's breathing intensity and rhythm. The audio output provides a continuous, natural-sounding auditory representation of breathing, which can be used for caregiver reassurance, sleep monitoring, or relaxation purposes. Abnormal events such as breathing cessation or muscle spasms may trigger audio alerts or spoken messages. Processed physiological metrics, including respiratory rate, heart rate, and sleep state, being transmitted to a cloud platform for long-term analysis. The system may also interface with smart lighting or external alerting devices to support hearing-impaired caregivers.
Owner:EMFIT CORP +1

An AI cloud-based cross-language call and external network video and audio translation method and system

The application discloses a kind of based on AI cloud cross-language conversation and external network video and audio translation method and system, comprising the following steps: in the exclusive foreign language access communication identification of intelligent communication terminal preset;When the external input source request of exclusive foreign language access communication identification is monitored to intelligent communication terminal, the transmission path of original audio signal is redirected to AI intelligent body cloud service module;AI intelligent body cloud service module receives original audio signal, determines source language type, and then executes text translation and mother language synthesis, generates target voice stream;Original audio signal is executed in cloud to delete instruction, and target voice stream is encapsulated into downstream data packet;Intelligent communication terminal receives downstream data packet, and outputs pure mother language voice through audio synthesizer.The AI intelligent body service architecture deployed in cloud and the system kernel of intelligent communication terminal are deeply coupled in the application, interception, redirection, translation and original sound are removed to audio stream.
Owner:JIANGSU YOUZHISHUN BREEDING TECHNOLOGY CO LTD

An audio synthesis method based on a vits model improvement and a storage medium

The application discloses an audio synthesis method based on a VITS model improvement and a storage medium, and belongs to the technical field of speech synthesis. The method comprises the following steps: obtaining text of to-be-synthesized audio data, and preprocessing the text; inputting the preprocessed text into a pre-trained adaptive speech synthesis model AdaVITS to perform audio synthesis; and obtaining generated audio data according to the output of the adaptive speech synthesis model AdaVITS. The adaptive speech synthesis model AdaVITS is based on a speech synthesis model VITS, a loss function of the speech synthesis model VITS is improved and increased to obtain a joint loss function of the adaptive speech synthesis model AdaVITS, and the joint loss function is optimally solved, so that the speech quality and the training efficiency are synergistically optimized.
Owner:JIANGSU ZHIHENG INFORMATION TECH SERVICES CO LTD

Audio synthesis method and apparatus, computer device, and storage medium

The application relates to an audio synthesis method and device, computer equipment and a storage medium. The method can be used in cloud technology, artificial intelligence, intelligent transportation, audio and video and the like, and comprises the following steps: performing audio track feature separation on source audio data and to-be-processed audio data to obtain source singing track features, to-be-processed singing track features and to-be-processed accompaniment track features; performing voiceprint feature extraction on the source singing track features to obtain source singing voiceprint features; performing attention processing on the source singing voiceprint features and the to-be-processed singing track features through an attention network, and in the process of attention processing, taking the attention processing result obtained by each attention layer in the attention network and the source singing voiceprint features as input data of the next attention layer to perform attention processing, so as to obtain target singing track features; and fusing the target singing track features and the to-be-processed accompaniment track features to obtain synthesized audio data. The method can improve the synthesis effect of audio.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Device testing method and apparatus, electronic device, and storage medium

The present disclosure discloses a device testing method and device, an electronic device, and a storage medium, relates to the technical field of artificial intelligence, and in particular to the technical field of voice testing and voice interaction. The specific implementation scheme is: in response to a test request, based on the device type information of a voice interaction device carried by the test request, performing wake-up testing on the voice interaction device to obtain a wake-up testing result; in the case where the wake-up testing result indicates that the wake-up is successful, determining target audio synthesis parameters; and based on the target audio synthesis parameters, using a text sample set carried by the test request to test the voice interaction device to obtain a test result of the voice interaction device.
Owner:APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD

Method for generating audio based on large model, electronic device, and storage medium

The present application provides a method for generating audio based on large model, an electronic device, and a storage medium, which relates to a technical field of artificial intelligence such as an audio synthesis and a large model. A specific implementation includes: obtaining a character that is generated in real time during a process of generating a text using a large model; obtaining an audio feature of each audio unit of the character sequentially by using a pre-trained audio generation model based on the character; the audio feature of the audio unit is a discretized audio feature, and the character includes audio features of a plurality of different audio units; synthesizing a corresponding audio by using a pre-trained vocoder based on the audio feature of each audio unit.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Audio synthesis method, computer device, storage medium, and program product

The application provides an audio synthesis method, a computer device, a storage medium, and a program product. The method comprises the following steps: performing feature fusion processing on phoneme feature information of a target text and label information corresponding to the target text to obtain target phoneme feature information; performing splicing processing on the target phoneme feature information and target pitch feature information to obtain spliced feature information; generating a predicted mel spectrum according to the spliced feature information; and performing conversion processing on the predicted mel spectrum by using a vocoder to obtain audio data matched with a score corresponding to the target text. By using the application, the cost of audio synthesis can be reduced, and the timbre stability of the synthesized audio data can be improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD