Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

131 results about "Audio synthesis" patented technology

Audio synthesis method, audio synthesis model training method, apparatus, electronic device, computer-readable storage medium, and computer program product

An audio synthesis method, an audio synthesis model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which relate to artificial intelligence technology. The method includes: invoking an audio synthesis model based on language information and preset style information of a target text to perform following processing, the audio synthesis model including a prior encoder and a waveform decoder: generating audio features corresponding to the target text based on the language information and the preset style information by using the prior encoder; performing normalizing flow processing on the audio features by using the prior encoder, to obtain a hidden variable of the target text; and performing waveform decoding on the hidden variable of the target text by using the waveform decoder, to obtain a synthetic waveform conforming to an audio style described in the preset style information and corresponding to the target text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Voice generation method and device

The embodiment of the invention provides a voice generation method and device, computer equipment, a computer readable storage medium and a computer program product, and belongs to the field of audio processing. The voice generation method comprises the following steps: acquiring a content text and an emotion description text; determining an emotion weight array according to the emotion description text; determining a target emotion vector according to the emotion weight array and a plurality of basic emotion vectors; the target emotion vector and the content text serve as model input, first target voice is generated through a pre-trained audio synthesis model, and the first target voice comprises text content in the content text and emotion features in the emotion description text. According to the technical scheme provided by the embodiment of the invention, emotion control on the first target voice can be realized by utilizing the emotion description text, so that the voice generation stability and the emotion controllability and accuracy in the voice generation process are improved.
Owner:SHANGHAI HODE INFORMATION TECH CO LTD

System and Method for Dynamic Interactive Storytelling Using Language Models and Generative Video and Audio Synthesis

A system and method are provided for dynamically generating interactive multimedia storytelling experiences using integrated artificial intelligence models. The system comprises a generative language model for producing narrative content in response to user input, a generative video synthesis module for visualizing story segments, and a generative audio synthesis module for producing synchronized speech, effects, and music. In alternative embodiments, a single multimodal generative model may perform both video and audio synthesis. A user interaction module accepts free-form input to evolve the story in real time, and a content generation coordinator manages orchestration, timing, and latency optimization between components. The system supports modular architecture, lip synchronization with character visuals, predictive pre-generation to reduce delay, personalization based on user profiles, and deployment across various platforms including desktop, mobile, and extended reality environments. The invention enables open-ended, user-driven narrative generation with seamless and adaptive audiovisual synthesis.
Owner:TERENNA BRIAN

Hearing aid health analysis method based on physiological signal monitoring function

The invention discloses a hearing aid health analysis method based on a physiological signal monitoring function. The method comprises the following steps: collecting physiological signal data of a user in real time through a multi-mode biosensor array integrated on a hearing aid body; the hearing aid body comprises a digital signal processor and a wireless communication module; the physiological signal data at least comprises a brain wave signal, a skin electric response signal and a temporal artery pulse signal; performing motion artifact elimination processing on the physiological signal data through an adaptive filtering algorithm; performing feature extraction on the preprocessed physiological signal by using an embedded health analysis engine to generate a composite health feature vector containing a sympathetic nerve activity index, cognitive load measurement and a cardiovascular function index; and dynamically generating a health risk assessment report based on an association relationship between the composite health feature vector and a preset hearing scene, the bone conduction vibration unit and the audio synthesis module of the hearing aid body output graded health alarm signals.
Owner:SHENZHEN XINZHENGYU TECH

Audio synthesis method and device, medium and equipment

The embodiment of the invention provides an audio synthesis method and device, a medium and equipment, and relates to the technical field of speech synthesis. The method comprises the steps of obtaining a target text to be subjected to audio synthesis; splitting the target text into a plurality of text units with complete semantics to obtain a task sequence; for the ith text unit in the task sequence, scheduling influence factors are obtained before the audio synthesis task is executed, the scheduling influence factors comprise application layer semantic information and / or equipment real-time performance indexes, and i is a positive integer; determining a target synthetic link from a cloud synthetic link and an end-side synthetic link based on the scheduling influence factors; and completing audio synthesis of the ith text unit through the target synthesis link. According to the scheme provided by the embodiment of the invention, in the audio synthesis process of the target text, adaptive switching between the cloud audio synthesis link and the end-side audio synthesis link can be realized in a low-delay manner.
Owner:IFLYTEK CO LTD

Interactive digital human generation method and system based on picture and audio synthesis

The invention discloses an interactive digital human generation method based on picture and audio synthesis, and belongs to the crossing field of artificial intelligence and computer graphics. According to the method, a full-link process of'feature extraction-model construction-emotion driving-real-time interaction-video output 'can be automatically completed only by uploading a character picture and a section of audio by a user: key point and semantic feature extraction is performed on the picture to obtain face / posture information; performing voice recognition, semantic analysis and emotion recognition on the audio to obtain a semantic tag and an emotion parameter; generating a personalized three-dimensional digital human based on the information, and driving the personalized three-dimensional digital human to generate facial expressions and limb actions which are synchronous with emotions; the user intention is analyzed in real time through natural language understanding and computer vision, and multi-modal interaction is achieved. According to the method, digital human creation can be completed without professional modeling and motion capture equipment, the generation cost is reduced, and the method can be widely applied to virtual anchors, online education, intelligent customer service and movie and television entertainment scenes.
Owner:NEW ONE (BEIJING) TECH CO LTD

Cross-culture melody feature extraction and musical instrument matching system based on deep learning

The invention discloses a cross-culture melody feature extraction and musical instrument matching system based on deep learning, and belongs to the technical field of music information processing and artificial intelligence. The system comprises an audio acquisition module, an audio preprocessing module, a melody feature extraction module, a cultural context understanding module, an emotion semantic analysis module, a musical instrument timbre database, a musical instrument matching recommendation module, a distributor scheme generation module, a user interaction module and an audio synthesis module. Multi-dimensional melody features are extracted through a multi-scale convolutional neural network and a bidirectional long-short term memory network, cultural semantic understanding is realized through reasoning on a music knowledge graph by using a graph neural network, and musical instrument matching is performed by using a multi-objective optimization algorithm in combination with three dimensions of timbre integrating degree, cultural consistency and emotional expressive power. According to the system, automation of the whole process from humming melody to musical instrument configuration is achieved, the efficiency of the instrument is improved by 400%, the accuracy rate reaches 92%, and an intelligent tool is provided for cross-culture music creation and national music modernization reorganization.
Owner:SHENZHEN UNIV

Feature semantic distinguishing and model single-step screening mimicry audio generation method

The invention discloses a mimicry audio generation method based on feature semantic distinguishing and model single-step screening, and belongs to the technical field of computer audio synthesis and signal processing. The implementation method comprises the following steps: 1, extracting Mel spectrum features from an audio sample, forming a frequency domain spectrogram of the audio through short-time Fourier transform, and performing frequency domain mapping and logarithmic compression on the frequency domain spectrogram by using a Mel filter bank to form a Mel spectrogram; 2, training an audio codec configured with a multi-stage residual quantizer by using fusion loss of semantic classification loss and supervised reconstruction loss to obtain audio features with semantic differentiation; 3, training the audio diffusion model with multi-task diffusion loss; and 4, performing single-step screening on the voice audio by using the trained audio diffusion model in combination with the language audio similarity to obtain a de-noised audio. Compared with the prior art, the technical problem of efficiently generating the high-quality audio consistent with the input semantics when the high-quality vivid scene sound effect is simulated is solved.
Owner:BEIJING INST OF TECH

Hearing aid health analysis method based on physiological signal monitoring function

The application discloses a hearing aid health analysis method based on physiological signal monitoring function, comprising the following steps: collecting physiological signal data of a user in real time through a multi-modal biological sensor array integrated in a hearing aid body; the hearing aid body comprises a digital signal processor and a wireless communication module; the physiological signal data at least comprises an electroencephalogram signal, a galvanic skin response signal and a temporal artery pulse signal; the physiological signal data is subjected to motion artifact elimination processing through an adaptive filtering algorithm; a preprocessed physiological signal is subjected to feature extraction through an embedded health analysis engine, and a composite health feature vector comprising a sympathetic nerve activity index, a cognitive load metric and a cardiovascular function index is generated; a health risk assessment report is dynamically generated based on the association between the composite health feature vector and a preset hearing scene; and a graded health alarm signal is output through a bone conduction vibration unit and an audio synthesis module of the hearing aid body.
Owner:SHENZHEN XINZHENGYU TECH

A method, system, and computer device for generating text-based videos.

This invention proposes a method, system, and computer device for generating text-based videos, relating to the technical field of video processing. The method includes: acquiring user input parameters; generating anchor point prompts and anchor point character diagrams based on the user input parameters; evaluating the anchor point character diagrams until they pass evaluation; extracting information from the anchor point prompts to obtain a storyboard; evaluating the storyboards until they pass evaluation; generating audio and storyboard video based on the storyboards; evaluating the storyboard video until it passes evaluation; post-processing the storyboard video to obtain a post-processed video; and combining the post-processed video with the audio to generate a target video. This invention effectively provides a closed-loop evaluation node for text-based video generation, improving the quality of generated text-based videos.
Owner:GUANGDONG HENGQIN SHUSHUSHUO STORY INFORMATION TECH CO LTD

Vocoder training method, audio synthesis method, medium, apparatus, and computing device

Embodiments of the present disclosure provide a vocoder training method. The vocoder training method comprises: obtaining a first fundamental frequency sequence of audio in an audio corpus; performing fundamental frequency perturbation processing on the first fundamental frequency sequence to obtain a second fundamental frequency sequence; performing mapping processing on the second fundamental frequency sequence to obtain a target tensor; and training the target tensor, an acoustic feature sequence corresponding to the audio, and the audio to obtain a vocoder used for audio synthesis. The method of the present disclosure improves the robustness of the vocoder to fundamental frequency prediction errors in actual application by introducing fundamental frequency perturbation, thereby significantly improving the accuracy and quality of audio synthesis and providing a better experience for users. In addition, embodiments of the present disclosure provide an audio synthesis method, a medium, a device and a computing device.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

A smart management system and method based on video privacy information identification

The application discloses a kind of wisdom management system and method based on video privacy information identification, it is related to information processing identification technical field, including: time stamp matching test is carried out to audio of different length, matching test model is established, audio in court video is extracted, audio is divided and optimization division result is obtained, each segment audio after division is converted into text separately, private content in text is identified, the location of private content in audio is found by time stamp matching and is carried out mute, complete audio is synthesized after mute processing, and the division process of subsequent audio to be muted is planned, after frame extraction is carried out to court video image, private information in each frame image is identified and is carried out coding, audio-video is synthesized after coding video and mute processing audio, the probability that mute is not complete even part of private information in audio is not muted due to that time stamp matching deviation is larger is reduced, the reliability of court audio private information processing is guaranteed.
Owner:JIANGSU XINSHIYUN TECH CO LTD

Sound scene enhancement device based on environmental perception

The utility model discloses a sound scene enhancing device based on environmental perception, and relates to the technical field of acoustics. The system comprises a microphone array, a signal conditioning circuit, an analog-to-digital converter, a DSP processor, a sound scene database memory, a master control MCU, an FPGA beam controller, an audio synthesis chip, a power amplifier, a loudspeaker array and a noise classification coprocessor, the microphone array is electrically connected with the signal conditioning circuit, and the signal conditioning circuit is electrically connected with the analog-to-digital converter. The analog-to-digital converter is electrically connected with the DSP processor, the analog-to-digital converter is electrically connected with the master control MCU, the DSP processor is electrically connected with the sound scene database memory, the DSP processor is electrically connected with the noise classification coprocessor, the noise classification coprocessor is electrically connected with the master control MCU, the master control MCU is electrically connected with the sound scene database memory, and the master control MCU is electrically connected with the FPGA beam controller. The sound scene can be automatically adjusted according to environment changes.
Owner:JIANGSU ACOUSTIC IND TECH INNOVATION CENT

Audio synthesis method, apparatus, device, computer-readable medium, and program product

Embodiments of the present disclosure disclose an audio synthesis method, device, equipment, computer readable medium and program product. A specific implementation of the method comprises: obtaining target audio corresponding to a first language and accent correction information corresponding to an audio synthesis scene, wherein the audio synthesis scene is a scene for synthesizing audio corresponding to a second language based on the target audio; extracting audio feature information corresponding to the target audio; and generating the synthesized audio corresponding to the second language by using a pre-trained speech synthesis large model according to the audio feature information and the accent correction information. The implementation is related to artificial intelligence, and by using the accent correction information, in the process of synthesizing the audio corresponding to the second language based on the target audio corresponding to the first language, the disturbance of the accent noise can be avoided, so that the synthesized audio without the accent noise can be accurately and efficiently generated.
Owner:BEIJING YIJING INFORMATION TECHNOLOGY CO LTD

Model training methods, devices, electronic equipment, computer-readable storage media, and computer program products

This application provides a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product. The method includes: extracting audio feature sequences from raw audio data; determining a first phoneme feature sequence based on the raw audio data and text data; aligning the first phoneme feature sequence based on the audio feature sequence to obtain a second phoneme feature sequence; configuring an omission flag in the second phoneme feature sequence based on text data to obtain a third phoneme feature sequence; structurally associating the audio feature sequence and the third phoneme feature sequence to obtain structured data; storing the structured data in a target storage file; and upon receiving a training instruction, reading the structured data from the target storage file and updating the model parameters of the audio synthesis model based on the data in the structured data that does not have an omission flag configured, thereby obtaining the trained audio synthesis model. This application can improve the training efficiency and accuracy of audio synthesis models.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Sound duplicating method, device and equipment and storage medium

The embodiment of the invention provides a sound copying method and device, equipment and a storage medium, which can be applied to scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving, audio and video, in the method, feature extraction is performed on reference audio in advance, and reference timbre features and reference rhythm features are obtained. And obtaining and storing a sound feature file of the reference audio based on the reference timbre feature and the reference rhythm feature, so that when the reference audio is used as input for multiple times of audio synthesis, only the sound feature file of the reference audio needs to be read in each time of audio synthesis, and based on the first text feature of the text to be synthesized and the sound feature file, the sound feature file of the reference audio is read. According to the method and the device, the synthetic audio corresponding to the to-be-synthesized text is generated without repeatedly reading the reference audio and repeatedly calculating the reference audio, so that the time consumption and resource consumption of sound copying are effectively reduced, and the waiting time of synthesizing the audio by using a sound copying model is also effectively reduced, thereby improving the use experience and enhancing the controllability of the system.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

AI digital human family education method and device based on large language model

The invention provides an AI digital human family education method and device based on a large language model, and relates to the technical field of artificial intelligence. The method comprises the following steps: according to student information, performing adaptive style generation by using a large language model to obtain a personalized interaction style; based on an educational resource database, according to the personalized interaction style, the student information and the student questions, a big language model is used for question answering, and a question answering text is obtained; based on a personalized interaction style, according to the question answering text, performing voice generation by using an Edge-TTS module to obtain a teaching audio; performing video generation by using an optimized video generation model according to the teaching audio and the digital human image picture to obtain a silent teaching video; and performing audio synthesis according to the silent teaching video and the teaching audio to obtain a teaching animation. The AI digital family education method based on artificial intelligence is high in accuracy, high in flexibility, high in real-time performance and high in interactivity.
Owner:HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Audio processing method, system and equipment based on multi-order frequency adjustment and medium

The invention relates to the technical field of audio processing based on multi-order frequency adjustment, and discloses an audio processing method, system and device based on multi-order frequency adjustment and a medium, and the method comprises the steps: obtaining an initial audio; frequency division is carried out on the initial audio frequency to obtain initial frequency band data, and the initial frequency band data comprise low-frequency data, intermediate-frequency data and high-frequency data; performing denoising processing on the initial frequency band data to obtain target frequency band data; performing audio synthesis on the target frequency band data to obtain a to-be-adjusted audio; and based on target cavity characteristic data of the sound production equipment, performing frequency adjustment on the to-be-adjusted audio by adopting at least two EQ adjustment modules for multi-order frequency adjustment to obtain a target audio suitable for the sound production equipment. Therefore, the man-made effect of the frequency band edge is effectively avoided, the frequency curve of the finally output audio is more continuous and smooth, and the audio better meets the hearing requirement of a user.
Owner:深圳市鸿宇光电有限公司

Audio playing method, electronic equipment, storage medium and computer program product

The invention discloses an audio playing method, electronic equipment, a storage medium and a computer program product, relates to the technical field of audio playing, and can improve a sound playing effect when the electronic equipment plays audio. At least four real loudspeakers are arranged in the electronic equipment, and the electronic equipment performs sound mixing processing on audio data to generate a first height sound channel audio signal, a first front sound channel audio signal, a first rear sound channel audio signal and a middle sound channel audio signal. And the electronic equipment generates a second front sound channel audio signal with a wider horizontal sound field and higher height perceptibility of a height sound source according to the first height sound channel audio signal and the first front sound channel audio signal. And the electronic equipment performs sound field broadening on the first rear sound channel audio signal, performs signal synthesis on the second front sound channel audio signal, a second rear sound channel audio signal obtained by sound field broadening and a middle sound channel audio signal, generates audio synthesis signals in one-to-one correspondence with the real loudspeakers, and controls the real loudspeakers to play the audio synthesis signals.
Owner:HONOR DEVICE CO LTD

Audio synthesis method, audio synthesis system, computer device and storage medium

The application relates to an audio synthesis method, an audio synthesis system, a computer device and a computer readable storage medium. The method comprises the following steps: acquiring a current buffer data amount of an audio buffer; the current buffer data amount is used for indicating a playing time length of accompaniment audio data in the audio buffer; based on the difference between the current buffer data amount and a reference buffer data amount, the data amount of the accompaniment audio data buffered in the audio buffer is adjusted to obtain an adjusted audio buffer; the accompaniment audio data is read out from the adjusted audio buffer, and the read accompaniment audio data is played; and when a user sings along with the played accompaniment audio data, the dry audio data of the user is recorded; the read accompaniment audio data and the recorded dry audio data are subjected to audio synthesis processing to obtain synthesized audio data. By adopting the method, the read accompaniment audio data and the recorded dry audio data can be ensured to correspond to each other, and the audio quality of the synthesized audio is improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

An audio cloning method based on cloud computing technology and related equipment

PendingCN122313989AAudio synthesisData center
This application provides an audio cloning method and related equipment based on cloud computing technology to ensure the security of cloned voice and prevent user privacy leaks. The method is applied to a cloud platform, which manages the infrastructure running cloud computing services. The infrastructure includes at least one data center, and each data center includes multiple servers. The method includes: acquiring original audio from a first tenant; encrypting the original audio to obtain encrypted audio; and storing the encrypted audio in the infrastructure. Acquiring a first text from the first tenant. Decrypting the encrypted audio to obtain the original audio. Synthesizing the original audio and the first text to obtain cloned audio, where the sound features of the cloned audio are identical to those of the original audio, and the content of the cloned audio is the content of the first text.
Owner:HUAWEI TECH CO LTD

An adaptive sleep environment system

PendingCN122297868AAudio synthesisNoise
This invention provides an adaptive sleep environment system, belonging to the field of smart home technology. It includes: a sound acquisition module for acquiring the wake-up voice of a monitored subject and the ambient noise within the sleep space of the monitored subject; a vital sign monitoring module for monitoring the sleep vital sign parameters of the monitored subject; a directional audio playback module for playing sleep-aiding audio; and a system host that, in response to the wake-up voice, generates reverse audio based on the ambient noise and synthesizes the reverse audio into the sleep-aiding audio. The directional audio playback module is also used to directionally play the synthesized reverse audio towards the direction of the monitored subject and adjust the volume of the sleep-aiding audio based on the sleep state determined by the sleep vital sign parameters. Beneficial effects: By directionally playing sleep-aiding audio, it avoids interference with other sleepers sharing the same space, thus improving their sleep environment.
Owner:SHANGHAI SIXTH PEOPLES HOSPITAL

Audio synthesis method, apparatus, device, and medium

Embodiments of the present application provide an audio synthesis method, device, equipment and medium, wherein the method comprises: obtaining audio key information used for synthesizing audio data; performing encoding processing on the audio key information to obtain audio attribute features, generating K candidate spectral features according to the audio attribute features and diffusion frequency information; K is a positive integer; obtaining time dimension information and frequency dimension information corresponding to the K candidate spectral features, performing sampling processing on the K candidate spectral features according to the time dimension information and the frequency dimension information to obtain K target spectral features; performing feature fusion processing on the K target spectral features to obtain a fusion spectral feature, and synthesizing the fusion spectral feature into target audio data. By using the embodiments of the present application, the quality of audio synthesis can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Dubbing method and device, electronic equipment and storage medium

The embodiment of the invention provides a dubbing method and device, electronic equipment and a storage medium, and relates to the technical field of data processing. The specific implementation scheme is as follows: obtaining a target line text and an original line audio; performing acoustic feature extraction on the original line audio to obtain at least one acoustic feature; based on the at least one acoustic feature, generating natural language description information about the acoustic feature of the original line audio; and based on the natural language description information, performing audio synthesis processing on the target line text to obtain a dubbing audio which has acoustic characteristics of the original line audio and belongs to the target language. Visibly, according to the scheme of the application, the dubbing efficiency for the target line text can be improved.
Owner:CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD

Multi-target 3D audio rapid synthesis method based on matrix sparse decomposition

The invention discloses a multi-target 3D audio rapid synthesis method based on matrix sparse decomposition. The multi-target 3D audio rapid synthesis method comprises the following steps: S1, performing PCA decomposition on an HRIR database for 3D audio synthesis to obtain a principal component score matrix and a principal component vector matrix; s2, performing sparse decomposition on the PCA principal component vector matrix obtained in the step S1 to obtain a kernel matrix and a sparse basis matrix; and S3, during multi-target 3D audio synthesis, selecting a principal component score corresponding to each target, sequentially multiplying sound source data to be processed by the principal component score matrix, the sparse basis matrix and the kernel matrix, and taking a trace to obtain processed 3D audio data. The method has the beneficial effects that the HRIR database is compressed under the condition that the reconstruction error is very small, and the calculation efficiency of multi-target 3D audio synthesis is greatly improved.
Owner:SHANGHAI AVIATION ELECTRIC

Self-adaptive audio coding automatic driving state auditory prompting method and self-adaptive audio coding automatic driving state auditory prompting system

The invention provides a self-adaptive audio coding automatic driving state auditory prompting method and system, and is applied to the technical field of intelligent driving. Internal operation state data and environment target data of an automatic driving system are collected in the vehicle operation process; based on the current decision-making result, decision-making related targets which directly influence the decision-making are screened out; global audio characteristic parameters are generated by performing time sequence modeling on the internal operation state and are used for expressing the overall operation situation of the automatic driving system; meanwhile, importance evaluation is conducted on decision-related targets, corresponding local audio attribute parameters are generated and used for highlighting auditory significance of the key targets, the global audio feature parameters serve as auditory backgrounds to be continuously output, the local audio attribute parameters serve as auditory foregrounds to be dynamically overlapped, and auditory prompts are formed through audio synthesis and playing. In this way, continuous and explainable auditory expression of the running state and decision causal information of the automatic driving system is achieved.
Owner:JILIN UNIVERSITY

Audio synthesis method, audio synthesis model training method and related devices

This application provides an audio synthesis method, an audio synthesis model training method, and related devices. The method includes: determining text features of a first text and audio data of a first object; performing a first encoding on the audio data to obtain the voice features of the first object, and performing a second encoding on the audio data to obtain the musical features of the audio data; performing attention processing on the musical features and text features to obtain a first feature; performing a fusion processing on the voice features, the first features, and the text features to obtain a second feature; and decoding the second feature to obtain synthesized audio. Through this application, the accuracy of audio synthesis can be improved, and synthesized audio consistent with the sound of the first object can be obtained.
Owner:MASHANG CONSUMER FINANCE CO LTD

Audio dialogue method and device, equipment and storage medium

The embodiment of the invention provides an audio dialogue method and device, equipment, a storage medium and a program product. The method comprises the following steps: encoding a first audio stream collected from an environment into an audio feature sequence by using a streaming audio encoder; generating a sequence of text units as a response to the first audio stream based on the system cue word and the audio feature sequence using a trained machine learning model; and in response to the sequence of text units satisfying the audio synthesis condition, using the streaming audio synthesizer to generate a second audio stream from the sequence of text units for playing. In this way, the audio conversation can be realized in a full duplex and streaming manner, and discrete audio coding does not need to be introduced in the audio generation process. This may improve the performance and efficiency of the audio dialog.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +2

Artificial intelligence-based audio generation methods, devices, equipment, and storage media

This application provides an audio generation method, apparatus, electronic device, and computer-readable storage medium based on artificial intelligence; it relates to artificial intelligence technology; the method includes: sampling multiple audio data of a target object to obtain reference audio data of the target object; performing audio encoding processing on the reference audio data of the target object to obtain a reference embedding vector of the reference audio data; performing timbre-based attention processing on the reference embedding vector of the reference audio data to obtain a timbre embedding vector of the target object; performing text encoding processing on the target text to obtain a content embedding vector of the target text; and performing synthesis processing based on the timbre embedding vector of the target object and the content embedding vector of the target text to obtain audio data that conforms to the timbre of the target object and corresponds to the target text. This application can improve the stability of audio synthesis.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Context aware sound field mapping system and method for enhancing display interaction

The invention provides a context-aware sound field mapping system and method for enhancing display interaction, and the system comprises a context analysis module which is used for analyzing the voice content to be broadcasted by AI, and extracting key semantics, emotion, intention and hidden scenes; the acoustic environment sensing module is used for analyzing acoustic characteristics of a real environment where a user is located through a microphone array; the background sound generation and selection module is used for generating or retrieving a most matched background sound material from a local or cloud sound database based on the analyzed context; and the audio synthesis and spatialization rendering module is used for intelligently mixing the dry sound signals of the AI with background sound, and applying a spatial audio technology, so that the finally output sound has a sense of direction, a sense of distance and a sense of environmental fusion. The method is used for dynamically generating and rendering the spatialized background sound matched with the dialogue context for artificial intelligence voice interaction, so that immersion of digital entity existence in a physical environment is created.
Owner:BEIJING LANGZHIWAN INTELLIGENT TECHNOLOGY CO LTD