Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

829 results about "Timbre" patented technology

In music, timbre (/ˈtæmbər, ˈtɪm-/ TAM-bər, TIM-, French: [tɛ̃bʁ]), also known as tone color or tone quality (from psychoacoustics), is the perceived sound quality of a musical note, sound or tone. Timbre distinguishes different types of sound production, such as choir voices and musical instruments, such as string instruments, wind instruments, and percussion instruments. It also enables listeners to distinguish different instruments in the same category (e.g., an oboe and a clarinet, both woodwind instruments).

Voice generation method and device based on multi-modal fusion, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology, cultural transmission and the like, and discloses a voice generation method based on multi-modal fusion, which comprises the following steps: collecting audio data to extract timbre features, and training a field feature timbre generation model; analyzing text semantic recognition emotion information, adjusting speech synthesis parameters, combining personalized information to construct a parameter mapping table, fusing to generate a synthesis control parameter sequence, aligning the synthesis control parameter sequence with character labels, visual elements and background music data, driving a domain feature timbre generation model, and generating synthesis data of synchronous speech, text, vision and music. Domain timbres are generated through timbre feature training, speech expression is optimized by combining semantic analysis and emotion recognition, user requirements are matched based on personalized information, and time alignment is performed by fusing text, vision and music data, so that the synthesized speech has domain features, emotion adaptability and personalization, and the speech quality is improved. And the voice immersion and the information transmission capability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Tone conversion method and device based on cultural semantics, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to the field of medical health, and discloses a timbre conversion method, device, equipment and medium based on cultural semanteme, the method comprises the following steps: constructing a cultural semantic timbre library comprising a semantic label and timbre characteristic parameter mapping relationship, the semantic label characterizing emotional semanteme of a target timbre, and the timbre characteristic parameter mapping relationship between the semantic label and the timbre characteristic parameter; the timbre characteristic parameters comprise a pitch range, rhythm rhythm and a harmonic structure; performing feature extraction based on text, image and audio multi-mode information to obtain semantic keywords, visual emotion features and audio acoustic features; performing attention weight fusion on the features through a multi-modal fusion deep learning model, and dynamically adjusting model parameters in combination with a semantic timbre library to generate a target timbre; and finally, intelligent conversion from the multi-mode information to the adaptive tone is realized. Through semantic-driven multi-modal feature collaborative optimization, the defect that timbre conversion machinery is stiff and lacks emotional expression is overcome, and the integrating degree of timbre expression and semantic scenes is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Cross-language voice migration synthesis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a cross-language voice migration synthesis method, device, equipment and medium. And training an acoustic model by adopting a hierarchical self-adaptive fine tuning strategy, fusing the phoneme sequence for reasoning and the tone mark to generate a representation sequence, and finally synthesizing a target language speech signal. According to the method, a sharing and separation parallel cross-language modeling structure is constructed, the accuracy of phoneme and tone modeling in a low-resource language is effectively improved, the target language acoustic model has higher generalization ability and migration efficiency in combination with a multi-stage self-adaptive fine tuning and training data enhancement strategy, and the method is suitable for being applied to the field of multi-language modeling. And finally, synchronous improvement of the voice naturalness and the tone fidelity is realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Speech cloning system and method fusing rhythm characteristics

The invention discloses a voice cloning system and method fusing rhythm characteristics, belongs to the technical field of voice synthesis and natural language, and is applied to the aspect of fine-grained rhythm control in zero-sample voice synthesis. The implementation method comprises the following steps of: 1, extracting rhythm features and audio features of an audio file, and further respectively acquiring pause features, speed features and tone features in the rhythm features by sequentially adopting transcriptional text inverse coding, syllable-level speed registration quantization and pitch sequence feature splicing modes; 2, fusing the features of the audio files in a manner of discarding feature screening without guidance of a classifier; 3, generating a target Mel spectrogram based on conditional flow matching; generating a target audio file from a to-be-cloned audio file through the trained voice cloning model controlled by the fusion rhythm; compared with the prior art, fine-grained rhythm control of tone and rhythm feature decoupling is realized in zero-sample speech synthesis, so that intonation accuracy based on a context scene is improved.
Owner:BEIJING INST OF TECH

Speech synthesis scheme and device, electronic equipment, storage medium and program product

The invention provides a voice synthesis scheme and device, electronic equipment, a storage medium and a program product, and relates to the technical field of voice processing, and the method comprises the steps: inputting a text content into an LLM semantic understanding module, and obtaining a deep semantic feature and multi-modal context semantic information corresponding to the text content; transmitting the deep semantic features to a neural codec, and outputting compressed acoustic features; inputting the compressed acoustic features into an acoustic modeling module, and outputting high-precision acoustic features; inputting the high-precision acoustic features and the multi-mode context semantic information into an emotion and rhythm control module, and inputting voice parameters with emotion and rhythm marks; and inputting the voice parameters with the emotion and rhythm marks and the reference audio into a tone migration module to obtain a synthetic voice of a tone corresponding to the reference audio.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Real-time sound duplicating method and system based on end-cloud fusion

The invention provides a real-time sound copying method and system based on end-cloud fusion. The method comprises the following steps: a cloud end carries out real-time tone copying and voice synthesis on a small amount of voice data of a user based on an AI large model; timbre samples are collected when a user registers voice audio data, and user timbre voice data of a preset text are synchronously generated by a large model and are used as fine tuning training data of an end-side voice synthesis model; user tone voice data of a preset text and voice audio data registered by a user are utilized to carry out migration fine tuning training on an end-side voice synthesis model to adapt to the personalized tone of the user, so that high-quality output of the end-side voice synthesis model is ensured, and personalized voice replication is realized; and issuing the trained end-side speech synthesis model to the user equipment, and independently completing speech replication in a network-free or weak network environment. According to the invention, through automatic generation of the user tone data and adaptive fine tuning of the model, the tone of the user is deployed to the end side after fine tuning, and high-quality and high-adaptability sound replication of end-cloud collaboration is realized.
Owner:PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD

Enhanced wireless communication handover management system

To effectively resolve the issue of inappropriate handovers during communication sessions, methods, systems, and machine-readable mediums which utilize speech features captured by microphones in the original wireless peripheral device and / or the wireless peripheral device to which the communication session is to be handed over to determine if a handover should proceed or should be reversed. Speech features refer to the composite attributes of spoken language that encompass both acoustic and linguistic features. Acoustic features characterize the sound properties of speech and include, but are not limited to, timbre, pitch, intonation, speaking rate, articulation, prosody, melody, spectral features, formant frequencies, and the like. Linguistic features pertain to the actual content conveyed, comprising words, phrases, syntax, and semantics. This includes the analysis of words, phrases, syntax, and semantics to understand the context and continuity of the conversation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Novel empty drum

The utility model relates to a novel empty drum. The novel empty drum comprises a drum body, the drum body comprises an upper drum body and a lower drum body which are fixedly connected, and eight peripheral tongues are arranged around the geometric center of the upper drum body; each peripheral tone tongue is provided with a first sub-tone tongue, a second sub-tone tongue, a third sub-tone tongue, a fourth sub-tone tongue and a fifth sub-tone tongue, and the first sub-tone tongue, the second sub-tone tongue, the third sub-tone tongue and the fourth sub-tone tongue are arranged close to the fifth sub-tone tongue; the geometric center of the lower drum body is provided with a timbre control hole and a rubber plug matched with the timbre control hole. According to the novel empty drum, the structure of the peripheral tongues is improved and designed, so that the tone of the empty drum in the playing process is richer, and the use experience of a user is improved.
Owner:惠州市广利润发科技有限公司

Multi-dimensional scoring auxiliary system and method in piano teaching

The invention relates to the technical field of piano teaching and computers, discloses a multi-dimensional scoring auxiliary system and method in piano teaching, and aims to solve the problems that existing piano teaching is single in evaluation dimension, one-sided in data acquisition and lagged in feedback. According to the method, a five-dimensional scoring model covering rhythm, strength, timbre, posture and fingering is constructed by synchronously collecting a key triggering time sequence, an audio frequency spectrum, limb posture and finger joint tracks, and quantitative scores are output through time alignment, feature extraction and neural network fusion calculation. The system comprises a key sensing module, an audio acquisition module, a posture capture module, a score calculation module and an AR visual feedback module, and supports teaching suggestion generation, historical trend analysis, difficulty self-adaption and multi-person comparison functions. Through multi-modal data closed-loop evaluation and intelligent intervention, the teaching accuracy and individuation level are remarkably improved, and scientific quantification and efficient improvement of the piano playing ability are achieved.
Owner:JILIN NORMAL UNIV

Intelligent identification and visual playing system for ancient music score

The invention relates to the technical field of ancient music score intelligent processing, and discloses an ancient music score intelligent identification and visual playing system, which comprises an image preprocessing module, a stroke enhancement module, a symbol segmentation module, a symbol identification module, a semantic analysis module, a visualization module, an acoustic synthesis module and the like. Through a specially designed image processing and symbol analysis method, the aging problem of the ancient music score can be accurately repaired, adhered symbols are separated, a complex structure is identified, music semantics are inferred in combination with a knowledge graph, and a multi-modal visualization effect and a high-sampling ancient charm tone are generated. The system supports feedback optimization and automatic workflow, significantly improves the efficiency and accuracy of ancient music score recognition, translation and presentation, and provides technical support for ancient music research and propagation.
Owner:ANHUI UNIV

Voice evaluation method and device, medium and program product

The invention provides a voice evaluation method, a voice evaluation device, a non-transitory storage medium and a computer program product. The voice evaluation method comprises the following steps: extracting a first person voice with a first tone based on a to-be-evaluated voice; converting the voice of the first person into voice with a second tone; and obtaining an evaluation result of the voice of the first person based on the converted voice with the second tone and the standard voice of the one or more accent.
Owner:NEW ORIENTAL EDUCATION & TECH GRP CO LTD

Video dubbing language conversion method and system and related equipment

The invention provides a video dubbing language conversion method, a video dubbing language conversion system and related equipment. The method comprises the following steps: acquiring audio track data from a video to be converted; carrying out human voice extraction on the audio track data and classifying according to roles to obtain a single speaker audio of each role; performing voice-to-text conversion on the single speaker audio of each role to obtain an original language copywriting of each role; performing sound cloning on the single speaker audio of each role to obtain a timbre model of each role; performing target language translation on the original language copywriting of each role to obtain a translated copywriting of each role; based on the translation copywriting of each role and the tone model of each role, performing text-to-voice conversion to obtain a translation audio of each role; and performing replacement of each role translation audio on the audio track data in the to-be-converted video to obtain a dubbing conversion video. According to the technical scheme, language video dubbing conversion combined with the tone of the speaker is achieved, the video is more diversified, and the user requirements can be better met.
Owner:SHENZHEN MAIFENG TECH CO LTD

Method for generating virtual character and interacting with user

PendingCN120495484AAnimationAnthropomorphaEngineering
The invention provides a method for generating a virtual character and interacting with a user, and solves the problems of insufficient personification degree and limited interaction capability of the virtual character in the prior art. The method comprises the steps of obtaining information of a virtual character to be generated, generating a cue word Prompt used for generating a virtual character image, generating a character image through the cue word, generating timbre, fixed verbal skill voice and video based on the virtual character image, and optimizing the timbre, the fixed verbal skill voice and the video; according to the virtual character and the large language model, interaction between the user and the virtual character is achieved, the anthropomorphic degree of the virtual character is improved, and the interaction ability of the virtual character is improved.
Owner:SHANGHAI 2345 NETWORK TECH

Solution for TTS (Tone To Send) in high-concurrency scene by clone timbre

The invention belongs to the technical field of text-to-speech conversion, and particularly relates to a solution for TTS (text-to-speech) in a high-concurrency scene through tone cloning. According to the method, the text can be generated and returned to the client by segmenting and parallelizing the text and reducing the delay of the first segment, compared with a traditional serial whole segment synthesis mode, the waiting time of a user is greatly shortened, the audio is returned after being generated, the waiting time of the user can be effectively shortened under high concurrency, the network bandwidth and the delay pressure are reduced, and the user experience is improved. The real-time performance is guaranteed, reconnection disasters are avoided, underlying model services can be fully utilized, the influence of the serial processing characteristic of a CosyVoice2 model is avoided, a client can immediately start to play a first segment of audio, even if subsequent segments are still synthesized, the continuous playing experience of previous content is not influenced, and when the concurrency is increased, the continuous playing experience of the previous content is not influenced. The model gateway can dynamically disperse requests to a plurality of GPU instances, and long queue waiting after a single-card video memory is fully occupied is avoided.
Owner:BEIJING HUAYUN WORLD TECH CO LTD

Flute ethnic music library construction method and system

ActiveCN120388550AElectrophonic musical instrumentsFlute repertoryTimbre
The invention relates to the technical field of music information management, in particular to a flute ethnic music library construction method and system, and the method comprises the following steps: obtaining flute track audio data, dividing paragraphs according to a time axis, extracting a tone waveform change track, carrying out the cumulative statistics of tone turning, intercepting melody main line information, and recognizing the rising and falling and continuous changes of musical scales. Defining path nodes, marking melody structures, comparing timbre fluctuation segments, recording style change positions, and generating style music library field combinations. According to the method, the audio data of the flute track is finely divided, the tone waveform change is counted, the key tone change in the track is captured, and the node division is performed through the directivity of the musical scale trend, so that the multi-dimensional visualization of the information is increased, and the structural characteristics of the melody are clearer; and national style sample tracks are called for structure comparison, and fragments with consistent styles are counted, so that the convenience and efficiency of searching and identifying specific style tracks by a user are greatly improved.
Owner:WEIFANG UNIVERSITY

Electronic piano accompaniment system with adaptive rhythm based on artificial intelligence

The invention provides a self-adaptive rhythm electronic piano accompaniment system based on artificial intelligence, and the system comprises a single-key triggering module which is used for indicating an accompaniment process; the speed-control signal conversion module is used for converting the speed signal acquired by the inductive sensor into a music control signal; the storage module is used for storing accompaniment note sequence data of the song; and the artificial intelligence accompaniment generation module adjusts the accompaniment rhythm according to the trigger signal and the playing speed, and generates accompaniment sound or harmony in real time. The system further comprises a pitch calculation unit, a timbre parameter adjustment module, a timbre generation module, a dynamic range adjustment module, a timbre smooth transition module, an automatic harmony generation module, a volume dynamic adjustment module and an audio output module. The control unit receives a user instruction and adjusts system setting parameters. Intelligentization and self-adaption of the accompaniment can be achieved, self-adjustment of the playing speed of the accompaniment is achieved, the playing speed of the accompaniment is kept consistent with the singing speed of the user, the playing flexibility and expressive force are improved, and the playing experience is improved.
Owner:陈必红

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

HiFi cable made of at least 24 layers of pure carbon with magnetic field protection

A metal-free signal cable made of at least 24 carbon rovings twisted together for use in music reproduction as a loudspeaker cable, small signal cable or bridge, which is kept free from external influences by spacers made of electromagnetically practically neutral aluminum-magnesium alloy. Even the highest-quality hi-fi systems sound significantly more energetic, clearer, more accurate, more tangible, and even more "in the room" when using my invention. Transients are much more realistic, the reproduced stage size increases significantly, the sound and positioning of the instruments involved are noticeably improved, and each individual note becomes significantly more subtle, complex, and rich in nuance. Even the smallest sound events become perceptible, which are lost with conventional solutions.
Owner:GESCHKE JAN HENDRIK

End side voice model deployment method and device, equipment and storage medium

The invention relates to the technical field of end-side model deployment, in particular to an end-side voice model deployment method and device, equipment and a storage medium. Comprising the steps of recording reference audio at an end side; extracting a reference semantic token and an embedded vector of the reference audio; obtaining a training text set, integrating and inputting the training text code, the embedded vector and the reference semantic token corresponding to each training text in the training text set into a preset language model, and outputting a comprehensive token sequence corresponding to the training text; performing Mel spectrum conversion on each comprehensive token sequence to obtain a first spectrum representation corresponding to the comprehensive token sequence; generating an audio signal corresponding to the training text according to each first spectrum representation, and integrating to obtain a training data set; inputting the training data set into a to-be-trained model, and training to obtain a lightweight voice model which only retains timbre modeling parameters; and updating the lightweight voice model to the end side to complete end side deployment. According to the invention, the voice synthesis model deployment of the end-side equipment can be realized.
Owner:SHENZHEN RAISOUND TECH

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-mode background sound condition input data; performing modal integrity detection on the multi-modal background sound condition input data to obtain a detection result; generating an environment background sound feature embedding vector according to a detection result; obtaining to-be-synthesized text data and speaker reference audio data, and performing feature extraction to obtain text semantic features and speaker timbre features; inputting the environment background sound feature embedded vector, the text semantic feature and the speaker timbre feature into an acoustic model to generate a Mel spectrum; and converting the Mel spectrum into a target voice waveform to obtain synthetic voice data. By implementing the method, scene requirements can be deeply matched, diversified scene types can be covered, accurate matching of background sounds and voice semantics is realized, and the technical scheme can be applied to the fields of finance and medical health.
Owner:PING AN TECH (SHENZHEN) CO LTD

Mediator timbre cloning method and system, electronic equipment and storage medium

The invention provides a mediator timbre cloning method and system, electronic equipment and a storage medium. The method comprises the following steps: selecting a tone of a mediator in response to a selection instruction input by a user; obtaining an input text; predicting a target audio feature vector of the input text by using an autoregression model; a target clustering center matched with the target audio feature vector is searched in an audio dictionary, the audio dictionary comprises a plurality of clustering clusters, and each clustering cluster comprises a plurality of audio feature vectors; and inputting the input text phonemes of the input text, the frequency characteristics of a reference audio and the target clustering center into the trained acoustic model to obtain a target output audio, the reference audio being a mediator audio corresponding to the tone of the mediator. According to the method, the mediation efficiency can be improved, the service consistency is ensured, the user experience is improved, and the method has better practicability and market competitiveness.
Owner:SHANGHAI JINQIAO YIFA INFORMATION TECH CO LTD

Personalized communication system fused with music emotion perception

The invention discloses a communication resource optimization method and system based on multi-source emotion perception, and aims to solve the problem that emotion analysis and communication resource allocation are disjointed in the prior art. The method comprises three core modules: a music emotion analysis module, a multi-source emotion perception module and an emotion-channel mapping engine module. The method comprises the following steps: firstly, constructing a music emotion vector by extracting a Mel-frequency cepstral coefficient, a rhythm characteristic and a timbre characteristic of a music signal; secondly, fusing the user voice fundamental frequency contour and the micro-expression dynamic characteristics to generate a user emotion vector; and finally, constructing an emotion-channel mapping engine based on deep reinforcement learning (DRL), and realizing dynamic optimization of a modulation coding scheme by defining a joint state space and designing a third-order reward function. According to the method, the emotion satisfaction degree is creatively used as a core index of resource allocation, the problem of emotion perception deficiency in traditional communication optimization is solved, and the user experience and the network resource utilization rate are remarkably improved.
Owner:马白雪

Multi-modal feature fusion high-quality intelligent sound ray editing method and device

The embodiment of the invention provides a multi-modal feature fusion high-quality intelligent sound ray editing method and device, and the method comprises the steps: determining an original timbre spectrum corresponding to an original audio according to a timbre conversion request for the original audio, and determining a target timbre spectrum corresponding to the timbre conversion request; determining a to-be-adjusted frequency band corresponding to the original audio based on the timbre conversion request, and replacing the original timbre spectrum by using the target timbre spectrum based on the to-be-adjusted frequency band to obtain a timbre control spectrum; and fusing the timbre control spectrum in an initial audio feature corresponding to the original audio to obtain a target audio feature, and generating a target audio corresponding to the target timbre spectrum based on the target audio feature. The generated target audio can realize more efficient, more accurate and more natural tone conversion on the premise of improving the audio emotion expressive force effect.
Owner:SHANGHAI LE ELEMENT WORLD TECH CO LTD

Personalized tone migration and synthesis method based on virtual singer

The invention relates to a personalized timbre migration and synthesis method based on a virtual singer, and the method comprises the steps: obtaining a reference singing audio of a target virtual singer, and extracting a timbre identity benchmark feature used for representing the timbre identity stability through timbre coding; obtaining personalized timbre migration demand information for the virtual singer, wherein the demand information comprises a to-be-migrated timbre attribute and a corresponding target change amplitude or change direction; calculating a compatibility score according to the timbre identity reference feature and a migration demand, generating a timbre migration risk indication value, and comparing the risk indication value with a preset threshold value or a preset threshold value interval to determine that the migration is high-risk migration or low-risk migration or conservative-risk migration; according to the method, the tone identity benchmark features in the reference singing audio are extracted, and the tone identity invariant set with stable transpitch and sounding intensity is further constructed, so that the core tone identity of the virtual singer is accurately described.
Owner:CHANGSHA NORMAL UNIV

Video data processing method and device and electronic equipment

The invention provides a video data processing method and apparatus, and an electronic device. The method comprises the steps of obtaining initial video data; the initial video data comprises initial video stream data and initial audio stream data corresponding to the initial video stream data; determining a role feature of at least one audio generation role in the initial video data; determining scene features of the initial video data; based on the role features, the scene features and the content translation text corresponding to the initial audio stream data, determining input data; inputting the input data into a preset voice generation model, and generating target audio stream data through the voice generation model; and generating target video data based on the initial video stream data and the target audio stream data. According to the mode, in the process of generating the translated audio corresponding to the video, tone cloning and emotion restoration of the translated audio are realized through the role features of the speaker of the original video and the scene features embodied by the video, the translation effect of the video is improved, and the video translation cost is reduced.
Owner:WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD

Voice conversion method and related equipment

The embodiment of the invention discloses a voice conversion method and related equipment. The related equipment can comprise a voice conversion device, electronic equipment, a computer program product and a computer readable storage medium. After at least one to-be-converted voice and a target timbre identifier corresponding to the to-be-converted voice are acquired, audio content features and acoustic features are extracted from the to-be-converted voice, and the target timbre feature corresponding to the to-be-converted voice is determined based on the target timbre identifier; extracting audio conversion features from the audio content features, extracting rhythm features from the acoustic features, fusing the target timbre features, the audio conversion features and the rhythm features to obtain target audio features, and then generating target voice corresponding to the target timbre identifier based on the target audio features; according to the scheme, the voice conversion accuracy can be improved. The embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method for training speech synthesis model, speech synthesis method, and electronic device

A method for training a speech synthesis model includes obtaining training data; obtaining an initial speech synthesis model; training a semantic encoding network and a semantic decoding network in the speech synthesis model respectively based on a style sample speech, a timbre sample speech, an input sample text, and an output sample speech in training samples of the training data, to obtain a trained speech synthesis model.
Owner:BAIDU INT TECH (SHENZHEN) CO LTD