Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

597 results about "Synthesis methods" patented technology

Synthesizers use various methods to generate electronic signals (sounds). Among the most popular waveform synthesis techniques are subtractive synthesis, additive synthesis, wavetable synthesis, frequency modulation synthesis, phase distortion synthesis, physical modeling synthesis and sample-based synthesis.

Cross-language voice migration synthesis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a cross-language voice migration synthesis method, device, equipment and medium. And training an acoustic model by adopting a hierarchical self-adaptive fine tuning strategy, fusing the phoneme sequence for reasoning and the tone mark to generate a representation sequence, and finally synthesizing a target language speech signal. According to the method, a sharing and separation parallel cross-language modeling structure is constructed, the accuracy of phoneme and tone modeling in a low-resource language is effectively improved, the target language acoustic model has higher generalization ability and migration efficiency in combination with a multi-stage self-adaptive fine tuning and training data enhancement strategy, and the method is suitable for being applied to the field of multi-language modeling. And finally, synchronous improvement of the voice naturalness and the tone fidelity is realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-agent-based material performance prediction and synthesis method and system

The invention relates to a multi-agent-based material performance prediction and synthesis system, and the system comprises a multi-agent data enhancement module which is configured to be used for firstly disassembling a complex problem into a plurality of subtasks, and then constructing a fine tuning data set comprising Sub-CoQ question and answer pairs by starting multi-source parallel retrieval; the multi-expert debate module is configured to be used for simulating decision conflicts of different roles in material engineering and generating a direct preference optimization DPO data set through debate; the training and verification module is configured to be used for training and verifying a large model MatMind in the field of materials by utilizing supervised fine tuning SFT and reinforcement learning RLHF based on the fine tuning data set and the DPO data set; and the material performance prediction and synthesis module is configured to be used for realizing intelligent recommendation of a material performance prediction and synthesis process by importing input parameters into the large model MatMind.
Owner:SHANGHAI INST OF CERAMIC CHEM & TECH CHINESE ACAD OF SCI

Sound duplicating and low-delay streaming speech synthesis method and system based on ultra-short sample

The invention provides a voice replication and low-delay streaming voice synthesis method and system based on an ultra-short sample, relates to the technical field of artificial intelligence, and is suitable for intelligent interaction, outbound service and multi-mode communication scenes. Deep personalized customization of intelligent voice interaction and real-time generation of ultra-low delay are realized, and bidirectional streaming interaction of a system level is supported, so that the fluency and response speed of dialogues are improved; an ultra-short sample sound duplicating module specially designed for processing ultra-short audio samples and a speech synthesis engine with ultra-low delay and bidirectional streaming output capability are integrated, and the capabilities of the ultra-short sample sound duplicating module and the speech synthesis engine are applied to a real-time and interactive intelligent speech interaction process to form a complete and efficient solution. The industrial pain point is directly solved in a targeted manner, and the method has important commercial application value.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Phoneme time axis driven high-definition video mouth shape automatic synthesis method

The invention discloses a high-definition video mouth shape automatic synthesis method driven by a phoneme time axis, and relates to the technical field of electrical digital data processing, and the method comprises the steps: 1, obtaining an input text or voice, employing a fixed phoneme set and a deterministic pronunciation rule to generate a phoneme sequence and a corresponding starting and ending time point, and building the phoneme time axis; outputting the phoneme time axis, the frame index mapping table and the target frame set; 2, on the basis of a fixed and limited mouth shape dictionary and a pronunciation part rule, three complementary derivation chains are established at the same time, and a mouth shape key frame sequence meeting phase continuity and boundary traceability is obtained; and step 3, extracting three-dimensional face feature points and a head posture track based on the original video, and outputting a video mouth shape synthesis result which is consistent with the resolution of the original video and is synchronous with a phoneme time axis according to the posture uniformization mouth shape sequence. According to the invention, the accuracy and reality of mouth shape video generation are obviously improved.
Owner:CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Speech synthesis method and device, vehicle and storage medium

The embodiment of the invention provides a voice synthesis method and device, a vehicle and a storage medium, and the method comprises the steps: obtaining multi-modal data which comprises a voice signal, a user instruction, a historical interaction log and bus data; determining model input parameters according to the multi-modal data; according to the model input parameters, the static corpus and a preset large model, generating a personalized script, the preset large model being used for dynamically generating the personalized script in combination with the model input parameters and the static corpus; and synthesizing the broadcast audio according to the personalized script. According to the invention, the technical problem that the voice synthesis technology in the related technology cannot adapt to the personalized demands of the user is solved.
Owner:GUANGZHOU AUTOMOBILE GROUP CO LTD

Method and apparatus for synthesizing spatially cross-correlated multi-point ground motions fitting response spectrum

Disclosed in the present invention are a method and apparatus for synthesizing spatially cross-correlated multi-point ground motions fitting a response spectrum. The method comprises: using an influence matrix method to obtain ground motion acceleration time histories fitting a target reaction spectrum with high precision, and calculating a Fourier amplitude spectrum and a phase spectrum thereof; and considering a calculation relationship between a power spectral density function and the Fourier amplitude spectrum, taking the power spectral density function of the obtained time histories as an initial value, using a random vibration theory and Cholesky decomposition to obtain spatially cross-correlated multi-point ground motion time histories, further taking into account a phase spectrum of the ground motion time histories fitting the target reaction spectrum with high precision, and performing correction and iteration to obtain spatially cross-correlated multi-point ground motions fitting a response spectrum. The present invention overcomes the shortcomings in current multi-point ground motion time history synthesis techniques, such as the difficulty in ensuring good fitting to a target response spectrum and the non-stationarity of synthesized ground motions, helping to improve the reliability of dynamic response analysis results for large-scale engineering structures under the action of multi-point ground motions.
Owner:JIANGNAN UNIV

Speech synthesis method and device based on optimization strategy algorithm, equipment and medium

The invention relates to the technical field of intelligent decision making, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a speech synthesis method, device, equipment and medium based on an optimization strategy algorithm, comprising: extracting a time sequence processing network unit and a data sampling scheduling unit in a speech synthesis model; mapping a denoising function in the time sequence processing network unit into a multi-step Markov decision function, and converting an ordinary differential equation in the data sampling scheduling unit into a multi-source stochastic differential equation; sampling multiple groups of independent audio tracks corresponding to the input text based on a multi-source stochastic differential equation; calculating a strategy gradient modulation factor by using an optimization strategy algorithm and a multi-step Markov decision function; optimizing strategy parameters of the speech synthesis model according to the strategy gradient modulation factor to obtain an optimized speech synthesis model; and obtaining a to-be-converted text, and synthesizing voice corresponding to the to-be-converted text by using the optimized voice synthesis model. And the speech synthesis accuracy is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Defect synthesis method of SMT element based on cGAN network driving

The invention provides a defect synthesis method of an SMT (Surface Mount Technology) element based on cGAN network driving, which comprises the following steps: acquiring an input condition, constructing an adversarial loss function, an L1 reconstruction loss function, a circulation loss function and a physical loss function, and carrying out weighted summation on the four loss functions to obtain a total loss function. Training the cGAN network according to the total loss function to obtain an SMT element defect synthesis model; generating a virtual defect sample by adopting an SMT element defect synthesis model; wherein the virtual defect sample is used for reducing the false detection rate of the SMT defect detection model. The virtual defect sample generated by the SMT element defect synthesis model has visual authenticity, and the defect form in the virtual defect sample and the real welding process have physical consistency. The virtual defect sample generated by the SMT element defect synthesis model is combined with the real defect sample, so that the trained SMT defect detection model is suitable for SMT defect detection under a small sample condition, and the SMT defect detection precision under a real processing scene is improved.
Owner:BEIJING DEZHI MATRIX TECHNOLOGY CO LTD +3

Audio synthesis method, audio synthesis model training method, apparatus, electronic device, computer-readable storage medium, and computer program product

An audio synthesis method, an audio synthesis model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which relate to artificial intelligence technology. The method includes: invoking an audio synthesis model based on language information and preset style information of a target text to perform following processing, the audio synthesis model including a prior encoder and a waveform decoder: generating audio features corresponding to the target text based on the language information and the preset style information by using the prior encoder; performing normalizing flow processing on the audio features by using the prior encoder, to obtain a hidden variable of the target text; and performing waveform decoding on the hidden variable of the target text by using the waveform decoder, to obtain a synthetic waveform conforming to an audio style described in the preset style information and corresponding to the target text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Deterministic material synthesis method and system for chiral intervention before nucleation

The invention aims to serve the national strategic emerging industry, and discloses a deterministic material synthesis method and a deterministic material synthesis system for pre-nucleation chiral intervention (PNChI for short), which are suitable for one-dimensional or quasi-one-dimensional material systems (such as carbon nanotubes, biomacromolecules and the like) with formalized chiral tags. According to the method, before irreversible closed nucleation of a precursor is completed, deterministic authorization of a target chiral index (n, m) is realized in a controllable time window through at least one engineering intervention mechanism (such as physical field regulation, catalytic interface design, information coding guidance and the like) based on distinguishable physical properties, and consistency is kept in subsequent growth. The system comprises a programmable intervention unit and an intelligent sequence control module, and does not depend on a separation or screening step after nucleation. The single chiral index abundance in the obtained material is greater than or equal to 95%, preferably greater than or equal to 98%. The statistical limitation of traditional chiral control is broken through, a general synthesis normal form from'instruction input 'to'structure output' is constructed, and the method is suitable for high-consistency application scenes such as semiconductors, quantum devices, catalysts and biological recognition.
Owner:SHAANXI TAIWAT THERMAL POWER TECHNOLOGY CO LTD

Speech video synthesis method and system

The invention provides a speech video synthesis method and system, and the method comprises the steps: inputting speech audio into a pre-trained audio-motion encoder, extracting a speech feature sequence from the speech audio, and generating a motion sequence of a face according to the speech feature sequence; inputting a single two-dimensional face picture shot from a first visual angle into a pre-trained picture encoder, and extracting three-dimensional identity features of an object contained in the face picture; inputting the three-dimensional identity features and the emotion labels into a pre-trained emotion mapping layer, and fusing to obtain emotion-containing three-dimensional identity features; the action sequence and the emotion-containing three-dimensional identity features are input into a video generation network based on a neural radiation field, a speech video of the object matched with the speech audio at a required second view angle is synthesized through the neural radiation field and camera parameters, and the second view angle can be adjusted within a preset range through the camera parameters.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Video synthesis method and system

The invention discloses a video synthesis method and system, and relates to the technical field of audio and video processing. A video synthesis system comprises a video source acquisition and processing module, a cross-modal semantic understanding module, an attention tensor generation module, a hierarchical progressive fusion module and a quality evaluation and optimization module. According to the method, the spatial-temporal joint features are extracted through the three-dimensional convolutional network, and the audio-visual cross-modal attention mechanism is constructed, so that the dynamic association strength of the audio event and the video content can be quantified, and the main body space mask can be generated, and therefore, the traditional isolated visual processing can be expanded into sound and picture semantic linkage understanding; in this way, deep guidance of multi-modal information on the synthesis process is achieved, and the synthesis effect of the video synthesis method and system is improved.
Owner:SUZHOU BROADCASTING SYST +1

Low-temperature synthesis equipment for reducing viscosity of slurry and synthesis method of low-temperature synthesis equipment

The invention provides low-temperature synthesis equipment for reducing slurry viscosity and a synthesis method thereof, and relates to the field of ceramic slurry synthesis, the low-temperature synthesis equipment comprises a synthesis assembly, the synthesis assembly comprises a driving mechanism and a dispersion mechanism, the effective use outer diameter of the dispersion mechanism can be freely regulated and controlled, so that the dispersion mechanism can adapt to dispersion cylinders with different inner diameters; the working position can be flexibly adjusted without replacing the dispersion disc, the gap between the dispersion disc and the barrel wall is reduced to eliminate edge stirring dead angles, and meanwhile, the two dispersion mechanisms rotate in opposite directions, so that the dispersion mechanisms are matched with different outer diameters, a complex three-dimensional flow field can be formed in the barrel, the slurry dispersion uniformity and the viscosity reduction efficiency are remarkably improved, and the dispersion efficiency is improved. The problems that a dispersing disc of a dispersing machine used for synthesizing ceramic slurry is of a fixed structure and can only adapt to a dispersing barrel with a certain specific inner diameter, the dispersing machine is mostly of a single-disc structure, dead angles are easily formed at the upper part and the bottom of a barrel body, and the synthesis effect is poor are solved.
Owner:杭州淮瓷科技有限公司

Voice synthesis method and device based on voice token fusion, equipment and medium

The invention relates to the technical field of speech semantics, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a speech synthesis method, device, equipment and medium based on speech token fusion. Performing text coding on the initial potential representation to obtain a target text feature; generating a semantic token corresponding to the initial text according to the target text feature, and performing time sequence alignment on the semantic token and the target text feature to obtain a target semantic token; obtaining user voice of a reference user, extracting timbre characteristics of the user voice, and generating a Mel spectrogram frame by frame according to the timbre characteristics and the target semantic token; and performing speech synthesis according to the Mel spectrogram to obtain target speech. The speech synthesis efficiency and quality can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech synthesis method, system and device, storage medium and program product

The invention provides a voice synthesis method, system and device, a storage medium and a program product, and relates to the technical field of artificial intelligence and voice processing, and the method comprises the steps: obtaining Mel spectrum data corresponding to to-be-synthesized voice text data; inputting the Mel spectrum data into a neural vocoder based on a selective state space model; and performing long sequence processing on the Mel spectrum data by using the selective state space model based on the neural vocoder to obtain synthesized audio data corresponding to the to-be-synthesized voice text data. According to the invention, the neural vocoder can be constructed based on the state space model for speech synthesis, the high-frequency reconstruction capability is improved, the loss of high-frequency details is avoided, and better synthetic tone quality is obtained.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Personalized tone migration and synthesis method based on virtual singer

The invention relates to a personalized timbre migration and synthesis method based on a virtual singer, and the method comprises the steps: obtaining a reference singing audio of a target virtual singer, and extracting a timbre identity benchmark feature used for representing the timbre identity stability through timbre coding; obtaining personalized timbre migration demand information for the virtual singer, wherein the demand information comprises a to-be-migrated timbre attribute and a corresponding target change amplitude or change direction; calculating a compatibility score according to the timbre identity reference feature and a migration demand, generating a timbre migration risk indication value, and comparing the risk indication value with a preset threshold value or a preset threshold value interval to determine that the migration is high-risk migration or low-risk migration or conservative-risk migration; according to the method, the tone identity benchmark features in the reference singing audio are extracted, and the tone identity invariant set with stable transpitch and sounding intensity is further constructed, so that the core tone identity of the virtual singer is accurately described.
Owner:CHANGSHA NORMAL UNIV

Description video synthesis method and device, equipment and medium

The invention discloses a commentary video synthesis method and device, equipment and a medium, and the method comprises the steps: obtaining a commentary content text of an original video, and generating a subtitle data set corresponding to the commentary content text and an audio data set of the subtitle data set, the subtitle data set comprises a plurality of subtitle display time periods arranged according to a time sequence and subtitle texts corresponding to the subtitle display time periods; determining a corresponding image display time period according to each subtitle display time period in the subtitle data set; corresponding to each image display time period, determining a matched image frame sub-sequence in a preset image material library according to a caption text in the corresponding caption display time period in the caption data set so as to construct a video image frame sequence; and synthesizing a corresponding target explanation video according to the subtitle data set, the audio data set and the video image frame sequence. According to the method, the subtitles, the audios and the corresponding image frame sequences accurately matched with the explanation content text are automatically generated, so that the production efficiency and quality of the explanation video are remarkably improved.
Owner:GUANGZHOU OVERSEAS KANGBAZI NETWORK TECHNOLOGY CO LTD

Method for training speech synthesis model, speech synthesis method, and electronic device

A method for training a speech synthesis model includes obtaining training data; obtaining an initial speech synthesis model; training a semantic encoding network and a semantic decoding network in the speech synthesis model respectively based on a style sample speech, a timbre sample speech, an input sample text, and an output sample speech in training samples of the training data, to obtain a trained speech synthesis model.
Owner:BAIDU INT TECH (SHENZHEN) CO LTD

Speech synthesis method, speech synthesis device, electronic equipment and storage medium

PendingCN120636366ASpeech recognitionSpeech synthesisSpeech segmentationSynthesis methods
The embodiment of the invention provides a speech synthesis method, a speech synthesis device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and is suitable for the field of financial science and technology and the field of digital medical treatment. The method comprises the steps of obtaining original voice, performing voice segmentation on the original voice to obtain candidate voice segments and speaker identifiers of the candidate voice segments, merging the candidate voice segments according to the speaker identifiers to obtain reference voice segments, screening the reference voice segments to obtain target voice segments, and performing voice recognition on the target voice segments to obtain voice recognition results. The method comprises the steps of obtaining a target text of a target voice segment, constructing a voice text pair according to the target voice segment and the target text, performing model updating on an original voice synthesis model according to the voice text pair to obtain a target voice synthesis model, and performing voice synthesis on a preset reference text through the target voice synthesis model, so that the voice synthesis quality can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Synchronous synthesis method and system for grounding current of station transformer

The invention belongs to the technical field of power protection, and particularly relates to a station transformer grounding current synchronous synthesis method and system, and the method comprises the steps: carrying out the grouping according to the number of a station transformer, synchronously collecting the grounding branch current of each station transformer to obtain an original signal, carrying out the double identification, and carrying out the homologous copying; generating two paths of independent signals of frequency detection and phase calibration; frequency detection signals are grouped according to double identifications, the independent actual frequency of each station transformer is obtained after frequency detection processing, and a same-frequency sine dynamic reference signal is generated based on the independent actual frequency; performing phase alignment and compensation on the phase calibration signal and the corresponding reference signal according to the double identifiers to obtain a calibrated signal; according to the method, the phase calibration reference of each station transformer is completely matched with the actual frequency of the station transformer, the reason of phase reference dislocation caused by frequency fluctuation is eliminated from the source, and the accuracy of synthesizing the total grounding current in the station is improved.
Owner:ZHENGZHOU UNIV

Audio synthesis method and device, medium and equipment

The embodiment of the invention provides an audio synthesis method and device, a medium and equipment, and relates to the technical field of speech synthesis. The method comprises the steps of obtaining a target text to be subjected to audio synthesis; splitting the target text into a plurality of text units with complete semantics to obtain a task sequence; for the ith text unit in the task sequence, scheduling influence factors are obtained before the audio synthesis task is executed, the scheduling influence factors comprise application layer semantic information and / or equipment real-time performance indexes, and i is a positive integer; determining a target synthetic link from a cloud synthetic link and an end-side synthetic link based on the scheduling influence factors; and completing audio synthesis of the ith text unit through the target synthesis link. According to the scheme provided by the embodiment of the invention, in the audio synthesis process of the target text, adaptive switching between the cloud audio synthesis link and the end-side audio synthesis link can be realized in a low-delay manner.
Owner:IFLYTEK CO LTD

High-precision residual current CT broken line detection and multi-channel synthesis method and system

The invention relates to a high-precision residual current CT disconnection detection and multi-channel synthesis method. The method comprises the following steps: reading an input value of a dial switch; confirming a multi-channel synthesis mode; confirming the number of actually used channels; sampling precision detection; the main control unit reads 16-channel voltage data in real time and records the 16-channel voltage data; carrying out broken line detection if the line is broken; performing multi-channel synthesis; performing effective value calculation to obtain a synthesized channel voltage effective value; performing effective value conversion; and storing the actual current value of the synthesis channel in a Flash storage area. The invention also discloses a high-precision residual current CT broken line detection and multi-channel synthesis system. The residual current sampling precision is improved, the current sampling precision is within the range of 0.5%, the broken line false alarm condition is reduced, accurate residual current data are provided for subsequent vector synthesis and data processing, and the residual current monitoring accuracy is improved; the device has a broken line monitoring function, can diagnose short line faults in time, has broken line detection-free time aiming at broken line false alarm, and has a broken line state recovery function to collect residual current again when a channel returns to normal.
Owner:ANHUI NANRUI JIYUAN POWER GRID TECH CO LTD

Speech synthesis method and related device

The invention provides a speech synthesis method and a related device, and relates to the technical field of speech synthesis. The speech synthesis method comprises the following steps: acquiring a first emotion feature of a target historical interaction speech; predicting a second emotion feature of a target voice to be generated according to the first emotion feature, a historical interaction text and a target text; wherein the historical interaction text is a text corresponding to the target historical interaction voice, the target text is a reply text generated based on an input voice in the latest round of voice interaction, and the target voice is a reply voice to be generated in the latest round of voice interaction; and generating the target voice according to the first emotion feature, the historical interaction text, the second emotion feature and the target text. According to the technical scheme provided by the invention, the problem that the emotional rhythm of the reply voice does not accord with the current context when the reply voice is generated based on the reply text in the prior art can be solved.
Owner:IFLYTEK CO LTD

Reverse synthesis method of carbon-based environmental functional material

The invention provides a reverse synthesis method of a carbon-based environmental functional material, and belongs to the technical field of computer material science. The method comprises the following steps: firstly, carrying out collaborative acquisition and joint data set construction of multi-source heterogeneous data of the carbon-based environmental functional material; building a reverse generation model of a carbon aerogel hole topological structure by means of a deep generative adversarial architecture, converting macroscopic performance indexes into high-dimensional constraint features by means of a performance coding network, and realizing directional generation of micro-pores by means of a self-adaptive modulation mechanism of a synthetic network. Meanwhile, an embedded physical attribute check and adversarial collaborative optimization strategy is implemented, a physical regression check head embedded in the authentication network is reused, and a self-supervised feedback closed loop without an agent model is constructed. And finally, completing process parameter inversion and intelligent decision recommendation based on multi-dimensional feature decoupling, and mapping the optimal microstructure features into an executable synthesis process formula through morphological quantitative verification and a process parameter inversion network.
Owner:QINGDAO UNIV OF TECH

Video synthesis method and device based on large language model, and medium

The invention relates to a video synthesis method and device based on a large language model and a medium, and the method comprises the steps: obtaining user input information which comprises a portrait image, reference voice and content reference information; performing multi-modal feature extraction on the input information to obtain multi-modal features including portrait image features, reference voice features and content reference information features; inputting the portrait image features into a digital human model constructed based on deep learning to generate an initial video; inputting the reference voice features and the content reference information features into a preset first large language model, and generating a video copywriting and split mirror information; performing matching in a local video material library based on the multi-modal features to obtain matched materials; inputting the split information and the matching material into a second large language model to generate a video content timetable; and performing video synthesis based on the video content timetable, the initial video and the video copywriting, and outputting a final synthesized video.
Owner:XIAMEN CHANGYUEYUAN INFORMATION TECHNOLOGY CO LTD

Speech synthesis method, diffusion model training method, device and equipment

The invention relates to the technical field of speech synthesis, is used for artificial intelligence scenes, financial business scenes and medical health business scenes, and provides a speech synthesis method, a diffusion model training method, a device and equipment, and the method comprises the steps: obtaining text information; carrying out coding processing on the text information based on a text coding sub-model of a diffusion model to obtain a vector sequence; based on an acoustic feature extraction sub-model, performing acoustic feature extraction processing on the vector sequence to obtain a first Mel spectrum; based on a context sensing sub-model, according to the vector sequence, performing text-voice alignment processing on the first Mel spectrum to obtain a second Mel spectrum; performing residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum based on a high-frequency compensation diffusion sub-model to obtain a target Mel spectrum; and based on the diffusion sub-model, according to the target Mel spectrum, determining the target voice so as to improve the voice synthesis effect and further promote the development of financial services and medical health services.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-dialect speech synthesis method and device, computer equipment and storage medium

The invention provides a multi-dialect speech synthesis method and device, computer equipment and a storage medium, and relates to the technical field of speech synthesis. According to the method, the input text is received, the language category corresponding to the input text is recognized, then the mapping rule corresponding to the language category is adopted to convert the input text into the standardized phoneme sequence, and it is ensured that phoneme expressions of different languages and dialects are consistent. The standardized phoneme sequence is input to the text coding layer, coding is performed through at least one dialect expert sub-network, learning and optimization are performed for specific features of different dialects, dialect feature vectors are obtained, and the naturalness and accuracy of speech synthesis are improved. The dialect feature vector is input into the acoustic decoder and converted into the corresponding voice waveform, and the multi-dialect synthetic voice is generated, so that the naturalness and accuracy of the dialect voice are ensured, and the multi-language adaptability of the voice synthesis model in the application scenes of financial consultation service, medical health inquiry and the like is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Direct digital synthesis method and system based on mixed fraction proportion

The invention discloses a direct digital synthesis method and system based on a mixed fraction proportion. The method comprises the following steps: determining an integer frequency control word, a numerator frequency control word and a denominator frequency control word based on a target frequency, a sampling frequency and a bit width of a phase accumulator; performing a fractional phase accumulation operation based on the sampling frequency, the molecular frequency control word, and the phase accumulator to generate a carry signal; performing integer phase carry and accumulation operation in parallel with the fractional phase accumulation operation according to the sampling frequency, the integer frequency control word and the phase accumulator to obtain an integer phase value; and acquiring a waveform digital amplitude corresponding to the integer phase value, and performing signal conversion according to the waveform digital amplitude to obtain an amplitude signal corresponding to the target frequency so as to improve the precision of the frequency synthesis signal.
Owner:CHENGDU WEIDE QINGYUN ELECTRONICS CO LTD

Speech synthesis method and device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to the fields of financial science and technology and medical health, and discloses a speech synthesis method, device, equipment and medium, and the method comprises the steps: obtaining original speech, and extracting high-dimensional features from the original speech to obtain high-dimensional speech features; inputting the high-dimensional language features into a pre-trained vector quantizer for discretization to obtain a plurality of discrete Tokens; a prediction Token sequence is generated through a TTS generator according to text information corresponding to the original voice and the multiple discrete Token, and the TTS generator is obtained by adopting a sample set to train and verify a large language model; and inputting the predicted Token sequence into a voice decoder for voice synthesis to obtain target voice. And the quality and the accuracy of the synthesized speech are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech synthesis method and device, electronic equipment and storage medium

The invention provides a speech synthesis method and device, electronic equipment and a storage medium, and relates to the technical field of speech synthesis, and the method introduces a target attribute text in a speech synthesis process, can support speech synthesis with an audio attribute corresponding to the target attribute text, and improves the speech synthesis efficiency. Therefore, the expressive force and rhythm of the target synthetic speech can be controlled according to the user demand, so that the target synthetic speech better meets the user demand, and the user experience is improved. A speech synthesis model is obtained through training of a text sample with an attribute tag, so that the speech synthesis model has an audio attribute control capability during speech synthesis, and audio attributes such as tone, speech style, emotional expression, human settings, mood, rhythm and the like of target synthesis speech can be controlled at the same time; and audio attributes such as language switching, environment sound effect, dialect generation and the like can be supported, and controllability is improved while the generation quality of the speech synthesis model is ensured.
Owner:IFLYTEK CO LTD