Voice generation method and device based on multi-modal fusion, equipment and medium
Through the speech generation method of multimodal fusion, the shortcomings of speech synthesis in the prior art in terms of timbre diversity, emotional expression accuracy, personalized customization capabilities and multimodal fusion are solved, and the generation of multimodal synthesis speech data is realized, which improves the speech expression ability and user experience.
Patent Information
- Application Number
- CN202510275169.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-27
AI Technical Summary
The existing speech synthesis technology has shortcomings in tone diversity, emotional expression accuracy, personalized customization capabilities and multimodal fusion, making it difficult to generate multimodal synthetic speech data that meets the needs of different fields.
The speech generation method based on multimodal fusion is adopted. By collecting audio data from the target field, extracting tone characteristics, and training the tone generation model of the domain feature; obtaining target text data, performing semantic analysis to identify emotional information, and adjusting speech synthesis parameters; obtaining personalized information, and constructing a personalized parameter mapping table; fusing these parameters and aligning them with the associated text annotations, visual element characteristics and background music rhythm data in time to generate multimodal synthesis speech data.
It has achieved the improvement of timbre diversity in speech synthesis, the improvement of emotional expression accuracy, the enhancement of personalized customization capabilities, and the synchronous presentation of multimodal information, which has enhanced the expression ability of speech and the immersion and personalized adaptability of information communication.
Smart Images

Figure CN120048243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a voice generation method, device, equipment and storage medium based on multimodal fusion. Background Art
[0002] Currently, voice synthesis technology is widely used in many fields such as education, medical care, finance and cultural dissemination. However, existing voice synthesis systems have significant deficiencies in terms of timbre diversity, emotional expression accuracy, personalized regulation ability and multimodal fusion, and it is difficult to meet the requirements of different application scenarios. Specifically, the existing technologies mainly have the following problems:
[0003] In the field of cultural dissemination, such as Buddhist voice synthesis, the timbre libraries equipped in the vast majority of products are extremely limited, mainly based on ordinary reading timbres. Users expect to hear timbres with characteristics such as compassion, solemnity, and peace. However, existing technologies are difficult to generate characteristic timbres that meet this requirement, resulting in the synthetic voice lacking appeal and being unable to meet the requirements of immersive experience. At the same time, the voice expression needs to reflect a sense of context. For example, chanting scriptures should have a solemn and respectful rhythm, historical explanations need to be vivid and fluctuating, and opera performances need to emphasize the singing rhythm. However, existing voice synthesis technologies cannot accurately simulate these contextual features, resulting in the generated voice lacking appeal and expressiveness. In addition, users with different beliefs and backgrounds have different preferences for religious voices and traditional cultural interpretations. For example, some users hope that the voice synthesis system can generate a chanting style with local characteristics, while some users hope to express it in modern standard language. However, existing voice synthesis systems cannot adaptively adjust parameters such as timbre, speech rate, and rhythm, resulting in a relatively single content presentation method and being difficult to meet personalized needs.
[0004] In the field of medical and health, the reading of medical texts requires a solemn, authoritative and soothing voice. However, existing voice synthesis technologies cannot provide matching timbres for different scenarios such as doctor-patient consultations, rehabilitation guidance, and psychological counseling, resulting in the synthetic voice lacking a sense of reality and professionalism and being difficult to establish doctor-patient trust. In addition, the communication between doctors and patients not only depends on the language content, but also needs to convey comfort, encouragement or warning information through the tone of voice. For example, in scenarios such as chronic disease management and psychological counseling, it is difficult for the voice synthesis system to generate an appropriate caring tone, which may lead to a rigid or cold patient experience and affect the communication effect. At the same time, due to the hearing loss of elderly patients, they need a slower speech rate and clearer pronunciation, while young users may prefer a faster speech rate and concise and clear expressions. However, existing voice synthesis technologies cannot adaptively adjust voice characteristics for different populations, resulting in understanding obstacles for some groups when using.
[0005] In the financial field, voice applications such as financial broadcasts, investment analysis, and risk warnings require steady, rational, and professional expressions. However, the voices synthesized by existing technologies often lack the intonation characteristics of the financial field, making it difficult to effectively convey market dynamics, analysis logic, and decision-making suggestions, thus affecting the user experience and information acceptance. At the same time, the demand for emotional expression varies greatly in different scenarios. For example, market condition broadcasts need to be neutral and stable, while financial training or investment strategy explanations require emotional expression to enhance persuasion. Existing voice synthesis technologies are difficult to automatically adjust emotional expression according to the text content, resulting in a single synthesized voice for financial information, unable to effectively convey important information or enhance the listeners' understanding. In addition, there are significant differences in the knowledge levels and needs of investors. For example, ordinary investors need more concise and easy-to-understand voice analysis, while professional investors are more concerned about data-driven in-depth interpretations. However, existing voice synthesis systems lack a user preference learning mechanism and are difficult to provide customized voice outputs for different investment groups, affecting the user experience. Summary of the Invention
[0006] The main objective of the present invention is to provide a voice generation method, device, equipment, and storage medium based on multimodal fusion, aiming to solve the technical problems that existing voice synthesis technologies have deficiencies in timbre diversity, emotional expression accuracy, personalized customization ability, and multimodal fusion, and are difficult to generate multimodal synthesized voice data that meets the requirements of different fields.
[0007] To achieve the above objective, the present invention provides a voice generation method based on multimodal fusion, including:
[0008] Collecting target field audio data and extracting timbre features from the target field audio data;
[0009] Training and generating a domain feature timbre generation model based on the timbre features;
[0010] Obtaining target text data and performing semantic parsing on the target text data to identify the emotional information of the target text data;
[0011] Adjusting the basic voice synthesis parameters according to the emotional information to generate an emotion adaptation parameter set;
[0012] Obtaining personalized information and constructing a personalized parameter mapping table based on the personalized information;
[0013] Fusing the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthesis control parameter sequence;
[0014] Aligning the synthesis control parameter sequence with the associated text annotations, target field visual element features, and background music rhythm data on the time axis to establish a parameter modal binding relationship table;
[0015] Drive the domain feature timbre generation model based on the parameter-modal binding relationship table to generate multi-modal synthesized speech data.
[0016] Furthermore, to achieve the above object, the present invention provides a speech generation device based on multi-modal fusion, including:
[0017] An audio data acquisition module, configured to acquire target domain audio data and extract timbre features from the target domain audio data;
[0018] A timbre training module, configured to train and generate a domain feature timbre generation model based on the timbre features;
[0019] A text semantic parsing module, configured to obtain target text data and perform semantic parsing on the target text data to identify the emotional information of the target text data;
[0020] An emotion parameter adjustment module, configured to adjust the basic speech synthesis parameters according to the emotional information to generate an emotion adaptation parameter set;
[0021] A personalized parameter modeling module, configured to obtain personalized information and construct a personalized parameter mapping table based on the personalized information;
[0022] A synthesis control parameter generation module, configured to fuse the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthesis control parameter sequence;
[0023] A multi-modal data synchronization module, configured to perform time-axis alignment on the synthesis control parameter sequence with associated text annotations, target domain visual element features, and background music rhythm data to establish a parameter-modal binding relationship table;
[0024] A multi-modal synthesized speech generation module, configured to drive the domain feature timbre generation model based on the parameter-modal binding relationship table to generate multi-modal synthesized speech data.
[0025] Furthermore, to achieve the above object, the present invention also provides a computer device, where the computer device includes a memory, a processor, and a speech generation program based on multi-modal fusion stored in the memory and executable on the processor. When the speech generation program based on multi-modal fusion is executed by the processor, the steps of the speech generation method based on multi-modal fusion as described above are implemented.
[0026] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, where a speech generation program based on multi-modal fusion is stored on the storage medium. When the speech generation program based on multi-modal fusion is executed by a processor, the steps of the speech generation method based on multi-modal fusion as described above are implemented.
[0027] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health, fintech, and cultural communication. It discloses a voice generation method based on multimodal fusion, including: collecting audio data in the target field and extracting timbre features; training a domain feature timbre generation model based on the timbre features; obtaining target text data and performing semantic parsing to identify emotional information; adjusting basic speech synthesis parameters to generate an emotion adaptation parameter set; obtaining personalized information and constructing a personalized parameter mapping table; fusing the emotion adaptation parameter set and the personalized parameter mapping table to generate a synthesis control parameter sequence; aligning the synthesis control parameter sequence with associated text annotations, visual element features in the target field, and background music rhythm data on the time axis to establish a parameter modality binding relationship table; driving the domain feature timbre generation model based on the parameter modality binding relationship table to generate multimodal synthesized speech data. By training the domain feature timbre generation model, the present invention enables the synthesized speech to have a domain-specific timbre, enhancing the diversity of timbres; combines semantic parsing and emotion recognition to adjust speech synthesis parameters to achieve emotion matching; constructs a parameter mapping table based on personalized information to generate personalized speech; fuses various modal data such as text, vision, and background music, and performs time axis alignment to synchronously present the content of each modality, enhancing the expressive ability of the speech and improving the immersion and personalized adaptability of information transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0029] Figure 1 is a schematic diagram of an application environment of a voice generation method based on multimodal fusion in an embodiment of the present invention;
[0030] Figure 2 is a schematic flowchart of an embodiment of the voice generation method based on multimodal fusion of the present invention;
[0031] Figure 3 is a schematic diagram of functional modules of a preferred embodiment of a voice generation device based on multimodal fusion of the present invention;
[0032] Figure 4 is a schematic diagram of a structure of a computer device in an embodiment of the present invention;
[0033] Figure 5 is another schematic diagram of a structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] The voice generation method based on multimodal fusion provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server through a network. The server can collect target domain audio data through the client, extract timbre features; train a domain feature timbre generation model based on the timbre features; obtain target text data, perform semantic parsing to identify emotional information; adjust basic speech synthesis parameters to generate an emotion adaptation parameter set; obtain personalized information, construct a personalized parameter mapping table; fuse the emotion adaptation parameter set and the personalized parameter mapping table to generate a synthesis control parameter sequence; align the synthesis control parameter sequence with associated text annotations, target domain visual element features, and background music rhythm data on the time axis to establish a parameter modality binding relationship table; drive the domain feature timbre generation model based on the parameter modality binding relationship table to generate multimodal synthesized voice data. The present invention trains a domain feature timbre generation model to make the synthesized voice have a domain-specific timbre, improving the diversity of timbres; combines semantic parsing and emotion recognition to adjust speech synthesis parameters to achieve emotion matching; constructs a parameter mapping table based on personalized information to generate personalized voices; fuses various modal data such as text, vision, and background music, and performs time axis alignment to synchronously present the content of each modality, enhancing the expression ability of the voice and improving the immersion and personalized adaptability of information transmission. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0036] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the voice generation method based on multimodal fusion provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0037] As Figure 2 shown, the voice generation method based on multimodal fusion proposed by the present invention includes the following steps:
[0038] S10, collect target domain audio data, and extract timbre features from the target domain audio data;
[0039] In this embodiment, target domain audio data is collected, and timbre features are extracted from the target domain audio data. Target domain audio data refers to audio materials with specific application scenarios, specific styles, or specific technical requirements, such as standard voices in the industry field, sound characteristics under specific cultural backgrounds, etc. In the process of extracting timbre features, multi-dimensional analysis of the audio is required, including frequency characteristics, time-domain characteristics, harmonic characteristics, etc.
[0040] In the field of Buddhism, target audio data usually includes types such as sermons by eminent monks, scripture recitations, ritual chantings, and Buddhist music performances. These audio data have unique speech rhythms and intonation characteristics. For example, the chanting voice of Tibetan Buddhism is steady and long, the Pali scripture recitation of Theravada Buddhism has unique intonations and scale structures, and the Buddhist chanting of Han Buddhism has melody. Therefore, when collecting audio data in the field of Buddhism, not only the speech content needs to be concerned, but also the compatibility between the cultural characteristics of the audio and religious rituals needs to be ensured.
[0041] The process of collecting audio data in the target field involves multiple links. First, the definition of the target field needs to be clarified, and appropriate audio data sources are selected according to the application scenario. For example, when constructing a speech synthesis system in the field of medical and health, audio such as medical conversations, doctor explanations, and health education can be selected; in the application of intelligent customer service in the financial field, audio data such as financial expert explanations and financial news broadcasts can be collected. These data can be sourced from public audio databases, professional recordings in specific fields, or on-site collection using high-precision recording equipment.
[0042] The collection methods of audio data can include online data crawling, real-time recording collection, database import, etc. Among them, online data crawling is suitable for obtaining open speech resources on the Internet, while real-time recording collection is suitable for extracting speech features in specific environments, such as recording doctors' diagnostic voices and bank customer service consultation voices. Database import is suitable for existing large-scale audio data sets, which are introduced into the system for subsequent analysis through format conversion, quality screening, etc.
[0043] The collected audio data needs to be preprocessed, including steps such as noise removal, signal enhancement, and volume normalization, to ensure the quality and consistency of the audio data. Noise removal can use methods such as adaptive filtering and spectral subtraction to remove background noise and improve the signal-to-noise ratio; signal enhancement can improve the clarity of the audio through methods such as automatic gain control and dynamic range compression; volume normalization can be based on the RMS energy normalization method to keep the volume of all audio samples consistent and avoid the influence of amplitude changes on subsequent feature extraction.
[0044] In the stage of timbre feature extraction, parameters that can characterize the timbre characteristics need to be obtained from the audio data. Common timbre features include fundamental frequency (F0), formant, harmonic-to-noise ratio (HNR), spectral envelope, etc. The fundamental frequency reflects pitch information and can be extracted through methods such as autoregressive analysis and zero-crossing rate calculation; the formant represents the resonance characteristics of pronunciation and can be extracted through linear predictive coding (LPC); the harmonic-to-noise ratio is used to measure the clarity of the audio and can be calculated based on time-frequency analysis methods.
[0045] In different application scenarios, the ways of data collection and timbre feature extraction can be adjusted according to actual needs. For example, in the field of Buddhism, highly sensitive microphones can be used to record Buddhist monks chanting scriptures in a temple environment, and timbre features can be extracted in combination with ambient sounds to enhance the immersion and atmosphere of the sound. In addition, audio separation technology can be applied to separate background sounds from the main audio track, making the timbre features of the chanting voice more pure. In the field of medical and health, automatic speech recognition (ASR) technology can be used to transcribe medical lectures in real time, and timbre modeling can be performed on the voices of different doctors to extract standard medical speech styles suitable for patient education. In the financial field, voice annotation tools can be used to perform hierarchical processing on financial broadcast audio to extract professional timbres suitable for applications such as market analysis and financial news interpretation.
[0046] The accuracy of timbre feature extraction can also be optimized under different environmental conditions. For example, in a high-noise environment, adaptive filtering technology can be used to reduce background noise interference and improve the purity of timbre features. In the processing of audio data with a low signal-to-noise ratio, short-time Fourier transform (STFT) or mel-frequency cepstral coefficients (MFCC) can be combined for feature enhancement to make timbre modeling more accurate. In addition, in order to meet the needs of different user groups, a multi-sample learning method can be adopted to train timbre models suitable for different accents and speaking speeds to ensure the generalization ability of timbre features.
[0047] Through multi-source data collection, the diversity and adaptability of timbre features are improved, ensuring that the voice synthesis system can generate timbre styles that meet application requirements in different fields. Through deep learning technology, the accuracy of timbre feature extraction is optimized, making the generated voice more in line with the expression habits of the target field. By combining various methods such as spectral analysis, formant analysis, and harmonic modeling, timbre features are made more expressive, achieving a high-fidelity voice synthesis effect. Through environment adaptation and data augmentation technology, the robustness of timbre features is enhanced, enabling the system to work stably in different noise environments.
[0048] S20. Based on the timbre features, train and generate a domain feature timbre generation model;
[0049] In this embodiment, a domain feature timbre generation model is trained based on timbre features. Timbre features are key parameters extracted from target domain audio data, including but not limited to fundamental frequency, formant distribution, harmonic structure, pitch fluctuation, speaking speed rhythm, etc. These features determine the unique style of speech, such as the solemnity and steadiness of Buddhist chanting, the clarity and authority of medical and health interpretation, and the professionalism and rigor of financial broadcasts.
[0050] To train a domain-specific feature timbre generation model, a high-quality training dataset needs to be established. This dataset should include a large amount of labeled audio data and cover different pronunciation styles, different timbre characteristics, and different background environments within the target domain. For example, in the Buddhist field, the dataset can include chanting audio of different sects and language families, and annotate its emotional color, speech rhythm, and rhythm characteristics. In the medical field, the dataset can include doctors' diagnostic voices, medical explanations, and gentle intonations suitable for patient education. In the financial field, the dataset can contain audio samples such as financial news broadcasts and market analyses, and annotate their timbre categories, such as the timbre of formal news broadcasts and in-depth market interpretation timbres.
[0051] Model training uses deep learning techniques and is based on model structures such as generative adversarial networks (GANs) or variational autoencoders (VAEs) to learn and generate speech data that conforms to the timbre characteristics of the target domain. The training process usually includes data preprocessing, feature encoding, model optimization, and other steps. First, in the data preprocessing stage, the audio samples are normalized, noise is removed, and the timbre characteristics are standardized to ensure that the model can learn stably. Second, in the feature encoding stage, convolutional neural networks (CNNs) or long short-term memory networks (LSTMs) are used to extract the high-dimensional features of the audio, and combined with the attention mechanism, so that the model can automatically focus on the key parts of the timbre characteristics. For example, when learning the timbre of Buddhist chanting, the attention mechanism will strengthen the low-frequency formants to enhance the thickness of the speech and make it more ritualistic and sacred. Finally, in the model optimization stage, through adversarial training or self-supervised learning methods, the generated timbre is made more natural and stable and can adapt to different target application scenarios.
[0052] In the applications of different fields, different technical strategies can be adopted for the training of the domain-specific feature timbre generation model. For example, in the Buddhist field, LSTM or Transformer architectures based on temporal modeling can be used to enable the model to learn the pitch changes and rhythm characteristics of chanting. In addition, environmental sound modeling techniques can be combined to incorporate background sounds such as temple bells and wooden fish sounds into the timbre generation process to enhance the immersion of the synthetic speech. In the medical and health field, a variational autoencoder model based on VAE can be adopted to learn the clarity, stability, and professionalism of doctors' voices, and adjust the model weights during the speech generation process to meet the auditory needs of different patient groups. In the financial field, an adversarial training strategy based on GAN can be adopted to make the generated speech more natural and avoid the stiffness and lack of emotion that are prone to occur in traditional TTS (text-to-speech) systems.
[0053] In addition, to enhance the controllability of the timbre, voice conversion technology (VC) can be combined to achieve speech synthesis in multiple styles within the same model framework. For example, in Buddhist speech synthesis, through the voice conversion module, the same synthesis model can generate timbres in different styles such as Tibetan Buddhist chanting, Han Buddhist sermonizing, and Theravada Buddhist Pali chanting. In the field of healthcare, within the same model framework, the speech rate and intonation can be adjusted to meet the requirements of different scenarios such as medical expert explanations and patient health consultations. In the financial field, the broadcasting style can be adjusted based on the habits of the target users, such as switching between fast-paced financial news broadcasts and in-depth financial interviews.
[0054] Example illustration: In the Buddhist field, the domain-specific timbre generation model trained can be used in scenarios such as chanting scriptures and Dharma assembly sermons. For example, when generating the timbre of Buddhist chanting, the low-frequency formants of the timbre can be enhanced to make the voice more solemn and heavy, so as to create the atmosphere of temple chanting. In the field of healthcare, this model can be used in scenarios such as popular science of medical knowledge and interpretation of doctor diagnoses. For example, when synthesizing the voice of medical experts, the stability and clarity of the timbre can be adjusted to make it easier for patients to understand medical terms. In the financial field, this model can be used in scenarios such as financial news broadcasts and interpretation of market data. For example, when synthesizing the voice of financial analysis, the clarity and rhythm of the speech can be enhanced to match the fast-paced information transmission requirements of the financial market.
[0055] By constructing a domain-specific timbre generation model, the timbre characteristics of the target domain can be highly restored during the speech synthesis process, making the synthesized speech more natural, fluent, and in line with industry requirements. Combining deep learning modeling and voice conversion technology can achieve timbre adjustments in multiple styles to meet the personalized needs of different application scenarios. By combining the attention mechanism and environmental sound modeling technology, the sense of context can be enhanced in the synthesized speech, making it more immersive. In addition, based on the training methods of generative adversarial networks or variational autoencoders, problems such as single timbre and lack of emotion in traditional TTS systems can be effectively avoided, improving the user's auditory experience and interaction satisfaction.
[0056] S30, obtain the target text data, and perform semantic parsing on the target text data to identify the emotional information of the target text data;
[0057] In this embodiment, the target text data is obtained and semantic parsing is performed to identify the emotional information of the text. The text data usually comes from domain-specific corpora, such as Buddhist scriptures, medical reports, financial analysis reports, etc. The obtained text needs to undergo data preprocessing, including noise removal, normalization, and structured annotation, for subsequent parsing work.
[0058] Semantic parsing is a process of deeply understanding text content based on natural language processing (NLP) technology, mainly including lexical analysis, syntactic analysis, semantic analysis, and sentiment recognition. First, lexical analysis is carried out to divide the target text data into basic language units such as words, phrases, or subwords, and perform part-of-speech tagging on them. Second, syntactic analysis is carried out. Using dependency syntax trees or syntactic constituent analysis methods, the syntactic structure of the text is parsed to understand the internal relationships of the sentence, such as the subject-predicate-object structure, modification relationships, etc. Then, semantic analysis is carried out. Combining pre-trained language models (such as BERT, GPT) and domain semantic knowledge bases, the word vector representation of the text is calculated to generate context-related semantic encoding vectors.
[0059] Based on the results of semantic analysis, a sentiment classification model is used to classify the text sentiment. Sentiment classification models usually adopt deep learning methods, such as bidirectional long short-term memory networks (BiLSTMs), convolutional neural networks (CNNs), or models based on the Transformer architecture. This model is trained with a large amount of labeled data and combines information such as sentiment words and syntactic structures to identify the sentiment type (such as solemn, merciful, serene, rigorous, professional, passionate, etc.) and sentiment intensity level expressed in the text.
[0060] To enhance the accuracy of sentiment recognition, a weighted statistical method can be combined to analyze the distribution of sentiment keywords in the text and quantify the sentiment intensity of the text. For example, in the field of Buddhism, if the text contains high-frequency words such as "mercy", "saving all sentient beings", "enlightenment", etc., it indicates that the overall sentiment of the text tends to be solemn and serene; in the field of healthcare, if the text involves words such as "recovery", "treatment plan", "efficient", etc., it indicates that the text may tend to be professionally rigorous in sentiment; in the financial field, if words such as "market volatility", "investment risk", "profit expectation" appear, it may indicate neutral or cautious sentiment characteristics.
[0061] Finally, the results of sentiment analysis are organized into structured sentiment information data, including sentiment categories (such as solemn, professional, rational), sentiment intensity levels (such as high, medium, low), context influence, etc., providing sentiment parameter inputs for subsequent speech synthesis.
[0062] In different fields, the technologies of semantic parsing and sentiment recognition can be adjusted according to actual needs. For example, in the field of Buddhism, the sentiment vocabulary list can be expanded based on the Buddhist semantic knowledge base, enabling the sentiment classification model to recognize the sentiment contained in specific religious terms. For example, the word "prajna" may not have an obvious sentiment in ordinary corpora, but in Buddhist texts, it represents wisdom and can be associated with "peace" and "rationality". By introducing large-scale training data of Buddhist canonical texts and combining language models such as BERT, the sentiment expressions in the Buddhist context can be parsed more accurately.
[0063] In the field of healthcare, semantic parsing can be combined with medical terminology libraries such as SNOMED CT and UMLS (Unified Medical Language System) to ensure that sentiment recognition can understand the professionalism of medical texts. For example, for a doctor's diagnosis report, the model needs to distinguish between objective descriptions (such as "the patient's blood sugar level has increased") and emotional expressions (such as "the condition needs to be closely monitored"), to avoid misjudging medical reports. At the same time, self-supervised learning methods can be used to train medical sentiment classification models to enhance the understanding ability of clinical texts.
[0064] In the financial field, semantic parsing can be combined with financial texts such as financial news and market analysis reports, and combined with time series data analysis to model the sentiment of the financial market. For example, when words such as "the market is warming up" and "the bull market is coming" appear in the news, the model can judge it as a positive sentiment; when the text involves "financial crisis" and "market crash" and other content, it can be marked as a negative sentiment. In addition, external data such as stock prices and social media sentiment indices can be combined to optimize text sentiment analysis in multiple dimensions.
[0065] Through semantic parsing and sentiment recognition of the target text data, the sentiment information of the text can be accurately extracted, and key sentiment parameters can be provided for subsequent speech synthesis. Combining deep learning technology can improve the accuracy of sentiment analysis and make the speech synthesis process more natural and vivid.
[0066] S40, adjust the basic speech synthesis parameters according to the sentiment information to generate a set of sentiment adaptation parameters;
[0067] In this embodiment, based on the sentiment information, the basic speech synthesis parameters are adjusted to generate a set of sentiment adaptation parameters. The basic speech synthesis parameters usually include core speech parameters such as fundamental frequency curve, speech rate interval, and spectral envelope shape. These parameters determine the pitch, rhythm, timbre and other characteristics of the synthesized speech. Adjusting these parameters enables the synthesized speech to more accurately express the emotional characteristics contained in the text.
[0068] First, obtain the basic speech synthesis parameters. These parameters come from a preset parameter library, which can be classified according to different languages, timbre styles and usage scenarios. For example, it includes Buddhist chanting timbre parameters, medical commentary timbre parameters, financial broadcast timbre parameters, etc. The basic speech parameters in the parameter library include:
[0069] Fundamental frequency curve (F0): Reflects the pitch fluctuations of speech and corresponds to the cadence of speech.
[0070] Speech rate interval: Refers to the pronunciation rhythm and pause time of speech, affecting the clarity and expression rhythm of speech.
[0071] Spectral envelope shape: Used to describe the formant characteristics of speech, affecting the brightness and resonance of timbre.
[0072] Then, adjust the basic speech synthesis parameters according to the emotional information. The adjustment method depends on the emotion category and the emotional intensity level:
[0073] Adjust the fundamental frequency curve based on the emotion category. Different emotion categories correspond to different pitch change patterns. For example, in the field of Buddhism, for texts expressing the emotion of "solemnity", the fundamental frequency curve may be adjusted to make the pitch relatively stable and reduce excessive fluctuations, while for texts expressing the emotion of "compassion", the fundamental frequency fluctuations may be slightly increased to make the speech more gentle; in the medical field, when expressing the emotion of "comfort", the fundamental frequency changes tend to be stable, avoiding abrupt pitch fluctuations to enhance the affinity; in the financial field, when expressing the emotion of "market analysis", the fundamental frequency is adjusted to have lower fluctuations to convey a sense of rationality and authority.
[0074] Adjust the speech rate interval based on the emotional intensity level. The higher the emotional intensity, the greater the variation in the speech rate. For example, when expressing "urgency", the speech rate can be increased, while when expressing the emotion of "steadiness", the speech rate can be decreased. In the context of Buddhist chanting, the speech rate of texts expressing the emotion of meditation and tranquility may be reduced, while texts expressing solemn preaching maintain a stable speech rate. In medical consultations, when doctors convey key information, the speech rate can be appropriately slowed down to ensure that patients can clearly understand. In the financial field, when reporting market dynamics, the speech rate can be increased to meet the information dissemination requirements of rapid transmission.
[0075] Adjust the spectral envelope shape to make the timbre fit the text emotion. By adjusting the formant distribution of the speech, optimize the brightness, softness or resonance of the timbre. For example, in the field of Buddhism, when the text expresses emotions such as "wisdom" and "serenity", the low-frequency components of the formants can be increased to make the speech sound more thick and steady; in the medical field, when doctors explain treatment plans, the formant frequency bands can be adjusted to make the speech clearer and more authoritative, while in the context of nursing consultations, the mid-frequency part can be enhanced to make the speech more affable. In the financial field, when reporting economic news, the formant distribution can be adjusted to make the timbre more penetrating to enhance the clarity of information transmission.
[0076] Finally, integrate the adjusted fundamental frequency curve, speech rate interval and spectral envelope shape to generate an emotion-adapted parameter set. This parameter set can be used for subsequent speech synthesis to ensure that the synthesized speech can accurately convey the emotional information of the target text.
[0077] By adjusting the basic speech synthesis parameters based on emotional information and generating an emotion-adapted parameter set, speech synthesis can more accurately express the emotional characteristics contained in the text. The speech parameters can be dynamically adjusted to make the synthesized speech more expressive and infectious, improving the user's immersive experience.
[0078] S50. Obtain personalized information and construct a personalized parameter mapping table based on the personalized information;
[0079] In this embodiment, a personalized parameter mapping table is constructed based on personalized information to meet the speech synthesis requirements of different users. The personalized information mainly includes user preference data, interaction behavior characteristics, historical usage patterns, etc., which can reflect the specific requirements of users during the speech synthesis process, such as speech rate, tone color, intonation, etc.
[0080] First, collect personalized information. Mainly through multimodal interaction methods such as a voice command recognition module, a gesture recognition module, and an eye movement tracking module, obtain user preference data. For example, the user can input "play slower" through voice, and the system will record this requirement; the user can manually adjust the tone color curve on the interface, and the system will record their adjustment preference; the eye movement tracking module can identify the changes in the user's focus when listening to the speech, so as to judge the user's acceptance of a specific speech style.
[0081] Second, perform data preprocessing on the personalized information. Since the preference data input by the user may be affected by noise, data cleaning and standardization are required. Noise suppression methods can be used to remove outliers, and through normalization processing, map data from different sources to a unified data range. For example, the tone color preferences of different users can be converted into quantifiable vector data through spectral analysis, so as to ensure the comparability of personalized information.
[0082] Then, perform statistical analysis on the preprocessed personalized information to extract user feature weights. A hidden Markov model can be used to model the user's interaction behavior and identify their long-term preferences. For example, if a user always selects a "deep" and "slow" speech style, the system will automatically assign higher weights to these parameters. In addition, based on the collaborative filtering module, match the parameter combinations of similar users from the personalized data of the existing user group to obtain the group template matching result, so as to enhance the accuracy of personalized parameter prediction.
[0083] Finally, construct a personalized parameter mapping table. Map the user feature weights to the basic speech synthesis parameters and store them using a multi-level data structure, including a static parameter layer (user's long-term preferences), a dynamic adjustment layer (parameters modified in real time), and an associated weight layer (weight distribution between parameters), ensuring that parameters at different levels can be dynamically adjusted to adapt to the changing personalized needs of users.
[0084] By constructing a personalized parameter mapping table based on personalized information, speech synthesis can be dynamically adjusted according to the specific needs of users, making the generated speech more in line with the user's auditory habits and emotional expectations. It can continuously optimize the speech synthesis quality and improve the user's immersion and satisfaction.
[0085] S60. Integrate the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthetic control parameter sequence;
[0086] In this embodiment, the emotion adaptation parameter set and the personalized parameter mapping table are integrated to generate a synthetic control parameter sequence, so that speech synthesis can not only accurately express the emotion information in the text, but also meet the auditory preferences and habits of individual users. This process involves multiple key technical links such as parameter matching, weighted calculation, and data format conversion to ensure that the integrated parameter set can accurately guide the speech synthesis model to generate personalized speech that meets user needs.
[0087] First, extract the key parameters from the emotion adaptation parameter set and the personalized parameter mapping table. The emotion adaptation parameter set includes the fundamental frequency curve, speech rate interval, and spectral envelope shape, which are used to adjust the emotional expression of speech; the personalized parameter mapping table includes personalized information such as user-preferred pitch, speech rate adjustment weight, and timbre style selection.
[0088] Then, perform parameter normalization processing. Since the emotion adaptation parameters and personalized parameters come from different sources, there may be problems with inconsistent numerical ranges. Therefore, normalization and standardization methods such as Min-Max normalization or Z-score normalization need to be used to convert all parameters to the same scale for subsequent calculations.
[0089] Next, calculate the fusion parameter weights. Based on the influence weights of different parameters in speech synthesis, the weighted average method or a deep learning model is used to calculate the final synthetic control parameters. For example, the adjustment of the fundamental frequency curve is affected by both the emotion category and the user's pitch preference, so a linear weighting formula or an attention mechanism can be used to dynamically calculate the final fundamental frequency regulation value.
[0090] Finally, convert the fused parameters into a synthetic control parameter sequence. This parameter sequence is stored in the form of a time-axis index, enabling it to adapt to the dynamically changing speech synthesis process. For example, during the reading of a long text, different paragraphs may correspond to different emotions and speech styles, so the parameter sequence should have a time dimension index to ensure that the speech synthesis model can accurately call the corresponding speech control parameters at different time points.
[0091] By integrating the emotion adaptation parameter set with the personalized parameter mapping table, a synthetic control parameter sequence is generated, enabling the speech synthesis system to precisely combine text emotion and user preferences and dynamically adjust speech expression. It can achieve advantages such as personalized adjustment, accurate emotion expression, and cross-scenario adaptation, making speech synthesis more natural, realistic, and enhancing the user's immersive experience.
[0092] S70. Align the synthetic control parameter sequence with the associated text annotation, the visual element features of the target domain, and the background music rhythm data along the time axis to establish a parameter-modal binding relationship table;
[0093] In this embodiment, the synthetic control parameter sequence is aligned with the associated text annotation, the visual element features of the target domain, and the background music rhythm data along the time axis to establish a parameter-modal binding relationship table, ensuring the synchronous matching of multi-modal information, enabling the speech synthesis process to dynamically adapt to the text, vision, and music rhythm, and achieving consistent multi-modal output.
[0094] First, extract the time information of each modal data. The synthetic control parameter sequence contains timestamp information to identify the action time point of each voice parameter. The associated text annotation is based on text time segmentation indexing, indicating the time position of the text content during the speech output. The visual element features of the target domain include key frame time indexes to determine the dynamic change moments of visual elements. The background music rhythm data is based on beat detection to identify the time positions of rhythm features such as music accents and beat changes.
[0095] Then, establish a time axis mapping relationship. Through the timestamp alignment algorithm, compare the time indexes of the synthetic control parameter sequence with the text time sequence, the visual time sequence, and the music rhythm time sequence to construct a time correspondence mapping. For example, use DTW (Dynamic Time Warping algorithm) to calculate the optimal matching path between each time sequence to ensure that different modal information can be synchronized. For example, during the speech synthesis process, if the voice parameter is adjusted to a low-frequency and steady timbre, the visual element may need to switch to a static image of a Buddha statue, and at the same time, the background music rhythm also needs to be reduced to match the overall atmosphere.
[0096] Next, integrate the multi-modal data to establish a parameter-modal binding relationship table. Convert the aligned time mapping relationship into a time sequence database, recording the voice parameters, text annotations, visual elements, and music rhythm features corresponding to each time index. Store it in a hierarchical structure, for example:
[0097] Static layer: Fixed text annotations and visual background information, such as the original text annotations of Buddhist scriptures, background images, etc.
[0098] Dynamic adjustment layer: Visual elements that change over time (such as Buddha statue animations, gesture dynamics) and music rhythm change points (such as beat changes, drum beat enhancements).
[0099] Weight layer: Weight distribution of different modal data to ensure the main expression direction during the speech synthesis process. For example, in the scene of chanting scriptures, the voice weight is the highest, while in the Dharma assembly scene, the visual dynamics may be dominant.
[0100] Finally, input the parameter-modal binding relationship table into the multimodal synthesis system, enabling it to dynamically load the corresponding text, visual, and music information according to the timeline index and apply it to the final speech synthesis process, so that the output speech, text, visual, and music are synchronized, achieving an immersive multimodal fusion experience.
[0101] In different application scenarios, the way of timeline alignment can be different.
[0102] In the Buddhist speech synthesis scenario, a method of matching the chanting rhythm can be adopted to synchronize the speech with the characteristics of the background Buddhist music, such as the drumbeats and bell sounds. For example, when chanting the Heart Sutra, the rhythm points of the background music match the pauses in the speech, making the entire chanting process more in line with the rhythm perception of the practitioners.
[0103] In the medical and health speech scenario, based on the speech broadcast rhythm of the doctor, the electronic medical record information or medical imaging data of the patient can be aligned on the timeline, so that when the doctor explains a certain disease, the patient can synchronously see the corresponding examination images or medical annotations on the screen, improving the understanding efficiency.
[0104] In the financial market interpretation scenario, combined with the changes in stock market data, the speech information of the financial news broadcast can be time-mapped with the stock K-line chart and trend change curve, and the tone and rhythm of the speech broadcast can be automatically adjusted at the key points of the market (such as the moment of sharp changes in rise and fall), making the financial information more intuitive and vivid.
[0105] Example illustration:
[0106] In the Buddhist field, when chanting Buddhist scriptures, the system can synchronously display the text content of the scriptures with the speech, and combine the Buddha statue animation and background Sanskrit sounds, enabling users to feel as if they are in a temple chanting scene, enhancing the immersion of cultivation. For example, when chanting the Great Compassion Mantra, the system can automatically switch the Buddha statue picture at the moment when "Avalokitesvara Bodhisattva" appears and adjust the music rhythm to make it more solemn.
[0107] In the medical and health field, when a doctor explains a patient's diagnosis report, the system can match the medical record data and imaging report in real time and synchronously display the explanations of relevant medical terms on the patient side to help the patient more intuitively understand their own health status. For example, when the doctor explains a certain liver ultrasound report, the system will automatically mark the relevant abnormal areas and adjust the rhythm of the speech broadcast to ensure that the patient can understand the key content.
[0108] In the financial field, during the market situation broadcast, the system can dynamically match relevant visual content such as K-line charts and trend analysis charts according to the speech broadcast rhythm of financial news, enabling users to simultaneously see the real-time changes in the stock market while listening to financial news, thereby improving the accuracy and experience of market interpretation. For example, when broadcasting the news "The Federal Reserve announces an interest rate cut", the system can synchronously display the change curve of the US dollar index and adjust the speech rhythm to make the information more impactful and expressive.
[0109] By constructing a parameter-modal binding relationship table, multi-modal temporal alignment of speech, text, vision, and music is achieved, enabling the synthesized speech to not only conform to the text semantics but also adapt to the visual and music rhythms, thereby significantly enhancing the integrated experience of hearing, vision, and context.
[0110] S80, based on the parameter-modal binding relationship table, drives the domain feature timbre generation model to generate multi-modal synthesized speech data.
[0111] In this embodiment, based on the parameter-modal binding relationship table, the domain feature timbre generation model is driven to generate multi-modal synthesized speech data, enabling the synthesized speech to achieve precise adjustment in terms of timbre, emotion, personalized features, and the matching of multi-modal information, enhancing the immersion and applicability of speech synthesis.
[0112] First, load the parameter-modal binding relationship table. In the previous step, the parameter-modal binding relationship table has been constructed, which stores the temporal alignment information between the synthesized control parameter sequence, associated text annotations, target domain visual element features, and background music rhythm data. When loading, the system parses this relationship table, extracts the speech control parameters, text annotations, visual features, and music rhythm information corresponding to each time point, and inputs these data into the domain feature timbre generation model.
[0113] Then, adjust the timbre generation parameters. According to the synthesized control parameter sequence in the parameter-modal binding relationship table, adjust the key parameters of speech synthesis:
[0114] Pitch control: Adjust the fundamental frequency to make the timbre characteristics of the speech match the timbre style of the target domain. For example, in the scene of Buddhist chanting, the fundamental frequency is relatively stable and has a resonance cavity effect to simulate the echo of temple chanting.
[0115] Speech rate control: Adjust the speech rate to make the synthesized speech conform to the semantic rhythm of the text content. For example, in paragraphs with strong emotional expressions, the speech rate is appropriately increased, while in content that needs to create a solemn atmosphere, the speech rate is appropriately slowed down.
[0116] Emotional expression: Combine the emotion adaptation parameter set to adjust the pitch, formants, and volume to ensure that the speech expression conforms to the emotional attributes of the text, such as different styles of compassion, solemnity, and profundity.
[0117] Personalized matching: Combine with the personalized parameter mapping table to make the speech synthesis result meet the needs of specific users, such as timbre preference, speech rate setting, pitch control, etc.
[0118] Next, synchronously generate multi-modal synthetic speech data. While generating the speech, the system automatically matches the text annotation, visual elements, and background music according to the parameter-modal binding relationship table to ensure the time synchronization of all modalities. For example:
[0119] When reciting Buddhist scriptures, the system can automatically display the scriptures, highlight important sentences in key paragraphs, and trigger visual animations, such as Buddha statue animations or candlelight changes, at appropriate time nodes.
[0120] In the medical explanation scenario, the system can display relevant pathological pictures or medical reports in real time according to the doctor's speech content to improve the patient's understanding.
[0121] In the financial information broadcast, the system can automatically switch the stock market K-line chart according to the market dynamics and adjust the chart animation according to the rhythm of the voice broadcast to make the data change more expressive.
[0122] Finally, output the multi-modal synthetic speech data. After the generated speech data is fused with the multi-modal information, a complete output data stream is formed, supporting multiple formats (such as MP3, WAV, AAC, etc.), and can be used for different terminal devices, such as mobile apps, smart speakers, AR / VR devices, etc., to achieve efficient information interaction and immersive experience.
[0123] By constructing a parameter-modal binding relationship table, multi-modal synchronization of speech synthesis with text, vision, and music is achieved, making the synthetic speech not only conform to the text content but also highly match the vision and sound effects, thus providing a more natural immersive experience. The timbre style, emotional expression, and personalized parameters can be adjusted according to different users and different scenarios to enhance the interaction effect, and it is widely applicable to fields such as Buddhist scripture recitation, medical explanation, and financial information broadcast.
[0124] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health, fintech, and cultural dissemination. It discloses a voice generation method based on multimodal fusion, including: collecting audio data in the target field and extracting timbre features; training a domain feature timbre generation model based on the timbre features; obtaining target text data and performing semantic parsing to identify emotional information; adjusting basic speech synthesis parameters to generate an emotion adaptation parameter set; obtaining personalized information and constructing a personalized parameter mapping table; fusing the emotion adaptation parameter set and the personalized parameter mapping table to generate a synthesis control parameter sequence; aligning the synthesis control parameter sequence with associated text annotations, visual element features in the target field, and background music rhythm data on the time axis to establish a parameter modality binding relationship table; driving the domain feature timbre generation model based on the parameter modality binding relationship table to generate multimodal synthesized speech data. By training the domain feature timbre generation model, the present invention enables the synthesized speech to have a domain-specific timbre, enhancing the diversity of timbres; combines semantic parsing and emotion recognition to adjust speech synthesis parameters to achieve emotion matching; constructs a parameter mapping table based on personalized information to generate personalized speech; fuses various modal data such as text, vision, and background music, and performs time axis alignment to synchronously present the content of each modality, enhancing the expressive ability of the speech and improving the immersion and personalized adaptability of information transmission.
[0125] In one embodiment, the above S10 includes:
[0126] S101, collecting speech samples in the target field including dialect voices and pronunciations of domain-specific terms;
[0127] S102, using an audio analysis tool to extract the fundamental frequency and formant distribution of the speech samples;
[0128] S103, collecting scene environmental sound effect data including mechanical operation sounds, natural background sounds, and indoor reverberation audio;
[0129] S104, performing spectral analysis on the environmental sound effect data to extract acoustic features from the environmental sound effect data;
[0130] S105, fusing the fundamental frequency, formant distribution, and acoustic features into timbre features.
[0131] In this embodiment, audio data in the target field is collected and timbre features are extracted therefrom to ensure that the speech synthesis result has domain features and the synthesized timbre better meets the style requirements of the application scenario. This process includes multiple technical key points, and each step involves different technical implementation methods.
[0132] First, collect voice samples in the target field. The core of timbre features comes from real audio data. It is necessary to collect voice data that conforms to the characteristics of the target field, including dialect voices, pronunciations of specialized terms, and timbre samples of different speakers.
[0133] In the application scenario of Buddhism, the voice samples mainly include the sermons of eminent monks, the chanting of believers, the recitation of mantras, etc., covering the voice styles of different sects such as Tibetan Buddhism, Theravada Buddhism, and Han Buddhism, so as to ensure that the synthesized voice can reproduce the unique timbre characteristics of Buddhist culture.
[0134] In the field of medical and health, the voice samples need to cover doctors' explanations, patients' descriptions of their conditions, pronunciations of medical terms, etc., to adapt to application scenarios such as medical lectures and disease popularization.
[0135] In the financial field, the timbre samples can include the voices of financial news announcers, covering content such as stocks, exchange rates, and market forecasts, to ensure that the synthesized voice meets the expression requirements of financial interpretation.
[0136] Then, use audio analysis tools to extract voice features. The collected audio data needs to be processed by professional audio analysis tools (such as Praat, MATLAB, Librosa) for feature extraction, including:
[0137] Fundamental frequency (F0): Determines the pitch range of the timbre and is used to analyze the natural vocal characteristics of speech. For example, the timbre of Buddhist chanting is usually lower and the fundamental frequency is relatively stable, while the audio of financial news broadcasts may have greater pitch fluctuations.
[0138] Formant distribution (Formants): Reflects the quality of speech and affects the clarity of the synthesized voice. For example, the formant distribution of Tibetan Buddhist chanting is relatively concentrated, with a strong sense of resonance, while the speech expressions in the medical field are usually clear and have weak low-frequency resonance.
[0139] Next, collect environmental sound effect data. Environmental sound effects are crucial for constructing a realistic timbre. The following types are mainly collected:
[0140] Mechanical operation sounds: Used to construct scenarios such as medical device explanations and industrial production voice synthesis, so that the synthesized voice conforms to the real environment.
[0141] Natural background sounds: Used to create an immersive experience. For example, in the field of Buddhism, collect temple bells, wind chimes, bird calls, etc.; in the medical and health scenario, it may involve the background sound of the consulting room, hospital broadcasts, etc.
[0142] Indoor reverberation: Collect the reverberation characteristics of different scenarios such as large halls, conference halls, and recording studios to make the voice synthesis more spatial. For example, Buddhist chanting is usually accompanied by echoes, while financial news broadcasts require a clean timbre without reverberation.
[0143] Subsequently, perform spectral analysis on the ambient sound effects to extract acoustic features. Use the Short-Time Fourier Transform (STFT) to analyze the audio data and extract the spectral energy distribution, enabling the system to distinguish between speech and ambient sound components. Use the Mel Frequency Cepstral Coefficients (MFCC) to analyze the timbre changes in different environments, ensuring that the timbre can adapt to different application scenarios. For example, in Buddhist speech synthesis, there are more low-frequency components in the ambient sound effects, and it is necessary to enhance the harmonic characteristics to improve the spatial sense of the timbre. Through impulse response modeling, simulate the acoustic characteristics of specific scenarios. For example, in the scenario of Buddhist chanting, reproduce the natural reverberation of the temple space.
[0144] Finally, fuse the timbre features to form training data. Integrate the fundamental frequency, formant distribution, and ambient sound effect features to form the final timbre feature data:
[0145] Adopt the method of feature vector splicing to convert various timbre parameters into a unified vector representation and input it into the neural network for training. Through dimension normalization, enable timbre data from different sources to be trained on the same scale, improving the generalization ability of the speech synthesis model. Combine speech enhancement technology to remove background noise and optimize the timbre, making the synthesized speech clearer and more natural.
[0146] This embodiment enables speech synthesis to possess real domain characteristics through multi-source data collection, professional timbre analysis, and ambient sound fusion. It can not only capture the unique timbre characteristics of the target domain but also simulate the timbre environment in specific scenarios, improving the immersion and realism of speech synthesis, and is applicable to multiple fields such as Buddhist chanting, medical interpretation, and financial news broadcasting.
[0147] In one embodiment, the above S30 includes:
[0148] S301, construct a domain semantic knowledge base pre-annotated with domain emotion words and the semantics of professional terms;
[0149] S302, obtain the target text data, and based on the annotation information in the domain semantic knowledge base, perform context semantic encoding on the target text data through a pre-trained language model to generate a semantic encoding vector;
[0150] S303, classify the semantic encoding vector through an emotion classification model to generate an emotion classification result;
[0151] S304, according to the emotion classification result and the distribution of the emotion keywords annotated in the domain semantic knowledge base in the target text data, analyze the emotion intensity level of the target text data through weighted statistical operations;
[0152] S305, generate emotion information metadata containing the emotion classification result and the emotion intensity level.
[0153] In this embodiment, target text data is obtained and semantically parsed to identify the emotional information in the text. This process needs to combine a semantic knowledge base, a deep learning model, and statistical analysis methods to ensure accurate extraction of the emotional features of the text.
[0154] First, a domain semantic knowledge base is constructed. This knowledge base contains emotional words, professional terms, and their semantic association relationships, enabling the system to accurately parse text emotions based on the language characteristics of a specific domain. For example:
[0155] In the field of Buddhism, the knowledge base needs to cover Buddhist-specific emotional words such as "compassion", "wisdom", "liberation", etc., and label their positive or negative emotional tendencies; in the field of medical and health, the emotional weights of words such as "ailment", "cure", "recovery" need to be labeled to accurately identify the emotional states in medical reports or patient condition descriptions; in the financial field, financial terms such as "market turmoil", "good news", "venture capital" need to be included and their emotional tendencies labeled to meet the emotional recognition requirements of financial news and market analysis texts.
[0156] Then, the target text data is obtained and semantically encoded. The text data can be sourced from user input, document parsing, database retrieval, etc., and semantic parsing is required after acquisition. A pre-trained language model (such as BERT, GPT, T5) is used to encode the text contextually to generate a semantic encoding vector, enabling the system to understand the overall semantics of the text. The pre-trained language model will combine the annotation information in the domain semantic knowledge base to improve the recognition ability for professional terms and specific expressions. For example:
[0157] In the field of Buddhism, identify the emotional dimensions (wisdom, cultivation) corresponding to "Six Paramitas".
[0158] In the field of medical and health, analyze the emotional tendency (positive) in "optimization of treatment plan".
[0159] In the financial field, analyze the emotion (negative) in "expectation of economic recession".
[0160] Next, the semantic encoding vector is classified through an emotion classification model. A deep learning emotion analysis model (such as LSTM, CNN+Attention, Transformer) is used to classify the text to generate an emotion classification result: the classification dimensions can include positive, neutral, negative, etc.; the sub-domains can be further divided, for example, the Buddhist field can be divided into enlightenment, solemnity, joy, piety, etc., the medical field can be divided into anxiety, hope, calmness, etc., and the financial field can be divided into optimism, prudence, panic, etc.
[0161] Then, based on the statistical analysis of sentiment keywords, calculate the sentiment intensity. Combine the distribution of sentiment keywords annotated in the domain semantic knowledge base in the text, and conduct weighted statistical analysis on the sentiment information. The calculation methods include:
[0162] Calculation of sentiment word weights: Calculate the overall sentiment intensity based on the density and weight of sentiment words in the text.
[0163] Sentiment co-occurrence relationship: Analyze the combination relationship of multiple sentiment words in the text and adjust the weights of the sentiment classification model.
[0164] Syntactic dependency analysis: Use NLP technology to parse the syntactic structure to ensure that sentiment judgments are based on complete semantic units rather than isolated words.
[0165] Finally, generate sentiment information metadata. Based on the above analysis, form structured sentiment information data, including: sentiment classification results (the main sentiment category expressed by the text); sentiment intensity level (such as a numerical sentiment distribution between 0 and 1); context influence factor (considering the adjustment of the sentiment tendency by the text context).
[0166] This metadata can be used for subsequent speech synthesis adjustment to make the generated speech more capable of expressing emotions.
[0167] In this embodiment, through the sentiment recognition technology based on semantic parsing, the sentiment information in the text is accurately obtained, enabling the speech synthesis to dynamically adjust the intonation, timbre, and speech rate during the expression process. It solves the problem of single emotion and lack of personalized expression in traditional speech synthesis, making the synthesized speech more in line with the emotional needs in specific scenarios.
[0168] In one embodiment, the above S40 includes:
[0169] S401, obtain basic speech synthesis parameters from a preset parameter library, where the basic speech synthesis parameters include fundamental frequency curve, speech rate interval, and spectral envelope shape;
[0170] S402, adjust the fundamental frequency curve in the basic speech synthesis parameters according to the sentiment category in the sentiment information;
[0171] S403, adjust the speech rate interval in the basic speech synthesis parameters according to the sentiment intensity level in the sentiment information;
[0172] S404, adjust the spectral envelope shape in the basic speech synthesis parameters according to the sentiment information;
[0173] S405, integrate the adjusted fundamental frequency curve, the adjusted speech rate interval, and the adjusted spectral envelope shape to generate the emotion adaptation parameter set.
[0174] In this embodiment, the process of adjusting the basic speech synthesis parameters to match the emotional information and generating an emotion-adaptive parameter set involves dynamically adjusting multiple core parameters of speech synthesis, making the synthesized speech more in line with the emotional expression of the text.
[0175] First, obtain the basic speech synthesis parameters from the preset parameter library. The speech synthesis system usually contains a set of basic speech synthesis parameter libraries, which store the audio parameters defaultly used by different speech synthesis models, including:
[0176] Fundamental frequency curve (F0 contour): Determines the pitch fluctuations of speech and affects the naturalness and emotional expression of speech. For example, a high fundamental frequency is usually associated with emotions such as excitement and agitation, while a low fundamental frequency usually corresponds to calm and steady expressions.
[0177] Speech rate and pause duration: Defines the rhythm of speech, including pauses between sentences, speech rate, etc. Adjusting the speech rate and pause duration can affect the expressiveness of speech. For example, slow speech is suitable for solemn or contemplative situations, while fast speech can be used for urgent or passionate expressions.
[0178] Spectral envelope shape: Affects the timbre characteristics of speech, adjusts formants and spectral characteristics, and can be used to simulate different speaking styles. For example, a relatively sharp spectral envelope can express anger, while a relatively rounded spectral envelope can convey gentle and serene emotions.
[0179] Then, adjust the fundamental frequency curve according to the emotion category in the emotional information. The emotion category (such as joy, sadness, anger, etc.) determines the change pattern of the fundamental frequency:
[0180] Pleasant / cheerful: The fundamental frequency range shifts upward, the pitch is more variable, and the speech appears lively and vivid.
[0181] Sad / low: The fundamental frequency decreases, the pitch is more monotonous, and the speech is softer and slower.
[0182] Angry / agitated: The fundamental frequency changes violently, the speech fluctuates greatly, showing tension or strong emotions.
[0183] Solemn / somber (suitable for Buddhist chanting): The fundamental frequency is stable, avoiding excessive pitch changes, making the speech more stable.
[0184] Next, adjust the speech rate and pause duration according to the emotion intensity level in the emotional information. The emotion intensity level determines the speed and pauses of the speech:
[0185] Strong emotions (such as anger, excitement): The speech rate speeds up, the pauses decrease, and the language rhythm is compact.
[0186] Weak emotions (such as calmness, contemplation): The speech rate slows down and the pauses increase, showing thinking and tranquility.
[0187] For example, for classic texts such as "Heart Sutra" and "Diamond Sutra", the speech rate can be appropriately slowed down and the pauses between paragraphs can be increased to give the audience sufficient time for understanding and contemplation. When the market fluctuates (such as a stock market crash), the speech rate can be appropriately increased to emphasize urgency; while when analyzing a stable market environment, the speech rate can be appropriately slowed down to make the information transmission clearer.
[0188] Then, adjust the spectral envelope shape according to the emotional information. The spectral envelope controls the formant characteristics of speech, and different adjustment methods can affect the emotional expression of speech:
[0189] Enhancement of high-frequency components: It can make the speech brighter and sharper, suitable for expressions of excitement and anger.
[0190] Enhancement of low-frequency components: It can make the speech more steady and heavy, suitable for solemn and majestic emotions.
[0191] For example, when simulating a monk chanting scriptures, the high-frequency energy can be reduced and the middle and low-frequency parts can be enhanced to make the voice more mellow and enhance the sense of immersion.
[0192] Finally, integrate the adjusted fundamental frequency curve, speech rate interval, and spectral envelope shape to generate an emotion-adaptive parameter set. This parameter set contains all the adjusted speech synthesis parameters and can be used for subsequent speech generation, enabling the synthesized speech to dynamically match the emotional expression of the text and improving the naturalness and expressiveness of the synthesized speech.
[0193] In this embodiment, by dynamically adjusting the speech synthesis parameters, the synthesized speech can accurately match the emotional expression of the text, improving the naturalness and expressiveness of speech synthesis. It can perform fine-grained parameter adjustment for different emotional categories and intensities, making the speech expression more realistic and credible, and is widely applicable to multiple fields such as Buddhism, healthcare, and finance to meet the personalized needs of different users for speech styles.
[0194] In one embodiment, the above S50 includes:
[0195] S501, collecting user preference data through a three-modal interaction interface composed of a voice command recognition module, a gesture recognition module, and an eye movement tracking module;
[0196] S502, performing noise suppression processing on the user preference data and generating a standardized preference vector using a standardization method;
[0197] S503. Extract the user operation behavior feature sequence from the standardized preference vector through a hidden Markov model, and match parameter combinations from the group preference template library based on the collaborative filtering module to obtain the group template matching result;
[0198] S504. Input the user operation behavior feature sequence and the group template matching result into a multi-level decision tree model, and apply cultural adaptation rules at the decision nodes to generate a parameter weight vector;
[0199] S505. Map the parameter weight vector and the basic speech synthesis parameters, and encode the mapping result into a hierarchical data structure including a static parameter layer, a dynamic adjustment layer, and an associated weight layer to generate the personalized parameter mapping table.
[0200] In this embodiment, the personalized speech synthesis system needs to adjust the style, emotional characteristics, and speech parameters of the synthesized speech according to the user's usage habits, behavior patterns, and preferences. To achieve this goal, the system uses a speech command recognition module, a gesture recognition module, and an eye movement tracking module to build a three-modal interaction interface to collect user preference data.
[0201] Speech command recognition module: The user can put forward personalized requirements for speech synthesis parameters through speech commands, such as "slower speed", "gentler intonation", or "change the voice color to male voice". The system uses automatic speech recognition (ASR) technology to recognize the user's speech commands and extract the key information therein.
[0202] Gesture recognition module: Use computer vision technology to recognize the user's gestures. For example, adjust the speech speed through a sliding gesture, or select a specific voice color through a tapping action. This method is applicable to the personalized adjustment of speech synthesis in intelligent devices (such as smart speakers, VR / AR devices).
[0203] Eye movement tracking module: Analyze the user's focus on the interaction interface through eye movement tracking technology to infer the user's preferences. For example, in the interface of multiple speech synthesis options, the user's fixation duration can reflect the interest in specific voice colors or speed settings, and thus be used to optimize parameter recommendations.
[0204] The collected original user preference data may contain noise or abnormal data, so noise suppression and normalization are required before use. Voice command data may be affected by environmental noise, gesture data may contain accidental touches, and eye movement data may be abnormal due to user blinking or line-of-sight deviation. The system uses signal processing methods (such as band-pass filtering, statistical denoising, etc.) to clean the data and remove invalid information. Since there may be significant differences in the interaction methods of different users, it is necessary to normalize the data so that it can be processed within a unified numerical range. For example, map the speech rate adjustment range to the interval [-1,1] to make the personalized preference data comparable and universal.
[0205] The normalized data is input into the Hidden Markov Model (HMM) for analyzing user behavior characteristics.
[0206] Hidden Markov Model (HMM): This model can capture the dynamic change patterns of user preferences, such as: the voice preferences of users in different scenarios (e.g., prefer slow reading when studying, and prefer fast reading when listening to the news). The continuous adjustment patterns of users for different adjustment options (e.g., users habitually adjust the timbre first and then the speech rate). The stability of user operations (e.g., whether the speech rate adjustment of a certain user tends to a specific value).
[0207] The user operation behavior feature sequence extracted based on the HMM is used to further match with the group template, so as to recommend parameter configurations that better meet the user's needs.
[0208] To improve the accuracy of recommendations, the system adopts a Collaborative Filtering (CF) module to compare the operation behavior of a single user with historical user data and find the parameter configurations of similar users.
[0209] The collaborative filtering matching methods can include:
[0210] User-based collaborative filtering: Find users with similar voice preferences and recommend the voice parameter settings of these users. For example, if multiple users choose a low timbre and slow reading in the Buddhist chanting scene, then similar parameter configurations will be preferentially recommended to new users in this scene.
[0211] Content-based collaborative filtering: Analyze the historical operations of the user himself and recommend the commonly used parameter configurations. For example, if a user has habitually adjusted the timbre to a deep and mellow style in the past, the system will preferentially match this type of timbre.
[0212] The matching results are used to generate the group template matching results, which serve as the input for the subsequent decision tree model.
[0213] To generate the final personalized parameters, the system uses a multi-level decision tree model for optimization and applies cultural adaptation rules during the decision-making process.
[0214] Multi-level decision tree model: This model calculates the personalized parameter weights suitable for the user based on the user's personalized preferences, group template matching results, and speech synthesis requirements. For example, if the user has a long-term preference for slow speech, the decision tree will gradually increase the slow weight and reduce the recommendation probability of high-speed speech. If the user uses speech synthesis in the field of Buddhism, the decision tree will prefer a timbre style with enhanced mid-low frequencies and stable fundamental frequency changes.
[0215] Cultural adaptation rules ensure that the personalized speech synthesis parameters conform to the user's cultural background. For example, in the field of Buddhism, the speech style should be more stable and solemn, rather than exaggerated and drastically changing. In the medical field, the speech should be more amiable and slow, rather than fast and high-pressure broadcasting.
[0216] Finally, the decision model generates a parameter weight vector and maps it to the basic speech synthesis parameters to form a personalized parameter mapping table. This mapping table uses a hierarchical data structure, including:
[0217] Static parameter layer: Stores the user's long-term preferences, such as the default values of timbre and speech rate.
[0218] Dynamic adjustment layer: Allows the user to make fine-tuning during real-time interaction.
[0219] Associated weight layer: Stores the mutual influence relationships between different parameters, such as how emotional parameters affect speech rate, timbre, etc.
[0220] In this embodiment, through multi-modal data collection, feature extraction, group template matching, and decision tree optimization, a personalized parameter mapping table is constructed, making the speech synthesis more in line with the user's needs. It can dynamically adapt to the speech preferences of different users, achieve more accurate personalized speech synthesis, and improve the user experience.
[0221] In one embodiment, the above S70 includes:
[0222] S701, processing the target text data through a predefined annotation module to generate associated text annotations, and processing the associated text annotations according to a preset segmentation strategy to generate a text time series;
[0223] S702, extracting target domain visual element features from the target domain image, and processing the target domain visual element features through a key frame detection module to generate a visual time series;
[0224] S703 extracts the background music rhythm data from the background music audio data, and processes the background music rhythm data through a beat detection module to generate a music rhythm time series;
[0225] S704 parses the timestamp information included in the synthesis control parameter sequence, and performs a comparison process on the timestamp information with the time information recorded in the text time series, visual time series, and music rhythm time series to generate a modal time correspondence mapping;
[0226] S705, based on a unified time reference, combines the modal time correspondence mapping, and assigns corresponding time identifiers to the data items in the synthesis control parameter sequence, associated text annotations, target domain visual element features, and background music rhythm data, and combines them to form a parameter-modal binding relationship table with unified time tags.
[0227] In this embodiment, in order to ensure the precise alignment of the content of the speech synthesis with the corresponding text annotations, it is necessary to first process the target text data to form a text time series.
[0228] The predefined annotation module is used to analyze the target text data, extract keywords, syntactic structures, prosody information, etc., and add semantic tags to the text. For example: in Buddhist applications, pause points and stress information for chanting scriptures such as the Heart Sutra and Diamond Sutra can be added. In medical and health applications, key medical terms in patient medical records or medical reports can be specially annotated. In the financial field, market terms and key data points can be annotated to make the speech synthesis emphasize market movement information.
[0229] The text data processed by the predefined annotation module needs to be segmented according to the time axis. Different text contents apply different segmentation strategies:
[0230] Segment by natural pauses: For example, in Buddhist chanting voices, it is divided according to syntactic structures and natural pause points to avoid syllable breaks affecting the listening experience.
[0231] Segment by punctuation marks: For news broadcasts and financial voice reports, they can be segmented according to punctuation marks to ensure logical coherence.
[0232] Segment by content hierarchy: In medical and health data, it can be segmented according to structured contents such as diagnoses, treatments, and suggestions in medical records, so that patients can listen to the information more hierarchically.
[0233] The generated text time series contains the start time and end time of each text segment, providing time identifiers for subsequent modal alignment.
[0234] The visual elements in the target domain image need to be synchronized with the speech content. For example, in Buddhist applications, key visual elements such as Buddha statues, instruments, and scriptures need to be synchronized with the explanation. In medical applications, CT images and medical imaging data need to be synchronized with the audio of the explanation. In the financial field, financial data such as K-line charts and trend analysis charts can be matched with speech synthesis.
[0235] Use computer vision methods to analyze the target domain image and extract key visual element features. For example, in the field of Buddhism, extract the features of elements such as Buddha statues, hand seals, and instruments. In medical imaging, extract the lesion area and medical scan level information. In financial data visualization, extract the key turning points of the candlestick chart and abnormal trading volume points.
[0236] Use inter-frame change detection algorithms (such as SIFT, SURF, and optical flow analysis) to detect key visual frames: in Buddhist teaching videos, identify changes in gestures and the turning of pages in scriptures. In medical data presentations, locate key change points in X-ray and MRI images. In financial data presentations, detect moments of drastic market fluctuations.
[0237] The generated visual time series contains the time points at which each visual element changes, ensuring a match with the speech synthesis.
[0238] In multimodal speech synthesis systems, the rhythm information of background music has an important influence on the rhythm control of speech. For example, in Buddhist applications, the background music for chanting scriptures usually contains bells and wooden fish sounds, and its rhythm needs to match the speech synthesis. In medical applications, soothing music can be used for meditation guidance, and the volume and rhythm need to be adjusted synchronously with the speech. In financial data broadcasts, background music with a strong rhythm can be used to enhance the sense of urgency of market analysis.
[0239] Music signal processing technology is used to extract the rhythm characteristics of background music: the spectrum is analyzed through Fourier transform to obtain the main frequency of the rhythm. The volume fluctuation is detected through the autoregressive model to identify the changes in music sections.
[0240] Use the Beat Tracking algorithm to obtain the beat time of background music: In Buddhist chanting applications, detect the time when the bell and wooden fish sound appear. In medical relaxation music, identify the beat of the music and slow down or speed up in sync with the voice. In financial data visualization, analyze the rhythm changes of background music to make information broadcasting more dynamic.
[0241] The generated music rhythm time series records the rhythm time points of the background music so that it can be synchronized with the speech synthesis content.
[0242] The synthetic control parameter sequence contains multiple timestamps, each of which indicates a key time point for speech synthesis. After parsing these timestamps, they are compared and processed with the time information in the text time series, visual time series, and music rhythm time series.
[0243] Using time alignment algorithms (such as DTW, HMM), calculate the time mapping relationship between each modal time series and the synthetic control parameter sequence. For example: in Buddhist applications, ensure that the bell sounds, gesture changes are synchronized with the chanting of scriptures. In medical applications, ensure that the page turning of MRI images is synchronized with the doctor's interpretation audio. In financial data broadcasts, ensure that the trend changes of the K-line chart are synchronized with the broadcast voice.
[0244] Finally, output the modal time corresponding mapping for time axis alignment.
[0245] Using linear time alignment (LTA) or non-linear time alignment (NLTA), unify the time reference for all modal data. For example: in Buddhist applications, adjust the bell sounds, wooden fish sounds, and voice to be synchronized at the same time point. In medical image interpretation, ensure that the audio content explained by the doctor is synchronized with the image page turning. In financial market broadcasts, adjust the audio and data charts to have consistent timelines.
[0246] Record the time tags of all modal data: voice time points; text time points; visual time points; background music time points.
[0247] Form a multi-modal data binding relationship table with unified time tags to ensure the synchronization of all modal information.
[0248] In this embodiment, through the multi-modal time alignment method, the synchronous mapping of multi-modal information such as voice, text, vision, and background music is realized. It can ensure that all modal information works together under a unified time axis, enabling users to obtain a more real and immersive experience when using, improving the usability and adaptability of speech synthesis, and being widely applicable to fields such as Buddhism, medical health, and finance.
[0249] In one embodiment, the above S80 includes:
[0250] S801, input the parameter modal binding relationship table into the domain feature timbre generation model;
[0251] S802, generate multi-modal synthetic speech data for synchronously displaying text annotations, visual elements of the target domain, and background music through the domain feature timbre generation model;
[0252] S803, according to the application scenario requirements, perform format conversion on the multi-modal synthetic speech data to generate target data suitable for the target terminal device.
[0253] In this embodiment, in a multimodal speech synthesis system, the parameter modal binding relationship table is used to provide the time axis alignment information required for the timbre generation process to ensure the synchronization of speech, text, visual elements and background music. When the relationship table is input into the domain feature timbre generation model, the following processing steps are required:
[0254] Read the time tags, speech synthesis parameters, text annotation information, visual element features, and background music rhythm information in the table. Analyze the time synchronization mapping relationship of different modes to ensure that the timbre model can trigger the corresponding speech synthesis operation at the appropriate time point. In Buddhist application scenarios, this relationship table can be used to ensure that the sound of bells, chanting rhythms, and visually displayed Buddha statues or instruments are consistent with the synthesized speech. In medical applications, this relationship table can be used to time match the doctor's interpretation of audio, medical imaging data, and background description text. In financial data broadcasting, ensure that the key market changes reported by voice are consistent with the chart data and sound effects.
[0255] Combined with the timbre feature data and the emotion adaptation parameters, the speech synthesis parameters are adjusted to make the synthesized timbre conform to the phonological characteristics of the target field. For example, in Buddhist applications, the timbre characteristics can be dynamically adjusted according to the chanting styles of different sects (Tibetan, Han, Southern), so that the generated speech can accurately restore the rhythmic characteristics of a specific sect. In medical applications, the clarity, speed, and pauses of the doctor's voice can be adjusted so that patients can more easily understand the medical interpretation. In financial market broadcasts, the tone and broadcast speed can be adjusted to adapt to different market sentiments, such as stable broadcasts or emergency market fluctuations.
[0256] The domain-specific timbre generation model is used to synthesize speech data with multimodal synchronization capabilities, and ensure that text, visual elements, and background music are synchronized with it.
[0257] Deep learning models (such as Tacotron and FastSpeech) are used in combination with domain-specific timbre data to train synthetic timbre, so that it has the proprietary phonological characteristics of the target domain. For example: in Buddhist applications, fundamental frequency regulation and resonance peak optimization technology are used to enable the synthesized chanting voice to accurately simulate the timbre characteristics of different monks, such as powerful and deep or soft and broad. In medical applications, emotional embedding technology is used to ensure that doctors' voice expressions are accurate and natural, so that patients can understand medical reports more clearly. In financial market broadcasts, rhythm adjustment algorithms are used to make the broadcast voice emphasize key information points when the market fluctuates violently, thereby enhancing the effect of information transmission.
[0258] During speech synthesis, it is necessary to display text annotation data synchronously according to the timestamp information: In Buddhist applications, text synchronization can be used to display scriptures and provide annotations, such as the Sanskrit, Pinyin, and modern Chinese interpretations of the Heart Sutra. In medical applications, text synchronization can be used to assist doctors in interpreting medical images. For example, lesion marks in MRI films can be automatically highlighted as the doctor explains. In the financial field, data indicators, stock price changes, etc. can be displayed synchronously on the screen when market news is broadcast.
[0259] Computer vision technology is used to synchronize visual elements with speech synthesis. For example, in Buddhist applications, temple scenes, scripture calligraphy displays, or image animations can be used to present Buddhist stories. In medical image interpretation applications, speech interpretation can synchronously trigger dynamic annotation of medical images. For example, when a doctor explains a lung X-ray, the corresponding area in the image is highlighted. In financial data broadcasts, K-line charts and index trend charts are dynamically updated with the broadcast voice.
[0260] Using background track mixing technology, the background music can be adjusted according to the rhythm of the speech to match the auditory experience: in Buddhist applications, the timing of the background bell and wooden fish sounds can be adjusted according to the rhythm of chanting. In medical applications, the background soothing music can be adjusted to match the rhythm of the interpretation to improve the patient's auditory comfort. In the financial field, the rhythm of the background music can be adjusted according to market sentiment, such as reducing the speed of the music when the market plummets to create a stable atmosphere.
[0261] Different terminal devices have different format requirements for audio, visual, and text data, so the generated multimodal synthesized speech data needs to be adapted and converted.
[0262] Adapt to different platform requirements: Generate MP3, AAC, WAV and other formats, suitable for mobile devices and streaming platforms. Use low bit rate encoding (such as Opus) to optimize file size, suitable for low bandwidth environments. Use high dynamic range encoding (such as FLAC) for high-quality voice storage, such as Buddhist voice archives.
[0263] Adapt to different terminals, such as: in mobile applications, use JSON or XML format to store text annotations for front-end rendering of synchronous subtitles. In smart speaker devices, use embedded TTS data format to achieve simultaneous voice and text broadcasting. In online education platforms, generate PDF or HTML format annotation documents for user convenience.
[0264] In addition, visual element adaptation can also be achieved:
[0265] In Buddhist applications: Use vector format to store Buddha images to ensure high-definition display across devices. Use WebP format to optimize image loading speed and is suitable for web page display.
[0266] In medical applications: Store medical images in DICOM format to ensure compatibility with hospital systems. Compress medical image videos using H.265 encoding to improve storage efficiency.
[0267] In financial applications: Store charts in SVG vector format to ensure clear data display.
[0268] Cross-platform optimization can also be achieved: Deployed in the cloud, WebRTC or WebSocket can be used to transmit multimodal data, supporting real-time interaction. On edge devices (such as smart speakers), use offline speech synthesis technology to reduce network dependence and improve response speed.
[0269] In this embodiment, the domain feature timbre generation model is driven by the parameter modality binding relationship table to generate multimodal synthesized speech data. It is possible to establish a time synchronization mapping between different modalities, enabling speech, text, visual elements, and background music to work together, improving the immersion and adaptability of multimodal synthesized speech. After format conversion, it can be widely applied to various platforms such as mobile devices, online streaming media, and smart devices, realizing cross-industry application promotion.
[0270] In one embodiment, a speech generation device based on multimodal fusion is provided. The speech generation device based on multimodal fusion corresponds one-to-one with the speech generation method based on multimodal fusion in the above embodiment. Refer to Figure 3 , Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the speech generation device based on multimodal fusion of the present invention. Audio data acquisition module 10, timbre training module 20, text semantic parsing module 30, emotion parameter adjustment module 40, personalized parameter modeling module 50, synthesis control parameter generation module 60, multimodal data synchronization module 70, and multimodal synthesized speech generation module 80. The detailed description of each functional module is as follows:
[0271] The audio data acquisition module 10 is used to acquire audio data in the target domain and extract timbre features from the audio data in the target domain;
[0272] The timbre training module 20 is used to train and generate a domain feature timbre generation model based on the timbre features;
[0273] The text semantic parsing module 30 is used to obtain target text data and perform semantic parsing on the target text data to identify the emotion information of the target text data;
[0274] The emotion parameter adjustment module 40 is used to adjust the basic speech synthesis parameters according to the emotion information to generate an emotion adaptation parameter set;
[0275] The personalized parameter modeling module 50 is used to obtain personalized information and construct a personalized parameter mapping table based on the personalized information;
[0276] The synthesis control parameter generation module 60 is used to fuse the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthesis control parameter sequence;
[0277] The multimodal data synchronization module 70 is used to perform timeline alignment on the synthesis control parameter sequence, the associated text annotation, the target domain visual element features, and the background music rhythm data, and establish a parameter-modal binding relationship table;
[0278] The multimodal synthetic speech generation module 80 is used to drive the domain feature timbre generation model based on the parameter-modal binding relationship table to generate multimodal synthetic speech data.
[0279] In one embodiment, the audio data acquisition module 10 is specifically used for:
[0280] Collect speech samples in the target domain, including dialect speech and audio of the pronunciation of target domain-specific terms;
[0281] Use an audio analysis tool to extract the fundamental frequency and formant distribution of the speech samples;
[0282] Collect scene environmental sound effect data, including mechanical operation sounds, natural background sounds, and indoor reverberation audio;
[0283] Perform spectral analysis on the environmental sound effect data to extract acoustic features from the environmental sound effect data;
[0284] Fuse the fundamental frequency, formant distribution, and acoustic features into timbre features.
[0285] In one embodiment, the text semantic parsing module 30 is specifically used for:
[0286] Construct a domain semantic knowledge base pre-annotated with domain emotion vocabulary and the semantics of professional terms;
[0287] Obtain target text data, and based on the annotation information in the domain semantic knowledge base, perform context semantic encoding on the target text data through a pre-trained language model to generate a semantic encoding vector;
[0288] Classify the semantic encoding vector through an emotion classification model to generate an emotion classification result;
[0289] According to the emotion classification result and the distribution of emotion keywords annotated in the domain semantic knowledge base in the target text data, analyze the emotion intensity level of the target text data through weighted statistical operations;
[0290] Generate emotional information metadata including the emotional classification result and the emotional intensity level.
[0291] In one embodiment, the emotional parameter adjustment module 40 is specifically configured to:
[0292] Obtain basic speech synthesis parameters from a preset parameter library, where the basic speech synthesis parameters include a fundamental frequency curve, a speech rate interval, and a spectral envelope shape;
[0293] Adjust the fundamental frequency curve in the basic speech synthesis parameters according to the emotional category in the emotional information;
[0294] Adjust the speech rate interval in the basic speech synthesis parameters according to the emotional intensity level in the emotional information;
[0295] Adjust the spectral envelope shape in the basic speech synthesis parameters according to the emotional information;
[0296] Integrate the adjusted fundamental frequency curve, the adjusted speech rate interval, and the adjusted spectral envelope shape to generate the emotion adaptation parameter set.
[0297] In one embodiment, the personalized parameter modeling module 50 is specifically configured to:
[0298] Collect user preference data through a three-modal interaction interface composed of a voice command recognition module, a gesture recognition module, and an eye movement tracking module;
[0299] Perform noise suppression processing on the user preference data and generate a standardized preference vector using a standardization method;
[0300] Extract a user operation behavior feature sequence from the standardized preference vector through a hidden Markov model, and match a parameter combination from a group preference template library based on a collaborative filtering module to obtain a group template matching result;
[0301] Input the user operation behavior feature sequence and the group template matching result into a multi-level decision tree model, and apply a cultural adaptation rule at the decision node to generate a parameter weight vector;
[0302] Map the parameter weight vector to the basic speech synthesis parameters, and encode the mapping result into a hierarchical data structure including a static parameter layer, a dynamic adjustment layer, and an associated weight layer to generate the personalized parameter mapping table.
[0303] In one embodiment, the multi-modal data synchronization module 70 is specifically configured to:
[0304] Process the target text data through a predefined annotation module to generate associated text annotations, and process the associated text annotations according to a preset segmentation strategy to generate a text time series;
[0305] Extract target domain visual element features from target domain images, and process the target domain visual element features through a key frame detection module to generate a visual time series;
[0306] Extract background music rhythm data from background music audio data, and process the background music rhythm data through a beat detection module to generate a music rhythm time series;
[0307] Parse the timestamp information included in the synthetic control parameter sequence, and perform comparison processing on the timestamp information with the time information recorded in the text time series, visual time series, and music rhythm time series to generate a modality time correspondence mapping;
[0308] Based on a unified time reference, combined with the modality time correspondence mapping, assign corresponding time identifiers to the data items in the synthetic control parameter sequence, associated text annotations, target domain visual element features, and background music rhythm data, and combine them to form a parameter modality binding relationship table with unified time tags.
[0309] In one embodiment, the multimodal synthesis speech generation module 80 is specifically configured to:
[0310] Input the parameter modality binding relationship table into a domain feature tone generation model;
[0311] Generate multimodal synthesis speech data for synchronously displaying text annotations, target domain visual elements, and background music through the domain feature tone generation model;
[0312] According to the application scenario requirements, perform format conversion on the multimodal synthesis speech data to generate target data suitable for target terminal devices.
[0313] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as shown in Figure 4As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media, and internal memory. The non-volatile storage media stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a voice generation method based on multimodal fusion.
[0314] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a voice generation method based on multimodal fusion
[0315] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0316] Collect audio data in the target domain, and extract timbre features from the target domain audio data;
[0317] Based on the timbre features, train and generate a domain feature timbre generation model;
[0318] Obtain target text data, and perform semantic parsing on the target text data to identify the emotional information of the target text data;
[0319] Adjust the basic speech synthesis parameters according to the emotional information to generate an emotion adaptation parameter set;
[0320] Obtain personalized information, and construct a personalized parameter mapping table based on the personalized information;
[0321] Fuse the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthesis control parameter sequence;
[0322] Align the synthesized control parameter sequence with the associated text annotation, the visual element features of the target domain, and the background music rhythm data on the time axis to establish a parameter-modal binding relationship table;
[0323] Drive the domain feature timbre generation model based on the parameter-modal binding relationship table to generate multi-modal synthesized speech data.
[0324] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0325] Collect the audio data of the target domain and extract the timbre features from the audio data of the target domain;
[0326] Train a domain feature timbre generation model based on the timbre features;
[0327] Obtain the target text data and perform semantic parsing on the target text data to identify the emotional information of the target text data;
[0328] Adjust the basic speech synthesis parameters according to the emotional information to generate an emotion-adapted parameter set;
[0329] Obtain personalized information and construct a personalized parameter mapping table based on the personalized information;
[0330] Fuse the emotion-adapted parameter set with the personalized parameter mapping table to generate a synthesized control parameter sequence;
[0331] Align the synthesized control parameter sequence with the associated text annotation, the visual element features of the target domain, and the background music rhythm data on the time axis to establish a parameter-modal binding relationship table;
[0332] Drive the domain feature timbre generation model based on the parameter-modal binding relationship table to generate multi-modal synthesized speech data.
[0333] It should be noted that for the functions or steps that can be implemented by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0334] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0335] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0336] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A speech generation method based on multimodal fusion, characterized in that: The following steps are involved: Collecting target domain audio data, and extracting timbre features from the target domain audio data; Based on the timbre features, training and generating a domain-featured timbre generation model; Acquire target text data, and perform semantic analysis on the target text data to identify emotional information of the target text data; Adjusting basic speech synthesis parameters according to the emotion information to generate an emotion adaptation parameter set; Acquire personalized information, and construct a personalized parameter mapping table based on the personalized information; Merging the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthetic control parameter sequence; Aligning the synthesis control parameter sequence with the associated text annotation, the target domain visual element features and the background music rhythm data on a time axis to establish a parameter modal binding relationship table; The domain characteristic timbre generation model is driven based on the parameter modal binding relationship table to generate multimodal synthesized speech data.
2. The method for generating speech based on multimodal fusion according to claim 1, characterized in that: Collecting target domain audio data and extracting timbre features from the target domain audio data includes: Collect speech samples in the target domain, including dialect speech and pronunciation audio of domain-specific terms; Extracting the pitch frequency and formant distribution of the speech sample using an audio analysis tool; Collect scene environment sound data including mechanical operation sound, natural background sound and indoor reverberation audio; Performing spectrum analysis on the environmental sound effect data to extract acoustic features from the environmental sound effect data; The fundamental frequency, formant distribution and acoustic features are integrated into timbre features.
3. The method for generating speech based on multimodal fusion according to claim 1, characterized in that: Acquiring target text data and performing semantic analysis on the target text data to identify emotional information of the target text data includes: Build a domain semantic knowledge base that is pre-labeled with domain sentiment vocabulary and professional terminology semantics; Obtain target text data, and based on the annotation information in the domain semantic knowledge base, perform context semantic encoding on the target text data through a pre-trained language model to generate a semantic encoding vector; Classifying the semantic coding vector through a sentiment classification model to generate a sentiment classification result; Analyzing the sentiment intensity level of the target text data through weighted statistical operations according to the sentiment classification results and the distribution of sentiment keywords annotated in the domain semantic knowledge base in the target text data; Generate emotional information metadata including the emotional classification result and the emotional intensity level.
4. The method for generating speech based on multimodal fusion according to claim 1, characterized in that: Adjusting basic speech synthesis parameters according to the emotion information to generate an emotion adaptation parameter set includes: Acquire basic speech synthesis parameters from a preset parameter library, wherein the basic speech synthesis parameters include a fundamental frequency curve, a speech rate interval, and a spectrum envelope morphology; According to the emotion category in the emotion information, adjusting the fundamental frequency curve in the basic speech synthesis parameters; According to the emotion intensity level in the emotion information, adjusting the speech rate interval in the basic speech synthesis parameter; Adjusting the spectrum envelope morphology in the basic speech synthesis parameters according to the emotional information; The adjusted fundamental frequency curve, the adjusted speech rate interval and the adjusted spectrum envelope morphology are integrated to generate the emotion adaptation parameter set.
5. The method for generating speech based on multimodal fusion according to claim 1, characterized in that: Acquiring personalized information and constructing a personalized parameter mapping table based on the personalized information includes: Collect user preference data through a three-modal interactive interface consisting of a voice command recognition module, a gesture recognition module, and an eye tracking module; Performing noise suppression processing on the user preference data and generating a standardized preference vector using a standardization method; Extracting a user operation behavior feature sequence from the standardized preference vector through a hidden Markov model, and matching parameter combinations from a group preference template library based on a collaborative filtering module to obtain a group template matching result; Inputting the user operation behavior feature sequence and the group template matching result into a multi-level decision tree model, and applying cultural adaptation rules at the decision nodes to generate parameter weight vectors; The parameter weight vector is mapped to the basic speech synthesis parameter, and the mapping result is encoded into a hierarchical data structure including a static parameter layer, a dynamic adjustment layer and an associated weight layer to generate the personalized parameter mapping table.
6. The method for generating speech based on multimodal fusion according to claim 1, characterized in that: The synthesis control parameter sequence is aligned with the associated text annotation, the target domain visual element characteristics and the background music rhythm data on the time axis to establish a parameter modal binding relationship table, including: Processing the target text data through a predefined annotation module to generate associated text annotations, and processing the associated text annotations according to a preset segmentation strategy to generate a text time series; Extracting target domain visual element features from the target domain image, and processing the target domain visual element features through a key frame detection module to generate a visual time series; Extracting background music rhythm data from background music audio data, and processing the background music rhythm data through a beat detection module to generate a music rhythm time series; Parsing the timestamp information contained in the synthesis control parameter sequence, comparing the timestamp information with the time information recorded in the text time sequence, the visual time sequence and the music rhythm time sequence, and generating a modal time correspondence mapping; Based on a unified time reference and in combination with the modal time correspondence mapping, corresponding time identifiers are assigned to the data items in the synthetic control parameter sequence and associated text annotations, target domain visual element features and background music rhythm data, and combined to form a parameter modal binding relationship table with a unified time label.
7. The method for generating speech based on multimodal fusion according to claim 1, characterized in that: The method drives the domain characteristic timbre generation model based on the parameter modal binding relationship table to generate multimodal synthesized speech data, including: Inputting the parameter modal binding relationship table into a domain characteristic timbre generation model; Generate multimodal synthesized speech data that synchronously displays text annotations, target domain visual elements, and background music through the domain characteristic timbre generation model; According to the application scenario requirements, the multimodal synthesized speech data is formatted to generate target data suitable for the target terminal device.
8. A speech generation device based on multimodal fusion, characterized in that: The speech generation device based on multimodal fusion comprises: An audio data acquisition module, used to acquire target domain audio data and extract timbre features from the target domain audio data; A timbre training module, used for training and generating a domain-feature timbre generation model based on the timbre features; A text semantic analysis module, used to obtain target text data and perform semantic analysis on the target text data to identify emotional information of the target text data; An emotion parameter adjustment module, used to adjust basic speech synthesis parameters according to the emotion information to generate an emotion adaptation parameter set; A personalized parameter modeling module, used to obtain personalized information and construct a personalized parameter mapping table based on the personalized information; A synthetic control parameter generation module, used to merge the emotion adaptation parameter set with the personalized parameter mapping table to generate a synthetic control parameter sequence; A multimodal data synchronization module is used to align the synthesis control parameter sequence with the associated text annotation, the target domain visual element characteristics and the background music rhythm data on the time axis, and establish a parameter modal binding relationship table; The multimodal synthesized speech generation module is used to drive the domain characteristic timbre generation model based on the parameter modal binding relationship table to generate multimodal synthesized speech data.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a speech generation program based on multimodal fusion stored in the memory and executable on the processor. When the speech generation program based on multimodal fusion is executed by the processor, the steps of the speech generation method based on multimodal fusion are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a speech generation program based on multimodal fusion, and when the speech generation program based on multimodal fusion is executed by the processor, the steps of the speech generation method based on multimodal fusion as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Financial service product personalized recommendation generation method and system based on AI
CN120612156A
Speech generation method and device of script, equipment and medium
CN120726988A
Sound duplicating and low-delay streaming speech synthesis method and system based on ultra-short sample
CN120748417A
Voice generation method and device, storage medium and electronic equipment
CN120808747A
Voice generation method based on multi-modal input and related equipment
CN121053996A