Voice broadcasting method and device and electronic equipment

Through electronic devices, the dialect type is determined based on the device location and user voice information, and the dialect synthesis model is used to broadcast dialect voice matching dialect type, which solves the problem of insufficient naturalness and accuracy of dialect voice synthesis in the prior art, and achieves high-quality dialect voice broadcast.

CN120375804APending Publication Date: 2025-07-25VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510493547.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Most of the existing screen reading products are broadcast in Mandarin, or the dialectical pronunciation synthesis is poor in nature and accuracy, which cannot meet the understanding needs of the elderly and people with lower education levels.

Method used

Through electronic devices, dialect types are determined based on device location information and user voice information, dialect synthesis model is used to synthesize and broadcast dialect speech matching dialect type, including building a data processing pipeline to process multiple dialect speech samples, adjusting the front-end processing module and acoustic model module of the Mandarin synthesis model, and training to obtain the dialect synthesis model.

Benefits of technology

It realizes high-quality dialect voice broadcasts, improves the accuracy and nature of dialect voice, increases the intimacy and belonging of voice broadcasts, and facilitates users' understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375804A_ABST
    Figure CN120375804A_ABST
Patent Text Reader

Abstract

The invention discloses a voice broadcasting method and device and electronic equipment, and belongs to the technical field of artificial intelligence. The voice broadcasting method comprises the steps that the dialect type of voice broadcasting is determined according to reference information, and the reference information comprises at least one of equipment position information and user voice information; inputting to-be-broadcasted text information into the dialect synthesis model, and outputting dialect voice matched with the dialect type; and broadcasting the dialect voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of artificial intelligence, and particularly relates to a voice broadcast method, device and electronic device. Background Art

[0002] Dialects are an important part of local culture. Using dialect broadcasts helps protect and spread local dialects. Moreover, dialect voice broadcasts are easier to understand. Especially for the elderly and people with a lower level of education, they may be more accustomed to dialects rather than Mandarin.

[0003] However, current screen reading products either only have the function of Mandarin broadcast or have poor naturalness and accuracy in dialect voice synthesis. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a voice broadcast method, device and electronic device that can achieve natural and accurate dialect voice broadcasts.

[0005] In a first aspect, the embodiments of this application provide a voice broadcast method, which includes: determining the dialect type of the voice broadcast according to reference information, where the reference information includes at least one of device location information and user voice information; inputting the text information to be broadcast into a dialect synthesis model to output a dialect voice matching the dialect type; and broadcasting the dialect voice.

[0006] In a second aspect, the embodiments of this application provide a voice broadcast device, which includes: a processing module for determining the dialect type of the voice broadcast according to reference information, where the reference information includes at least one of device location information and user voice information; the processing module is further configured to input the text information to be broadcast into a dialect synthesis model to output a dialect voice matching the dialect type; and a broadcast module for broadcasting the dialect voice.

[0007] In a third aspect, the embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the voice broadcast method as in the first aspect are implemented.

[0008] In a fourth aspect, the embodiments of this application provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the voice broadcast method as in the first aspect are implemented.

[0009] In a fifth aspect, the embodiments of this application provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the steps of the voice broadcast method as in the first aspect.

[0010] In a sixth aspect, an embodiment of the present application provides a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement the steps of the voice broadcast method as in the first aspect.

[0011] In the voice broadcast method provided by the embodiment of the present application, according to reference information, a dialect type of the voice broadcast is determined. The reference information includes at least one of device location information and user voice information; the text information to be broadcast is input into a dialect synthesis model, and a dialect voice matching the dialect type is output; and the dialect voice is broadcast. Through the above voice broadcast method, based on reference information such as device location information and user voice information, the dialect type of the voice broadcast is determined, and then a dialect synthesis model is used to synthesize and broadcast a dialect voice matching the dialect type according to the text information to be broadcast. In this way, the dialect synthesis model can accurately generate a dialect voice that conforms to the pronunciation characteristics and prosody features of the corresponding dialect type, ensuring the accuracy and naturalness of the dialect voice broadcast, achieving high-quality dialect voice broadcast, and realizing natural and accurate dialect voice broadcast. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 A flowchart of the voice broadcast method provided by some embodiments of the present application;

[0013] Figure 2 A schematic diagram of an operation interface of the voice broadcast method provided by some embodiments of the present application;

[0014] Figure 3 A schematic diagram of the principle of the voice broadcast method provided by some embodiments of the present application;

[0015] Figure 4 A schematic diagram of the principle of the voice broadcast method provided by some embodiments of the present application;

[0016] Figure 5 A flowchart of the voice broadcast method provided by some embodiments of the present application;

[0017] Figure 6 A block diagram of the structure of the voice broadcast device provided by some embodiments of the present application;

[0018] Figure 7 A block diagram of the structure of an electronic device provided by some embodiments of the present application;

[0019] Figure 8 A schematic diagram of the hardware structure of an electronic device provided by some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0021] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the related objects before and after.

[0022] The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application. The terms related to the embodiments of the present application are explained below.

[0023] Model: A tool that uses mathematics, physics, computer, etc. to describe and simulate the structure, behavior, function, etc. of the research object, and can independently complete specific tasks such as dialect speech synthesis tasks.

[0024] Model training: The process of optimizing and adjusting the model by using a large amount of data so that the model can better complete specific tasks.

[0025] Dialect: A language variant used within a specific geographical area or among a specific group of people, which has certain differences from other variants of the same language.

[0026] Dialect synthesis model: A speech synthesis model that supports accurate synthesis of dialect voices of multiple dialect types.

[0027] Speech synthesis or speech broadcast: A technology that converts written text into voice output. The main goal of this technology is to replicate the natural pronunciation of human language so that the machine can convey information in spoken form.

[0028] Multimodal interaction: An interaction method involving multiple perceptual modalities, which improves the naturalness and efficiency of human-computer interaction by combining multiple input and output modes. The core idea of multimodal interaction is to combine two or more data input modes to enhance the system's ability to understand and respond to users.

[0029] Data processing pipeline: An automated and modular process system formed by connecting multiple steps of data processing in sequence.

[0030] Voice replication: A technology that captures and replicates an individual's voice characteristics, such as timbre, intonation, rhythm, etc., and then generates natural speech highly similar to the original speaker's voice.

[0031] Phoneme: The smallest speech unit that can distinguish meaning in speech.

[0032] Phoneme encoder: A component in a speech processing system used to encode phonemes in speech.

[0033] Mel spectrogram: A spectral representation that conforms to the auditory characteristics of the human ear and can better reflect the acoustic characteristics of speech.

[0034] Learning rate: An important hyperparameter in model supervised learning and deep learning processes.

[0035] Hallucination rate: In fields such as natural language processing, it refers to the proportion of content in the text generated by the model that is inconsistent with facts or unreasonable, reflecting the reliability and accuracy of the content generated by the model.

[0036] The following combines the accompanying drawings and details the voice broadcast method provided by the embodiments of the present application through specific embodiments and their application scenarios.

[0037] As Figure 1 shown, the embodiments of the present application provide a voice broadcast method, which may include the following steps 102 to 106:

[0038] Step 102: Determine the dialect type of the voice broadcast according to the reference information.

[0039] The voice broadcast method proposed by the embodiments of the present application is executed by an electronic device, which may specifically be a smart electronic device such as a smartphone, a tablet computer, a laptop computer, and a smart watch, without specific limitation here.

[0040] Among them, the above reference information includes at least one of device location information and user voice information.

[0041] Among them, the device location information may specifically be the GPS (Global Positioning System) location information of the electronic device.

[0042] Furthermore, the above dialect types include but are not limited to: Cantonese, Wu dialect, Hakka, Zhuang language, Miao language, etc., without specific limitation here.

[0043] Specifically, in the voice broadcast method provided in the embodiments of the present application, when the electronic device performs a dialect voice broadcast, the electronic device first determines the dialect type of the voice broadcast according to reference information, such as device location information and user voice information. In this way, based on multi-modal information, determining the dialect type of the voice broadcast can flexibly select different dialects for broadcast according to user needs or scenarios, meeting the diverse voice broadcast requirements.

[0044] In the actual application process, the user can trigger the electronic device to perform a dialect voice broadcast through touch input on the electronic device.

[0045] Exemplarily, as Figure 2 shown, a voice input control 602 is added to the screen reader product AI (Artificial Intelligence) assistive reading. After the user enables the screen reader product, the electronic device floatingly displays the application icon of the screen reader product on the current display interface. When the user opens any page of the electronic device, the user can click on the application icon of the screen reader product to expand the application icon, and then trigger the electronic device to perform a dialect voice broadcast by touching and clicking "Reading List" or "Reading This Page", or input voice through the voice input control 602 to trigger the electronic device to perform a dialect voice broadcast in a voice manner.

[0046] Step 104: Input the text information to be broadcast into the dialect synthesis model, and output the dialect voice matching the dialect type.

[0047] Among them, the dialect synthesis model is obtained by adjusting and training the Mandarin synthesis model. Adjusting and training the dialect synthesis model based on the Mandarin synthesis model can reduce the workload and cost of building a dialect synthesis model from scratch by leveraging the existing basis of the Mandarin synthesis model, improving the development efficiency of the dialect synthesis model.

[0048] Furthermore, the dialect synthesis model supports accurate synthesis of dialect voices of multiple dialect types.

[0049] Furthermore, the text information to be broadcast is the text selected by the user that needs to be broadcast in dialect, such as the text displayed on the current page of the electronic device.

[0050] Specifically, in the voice broadcast method provided in the embodiments of the present application, after determining the dialect type of the voice broadcast, the text information to be broadcast selected by the user is input into the trained dialect synthesis model, and the dialect synthesis model synthesizes and outputs the dialect voice matching the dialect type. In this way, using the dialect synthesis model to synthesize the dialect voice can accurately generate the dialect voice that conforms to the pronunciation characteristics and prosody features of the corresponding dialect type, ensuring the accuracy and naturalness of the dialect voice and realizing high-quality dialect voice output.

[0051] It should be noted that before using the dialect synthesis model to synthesize dialect speech matching the dialect type, it is necessary to determine whether the dialect synthesis model supports synthesizing the dialect speech of this dialect type. If the dialect synthesis model supports synthesizing the dialect speech of the above dialect type, the dialect synthesis model is used to synthesize the dialect speech of this dialect type. If the dialect synthesis model does not support synthesizing the dialect speech of the above dialect type, the dialect synthesis model is used to synthesize Mandarin speech.

[0052] In addition, in the actual application process, an emotional expression function can also be added to the dialect synthesis model according to user needs to further improve the naturalness and intelligence of voice broadcasting using the dialect synthesis model.

[0053] Step 106: Broadcast dialect speech.

[0054] Specifically, in the voice broadcasting method provided in the embodiments of the present application, after using the dialect synthesis model to synthesize and output dialect speech matching the dialect type, the dialect speech is broadcast to the user to read the text information selected by the user according to the determined dialect type, increasing the intimacy and sense of belonging of the voice broadcasting and facilitating the user to more clearly understand the target text.

[0055] The voice broadcasting method provided in the embodiments of the present application determines the dialect type of voice broadcasting according to reference information, where the reference information includes at least one of device location information and user voice information; inputs the text information to be broadcast into the dialect synthesis model, and outputs dialect speech matching the dialect type; broadcasts the dialect speech. Through the above voice broadcasting method, based on reference information such as device location information and user voice information, the dialect type of voice broadcasting is determined, and then the dialect synthesis model is used to synthesize and broadcast dialect speech matching the dialect type according to the text information to be broadcast. In this way, the dialect synthesis model can accurately generate dialect speech that conforms to the pronunciation characteristics and prosody features of the corresponding dialect type, ensuring the accuracy and naturalness of dialect speech broadcasting, realizing high-quality dialect speech broadcasting, and realizing natural and accurate dialect speech broadcasting.

[0056] In the embodiments of the present application, before step 104, the above voice broadcasting method further includes the following steps 108 to 112:

[0057] Step 108: Process multiple dialect speech samples through a data processing pipeline to obtain multiple effective audio segments with the same sampling rate.

[0058] Among them, the multiple effective audio segments are at least two or more effective audio segments.

[0059] Specifically, in the voice broadcast method provided in the embodiments of the present application, before synthesizing and outputting dialect voice using a dialect synthesis model, a data processing pipeline is also constructed, and a preset plurality of dialect voice samples are batch-processed through the data processing pipeline to obtain a plurality of effective audio segments with the same sampling rate and high quality for model training in subsequent steps. In this way, by batch-processing data through the data processing pipeline, the data format can be unified, noise can be eliminated, the quality and usability of the data can be improved, and a high-quality data foundation can be provided for subsequent model training.

[0060] Among them, the above-mentioned dialect voice samples are obtained by users from open-source data sets, corpora, and third-party dialect databases, and dialect text samples corresponding to the dialect voice samples can also be obtained simultaneously. For example, obtain dialect voice samples and corresponding dialect text samples from the recording library of the "Atlas of Chinese Dialects", the open-source database of Mozilla Common Voice, etc., or cooperate with universities or dialect radio stations specializing in various dialects, and obtain high-quality dialect voice samples and corresponding dialect text samples from the dialect databases of the cooperative universities or cooperative radio stations. In this way, on the one hand, a large number of publicly available voice resources can be fully utilized without starting from scratch, saving manpower and time costs, quickly expanding the data reserve, and covering a variety of dialect scenarios and expressions; on the other hand, the data obtained in the above manner is usually professionally sorted and labeled, with high data quality and strong standardization, and can provide reliable and standard data support for subsequent model training.

[0061] Furthermore, the above-mentioned dialect voice samples can also be obtained by users by recording voices according to preset dialect text samples. For example, in a specified region, select specified household register personnel to read dialects according to preset dialect text samples to collect dialect voice samples of the specified dialect text samples. In this way, precise collection can be carried out for specific requirements, specific dialect regions, or specific populations, making the obtained dialect voice samples more suitable for actual application scenarios and meeting the needs of research on specific dialect pronunciation, vocabulary, and other features and model training.

[0062] Furthermore, the above-mentioned dialect voice samples can also be obtained by users by voice replication according to preset dialect text samples. For example, for some minority dialects, the available dialect voice samples are limited. At this time, specifically, existing open-source voice replication models can be used to synthesize dialect voice samples. In this way, when real data is insufficient, supplementary data can be generated through technical means, breaking through the actual collection limitations, increasing data diversity, and also being able to flexibly control voice features and simulate dialect voices in different styles and tones.

[0063] In this way, dialect speech samples are obtained through various methods. On the one hand, it ensures that the collected dialect speech samples have a wide coverage range and diverse types, enabling the dialect synthesis model to learn more comprehensive dialect features and enhancing the generalization ability of the dialect synthesis model to accurately synthesize dialect speech in different scenarios. On the other hand, rich and high-quality dialect speech samples, as the basis for model training, can make the model training more sufficient, better fit the dialect speech rules, improve the accuracy, naturalness, and fluency of dialect synthesis, and optimize the model training effect.

[0064] In the actual application process, after obtaining the dialect speech samples, the dialect type corresponding to each dialect speech sample can also be labeled by manually adding tags. In this way, the dialect speech samples have clear attribute identifiers, facilitating classification management and subsequent use.

[0065] Step 110: Adjust the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model according to the dialect text sample, reference dialect phonemes, and reference dialect tones corresponding to the dialect speech sample to obtain an initial model for dialect synthesis.

[0066] It can be understood that the audio processing architecture of the speech synthesis model is as Figure 3 shown. Taking the original audio signal and audio spectrogram as the speech input, it receives text encoding, timbre, language, or dialect and other information as conditional inputs through a large language model. Further, the audio spectrogram is encoded through frequency-domain encoding to extract frequency-domain features, and the original audio signal is feature-extracted from the time dimension through time-domain encoding to provide audio features in different dimensions for subsequent processing. Further, the encoded audio features are converted into discrete vector representations through vector quantization and audio tokens are generated. Then, the audio tokens are processed by an audio tokenizer for the training and inference stages. Further, the vector representations after vector quantization are decoded by a decoder to obtain an audible audio signal that meets the requirements. Based on this, the main difference between dialect speech synthesis and Mandarin speech synthesis lies in the front-end processing module and the acoustic model module.

[0067] Therefore, in the speech broadcast method provided in the embodiments of the present application, before synthesizing and outputting dialect speech using the dialect synthesis model, the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model are adjusted according to the dialect text sample, reference dialect phonemes, and reference dialect tones corresponding to the obtained dialect speech sample to obtain an initial model for dialect synthesis. In this way, on the one hand, it can optimize the model structure according to the dialect speech characteristics to make it more suitable for the dialect speech synthesis task, such as adapting to dialect pronunciation rules, prosody characteristics, etc.; on the other hand, it is convenient to rely on the existing basis of the Mandarin synthesis model to reduce the workload and cost of building a dialect synthesis model from scratch.

[0068] Step 112: Train the initial model based on multiple valid audio segments and the corresponding dialect text samples and dialect types for each valid audio segment to obtain a dialect synthesis model.

[0069] Specifically, in the voice broadcast method provided in the embodiments of the present application, after adjusting the Mandarin synthesis model to obtain an initial model for dialect synthesis, the initial model is trained based on the processed multiple valid audio segments and the corresponding dialect text samples and dialect types for each valid audio segment to obtain a dialect synthesis model. In this way, an efficient migration from the Mandarin synthesis model to the dialect synthesis model is achieved, reducing the cost and difficulty of developing a completely new dialect synthesis model. At the same time, leveraging the existing advantages of the Mandarin synthesis model improves the processing performance and stability of the dialect synthesis model.

[0070] In the above embodiments provided by the present application, before inputting the text information to be broadcast into the dialect synthesis model and outputting the dialect voice matching the dialect type, the following steps are performed: processing multiple dialect voice samples through a data processing pipeline to obtain multiple valid audio segments with the same sampling rate; adjusting the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model according to the dialect text samples, reference dialect phonemes, and reference dialect tones corresponding to the dialect voice samples to obtain an initial model for dialect synthesis; training the initial model based on multiple valid audio segments and the corresponding dialect text samples and dialect types for each valid audio segment to obtain a dialect synthesis model. In this way, an efficient migration from the Mandarin synthesis model to the dialect synthesis model is achieved, improving the development efficiency of the dialect synthesis model. At the same time, leveraging the existing advantages of the Mandarin synthesis model improves the processing performance and stability of the dialect synthesis model.

[0071] In the embodiments of the present application, the above step of processing multiple dialect voice samples through a data processing pipeline may specifically include the following steps 114 to 120:

[0072] Step 114: Delete the invalid audio segments in each dialect voice sample and retain the valid audio segments in each dialect voice sample.

[0073] Among them, there is no audio signal in the invalid audio segment, and there is an audio signal in the valid audio segment.

[0074] Specifically, in the voice broadcast method provided in the embodiments of the present application, after obtaining the dialect voice samples, delete the invalid audio segments in each dialect voice sample and retain the valid audio segments in each dialect voice sample to remove the redundant data in the dialect voice samples, making the dialect voice samples more accurate, facilitating subsequent processing and analysis, and also reducing the computing resources occupied by the redundant data.

[0075] Step 116: According to the audio spectrum information of each dialect voice sample, filter out the dialect voice samples with audio quality greater than the quality threshold from all the dialect voice samples after deleting the invalid audio segments.

[0076] Specifically, in the voice broadcast method provided in the embodiments of the present application, for all the dialect voice samples after deleting the invalid audio segments, the audio spectrum information of each dialect voice sample is also detected. Then, according to the audio spectrum information of each dialect voice sample, filter out the dialect voice samples with audio quality greater than the quality threshold from all the dialect voice samples after deleting the invalid audio segments. In this way, the audio quality of the dialect voice samples used for subsequent model training is ensured.

[0077] Step 118: Perform noise reduction processing on each of the filtered dialect voice samples with audio quality greater than the quality threshold.

[0078] Specifically, in the voice broadcast method provided in the embodiments of the present application, after filtering out the dialect voice samples with audio quality greater than the quality threshold, perform noise reduction processing on each of the filtered dialect voice samples with audio quality greater than the quality threshold to remove the interference information in each of the filtered dialect voice samples, and further improve the audio quality of the dialect voice samples used for subsequent model training.

[0079] Step 120: Adjust the audio sampling rate of each of the noise-reduced dialect voice samples to the reference sampling rate.

[0080] Specifically, in the voice broadcast method provided in the embodiments of the present application, after performing noise reduction processing on the dialect voice samples, the audio sampling rate of each of the noise-reduced dialect voice samples is also adjusted to the reference sampling rate to unify the sampling rates of the dialect voice samples from different sources, make them conform to the standard format required for model training, avoid compatibility problems caused by sampling rate differences, and ensure the stability and consistency of model training.

[0081] That is, in the voice broadcast method provided in the embodiments of the present application, the processing algorithms included in the above data processing pipeline are specifically as shown in Table 1 below.

[0082] Table 1

[0083]

[0084]

[0085] Among them, the NNVAD (Neural Network-based Voice Activity Detection) algorithm is used to cut the dialect speech samples according to the presence or absence of audio signals in each audio segment of the dialect speech samples, so as to separate the valid audio segments and invalid audio segments in the dialect speech samples, and remove the invalid audio segments. The high-sampling spectrum detection algorithm is used to detect the audio spectrum information, so as to analyze features such as the frequency components of the audio and judge the audio characteristics. The audio quality detection algorithm uses the PESQ (Perceptual Evaluation od Speech Quality) and SISDR (Scale-Invariant Signal-to-Distortion Ratio) metrics to evaluate the audio quality, so as to measure the clarity, intelligibility, etc. of the speech. The background music detection algorithm is based on the sound event detection technology to identify whether there is noise information and interference information in the audio. The resampling algorithm is used to implement the sampling rate conversion to unify the audio sampling rate so that the audio sampling rate meets the requirements of subsequent processing or applications.

[0086] In the above embodiments provided by the present application, during the process of processing multiple dialect speech samples through the data processing pipeline, the invalid audio segments in each dialect speech sample are deleted, and the valid audio segments in each dialect speech sample are retained, where there is no audio signal in the invalid audio segments and there is an audio signal in the valid audio segments; according to the audio spectrum information of each dialect speech sample, from all the dialect speech samples from which the invalid audio segments have been deleted, the dialect speech samples with an audio quality greater than the quality threshold are screened out; noise reduction processing is performed on each of the screened dialect speech samples with an audio quality greater than the quality threshold; the audio sampling rate of each of the dialect speech samples after the noise reduction processing is adjusted to the reference sampling rate. In this way, on the one hand, the processed dialect speech samples are more standardized and of higher quality, which can reduce the computational overhead and errors caused by data quality problems during the subsequent model training process, facilitate accelerating the model training speed, and improve the model training efficiency; on the other hand, it can provide high-quality and standardized data input for the model, enabling the dialect synthesis model to learn more accurate dialect speech features, thereby improving the accuracy, naturalness, and stability of the model's synthesized dialect speech and improving the model performance.

[0087] In the embodiments of the present application, the steps of adjusting the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model according to the corresponding dialect text samples, reference dialect phonemes, and reference dialect tones of the dialect speech samples may specifically include the following steps 122 to step 130:

[0088] Step 122: construct a dialect dictionary based on the dialect text samples corresponding to the dialect speech samples; and load the dialect dictionary into the front-end processing module of the Mandarin synthesis model.

[0089] Specifically, in the speech broadcast method provided in the embodiment of the present application, in the process of adjusting the Mandarin synthesis model, a dialect dictionary is constructed based on the dialect text samples corresponding to the dialect speech samples, and then the dialect dictionary is loaded into the front-end processing module of the Mandarin synthesis model to lay an accurate acoustic parameter foundation for subsequent dialect speech generation, so that the dialect synthesis model can recognize and process dialect vocabulary, solve the problem that the unique vocabulary in the dialect cannot be accurately processed in the Mandarin synthesis model, and improve the dialect synthesis model's ability to parse dialect text.

[0090] It is understandable that the front-end processing module in the speech synthesis model is the core link to ensure the naturalness and accuracy of the synthesized speech. For dialect speech synthesis, since dialects generally have complex pronunciation rules and a large number of unique words such as Cantonese "咁" and Wu dialect "农", the front-end processing module needs to accurately map the sound, form and meaning relationship of dialect words through the dialect dictionary, adapt to the dialect-specific word order logic, and thus convert the original text into a phoneme sequence that conforms to the acoustic laws of the dialect. If this step is missing, the synthesized dialect speech may have a "Mandarin accent".

[0091] The dialect dictionary includes a plurality of Mandarin texts, a plurality of dialect texts, and a correspondence between each Mandarin text and the dialect text. For example, as shown in Table 2 below, the Mandarin text "this" corresponds to the dialect text "勒个", the Mandarin text "什么" corresponds to the dialect text "什子", the Mandarin text "女孩" corresponds to the dialect text "女娃娃", the Mandarin text "婆" corresponds to the dialect text "婆子妈", and the Mandarin text "婆" corresponds to the dialect text "婆子妈".

[0092] Table 2

[0093] Mandarin text Dialect text This This one What What Girl Little girl Grandma Mother-in-law

[0094] Furthermore, for dialects whose grammatical structures differ significantly from Mandarin, such as Cantonese, the dialect dictionary may also include a grammatical rule library of typical sentences. For example, for Cantonese, "你先走" corresponds to "你行先" to facilitate text replacement of typical sentences.

[0095] In actual application, the specific content of the dialect dictionary can be obtained by asking local people for help, translating according to a regular dialect dictionary, etc., and no specific restrictions are made here.

[0096] In addition, due to the wide generalization of dialects, there are more or less differences in pronunciation in different cities and counties. To ensure the stability of synthesized dialects, when constructing a dialect dictionary, as shown in Table 3 below, the regional requirements of different dialect types can be restricted, and a small range of counties or cities can be selected as the regional basis for different dialect types.

[0097] Table 3

[0098] Dialect type Regional requirement Zhuang language City A or City B Miao language County D, City C Hakka dialect City E Southern Min dialect City F or City G

[0099] Step 124: Construct a dialect language matrix according to the mapping relationship between dialect types and discrete encodings.

[0100] Specifically, in the voice broadcast method provided in the embodiments of the present application, it is expected that the dialect synthesis model can achieve mixed training of multiple dialects. Therefore, the concept of discrete encoding is introduced, and a mapping relationship between different dialect types and discrete encodings is established to assign a unique discrete encoding to each dialect type. Further, the discrete encoding is added to the training, and the discrete encoding is mapped to a dialect feature vector through a learnable embedding layer to construct a dialect language matrix according to the mapping relationship between dialect types and discrete encodings. In this way, on the one hand, it can provide a way for the model to quantitatively represent dialect characteristics, enabling the model to more accurately capture and learn the unique characteristics of dialects in terms of pronunciation, vocabulary, etc. On the other hand, through discrete encoding, the training and synthesis of each dialect type are controlled, realizing the unified training and flexible control of the multi-dialect synthesis model.

[0101] Among them, specifically, the mapping from dialect type to discrete encoding can be established through the following formula (1):

[0102] L = {l k |l k ∈Z} (1)

[0103] Among them, L represents the dialect type, l k represents the discrete encoding of the k-th dialect type, and Z represents the set of integers. For example,

[0104] the discrete encoding of Minnan dialect is 0, the discrete encoding of Zhuang language is 1, and the discrete encoding of Wu dialect is 2.

[0105] On this basis, the dialect language matrix can be specifically determined through the following formula (2):

[0106]

[0107] Among them, E lang represents the dialect language matrix, d l represents the dialect feature dimension, and R represents the real number matrix.

[0108] Step 126: Take the intersection of the first dialect phoneme set and the Mandarin phoneme set to obtain the common phoneme set of the dialect and Mandarin; take the difference set of the first dialect phoneme set and the common phoneme set to obtain the unique phoneme set of the dialect; take the union of the common phoneme set and the unique phoneme set to obtain the second dialect phoneme set.

[0109] Among them, the phonemes in the common phoneme set can be used both in Mandarin and in the dialect, and are called common phonemes, while the phonemes in the unique phoneme set can only be used in the dialect, and are called dialect-specific phonemes.

[0110] For example, for the four phonemes n, i3, o3, and ng in Table 4 below, since n, i3, and o3 can all be found in the Mandarin phonemes, and ng cannot be found in the Mandarin phonemes, therefore, n, i3, and o3 are the common phonemes of Mandarin and the dialect, and ng is the dialect-specific phoneme of the dialect.

[0111] Table 4

[0112] Text Mandarin phoneme Dialect phoneme I w o3 ng o3 You ni3 ni3

[0113] Specifically, in the voice broadcast method provided in the embodiments of the present application, during the process of adjusting the Mandarin synthesis model, the intersection of the first dialect phoneme set and the Mandarin phoneme set is also taken to obtain the common phoneme set of the dialect and Mandarin, and then the difference set of the first dialect phoneme set and the common phoneme set is taken to obtain the unique phoneme set of the dialect. Further, the union of the common phoneme set and the unique phoneme set is taken to obtain the second dialect phoneme set.

[0114] Among them, compared with the first dialect phoneme set, for the common phonemes that can be used both in Mandarin and in the dialect, the Mandarin phonemes are still used in the second dialect phoneme set, while for the dialect-specific phonemes that can only be used in the dialect, they are added and extended to modify the phoneme encoder. In this way, for the dialect-specific phonemes that do not exist in Mandarin, such as the Cantonese stop codas [-p / -t / -k], the Wu dialect voiced consonants, the Minnan dialect nasalized vowels, etc., accurate modeling can be achieved through phoneme system expansion and phoneme encoder reconstruction, so as to realize multi-dialect mixed training, which not only ensures the dialect characteristics, but also can ensure the pronunciation stability of dialect synthesis by means of Mandarin phonemes, avoiding the problem of poor stability caused by insufficient data volume.

[0115] In the actual application process, specifically, the second dialect phoneme set can be obtained through the following formulas (3) to (7).

[0116] P = {p1, L, p m} (3)

[0117] Q = {q1, L, q n} (4)

[0118] W = PIQ(5)

[0119] D = Q \ W(6)

[0120] Q' = WUD(7)

[0121] Wherein, P represents the set of Mandarin phonemes, and p m represents the m-th Mandarin phoneme, Q represents the set of first-dialect phonemes, and q n represents the n-th dialect phoneme, W represents the set of common phonemes between Mandarin and the dialect, D represents the set of unique phonemes of the dialect, and Q' represents the set of second-dialect phonemes.

[0122] Step 128: Combine the phoneme embedding sub-matrices corresponding to the set of common phonemes and the set of unique phonemes in the second-dialect phoneme set respectively to obtain the first phoneme embedding matrix; combine the dialect language matrix and the first phoneme embedding matrix to obtain the second phoneme embedding matrix.

[0123] Specifically, in the voice broadcast method provided in the embodiments of the present application, after obtaining the second-dialect phoneme set, combine the phoneme embedding sub-matrices corresponding to the set of common phonemes and the set of unique phonemes in the second-dialect phoneme set respectively to obtain the first phoneme embedding matrix, and then combine the dialect language matrix and the first phoneme embedding matrix to obtain the second phoneme embedding matrix, so as to adjust the original phoneme embedding matrix in the Mandarin synthesis model to the second phoneme embedding matrix, thereby realizing the fusion of dialect embedding and phoneme embedding for training in subsequent model training.

[0124] In the actual application process, the first phoneme embedding matrix can be specifically determined by the following formula (8).

[0125] E = [E shared ; E dialect (8)

[0126] Wherein, E represents the first phoneme embedding matrix, and E shared represents the phoneme embedding sub-matrix corresponding to the set of common phonemes, E shared is inherited from the pre-trained model, E dialect represents the phoneme embedding sub-matrix corresponding to the set of unique phonemes, and E dialect is obtained by random initialization.

[0127] Furthermore, the second phoneme embedding matrix can be determined by the following formula (9):

[0128] E' = [E lang ; E] (9)

[0129] Wherein, E' represents the second phoneme embedding matrix.

[0130] Step 130: Copy the four-tone embedding vectors in the first tone embedding matrix of the acoustic model module, and merge the first tone embedding matrix and the copied four-tone embedding vectors to obtain a second tone embedding matrix.

[0131] Among them, the second tone embedding matrix includes the first type of tone embedding vectors and the second type of tone embedding vectors. Mandarin tones include the first type of tones, and dialect tones include the first type of tones and the second type of tones.

[0132] It can be understood that there may be multiple tones in a dialect. For example, Cantonese has nine tones. This application adjusts and trains a dialect synthesis model based on an existing Mandarin synthesis model. However, the Mandarin synthesis model usually has only four tones. When adjusting and training the model, it is necessary to handle the tones in the dialect that exceed the four tones of Mandarin.

[0133] Therefore, in the voice broadcast method provided in the embodiments of this application, during the process of adjusting the Mandarin synthesis model, for the acoustic model module in the Mandarin synthesis model, the original first tone embedding matrix in the acoustic model module is adjusted according to the reference dialect tones. Specifically, retain the four tone embedding vectors in the first tone embedding matrix that are originally in Mandarin, and copy the four-tone embedding vectors in the first tone embedding matrix that are originally in Mandarin as the initial embedding vectors for the tones in the dialect that exceed the four tones of Mandarin. On this basis, merge the first tone embedding matrix and the copied four-tone embedding vectors to obtain a second tone embedding matrix, and then perform subsequent training fine-tuning based on this second tone embedding matrix, rather than randomly initializing, so as to realize the expansion from the four tones of Mandarin to the multiple tones of the dialect and realize the expansion of the tone embedding layer in the acoustic model module. In this way, on the one hand, the model can better adapt to the diverse tone features in the dialect and enhance the model's ability to model different dialect tones; on the other hand, starting from the four-tone embedding vectors of the Mandarin synthesis model, there is no need to train from scratch, which can converge faster and requires less data compared to random initialization.

[0134] In the actual application process, taking Cantonese as an example, the first tone embedding matrix can be adjusted specifically through the following formula (11):

[0135]

[0136] Among them, E extended represents the second tone embedding matrix, E base represents the original first tone embedding matrix of the Mandarin synthesis model, E base ∈ R 4×d , represents the four-tone embedding vectors in the first tone embedding matrix, E newRepresents the tone embedding matrix for tones in the dialect that exceed the four tones of Mandarin. d represents the embedding dimension. 4 corresponds to the four tones of Mandarin, and 9 corresponds to the 9 tones of Cantonese.

[0137] In the above embodiments provided by this application, according to the dialect text samples corresponding to the dialect voice samples, a dialect dictionary is constructed; the dialect dictionary is loaded into the front-end processing module of the Mandarin synthesis model. The dialect dictionary includes multiple Mandarin texts, multiple dialect texts, and the corresponding relationship between each Mandarin text and the dialect text; according to the mapping relationship between the dialect type and the discrete coding, a dialect language matrix is constructed; the intersection of the first dialect phoneme set and the Mandarin phoneme set is taken to obtain the common phoneme set of the dialect and Mandarin; the difference set of the first dialect phoneme set and the common phoneme set is taken to obtain the unique phoneme set of the dialect; the union of the common phoneme set and the unique phoneme set is taken to obtain the second dialect phoneme set; the phoneme embedding sub-matrices corresponding to the common phoneme set and the unique phoneme set in the second dialect phoneme set are merged to obtain the first phoneme embedding matrix; the dialect language matrix and the first phoneme embedding matrix are merged to obtain the second phoneme embedding matrix; the four-tone embedding vectors in the first tone embedding matrix of the acoustic model module are copied, and the first tone embedding matrix and the copied four-tone embedding vectors are merged to obtain the second tone embedding matrix; wherein, the second tone embedding matrix includes the first type of tone embedding vectors and the second type of tone embedding vectors. The Mandarin tones include the first type of tones, and the dialect tones include the first type of tones and the second type of tones. In this way, on the one hand, through the targeted adjustment of the front-end processing module and the acoustic model module, the dialect synthesis model can more accurately learn the characteristics of the vocabulary, pronunciation, tones, etc. of the dialect, so as to synthesize more compliant dialect voices, improving the naturalness and accuracy of dialect voice synthesis; on the other hand, the adaptability of the model is enhanced, enabling the model to better handle the synthesis tasks of different dialect types.

[0138] In the embodiments of this application, the above step 112 may specifically include the following steps 112a to 112d:

[0139] Step 112a: Construct model training data according to multiple valid audio segments and the corresponding dialect text samples and dialect types for each valid audio segment.

[0140] Specifically, in the voice broadcast method provided by the embodiments of this application, after processing the dialect voice samples to obtain multiple valid audio segments, model training data is constructed according to the multiple valid audio segments and the corresponding dialect text samples and dialect types for each valid audio segment. In this way, the training requirements of the dialect synthesis model can be accurately matched, ensuring the strong pertinence of the training data, and enabling the model to learn accurate dialect voice features.

[0141] In the actual application process, the data structure of the model training data may specifically be:

[0142] <Dialect language_id> <text> <phone> <audio token>。

[0143] Among them, <dialect language_id> is the unique discrete code assigned to each dialect type; <text>Refers to a dialect text sample; <phone>Refers to the phonemes corresponding to the dialect text sample, which can be obtained through the front-end processing module; <audiotoken>It is the result of mapping a continuous audio signal to a discrete symbol space through vector quantization, which can be implemented through VQVAE (Vector Quantized Variational Auto-Encoder) or open source tools.

[0144] In actual application, information such as speaker conversion labels and timestamps may also be added to the model training data, and no specific restrictions are made here.

[0145] Step 112b: Input the model training data into the initial model and output the predicted speech data.

[0146] Specifically, in the speech broadcast method provided in the embodiment of the present application, after constructing the model training data, the model training data is input into the initial model so that it outputs the predicted speech data. In this way, the model can process the input data based on the optimized structure and parameters, and initially output the dialect speech, so as to provide a basis for subsequent evaluation and optimization.

[0147] The processing flow of the initial model can be specifically described as follows: Figure 4 As shown, specifically, the initial model obtains the text information to be converted into the dialect speech, and converts the input text into the corresponding dialect text with the help of the dialect dictionary. For example, taking Cantonese as an example, the input text "是不是" is converted into "是不就是", the input text "的" is converted into "ސ", the input text "了" is converted into "咗", the input text "沒" is converted into "冇", and the input text "在" is converted into "喺". Further, the initial model further converts the dialect text into a phoneme sequence through phoneme conversion. Further, the initial model predicts the tone corresponding to each phoneme according to the phoneme sequence, determines the tone characteristics of the dialect pronunciation, and generates a corresponding feature vector based on the phoneme sequence and the predicted tone, and then uses the acoustic model to process the feature vector and map it into a Mel spectrum. Further, the initial model converts the Mel spectrum into the final dialect speech through a vocoder to achieve the conversion from spectrum information to audible speech.

[0148] Step 112c: Perform weighted calculation on the stop control loss value, text alignment loss value, and semantic quality loss value between the predicted speech data and the model training data to obtain a model loss value.

[0149] Among them, the stop control loss value is determined according to the true stop situation of the speech in the model training data and the predicted stop situation of the predicted speech data. The stop control loss value is used to control the stop timing of speech generation and solve the problem of the uncertainty of the generated sequence length; the text alignment loss value is used to ensure the correspondence between speech and text, and is obtained by calculating the cross-entropy loss between the predicted speech data and the phonemes of the model training data; the semantic quality loss value is used to ensure the quality of speech synthesis, and is obtained by calculating the <audio token>It is obtained from the cross-entropy loss between them.

[0150] Specifically, in the voice broadcast method provided by the embodiments of the present application, after the initial model outputs the predicted voice data, a weighted calculation is performed on the stop control loss value, text alignment loss value, and semantic quality loss value between the predicted voice data and the model training data to obtain the model loss value. In this way, from multiple dimensions such as the start and stop control of voice generation, the alignment degree between text and voice, and the semantic expression quality, the difference between the model output and the real data is comprehensively measured, and the problems existing in the model when synthesizing dialect voices are accurately located, so that the synthesized dialect voices are more in line with the actual needs in terms of start and stop control, text-voice alignment, semantic expression, etc., significantly improving the accuracy, naturalness, and semantic rationality of the synthesized voice.

[0151] In the actual application process, the model loss value can be specifically calculated by the following formula (13):

[0152] Total loss=20×GATE loss+Textloss+Semantic loss (13)

[0153] Among them, Total loss represents the model loss value, GATE loss represents the stop control loss value, Text loss represents the text alignment loss value, and Semantic loss represents the semantic quality loss value. In this way, using the three loss values to train the model improves the stability and accuracy of the model's voice synthesis, and increases the coefficient of the stop control loss value, reducing the "hallucination rate" of voice synthesis.

[0154] Step 112d: According to the model loss value, iteratively update the model parameters of the initial model to obtain a dialect synthesis model with a converged model loss value.

[0155] Specifically, in the voice broadcast method provided by the embodiments of the present application, after calculating the model loss value, according to the model loss value, iteratively update the model parameters of the initial model until a dialect synthesis model with a converged model loss value is obtained, and the training ends. In this way, through multi-dimensional loss evaluation and iterative optimization, the model can perform more stably when processing different dialect voice synthesis tasks, reducing the probability of errors or abnormal outputs, and improving the reliability of the model in actual applications.

[0156] Among them, in the initial model, for each training sample, the learning rate of the phonemes in the common phoneme set is less than the learning rate of the phonemes in the specific phoneme set, so that the model can more efficiently learn the pronunciation rules unique to the dialect, accelerate the learning speed of new phonemes, and at the same time take into account the existing knowledge of Mandarin phonemes.

[0157] Further, the learning rate of the first type of tone is less than that of the second type of tone. In this way, the original four tones in Mandarin can reach a similar convergence speed with the newly added tones in the dialect, and the model can focus on learning the newly added tones in the dialect, improving the synthesis accuracy of the complex dialect tones by the model.

[0158] In this way, training the model based on the hierarchical learning rate enables different parts of the model to update parameters at different speeds according to their own learning needs, improving the model training efficiency and accelerating the model convergence.

[0159] In the actual application process, the learning rates of the common phonemes and the dialect-specific phonemes can be specifically set through the following formula (10):

[0160]

[0161] where lr shared represents the learning rate of the common phonemes, lr dialect represents the learning rate of the dialect-specific phonemes, x represents the phoneme, and e is the natural constant, approximately 2.71828.

[0162] Further, the learning rates of the first type of tone and the second type of tone are set through the following formula (12):

[0163]

[0164] where represents the learning rate of the first type of tone, represents the learning rate of the second type of tone.

[0165] In the above embodiments provided by the present application, according to multiple valid audio segments and the corresponding dialect text samples and dialect types of each valid audio segment, model training data is constructed; the model training data is input into the initial model to output predicted speech data; the stop control loss value, text alignment loss value, and semantic quality loss value between the predicted speech data and the model training data are weighted and calculated to obtain the model loss value; according to the model loss value, the model parameters of the initial model are iteratively updated to obtain a dialect synthesis model with the model loss value converged; wherein, in the initial model, the learning rate of the phonemes in the common phoneme set is less than that of the phonemes in the specific phoneme set, and the learning rate of the first type of tone is less than that of the second type of tone. In this way, through multi-dimensional loss evaluation and iterative optimization, the accuracy, naturalness, and semantic rationality of the synthesized speech are improved, and the reliability of the model in actual applications is enhanced.

[0166] In the embodiments of the present application, the above step 102 may specifically include the following step 102a or step 102b or step 102c:

[0167] Step 102a: Determine the dialect type of the voice broadcast according to the device location information.

[0168] Specifically, in the voice broadcast method provided in the embodiments of the present application, when the electronic device performs a dialect voice broadcast, if the electronic device has enabled the location permission, the dialect type of the voice broadcast is determined according to the device location information. In this way, the strong correlation between the region and the dialect can be utilized to quickly locate the dialect type that the user wants to voice broadcast.

[0169] In the actual application process, specifically, the dialect type can be determined based on the device location information according to the corresponding relationship shown in Table 5 below:

[0170] Table 5

[0171] Device location Dialect type City A or City B Zhuang language County D, City C Miao language City E Hakka dialect City F or City G Southern Min dialect City H Wu dialect

[0172] Step 102b: Identify the key information in the user voice information, and determine the dialect type of the voice broadcast according to the key information.

[0173] Specifically, in the voice broadcast method provided in the embodiments of the present application, in the case where the electronic device triggers a dialect voice broadcast based on the user voice information, the electronic device can identify the text corresponding to the user voice information through an open-source ASR (Automatic Speech Recognition) model, perform key information matching, and then determine the dialect type of the voice broadcast according to the key information. For example, if the user voice information contains the word "Hakka", the dialect type of the voice broadcast is determined to be Hakka. In this way, the dialect type of the voice broadcast can be accurately judged, and it can be accurately recognized even when used across regions, enhancing the adaptability of the model to different usage scenarios.

[0174] Step 102c: Input the user voice information into a dialect recognition model, identify the acoustic features and text features of the user voice information, and output the dialect type of the voice broadcast according to the acoustic features and text features of the user voice information.

[0175] Among them, the above-mentioned dialect recognition model can automatically identify the dialect type based on the input user voice information.

[0176] Further, the above-mentioned acoustic features may include the Mel spectrum, fundamental frequency, tone, and phoneme pronunciation characteristics of the voice, etc., which are not specifically limited here.

[0177] Further, the above-mentioned text features may include the text corresponding to the voice, dialect vocabulary, grammar, etc., which are not specifically limited here.

[0178] Specifically, in the voice broadcast method provided in the embodiments of the present application, when the electronic device triggers the dialect voice broadcast based on the user voice information, the electronic device can also identify the text features of the user voice information through an open-source ASR model, and then use the trained dialect recognition model to identify the acoustic features of the user voice information, and determine the dialect type of the voice broadcast according to the acoustic features and text features of the user voice information. In this way, by comprehensively judging multi-dimensional information, complex voice situations can be processed, and the accuracy and reliability of dialect type judgment are improved.

[0179] In the above embodiments provided by the present application, the dialect type of the voice broadcast is determined according to the device location information; or, the key information in the user voice information is identified, and the dialect type of the voice broadcast is determined according to the key information; or, the user voice information is input into the dialect recognition model, the acoustic features and text features of the user voice information are identified, and the dialect type of the voice broadcast is output according to the acoustic features and text features of the user voice information. In this way, the adaptive dialect type switching based on multi-modal interaction is realized, and the intelligence and convenience of dialect switching are improved.

[0180] In the embodiments of the present application, the above step 102 may specifically include the following steps 102d to 102g:

[0181] Step 102d: Determine the first dialect type according to the device location information.

[0182] In the voice broadcast method provided in the embodiments of the present application, during the process of determining the dialect type of the voice broadcast, the first dialect type will be determined according to the device location information.

[0183] Step 102e: Determine the second dialect type according to the key information in the user voice information.

[0184] In the voice broadcast method provided in the embodiments of the present application, during the process of determining the dialect type of the voice broadcast, the second dialect type will also be determined according to the key information in the user voice information.

[0185] Step 102f: Input the user voice information into the dialect recognition model and output the third dialect type.

[0186] In the voice broadcast method provided in the embodiments of the present application, during the process of determining the dialect type of the voice broadcast, the user voice information will also be input into the dialect recognition model and the third dialect type will be output.

[0187] Step 102g: In the case where at least two of the first dialect type, the second dialect type, and the third dialect type are the same, determine the at least two same dialect types as the dialect type of the voice broadcast.

[0188] Specifically, in the voice broadcast method provided by the embodiments of the present application, after determining the first dialect type, the second dialect type, and the third dialect type by the above three methods, when there are at least two identical dialect types determined by the methods, the at least two identical dialect types are determined as the dialect type for voice broadcast; otherwise, Mandarin is used for voice broadcast. In this way, different determination methods can complement each other, adapt to different application scenarios and user input situations, and effectively determine the dialect type in various environments.

[0189] In the above embodiments provided by the present application, the first dialect type is determined according to the device location information; the second dialect type is determined according to the key information in the user voice information; the user voice information is input into the dialect recognition model to output the third dialect type; when there are at least two identical dialect types among the first dialect type, the second dialect type, and the third dialect type, the at least two identical dialect types are determined as the dialect type for voice broadcast. In this way, by combining different methods to determine the dialect type for voice broadcast, the robustness and accuracy of dialect type recognition are improved.

[0190] In the embodiments of the present application, the step of outputting the dialect type for voice broadcast according to the acoustic features and text features of the user voice information may specifically include the following steps 132 and 134:

[0191] Step 132: Based on the cross-attention mechanism, construct a query vector according to the acoustic features of the user voice information, and construct a value vector and a key vector according to the text features of the user voice information.

[0192] Specifically, in the voice broadcast method provided by the embodiments of the present application, when the electronic device triggers dialect voice broadcast based on the user voice information, the electronic device can use the dialect recognition model to perform splicing and fusion processing on the acoustic features and text features of the user voice information, and based on the cross-attention mechanism, let the acoustic features and text features guide each other. Specifically, based on the cross-attention mechanism, a query vector is constructed according to the acoustic features of the user voice information, and a value vector and a key vector are constructed according to the text features of the user voice information, so that data in different modalities can dynamically focus on the key information of each other, so that in the dialect classification task, the dialect recognition model can automatically discover which lexical features are related to specific pronunciation patterns. In this way, an association can be established between the acoustic features and the text features, enabling the model to better understand the mutual relationship between the acoustic and text information in the speech, mining more discriminative dialect features, and enhancing the depth and accuracy of feature processing.

[0193] Among them, the core mechanism of the cross-attention mechanism can be represented by the following formula (14):

[0194]

[0195] Among them, Q represents the query vector, V represents the value vector, K represents the key vector, d represents the feature dimension scaling factor, and T represents the matrix transpose. In this way, based on the information of speech and text multimodality, the classification accuracy of dialect recognition is improved, and the optimal dialect recognition model can be obtained.

[0196] Step 134: Output the dialect type of the voice broadcast according to the query vector, value vector, and key vector.

[0197] Specifically, in the voice broadcast method provided in the embodiments of the present application, after obtaining the above query vector, value vector, and key vector, the dialect recognition model is used to recognize the dialect type of the user's voice information according to the query vector, value vector, and key vector, and the recognition result is output to obtain the dialect type of the voice broadcast.

[0198] In the above embodiments provided by the present application, based on the cross-attention mechanism, according to the acoustic features of the user's voice information, a query vector is constructed, and according to the text features of the user's voice information, a value vector and a key vector are constructed; according to the query vector, value vector, and key vector, the dialect type of the voice broadcast is output. In this way, based on the cross-attention mechanism, the text features are used to assist in multi-dialect recognition, and the multi-dimensional information is comprehensively used to judge the dialect type, which can handle complex voice situations and improve the accuracy and reliability of dialect type judgment.

[0199] To sum up, as Figure 5 shown, the voice broadcast method provided by the embodiments of the present application may specifically include the following steps 202 to step 226:

[0200] Step 202: Add a voice input control in the screen reader product for the user to perform voice input.

[0201] Step 204: Obtain various dialect text samples and corresponding dialect voice samples in multiple ways, and label the dialect types of the dialect voice samples.

[0202] Step 206: Use the voice replication technology to generate high-quality dialect voice samples for rare minority dialects.

[0203] Step 208: Construct a data processing pipeline to batch process the dialect voice samples to construct model training data.

[0204] Step 210: Construct a dialect dictionary and adjust the front-end processing module of the Mandarin generation model according to the dialect grammar structure.

[0205] Step 212: Adjust the phoneme embedding matrix of the Mandarin synthesis model to a second phoneme embedding matrix, and design a dialect phoneme encoder to realize multi-dialect mixed training.

[0206] Step 214: Initialize the tone embedding vector of the dialect by copying the four-tone embedding vector, adjust the first tone embedding matrix of the acoustic model module to the second tone embedding matrix, and set the hierarchical learning rate to achieve the expansion of Mandarin four tones to dialect multi-tones.

[0207] Step 216: Combine multimodal interaction methods to adaptively determine the dialect type of the voice broadcast.

[0208] Step 218: Determine the first dialect type according to the device location information.

[0209] Step 220: Determine the second dialect type according to the key information in the user voice information.

[0210] Step 222: Use the dialect recognition model, based on the cross-attention mechanism, to determine the third dialect type according to the acoustic features and text features of the user voice information.

[0211] Step 224: In the case where there are at least two identical dialect types among the first dialect type, the second dialect type, and the third dialect type, determine the at least two identical dialect types as the dialect type of the voice broadcast.

[0212] Step 226: Perform dialect broadcast according to the dialect type and the dialect synthesis model.

[0213] In the voice broadcast method provided by the embodiment of the present application, the execution subject may be a voice broadcast device. In the embodiment of the present application, taking the voice broadcast device executing the above voice broadcast method as an example, the voice broadcast device provided by the embodiment of the present application is described.

[0214] As Figure 6 shown, the embodiment of the present application provides a voice broadcast device 300, which may include the following processing module 302 and broadcast module 304.

[0215] The processing module 302 is configured to determine the dialect type of the voice broadcast according to the reference information, where the reference information includes at least one of the device location information and the user voice information;

[0216] The processing module 302 is further configured to input the text information to be broadcast into the dialect synthesis model and output the dialect voice matching the dialect type;

[0217] The broadcast module 304 is configured to broadcast the dialect voice.

[0218] The voice broadcast device 300 provided by the embodiment of the present application determines the dialect type of voice broadcast according to reference information, where the reference information includes at least one of device location information and user voice information; inputs the text information to be broadcast into a dialect synthesis model, and outputs a dialect voice matching the dialect type; and broadcasts the dialect voice. Through the above voice broadcast device 300, based on reference information such as device location information and user voice information, the dialect type of voice broadcast is determined, and then the dialect synthesis model is used to synthesize and broadcast a dialect voice according to the text information to be broadcast. In this way, the dialect synthesis model can accurately generate a dialect voice that conforms to the pronunciation characteristics and prosody features of the corresponding dialect type, ensuring the accuracy and naturalness of the dialect voice broadcast, achieving high-quality dialect voice broadcast, and realizing natural and accurate dialect voice broadcast.

[0219] In the embodiment of the present application, the processing module 302 is further configured to: process a plurality of dialect voice samples through a data processing pipeline to obtain a plurality of effective audio segments with the same sampling rate; adjust the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model according to the dialect text samples, reference dialect phonemes, and reference dialect tones corresponding to the dialect voice samples to obtain an initial model for dialect synthesis; and train the initial model according to the plurality of effective audio segments and the dialect text samples and dialect types corresponding to each effective audio segment to obtain a dialect synthesis model.

[0220] In the above embodiment provided by the present application, a plurality of dialect voice samples are processed through a data processing pipeline to obtain a plurality of effective audio segments with the same sampling rate; the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model are adjusted according to the dialect text samples, reference dialect phonemes, and reference dialect tones corresponding to the dialect voice samples to obtain an initial model for dialect synthesis; and the initial model is trained according to the plurality of effective audio segments and the dialect text samples and dialect types corresponding to each effective audio segment to obtain a dialect synthesis model. In this way, an efficient migration from the Mandarin synthesis model to the dialect synthesis model is realized, the development efficiency of the dialect synthesis model is improved, and at the same time, the existing advantages of the Mandarin synthesis model can be utilized to improve the processing performance and stability of the dialect synthesis model.

[0221] In the embodiment of the present application, the processing module 302 is specifically configured to: delete the invalid audio segments in each dialect speech sample and retain the valid audio segments in each dialect speech sample, where there is no audio signal in the invalid audio segments and there is an audio signal in the valid audio segments; according to the audio spectrum information of each dialect speech sample, screen out the dialect speech samples with audio quality greater than the quality threshold from all the dialect speech samples after deleting the invalid audio segments; perform noise reduction processing on each of the screened dialect speech samples with audio quality greater than the quality threshold; and adjust the audio sampling rate of each of the dialect speech samples after the noise reduction processing to the reference sampling rate.

[0222] In the above embodiment provided by the present application, in the process of processing multiple dialect speech samples through the data processing pipeline, the invalid audio segments in each dialect speech sample are deleted, and the valid audio segments in each dialect speech sample are retained, where there is no audio signal in the invalid audio segments and there is an audio signal in the valid audio segments; according to the audio spectrum information of each dialect speech sample, screen out the dialect speech samples with audio quality greater than the quality threshold from all the dialect speech samples after deleting the invalid audio segments; perform noise reduction processing on each of the screened dialect speech samples with audio quality greater than the quality threshold; and adjust the audio sampling rate of each of the dialect speech samples after the noise reduction processing to the reference sampling rate. In this way, on the one hand, the processed dialect speech samples are more standardized and of higher quality, which can reduce the computational overhead and errors caused by data quality problems in the subsequent model training process, facilitate accelerating the model training speed, and improve the model training efficiency; on the other hand, it can provide high-quality and standardized data input for the model, enable the dialect synthesis model to learn more accurate dialect speech features, thereby improving the accuracy, naturalness and stability of the model's synthesized dialect speech, and enhancing the model performance.

[0223] In the embodiment of the present application, the processing module 302 is specifically configured to: construct a dialect dictionary according to the dialect text sample corresponding to the dialect voice sample; load the dialect dictionary into the front-end processing module of the Mandarin synthesis model, where the dialect dictionary includes multiple Mandarin texts, multiple dialect texts, and the corresponding relationship between each Mandarin text and the dialect text; construct a dialect language matrix according to the mapping relationship between the dialect type and the discrete encoding; take the intersection of the first dialect phoneme set and the Mandarin phoneme set to obtain the common phoneme set of the dialect and Mandarin; take the difference set of the first dialect phoneme set and the common phoneme set to obtain the unique phoneme set of the dialect; take the union of the common phoneme set and the unique phoneme set to obtain the second dialect phoneme set; merge the phoneme embedding sub-matrices corresponding to the common phoneme set and the unique phoneme set in the second dialect phoneme set to obtain the first phoneme embedding matrix; merge the dialect language matrix and the first phoneme embedding matrix to obtain the second phoneme embedding matrix; copy the four-tone embedding vectors in the first tone embedding matrix of the acoustic model module, and merge the first tone embedding matrix and the copied four-tone embedding vectors to obtain the second tone embedding matrix; where the second tone embedding matrix includes the first type of tone embedding vector and the second type of tone embedding vector, the Mandarin tone includes the first type of tone, and the dialect tone includes the first type of tone and the second type of tone.

[0224] In the above embodiments provided by the present application, a dialect dictionary is constructed according to the dialect text samples corresponding to the dialect voice samples; the dialect dictionary is loaded into the front-end processing module of the Mandarin synthesis model, and the dialect dictionary includes multiple Mandarin texts, multiple dialect texts, and the corresponding relationship between each Mandarin text and the dialect text; according to the mapping relationship between the dialect type and the discrete coding, a dialect language matrix is constructed; the intersection of the first dialect phoneme set and the Mandarin phoneme set is taken to obtain the common phoneme set of the dialect and Mandarin; the difference set of the first dialect phoneme set and the common phoneme set is taken to obtain the unique phoneme set of the dialect; the union of the common phoneme set and the unique phoneme set is taken to obtain the second dialect phoneme set; the phoneme embedding sub-matrices corresponding to the common phoneme set and the unique phoneme set in the second dialect phoneme set are merged to obtain the first phoneme embedding matrix; the dialect language matrix and the first phoneme embedding matrix are merged to obtain the second phoneme embedding matrix; the four-tone embedding vectors in the first tone embedding matrix of the acoustic model module are copied, and the first tone embedding matrix and the copied four-tone embedding vectors are merged to obtain the second tone embedding matrix; wherein, the second tone embedding matrix includes the first type of tone embedding vectors and the second type of tone embedding vectors, the Mandarin tones include the first type of tones, and the dialect tones include the first type of tones and the second type of tones. In this way, on the one hand, through the targeted adjustment of the front-end processing module and the acoustic model module, the dialect synthesis model can more accurately learn the characteristics of the dialect's vocabulary, pronunciation, tones, etc., so as to synthesize more compliant dialect voices, improving the naturalness and accuracy of dialect voice synthesis; on the other hand, the model adaptability is enhanced, enabling the model to better handle the synthesis tasks of different dialect types.

[0225] In the embodiments of the present application, the processing module 302 is specifically configured to: construct model training data according to multiple valid audio segments and the corresponding dialect text samples and dialect types of each valid audio segment; input the model training data into the initial model to output predicted speech data; perform weighted calculation on the stop control loss value, text alignment loss value, and semantic quality loss value between the predicted speech data and the model training data to obtain a model loss value; according to the model loss value, iteratively update the model parameters of the initial model to obtain a dialect synthesis model with convergent model loss value; wherein, in the initial model, the learning rate of the phonemes in the common phoneme set is less than the learning rate of the phonemes in the unique phoneme set, and the learning rate of the first type of tones is less than the learning rate of the second type of tones.

[0226] In the above embodiments provided by the present application, model training data is constructed according to multiple valid audio segments and the corresponding dialect text samples and dialect types of each valid audio segment; the model training data is input into the initial model to output predicted speech data; a weighted calculation is performed on the stop control loss value, text alignment loss value, and semantic quality loss value between the predicted speech data and the model training data to obtain a model loss value; according to the model loss value, the model parameters of the initial model are iteratively updated to obtain a dialect synthesis model with convergent model loss value; wherein, in the initial model, the learning rate of the phonemes in the common phoneme set is less than the learning rate of the phonemes in the unique phoneme set, and the learning rate of the first type of tones is less than the learning rate of the second type of tones. In this way, through multi-dimensional loss evaluation and iterative optimization, the accuracy, naturalness, and semantic rationality of the synthesized speech are improved, and the reliability of the model in practical applications is enhanced.

[0227] In the embodiments of the present application, the processing module 302 is specifically configured to: determine the dialect type of the voice broadcast according to the device location information; or, identify the key information in the user voice information, and determine the dialect type of the voice broadcast according to the key information; or, input the user voice information into the dialect recognition model, identify the acoustic features and text features of the user voice information, and output the dialect type of the voice broadcast according to the acoustic features and text features of the user voice information.

[0228] In the above embodiments provided by the present application, the dialect type of the voice broadcast is determined according to the device location information; or, the key information in the user voice information is identified, and the dialect type of the voice broadcast is determined according to the key information; or, the user voice information is input into the dialect recognition model, the acoustic features and text features of the user voice information are identified, and the dialect type of the voice broadcast is output according to the acoustic features and text features of the user voice information. In this way, the adaptive dialect type switching based on multi-modal interaction is realized, and the intelligence and convenience of the dialect switching are improved.

[0229] In the embodiments of the present application, the processing module 302 is specifically configured to: determine the first dialect type according to the device location information; determine the second dialect type according to the key information in the user voice information; input the user voice information into the dialect recognition model to output the third dialect type; in the case that at least two of the first dialect type, the second dialect type, and the third dialect type are the same, determine the at least two same dialect types as the dialect type of the voice broadcast.

[0230] In the above embodiments provided by the present application, the first dialect type is determined according to the device location information; the second dialect type is determined according to the key information in the user voice information; the user voice information is input into the dialect recognition model, and the third dialect type is output; when there are at least two identical dialect types among the first dialect type, the second dialect type, and the third dialect type, the at least two identical dialect types are determined as the dialect type for voice broadcast. In this way, by combining different methods to determine the dialect type for voice broadcast, the robustness and accuracy of dialect type recognition are improved.

[0231] In the embodiments of the present application, the processing module 302 is specifically configured to: based on the cross-attention mechanism, construct a query vector according to the acoustic features of the user voice information, and construct a value vector and a key vector according to the text features of the user voice information; and output the dialect type for voice broadcast according to the query vector, the value vector, and the key vector.

[0232] In the above embodiments provided by the present application, based on the cross-attention mechanism, a query vector is constructed according to the acoustic features of the user voice information, and a value vector and a key vector are constructed according to the text features of the user voice information; and the dialect type for voice broadcast is output according to the query vector, the value vector, and the key vector. In this way, based on the cross-attention mechanism, the text features are used to assist in multi-dialect recognition, and multi-dimensional information is comprehensively used to judge the dialect type, which can handle complex voice situations and improve the accuracy and reliability of dialect type judgment.

[0233] The voice broadcast device 300 in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0234] The voice broadcast device 300 in the embodiments of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0235] The voice broadcast device 300 provided in the embodiments of the present application can implement Figure 1 and Figure 5 each process implemented in the method embodiments. To avoid repetition, it will not be elaborated here.

[0236] Optionally, as Figure 7 shown, the embodiments of the present application further provide an electronic device 400, including a processor 402 and a memory 404. A program or instruction that can run on the processor 402 is stored on the memory 404. When the program or instruction is executed by the processor 402, it implements each step of the above-mentioned voice broadcast method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0237] It should be noted that the electronic device in the embodiments of the present application includes the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0238] Figure 8 It is a schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0239] The electronic device 500 includes but is not limited to: a radio frequency unit 501, a network module 502, an audio output unit 503, an input unit 504, a sensor 505, a display unit 506, a user input unit 507, an interface unit 508, a memory 509, and a processor 510, etc.

[0240] Those skilled in the art can understand that the electronic device 500 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 510 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 8 The structure of the electronic device shown in

[0241] does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0242] The processor 510 is used to determine the dialect type of the voice broadcast according to the reference information, and the reference information includes at least one of device location information and user voice information.

[0243] The audio output unit 503 is used to broadcast dialect voices.

[0244] In the embodiment of the present application, according to the reference information, the dialect type of the voice broadcast is determined. The reference information includes at least one of device location information and user voice information; the text information to be broadcast is input into the dialect synthesis model, and a dialect voice matching the dialect type is output; the dialect voice is broadcast. In the embodiment of the present application, based on reference information such as device location information and user voice information, the dialect type of the voice broadcast is determined, and then the dialect synthesis model is used to synthesize and broadcast the dialect voice according to the text information to be broadcast. In this way, the dialect synthesis model can accurately generate dialect voices that conform to the pronunciation characteristics and prosody features of the corresponding dialect type, ensuring the accuracy and naturalness of the dialect voice broadcast, achieving high-quality dialect voice broadcast, and realizing natural and accurate dialect voice broadcast.

[0245] Optionally, the processor 510 is further used to: process multiple dialect voice samples through a data processing pipeline to obtain multiple effective audio segments with the same sampling rate; adjust the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model according to the dialect text samples corresponding to the dialect voice samples, the reference dialect phonemes, and the reference dialect tones to obtain an initial model for dialect synthesis; train the initial model according to the multiple effective audio segments and the dialect text samples and dialect types corresponding to each effective audio segment to obtain a dialect synthesis model.

[0246] In the above embodiment provided by the present application, multiple dialect voice samples are processed through a data processing pipeline to obtain multiple effective audio segments with the same sampling rate; the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model are adjusted according to the dialect text samples corresponding to the dialect voice samples, the reference dialect phonemes, and the reference dialect tones to obtain an initial model for dialect synthesis; the initial model is trained according to the multiple effective audio segments and the dialect text samples and dialect types corresponding to each effective audio segment to obtain a dialect synthesis model. In this way, an efficient migration from the Mandarin synthesis model to the dialect synthesis model is achieved, improving the development efficiency of the dialect synthesis model. At the same time, the existing advantages of the Mandarin synthesis model can be utilized to improve the processing performance and stability of the dialect synthesis model.

[0247] Optionally, the processor 510 is specifically configured to: delete the invalid audio segments in each dialect speech sample and retain the valid audio segments in each dialect speech sample, where there is no audio signal in the invalid audio segments and there is an audio signal in the valid audio segments; according to the audio spectrum information of each dialect speech sample, filter out the dialect speech samples with audio quality greater than the quality threshold from all the dialect speech samples after deleting the invalid audio segments; perform noise reduction processing on each of the filtered dialect speech samples with audio quality greater than the quality threshold; and adjust the audio sampling rate of each of the dialect speech samples after noise reduction processing to the reference sampling rate.

[0248] In the above embodiments provided by the present application, during the process of processing multiple dialect speech samples through the data processing pipeline, the invalid audio segments in each dialect speech sample are deleted, and the valid audio segments in each dialect speech sample are retained, where there is no audio signal in the invalid audio segments and there is an audio signal in the valid audio segments; according to the audio spectrum information of each dialect speech sample, the dialect speech samples with audio quality greater than the quality threshold are filtered out from all the dialect speech samples after deleting the invalid audio segments; noise reduction processing is performed on each of the filtered dialect speech samples with audio quality greater than the quality threshold; and the audio sampling rate of each of the dialect speech samples after noise reduction processing is adjusted to the reference sampling rate. In this way, on the one hand, the processed dialect speech sample audio is more standardized and of higher quality, which can reduce the computational overhead and errors caused by data quality problems during the subsequent model training process, facilitate accelerating the model training speed, and improve the model training efficiency; on the other hand, it can provide high-quality and standardized data input for the model, enable the dialect synthesis model to learn more accurate dialect speech features, thereby enhancing the accuracy, naturalness, and stability of the model's synthesized dialect speech, and improving the model performance.

[0249] Optionally, the processor 510 is specifically configured to: construct a dialect dictionary according to the dialect text sample corresponding to the dialect voice sample; load the dialect dictionary into the front-end processing module of the Mandarin synthesis model, where the dialect dictionary includes multiple Mandarin texts, multiple dialect texts, and the corresponding relationship between each Mandarin text and the dialect text; construct a dialect language matrix according to the mapping relationship between the dialect type and the discrete encoding; take the intersection of the first dialect phoneme set and the Mandarin phoneme set to obtain the common phoneme set of the dialect and Mandarin; take the difference set of the first dialect phoneme set and the common phoneme set to obtain the unique phoneme set of the dialect; take the union of the common phoneme set and the unique phoneme set to obtain the second dialect phoneme set; merge the phoneme embedding sub-matrices corresponding to the common phoneme set and the unique phoneme set in the second dialect phoneme set to obtain the first phoneme embedding matrix; merge the dialect language matrix and the first phoneme embedding matrix to obtain the second phoneme embedding matrix; copy the four-tone embedding vectors in the first tone embedding matrix of the acoustic model module, and merge the first tone embedding matrix and the copied four-tone embedding vectors to obtain the second tone embedding matrix; where the second tone embedding matrix includes the first type of tone embedding vectors and the second type of tone embedding vectors, the Mandarin tones include the first type of tones, and the dialect tones include the first type of tones and the second type of tones.

[0250] In the above embodiments provided by the present application, a dialect dictionary is constructed according to the dialect text samples corresponding to the dialect voice samples; the dialect dictionary is loaded into the front-end processing module of the Mandarin synthesis model, and the dialect dictionary includes multiple Mandarin texts, multiple dialect texts, and the corresponding relationship between each Mandarin text and the dialect text; according to the mapping relationship between the dialect type and the discrete encoding, a dialect language matrix is constructed; the intersection of the first dialect phoneme set and the Mandarin phoneme set is taken to obtain the common phoneme set of the dialect and Mandarin; the difference set of the first dialect phoneme set and the common phoneme set is taken to obtain the unique phoneme set of the dialect; the union of the common phoneme set and the unique phoneme set is taken to obtain the second dialect phoneme set; the phoneme embedding sub-matrices corresponding to the common phoneme set and the unique phoneme set in the second dialect phoneme set are merged to obtain the first phoneme embedding matrix; the dialect language matrix and the first phoneme embedding matrix are merged to obtain the second phoneme embedding matrix; the four-tone embedding vectors in the first tone embedding matrix of the acoustic model module are copied, and the first tone embedding matrix and the copied four-tone embedding vectors are merged to obtain the second tone embedding matrix; wherein, the second tone embedding matrix includes the first type of tone embedding vectors and the second type of tone embedding vectors, the Mandarin tones include the first type of tones, and the dialect tones include the first type of tones and the second type of tones. In this way, on the one hand, through the targeted adjustment of the front-end processing module and the acoustic model module, the dialect synthesis model can more accurately learn the characteristics of the dialect's vocabulary, pronunciation, tones, etc., so as to synthesize more compliant dialect voices, improving the naturalness and accuracy of dialect voice synthesis; on the other hand, the adaptability of the model is enhanced, enabling the model to better handle the synthesis tasks of different dialect types.

[0251] Optionally, the processor 510 is specifically configured to: construct model training data according to multiple valid audio segments and the corresponding dialect text samples and dialect types of each valid audio segment; input the model training data into the initial model to output predicted speech data; perform weighted calculation on the stop control loss value, the text alignment loss value, and the semantic quality loss value between the predicted speech data and the model training data to obtain a model loss value; according to the model loss value, iteratively update the model parameters of the initial model to obtain a dialect synthesis model with a converged model loss value.

[0252] In the above embodiments provided by the present application, model training data is constructed based on multiple valid audio segments, and the corresponding dialect text samples and dialect types of each valid audio segment; the model training data is input into the initial model to output predicted speech data; a weighted calculation is performed on the stop control loss value, text alignment loss value, and semantic quality loss value between the predicted speech data and the model training data to obtain a model loss value; according to the model loss value, the model parameters of the initial model are iteratively updated to obtain a dialect synthesis model with convergent model loss value. In this way, through multi-dimensional loss evaluation and iterative optimization, the accuracy, naturalness, and semantic rationality of the synthesized speech are improved, and the reliability of the model in practical applications is enhanced.

[0253] Optionally, the processor 510 is specifically configured to: determine the dialect type of the voice broadcast according to the device location information; or, identify the key information in the user voice information, and determine the dialect type of the voice broadcast according to the key information; or, input the user voice information into the dialect recognition model, identify the acoustic features and text features of the user voice information, and output the dialect type of the voice broadcast according to the acoustic features and text features of the user voice information.

[0254] In the above embodiments provided by the present application, the dialect type of the voice broadcast is determined according to the device location information; or, the key information in the user voice information is identified, and the dialect type of the voice broadcast is determined according to the key information; or, the user voice information is input into the dialect recognition model, the acoustic features and text features of the user voice information are identified, and the dialect type of the voice broadcast is output according to the acoustic features and text features of the user voice information. In this way, the adaptive dialect type switching based on multi-modal interaction is realized, and the intelligence and convenience of the dialect switching are improved.

[0255] Optionally, the processor 510 is specifically configured to: determine the first dialect type according to the device location information; determine the second dialect type according to the key information in the user voice information; input the user voice information into the dialect recognition model to output the third dialect type; in the case where at least two of the first dialect type, the second dialect type, and the third dialect type are the same, determine the at least two same dialect types as the dialect type of the voice broadcast.

[0256] In the above embodiments provided by the present application, the first dialect type is determined according to the device location information; the second dialect type is determined according to the key information in the user voice information; the user voice information is input into the dialect recognition model to output the third dialect type; in the case where at least two of the first dialect type, the second dialect type, and the third dialect type are the same, determine the at least two same dialect types as the dialect type of the voice broadcast. In this way, by combining different methods to determine the dialect type of the voice broadcast, the robustness and accuracy of the dialect type recognition are improved.

[0257] Optionally, the processor 510 is specifically configured to: based on the cross-attention mechanism, construct a query vector according to the acoustic features of the user voice information, and construct a value vector and a key vector according to the text features of the user voice information; and output the dialect type of the voice broadcast according to the query vector, the value vector, and the key vector.

[0258] In the above embodiments provided by the present application, based on the cross-attention mechanism, a query vector is constructed according to the acoustic features of the user voice information, and a value vector and a key vector are constructed according to the text features of the user voice information; and the dialect type of the voice broadcast is output according to the query vector, the value vector, and the key vector. In this way, based on the cross-attention mechanism, the text features are used to assist in multi-dialect recognition, and the dialect type is judged by integrating multi-dimensional information, which can handle complex voice situations and improve the accuracy and reliability of dialect type judgment.

[0259] It should be understood that in the embodiments of the present application, the input unit 504 may include a Graphics Processing Unit (GPU) 5041 and a microphone 5042. The graphics processor 5041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 506 may include a display panel 5061, and the display panel 5061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 507 includes at least one of a touch panel 5071 and other input devices 5072. The touch panel 5071 is also called a touch screen. The touch panel 5071 may include two parts: a touch detection device and a touch controller. The other input devices 5072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0260] The memory 509 can be used to store software programs and various data. The memory 509 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 509 may include a volatile memory or a non-volatile memory, or the memory 509 may include both a volatile and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 509 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0261] The processor 510 may include one or more processing units; optionally, the processor 510 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 510 either.

[0262] The embodiments of the present application also provide a readable storage medium. A program or instructions are stored on the readable storage medium. When the program or instructions are executed by a processor, each process of the above embodiments of the voice broadcast method is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0263] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disk or optical disc, etc.

[0264] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the voice broadcast method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0265] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0266] The embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned embodiment of the voice broadcast method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0267] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0268] Through the description of the above embodiments, those skilled in the art can clearly understand that the method of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0269] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit of the present application and the scope protected by the claims, can still make many forms, all of which fall within the protection scope of the present application.< / audio> < / audiotoken> < / phone> < / text> < / audio> < / phone> < / text>

Claims

1. A voice broadcast method, characterized in that Including: Determine the dialect type of the voice broadcast according to the reference information, where the reference information includes at least one of device location information and user voice information; Input the text information to be broadcast into the dialect synthesis model, and output the dialect voice matching the dialect type; Broadcast the dialect voice.

2. The voice broadcast method according to claim 1, wherein Before inputting the text information to be broadcast into the dialect synthesis model and outputting the dialect voice matching the dialect type, the voice broadcast method further includes: Process multiple dialect voice samples through a data processing pipeline to obtain multiple effective audio segments with the same sampling rate; According to the dialect text samples, reference dialect phonemes, and reference dialect tones corresponding to the dialect voice samples, adjust the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model to obtain an initial model for dialect synthesis; Train the initial model according to multiple effective audio segments and the dialect text samples and dialect types corresponding to each effective audio segment to obtain a dialect synthesis model.

3. The voice broadcast method according to claim 2, wherein The processing of multiple dialect voice samples through the data processing pipeline includes: Delete the invalid audio segments in each dialect voice sample and retain the effective audio segments in each dialect voice sample, where there is no audio signal in the invalid audio segments and there is an audio signal in the effective audio segments; According to the audio spectrum information of each dialect voice sample, screen out the dialect voice samples with audio quality greater than the quality threshold from all the dialect voice samples after deleting the invalid audio segments; Perform noise reduction processing on each of the screened-out dialect voice samples with audio quality greater than the quality threshold; Adjust the audio sampling rate of each of the noise-reduced dialect voice samples to the reference sampling rate.

4. The voice broadcast method according to claim 2, characterized in that The adjustment of the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model according to the dialect text samples, reference dialect phonemes, and reference dialect tones corresponding to the dialect voice samples includes: Construct a dialect dictionary according to the dialect text samples corresponding to the dialect voice samples; load the dialect dictionary into the front-end processing module of the Mandarin synthesis model, where the dialect dictionary includes multiple Mandarin texts, multiple dialect texts, and the corresponding relationship between each Mandarin text and the dialect text; Construct a dialect language matrix according to the mapping relationship between the dialect type and the discrete coding; Take the intersection of the first dialect phoneme set and the Mandarin phoneme set to obtain the common phoneme set of the dialect and Mandarin; take the difference set of the first dialect phoneme set and the common phoneme set to obtain the unique phoneme set of the dialect; take the union of the common phoneme set and the unique phoneme set to obtain the second dialect phoneme set; Merge the phoneme embedding sub-matrices corresponding to the common phoneme set and the unique phoneme set in the second dialect phoneme set to obtain the first phoneme embedding matrix; merge the dialect language matrix and the first phoneme embedding matrix to obtain the second phoneme embedding matrix; Copy the four-tone embedding vectors in the first tone embedding matrix of the acoustic model module, and merge the first tone embedding matrix and the copied four-tone embedding vectors to obtain the second tone embedding matrix; Among them, the second tone embedding matrix includes a first type of tone embedding vector and a second type of tone embedding vector. Mandarin tones include the first type of tone, and dialect tones include the first type of tone and the second type of tone.

5. The voice broadcast method according to claim 4, characterized in that, Training the initial model according to the multiple effective audio segments and the corresponding dialect text samples and dialect types of each effective audio segment to obtain a dialect synthesis model, including: Constructing model training data according to the multiple effective audio segments and the corresponding dialect text samples and dialect types of each effective audio segment; Inputting the model training data into the initial model to output predicted speech data; Performing weighted calculation on the stop control loss value, text alignment loss value, and semantic quality loss value between the predicted speech data and the model training data to obtain a model loss value; Iteratively updating the model parameters of the initial model according to the model loss value to obtain a dialect synthesis model with converged model loss value; Among them, in the initial model, the learning rate of the phonemes in the common phoneme set is less than the learning rate of the phonemes in the unique phoneme set, and the learning rate of the first type of tone is less than the learning rate of the second type of tone.

6. The voice broadcast method according to any one of claims 1 to 5, characterized in that Determining the dialect type of the voice broadcast according to the reference information, including: Determining the dialect type of the voice broadcast according to the device location information; Or, identifying the key information in the user voice information, and determining the dialect type of the voice broadcast according to the key information; Or, inputting the user voice information into a dialect recognition model, identifying the acoustic features and text features of the user voice information, and outputting the dialect type of the voice broadcast according to the acoustic features and text features of the user voice information.

7. The voice broadcast method according to any one of claims 1 to 5, characterized in that Determining the dialect type of the voice broadcast according to the reference information, including: Determining a first dialect type according to the device location information; Determining a second dialect type according to the key information in the user voice information; Inputting the user voice information into a dialect recognition model to output a third dialect type; In the case where at least two of the first dialect type, the second dialect type, and the third dialect type are the same, determining the at least two same dialect types as the dialect type of the voice broadcast.

8. The voice broadcast method according to claim 6, wherein Outputting the dialect type of the voice broadcast according to the acoustic features and text features of the user voice information, including: Based on the cross-attention mechanism, constructing a query vector according to the acoustic features of the user voice information, and constructing a value vector and a key vector according to the text features of the user voice information; Outputting the dialect type of the voice broadcast according to the query vector, the value vector, and the key vector.

9. A voice broadcast device, characterized in that, Including: A processing module for determining the dialect type of the voice broadcast according to reference information, where the reference information includes at least one of device location information and user voice information; The processing module is further configured to input the text information to be broadcast into the dialect synthesis model and output dialect speech matching the dialect type; A broadcast module for broadcasting the dialect speech.

10. The voice broadcast device according to claim 9, wherein, The processing module is further configured to: Processing multiple dialect speech samples through a data processing pipeline to obtain multiple effective audio segments with the same sampling rate; Adjust the module parameter values of the front-end processing module and the acoustic model module of the Mandarin synthesis model according to the dialect text sample, reference dialect phonemes, and reference dialect tones corresponding to the dialect voice sample, to obtain an initial model for dialect synthesis; Train the initial model according to the multiple effective audio segments and the dialect text samples and dialect types corresponding to each effective audio segment, to obtain a dialect synthesis model.

11. The voice broadcast device according to claim 10, wherein The processing module is specifically configured to: Delete the invalid audio segments in each dialect voice sample, and retain the effective audio segments in each dialect voice sample, where there is no audio signal in the invalid audio segments, and there is an audio signal in the effective audio segments; According to the audio spectrum information of each dialect voice sample, filter out the dialect voice samples with audio quality greater than the quality threshold from all the dialect voice samples after deleting the invalid audio segments; Perform noise reduction processing on each of the filtered dialect voice samples with audio quality greater than the quality threshold; Adjust the audio sampling rate of each of the dialect voice samples after noise reduction processing to the reference sampling rate.

12. The voice broadcast device according to claim 10, characterized in that, The processing module is specifically configured to: Construct a dialect dictionary according to the dialect text sample corresponding to the dialect voice sample; load the dialect dictionary into the front-end processing module of the Mandarin synthesis model, where the dialect dictionary includes multiple Mandarin texts, multiple dialect texts, and the corresponding relationship between each Mandarin text and the dialect text; Construct a dialect language matrix according to the mapping relationship between the dialect type and the discrete encoding; Take the intersection of the first dialect phoneme set and the Mandarin phoneme set to obtain the common phoneme set of the dialect and Mandarin; Take the difference set of the first dialect phoneme set and the common phoneme set to obtain the unique phoneme set of the dialect; Take the union of the common phoneme set and the unique phoneme set to obtain the second dialect phoneme set; Merge the phoneme embedding sub-matrices corresponding to the common phoneme set and the unique phoneme set in the second dialect phoneme set to obtain a first phoneme embedding matrix; Merge the dialect language matrix and the first phoneme embedding matrix to obtain a second phoneme embedding matrix; Copy the four-tone embedding vectors in the first tone embedding matrix of the acoustic model module, and merge the first tone embedding matrix and the copied four-tone embedding vectors to obtain a second tone embedding matrix; Wherein, the second tone embedding matrix includes first-type tone embedding vectors and second-type tone embedding vectors, the Mandarin tones include first-type tones, and the dialect tones include the first-type tones and second-type tones.

13. The voice broadcast device according to claim 12, wherein The processing module is specifically configured to: Construct model training data according to the multiple effective audio segments and the dialect text samples and dialect types corresponding to each effective audio segment; Input the model training data into the initial model to output predicted speech data; Perform weighted calculation on the stop control loss value, text alignment loss value, and semantic quality loss value between the predicted speech data and the model training data to obtain a model loss value; According to the model loss value, iteratively update the model parameters of the initial model to obtain a dialect synthesis model with convergent model loss value; Among them, in the initial model, the learning rate of the phonemes in the set of common phonemes is less than the learning rate of the phonemes in the set of unique phonemes, and the learning rate of the first type of tone is less than the learning rate of the second type of tone.

14. The voice broadcast device according to any one of claims 9 to 13, characterized in that The processing module is specifically configured to: Determine the dialect type of the voice broadcast according to the device location information; Alternatively, identify the key information in the user voice information, and determine the dialect type of the voice broadcast according to the key information; Alternatively, input the user voice information into a dialect recognition model, identify the acoustic features and text features of the user voice information, and output the dialect type of the voice broadcast according to the acoustic features and text features of the user voice information.

15. The voice broadcast device according to any one of claims 9 to 13, characterized in that, The processing module is specifically configured to: Determine a first dialect type according to the device location information; Determine a second dialect type according to the key information in the user voice information; Input the user voice information into a dialect recognition model and output a third dialect type; In the case that at least two of the first dialect type, the second dialect type, and the third dialect type are the same, determine the at least two same dialect types as the dialect type of the voice broadcast.

16. The voice broadcast device according to claim 14, characterized in that The processing module is specifically configured to: Based on the cross-attention mechanism, construct a query vector according to the acoustic features of the user voice information, and construct a value vector and a key vector according to the text features of the user voice information; Output the dialect type of the voice broadcast according to the query vector, the value vector, and the key vector.

17. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the voice broadcast method described in any one of claims 1 to 8 are implemented.