Speech generation method and apparatus based on staggered sequences, device, and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]本发明的主要目的在于提供一种基于交错序列的语音生成方法、装置、设备及存储介质,旨在解决现有技术难以在多说话人对话场景中对语音内容、说话人身份、方言差异及副语言特征进行统一建模与协同生成,导致语音交互缺乏连续性、一致性及多维表达能力的技术问题
[0010] Beneficial Effects: This invention relates to the field of speech synthesis technology and discloses a speech generation method, apparatus, device, and medium based on interleaved sequences. The method includes: acquiring original multi-speaker audio materials; extracting dialogue sample units and aligning them with timestamps to construct a multi-channel aligned metadata dataset; converting the dialogue sample units into a set of discrete identifiers, including speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, paralinguistic discrete identifiers, and dialect discrete identifiers; inserting paralinguistic discrete identifiers into text discrete identifiers and concatenating various discrete identifiers according to pronunciation timing to generate an interleaved sequence; performing supervised fine-tuning of a pre-trained language generation network based on the interleaved sequence and performing random discarding to obtain a target language generation network; receiving generation condition prompts and concatenating dialect guidance statements when dialect indication information is available; inputting the generation condition prompts into the target language generation network to generate target text discrete identifiers and target speech discrete identifiers, and decoding them into target audio waveforms. This invention can be applied to business scenarios such as fintech and healthcare. By constructing interleaved sequences, it achieves a unified representation of speech, speaker, text, paralinguistics, and dialect information. Combined with random discarding and dialect guidance mechanisms, it enables multi-dimensional speech information to work synergistically during the generation process, thereby improving the continuity, consistency, and expressiveness of multi-speaker speech.
Smart Images

Figure CN122531353A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech synthesis technology, and in particular to a speech generation method, apparatus, device, and medium based on interleaved sequences. Background Technology
[0002] In the development of speech generation technology, most existing technologies are based on single-speaker continuous speech modeling, focusing on text-to-speech mapping capabilities. However, in multi-speaker dialogue scenarios, it is difficult to achieve continuous expression of content, tone, and context across speakers. Furthermore, the ability to express paralinguistic features is insufficient, failing to effectively integrate non-linguistic information such as laughter, pauses, and sighs, resulting in a lack of realistic interactivity in the generated speech. In addition, in scenarios involving multiple dialects, existing technologies lack effective modeling methods for dialect differences and homonyms, easily leading to pronunciation deviations or intonation distortions. Moreover, in the data processing stage, multi-speaker speech data in complex environments generally suffers from overlapping speech, background noise, and incomplete annotations, making it difficult for existing technologies to achieve high-quality sample construction and multi-dimensional information alignment, thus affecting the overall quality of speech generation.
[0003] In the fintech sector, speech generation technology is widely used in scenarios such as intelligent customer service, intelligent outbound calling, voice broadcasting, and investment advisory services. However, existing systems are typically built on a single-speaker model, making it difficult to support multi-role conversational speech generation. This results in a lack of coherence and interactivity in complex business processes (such as risk assessment communication, loan approval interaction, and wealth management product recommendations). Furthermore, insufficient modeling capabilities for user emotional changes, tone of voice, and non-verbal cues render speech interactions lacking nuance in scenarios such as complaint handling, risk warnings, and anti-fraud communication. In addition, when serving users in multiple regions, the differences in speech between different dialects cannot be accurately distinguished and expressed, impacting the localized service experience. At the data level, because financial speech data often contains environmental noise, multi-person interactions, and incomplete annotations, existing technologies struggle to effectively align and unify speech, text, and related auxiliary information, hindering the application of speech generation systems in complex financial business scenarios.
[0004] In the healthcare sector, speech generation technology is primarily applied to scenarios such as intelligent consultation, health follow-up, voice broadcasting, and assisted diagnostic interactions. However, existing technologies also suffer from insufficient multi-speaker interaction capabilities, making it difficult to realistically simulate the multi-turn communication characteristics of doctor-patient dialogues. Furthermore, support for emotional expression and paralinguistic features is limited, failing to effectively reflect tone changes, pauses, and emotional feedback in doctor-patient communication, thus affecting the naturalness and credibility of the interaction. In multi-regional healthcare service scenarios, significant dialect differences make it difficult for existing systems to accurately handle speech expressions and pronunciation variations across different dialects, impacting the accuracy of information transmission. Moreover, medical speech data often originates from complex sources, including background noise, simultaneous speech from multiple speakers, and incomplete annotations. Existing technologies lack the capabilities for data filtering, alignment, and multi-dimensional information fusion, making it difficult to support the construction of high-quality speech generation models. Summary of the Invention
[0005] The main objective of this invention is to provide a speech generation method, apparatus, device, and storage medium based on interleaved sequences, aiming to solve the technical problem that existing technologies are unable to uniformly model and collaboratively generate speech content, speaker identities, dialect differences, and paralinguistic features in multi-speaker dialogue scenarios, resulting in a lack of continuity, consistency, and multi-dimensional expressive capabilities in speech interaction.
[0006] To achieve the above objectives, the present invention provides a speech generation method based on interleaved sequences, comprising: Obtain original multi-speaker audio material, extract dialogue sample units from the original multi-speaker audio material, and align the dialogue sample units with timestamps to construct a multi-channel aligned metadata dataset; The dialogue sample units in the multi-channel aligned metadata set are converted into a discrete identifier set, which includes speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, paralinguistic discrete identifiers, and dialect discrete identifiers. The sub-language discrete identifier is inserted into the text discrete identifier, and the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier and the speech discrete identifier are concatenated according to the pronunciation time sequence to obtain an interleaved sequence; The interleaved sequence is input into the pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech identifiers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network. Receive a generation condition prompt containing text to be generated and speaker prompt information. When the generation condition prompt has dialect indication information, obtain a dialect guidance statement according to the dialect indication information and concatenate the dialect guidance statement to the beginning of the text to be generated. The generation condition prompts are input into the target language generation network, which outputs a discrete target text identifier in an autoregressive manner, outputs a corresponding discrete target speech identifier based on the discrete target text identifier, and decodes the discrete target speech identifier into a target audio waveform.
[0007] Furthermore, to achieve the above objectives, the present invention provides a speech generation apparatus based on interleaved sequences, comprising: A multi-channel alignment construction module is used to acquire original multi-speaker audio materials, extract dialogue sample units from the original multi-speaker audio materials, and perform time stamp alignment on the dialogue sample units to construct a multi-channel alignment metadata dataset. The discrete identifier generation module is used to convert the dialogue sample units in the multi-channel aligned metadata set into a discrete identifier set, wherein the discrete identifier set includes speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, non-language discrete identifiers, and dialect discrete identifiers. An interleaved sequence construction module is used to insert the sub-language discrete identifier into the text discrete identifier, and to concatenate the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier and the speech discrete identifier according to the pronunciation time sequence to obtain an interleaved sequence; A network training module is used to input the interleaved sequence into a pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech identifiers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network. The condition prompt construction module is used to receive a generation condition prompt containing the text to be generated and speaker prompt information. When the generation condition prompt has dialect indication information, the module obtains a dialect guidance statement based on the dialect indication information and concatenates the dialect guidance statement to the beginning of the text to be generated. The speech generation and decoding module is used to input the generation condition prompts into the target language generation network, output the target text discrete identifier in an autoregressive manner, output the corresponding target speech discrete identifier based on the target text discrete identifier, and decode the target speech discrete identifier into a target audio waveform.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an interleaved sequence-based speech generation program stored in the memory and executable on the processor, wherein the interleaved sequence-based speech generation program, when executed by the processor, implements the steps of the interleaved sequence-based speech generation method as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an interleaved sequence-based speech generation program, which, when executed by a processor, implements the steps of the interleaved sequence-based speech generation method as described above.
[0010] Beneficial Effects: This invention relates to the field of speech synthesis technology and discloses a speech generation method, apparatus, device, and medium based on interleaved sequences. The method includes: acquiring original multi-speaker audio materials; extracting dialogue sample units and aligning them with timestamps to construct a multi-channel aligned metadata dataset; converting the dialogue sample units into a set of discrete identifiers, including speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, paralinguistic discrete identifiers, and dialect discrete identifiers; inserting paralinguistic discrete identifiers into text discrete identifiers and concatenating various discrete identifiers according to pronunciation timing to generate an interleaved sequence; performing supervised fine-tuning of a pre-trained language generation network based on the interleaved sequence and performing random discarding to obtain a target language generation network; receiving generation condition prompts and concatenating dialect guidance statements when dialect indication information is available; inputting the generation condition prompts into the target language generation network to generate target text discrete identifiers and target speech discrete identifiers, and decoding them into target audio waveforms. This invention can be applied to business scenarios such as fintech and healthcare. By constructing interleaved sequences, it achieves a unified representation of speech, speaker, text, paralinguistics, and dialect information. Combined with random discarding and dialect guidance mechanisms, it enables multi-dimensional speech information to work synergistically during the generation process, thereby improving the continuity, consistency, and expressiveness of multi-speaker speech. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a speech generation method based on interleaved sequences according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the speech generation method based on interleaved sequences of the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the speech generation device based on interleaved sequences of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The speech generation method based on interleaved sequences provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain original multi-speaker audio materials from the client, extract dialogue sample units, and perform timestamp alignment to construct a multi-channel aligned metadata dataset. The dialogue sample units are converted into a set of discrete identifiers, including speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, paralinguistic discrete identifiers, and dialect discrete identifiers. Paralinguistic discrete identifiers are inserted into text discrete identifiers, and various discrete identifiers are concatenated according to pronunciation sequence to generate interleaved sequences. Based on the interleaved sequences, the pre-trained language generation network is supervised and fine-tuned, and random dropout processing is performed to obtain the target language generation network. Generation condition prompts are received, and dialect guidance statements are concatenated when dialect indication information is present. The generation condition prompts are input into the target language generation network to generate target text discrete identifiers and target speech discrete identifiers, which are then decoded into target audio waveforms. This invention can be applied to business scenarios such as fintech and healthcare. It achieves a unified representation of speech, speaker, text, paralinguistics, and dialect information by constructing interleaved sequences. Combined with random dropout processing and dialect guidance mechanisms, it enables multi-dimensional speech information to work synergistically during generation, thereby improving the continuity, consistency, and expressiveness of multi-speaker speech. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the speech generation method based on interleaved sequences provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the speech generation method based on interleaved sequences proposed in this invention includes the following steps: S10, acquire the original multi-speaker audio material, extract dialogue sample units from the original multi-speaker audio material, and perform time stamp alignment on the dialogue sample units to construct a multi-channel aligned metadata dataset; In this embodiment, the original multi-speaker audio material represents a continuous set of speech data containing the vocal content of multiple speakers. The data sources can be financial customer service call records, outbound voice call recordings, intelligent investment advisor interactive voice, and health consultation voice data. During the acquisition process, speech data is received through audio acquisition devices or data interfaces, and the audio undergoes unified processing, including sampling rate normalization, channel consistency, and encoding format conversion, to ensure that the data structure remains consistent in subsequent processing stages.
[0016] A dialogue sample unit represents a semantically complete segment obtained from continuous speech, with each segment corresponding to a single speaker's vocal range or a range alternating between multiple speakers. In implementation, energy detection and speech activity detection are performed on the audio signal to identify valid speech ranges. Segmentation is then performed by combining speech pause information or speaker variation characteristics, dividing the continuous audio into multiple speech segments. Based on this, adjacent speech segments are combined according to semantic continuity to form a dialogue sample unit with complete semantic expression. A dialogue sample unit can contain speech segments and corresponding text content or business information, such as customer inquiry statements and risk warning statements in a financial scenario.
[0017] Timestamp alignment refers to uniformly time-marking different data within a dialogue sample unit, establishing a correspondence between speech data, text content, and other information on the same timeline. In implementation, the speech signal is divided into a continuous frame sequence, and each frame is assigned a time stamp. Simultaneously, the text content is mapped to the corresponding audio frame interval through alignment processing, creating a correspondence between text units and speech frames. If other related information exists, it is mapped to the unified timeline based on its time position, thus forming a complete time correspondence structure.
[0018] The multi-channel aligned metadata dataset represents a data set structure built on a unified timeline. Channels represent the parallel organization of different data types, including voice channels, text channels, and supplementary information channels. These channels are linked through time stamps. During construction, multiple dialogue sample units that have completed timestamp alignment are aggregated and organized according to a unified data structure. This ensures that each dialogue sample unit contains multi-dimensional data content within its corresponding time range, thereby achieving synchronous expression and unified management of multi-dimensional information.
[0019] During the audio acquisition phase, historical speech data can be processed using offline batch import or continuously received and cached using real-time streaming acquisition. In the dialogue sample unit segmentation process, a fixed time window segmentation strategy can be used, or dynamic segmentation can be performed by combining speech activity detection results with speaker change information. Different granularities of speech segmentation can be achieved by adjusting segmentation thresholds and pause time parameters.
[0020] In timestamp alignment, frame-level alignment can be used to divide audio into equally spaced frame sequences, or semantic unit-based alignment can be used to map text content to speech segments. Different precision requirements can be adapted by adjusting frame length, alignment accuracy, and time offset parameters. During the construction of the multi-channel alignment metadata dataset, structured data storage can be used to record data and timestamps for each channel, or serialization can be used to uniformly encode multi-channel data, adapting to different business environments through field organization methods.
[0021] This embodiment divides the audio materials of multiple speakers into dialogue sample units and performs timestamp alignment, so that the voice data and text and related information are established in a unified time dimension. Furthermore, a multi-channel aligned metadata dataset is constructed to realize the synchronous organization and unified expression of multi-dimensional data, thereby improving the correlation accuracy and consistency between data.
[0022] S20, the dialogue sample unit in the multi-channel aligned metadata set is converted into a discrete identifier set, the discrete identifier set including speech discrete identifier, speaker discrete identifier, text discrete identifier, sub-language discrete identifier and dialect discrete identifier; In this embodiment, the multi-channel aligned metadata dataset represents a dataset whose temporal correspondence has been organized. Each dialogue sample unit contains speech content, text content, and identity and semantic tagging information associated with the speech content. Converting to a discrete tag set represents the unified transformation of continuous or descriptive data into an indexable, sortable, and combinable discrete representation, enabling different types of information to be organized within the same representation space. This approach is necessary because multi-speaker speech content, speaker identity information, text content, paralinguistic information, and dialect information have structural differences in their original forms. Speech is a continuous time-domain signal, text is a symbolic sequence, and identity and event information are tagged data. Without prior discretization, it is difficult to express and associate them within a unified sequence.
[0023] When the speech content in a dialogue sample unit is converted into a discrete speech identifier, the processing object is the continuous acoustic signal corresponding to a short speech segment. The implementation first performs acoustic analysis on the speech content to extract acoustic representations related to speech spectrum changes, and then inputs these acoustic representations into a discrete speech coding structure. The discrete speech coding structure can include an acoustic front-end, a feature compression coding layer, and a quantization layer group. The acoustic front-end is responsible for converting continuous speech into frame-level acoustic representations. The feature compression coding layer is responsible for compressing redundant acoustic information while preserving speech style, prosody, and articulation structure. The quantization layer group is responsible for mapping the continuous representation to multiple discrete codebook indices. Each index position outputs a discrete result, and multiple indices are combined to form the discrete speech identifier. This results in a discrete speech identifier that retains the reproducible acoustic information in the speech content while compressing the continuous waveform into a serializable representation. For recordings of account consultations, repayment reminders, and risk disclosures in fintech businesses, discrete voice identifiers can carry information such as amount announcements, product name announcements, and tone changes; for recordings of health consultations, rehabilitation guidance, and medication reminders in healthcare businesses, discrete voice identifiers can carry information such as symptom descriptions, pauses, and question-and-answer tone.
[0024] When speaker information in dialogue sample units is converted into discrete speaker identifiers, the processing object is the identity information that can distinguish different speakers. The speaker information can come from existing speaker tags in the dialogue sample units or from identity mapping results formed based on voiceprint analysis. In the implementation, speaker information is input into the identity mapping unit, which assigns a unique discrete index to each different speaker. If the input is voiceprint features, the identity representation is first extracted through the voiceprint coding structure, and then the discrete index is assigned through the index mapping table. If the input is an existing tag, the discrete speaker identifier is directly output based on the correspondence between the tag and the index. This processed discrete speaker identifier no longer relies on continuous voiceprint vectors but uses stable discrete identity codes to represent different speakers, facilitating subsequent unified organization with text and speech.
[0025] When the text content in the dialogue sample unit is converted into discrete text identifiers, the processing object is the transcribed text sequence. The purpose of text content conversion is to transform natural language expressions into computable symbol sequences. In the implementation, the text content is tokenized by splitting it into character-level, phrase-level, or syllable-level identifiers based on a pre-defined identifier table. Each identifier is then mapped to a corresponding index to obtain discrete text identifiers. The text tokenization process is not limited to a single splitting granularity. For terms such as fund names, wealth management product names, credit limits, repayment plans, and risk control prompts in fintech businesses, a business terminology-first approach to maintain completeness can be adopted to reduce semantic drift caused by excessive segmentation of financial keywords. For content such as symptom names, health indicator names, and lifestyle recommendations in healthcare businesses, a medical terminology-first approach to maintain completeness can be adopted to reduce semantic fragmentation of health terms during the discretization process.
[0026] When paralinguistic information in dialogue sample units is converted into discrete paralinguistic identifiers, the processing object is non-linguistic expressive information that exists parallel to linguistic content, such as laughter, sighs, inhalations, coughs, prolonged pauses, and interruptions in tone. The source of paralinguistic information can be the paralinguistic event annotation results in the dialogue sample units, or the event category results output by the paralinguistic recognition structure. In the implementation, the paralinguistic information is input into the event mapping unit, and a dedicated discrete index is assigned to each type of paralinguistic event. One paralinguistic event corresponds to one discrete label, and multiple paralinguistic events are mapped according to their temporal positions to obtain discrete paralinguistic identifiers. After this processing, the paralinguistic information no longer remains only at the descriptive level but possesses a symbolic form that can directly participate in sequence expression, thus providing a foundation for the subsequent collaborative expression of linguistic and non-linguistic content.
[0027] When dialect information in dialogue sample units is converted into discrete dialect identifiers, the processing object is dialect category information that can characterize regional speech differences. Dialect information can come from dialect annotation results or from category results output by the dialect recognition structure. In the implementation, dialect information is input into the dialect mapping unit, and a corresponding index is established for each dialect category to obtain discrete dialect identifiers. Discrete dialect identifiers and discrete text identifiers are independent of each other at the representation level. The discrete dialect identifiers are responsible for the category annotation function of pronunciation style and intonation differences, while the discrete text identifiers are responsible for the semantic content expression function. After this separation processing, the textual content of the text and the pronunciation attributes of the dialect can enter the discrete representation space separately, which helps to avoid the confusion of homographs in different dialects.
[0028] The discrete identifier set represents a data set formed by uniformly organizing discrete identifiers for speech, speaker, text, paralinguistics, and dialect. In implementation, a set record unit is established for each dialogue sample unit, writing the five types of discrete identifiers into their corresponding fields and saving the correspondences between them. The discrete identifier set is not a simple concatenation result, but a unified representation structure with associative relationships, ensuring that the speech content, text content, speaker, paralinguistic information, and dialect attributes within the same dialogue sample unit are recorded synchronously. After this processing, data from different sources and with different structures are compressed into a unified discrete representation structure, providing consistent input for subsequent expression and computation.
[0029] In one implementation, the speech content in the dialogue sample unit is discretized using a multi-codebook approach. The acoustic front-end converts the speech content into a time-frequency representation, the coding layer compresses this representation, and the quantization layer outputs multiple discrete indices using a multi-level quantization method. These discrete indices are combined to form a discrete speech identifier. Speaker information is processed using a label mapping method, text content is processed using a business terminology-first text tokenization method, and paralinguistic and dialect information is processed using a classification label-to-index method. Finally, all results are written into a discrete identifier set.
[0030] In another implementation, speech content is processed using residual quantization. The continuous acoustic representation passes through multiple quantization layers sequentially, with each layer outputting a set of discrete indices. The outputs of multiple quantization layers collectively represent the speech content. Text content employs a hybrid tagging method combining character-level and business phrase tags. In fintech business, the integrity of financial phrases such as credit limits, repayment periods, fund codes, and risk levels is preserved; in healthcare business, the integrity of health phrases such as symptom names, rehabilitation action names, and lifestyle recommendations is preserved. Speaker information, paralinguistic information, and dialect information are all processed using independent mapping tables to ensure that different types of information are not confused during discrete representation.
[0031] Discrete identifier sets can also be constructed using a business-scenario-adaptive approach. In fintech, inputs can be customer consultation recordings, customer service response recordings, risk warning texts, and speaker identity tags, while outputs can be corresponding discrete voice identifiers, discrete speaker identifiers, discrete text identifiers, non-language discrete identifiers, and dialect discrete identifiers. In healthcare, inputs can be health consultation voice recordings, symptom description texts, emotion prompts, and dialect tags, while outputs can be the corresponding five types of discrete identifiers. By adjusting the granularity of text tokenization, the number of voice quantization layers, and the granularity of tag mapping, the different requirements for expression accuracy and processing efficiency in different business environments can be adapted.
[0032] This embodiment converts the speech content, speaker information, text content, paralinguistic information, and dialect information in the dialogue sample unit into discrete speech identifiers, discrete speaker identifiers, discrete text identifiers, discrete paralinguistic identifiers, and discrete dialect identifiers, respectively, and further forms a discrete identifier set. This unifies and compresses the originally heterogeneous data into an organized, indexable, and associative discrete representation structure, thereby improving the consistency of multidimensional data representation, enhancing the correspondence accuracy between different information, and providing stable input for subsequent unified processing.
[0033] S30, insert the sub-language discrete identifier into the text discrete identifier, and concatenate the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier and the speech discrete identifier according to the pronunciation time sequence to obtain an interleaved sequence; In this embodiment, the insertion of non-linguistic discrete identifiers into text discrete identifiers refers to embedding previously independent non-linguistic event representations into the text expression sequence. This allows the text expression to not only carry semantic content but also simultaneously carry communicative information such as laughter, sighs, pauses in breathing, coughs, and drawn-out tones. The source of the non-linguistic discrete identifiers is the non-linguistic event category results already formed in the preceding processing, while the source of the text discrete identifiers is the discrete expression formed after the transcribed text has been tokenized. The purpose of the insertion is to maintain the consistency between the non-linguistic information and the text semantics in terms of temporal position, preventing the non-linguistic information from existing independently of the specific text position. In implementation, based on the occurrence position of the non-linguistic event on the timeline, the non-linguistic discrete identifier is written into the corresponding position between or within the text discrete identifiers corresponding to that time position. This ensures that a text expression unit in the discrete sequence contains both semantic content and non-linguistic prompts occurring synchronously with that semantic segment. The resulting text discrete identifiers are no longer pure text sequences but composite expression sequences with interactive tone and non-linguistic attributes. In fintech business scenarios involving risk disclosure, profit description, overdue reminders, and anti-fraud verification, this insertion method can bind communication traces such as pauses, emphasis, and hesitation with textual content such as amounts, terms, rates, and risk levels. In healthcare business scenarios involving consultation responses, rehabilitation reminders, and health education, this insertion method can bind expressions such as sighs, pauses, and inhalations with textual content such as symptom descriptions, medication reminders, and behavioral suggestions.
[0034] The discrete identifiers of speaker, dialect, text, and speech are concatenated according to the order of pronunciation, representing the organization of multiple categories of discrete information into a unified sequence based on the actual speech expression order. The discrete identifier of speaker serves to distinguish the speaker, the discrete identifier of dialect serves to mark regional pronunciation differences, the discrete identifier of text serves to represent the composite semantic and paralinguistic content, and the discrete identifier of speech serves to represent the reproducible acoustic information. Concatenation according to the order of pronunciation is not a simple stacking by category, but rather an organization based on the chronological relationship in the actual pronunciation process, enabling the discrete sequence to reflect who is speaking, in what dialect, what is said, and the corresponding acoustic expression. In implementation, a single speech segment can be considered a pronunciation unit. Within each pronunciation unit, the discrete identifier of speaker is written first, followed by the discrete identifier of dialect, then the discrete identifier of text with embedded paralinguistic information, and finally the discrete identifier of speech corresponding to the text content. Multiple pronunciation units are then arranged sequentially according to the original speech time order, thus forming a complete interleaved sequence. The significance of interleaved sequences lies in unifying previously scattered identity information, dialect information, text information, paralinguistic information, and acoustic information into a single sequential expression structure. This maintains the adjacency of multidimensional information occurring within the same time period and ensures traceability of transitions between different time periods. In fintech, when customers inquire about profit ranges, customer service provides risk warnings, or customers reconfirm redemption rules, the speakers, dialects, text, and voice information from different rounds can form a continuous interleaved sequence based on the order of pronunciation. Similarly, in healthcare, when users describe discomfort, the system provides health advice, or users supplement details, the expressions from different rounds can also be uniformly organized into an interleaved sequence. Once the interleaved sequence is formed, the discrete identifier at any position in the sequence not only retains its own semantics but also carries the contextual relationship of its preceding and following positions, thereby enhancing the organizeability of continuous multi-speaker expressions.
[0035] Paralinguistic discrete identifiers can only acquire complete semantic and non-linguistic expressive capabilities after first entering the text discrete identifiers. Only after this internal expansion, and then concatenated with speaker discrete identifiers, dialect discrete identifiers, and speech discrete identifiers according to pronunciation sequence, can the text discrete identifiers form an interlaced sequence that truly reflects the multi-speaker communication process. Without this insertion action, paralinguistic information will be separated from its semantic location; without sequential concatenation according to pronunciation sequence, the correspondence between identity, dialect, text, and speech will lose its sequential constraint.
[0036] In one implementation, the insertion position of the paralinguistic discrete identifier is determined based on the time alignment result. The time alignment result provides the corresponding intervals of the paralinguistic event and the text content on a unified time axis. The closest text discrete identifier position is selected based on the start and end positions of the interval, and the paralinguistic discrete identifier is written before, after, or within that position. For paralinguistic events with short durations, a single-point insertion method can be used; for paralinguistic events with longer durations, a multi-point insertion method within the interval can be used to maintain temporal coverage of the paralinguistic expression in the text sequence. After the text discrete identifier completes the paralinguistic embedding, it is then concatenated with the speaker discrete identifier, dialect discrete identifier, and speech discrete identifier in the order of articulation units to form a complete interleaved sequence.
[0037] In another implementation, the insertion position of the paralinguistic discrete markers is determined jointly by the text segmentation results and the boundaries of the pronunciation segments. If the paralinguistic event occurs at the beginning of a sentence, the paralinguistic discrete marker is placed before the corresponding text discrete marker; if the paralinguistic event occurs in the middle of a sentence, the paralinguistic discrete marker is placed between adjacent text discrete markers; if the paralinguistic event occurs at the end of a sentence, the paralinguistic discrete marker is placed after the corresponding text discrete marker. Speaker discrete markers and dialect discrete markers are placed at the beginning according to a fixed order of pronunciation units, text discrete markers are located in the middle, and speech discrete markers are located at the end. This sequential organization maintains consistency in pronunciation subject, dialect attributes, content expression, and acoustic expression.
[0038] An organizational approach adapted to the business environment can also be adopted. In fintech businesses, text discrete identifiers containing amount announcements, risk disclosures, return comparisons, and repayment reminders can maintain the integrity of business keywords. Then, secondary language discrete identifiers are inserted in corresponding positions to preserve communication traces such as pauses, emphasis, and hesitation. Subsequently, the speaker discrete identifiers corresponding to customer service and customer identities, as well as dialect discrete identifiers corresponding to regional pronunciation differences, are concatenated with the text discrete identifiers and speech discrete identifiers in chronological order. In healthcare businesses, text discrete identifiers containing symptom descriptions, health advice, and rehabilitation tips can prioritize maintaining the integrity of health terminology before inserting secondary language discrete identifiers to preserve expressions such as rapid breathing, sighs, and pauses. These are then concatenated with the speaker discrete identifier, dialect discrete identifier, and speech discrete identifier to form a unified sequence. By adjusting the insertion granularity of secondary language discrete identifiers, the segmentation granularity of text discrete identifiers, and the concatenation granularity of pronunciation units, different requirements for expression precision and sequence length in different business environments can be accommodated.
[0039] This embodiment inserts the discrete identifiers of the paralinguistic language into the discrete identifiers of the text, establishing a positional correspondence between non-linguistic events and semantic content in the same discrete expression. Then, according to the pronunciation sequence, the discrete identifiers of the speaker, dialect, text, and speech are concatenated into an interleaved sequence, so that identity information, regional pronunciation information, semantic information, paralinguistic information, and acoustic information are synchronously organized in a unified temporal structure, thereby enhancing the correspondence accuracy and expression consistency between multidimensional information.
[0040] S40, the interleaved sequence is input into the pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech identifiers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network. In this embodiment, the interleaved sequence represents unified temporal data that has been organized with multidimensional information. The sequence contains speaker information, dialect information, text information, paralinguistic information, and speech information, all arranged in the order of pronunciation. Inputting the interleaved sequence into a pre-trained language generation network for supervised fine-tuning means using the interleaved sequence as training input to readjust the parameters of the pre-trained language generation network for the target task, enabling the network to shift from general language generation capabilities to multi-speaker speech expression capabilities. The pre-trained language generation network can employ a sequence modeling network based on an autoregressive prediction structure, including an input embedding layer, a positional encoding layer, a multi-layer contextual modeling layer, and an output prediction layer. The input embedding layer converts different discrete identifiers in the interleaved sequence into a unified-dimensional vector; the positional encoding layer writes temporal positional information into the embedding vector; the multi-layer contextual modeling layer captures speaker switching, semantic continuity, dialect attribute changes, and paralinguistic insertion relationships over a long time span; and the output prediction layer provides a conditional distribution for the discrete result at the next position. The context modeling layer can be composed of cascaded attention units and feedforward transformation units, or it can be composed of a hybrid of gated temporal units and attention units. Residual connections and normalization structures maintain feature propagation stability between layers. Supervised fine-tuning essentially involves inputting the interleaved sequence position-by-position into the pre-trained language generation network. At each prediction position, the output is constrained by the preceding discrete information, and then the output is compared with the corresponding true discrete result. The network parameters are then adjusted in reverse based on the error.
[0041] The goal of supervised fine-tuning is to improve the conditional probability of the true discrete result at the current position in the output prediction layer, given the interleaved sequence conditions preceding the current position, while maintaining consistency in speaker expression, dialect attributes, paralinguistic events, and speech reconstruction. The corresponding optimization objective is composed of speech prediction cross-entropy error terms, speaker consistency error terms, paralinguistic event classification error terms, dialect classification error terms, and reconstruction error terms. The speech prediction cross-entropy error term constrains the prediction accuracy of the discrete result at the current position; the speaker consistency error term constrains the consistency between the output representation and the current speaker's identity representation; the paralinguistic event classification error term constrains the consistency between the paralinguistic category and the sample label; the dialect classification error term constrains the consistency between the pronunciation attribute and the dialect label; and the reconstruction error term constrains the distance between the output representation and the reference acoustic representation. These error terms are weighted and summed to form a joint optimization objective, which is then used for backpropagation to update the network parameters.
[0042] The current prediction position represents the discrete position being predicted during training, and its definition depends on the temporal arrangement of the interleaved sequence itself. The discrete speech markers preceding the current prediction position represent discrete expressions that already exist within the preceding context and belong to the speech category at the current training moment. Random discarding involves artificially masking a portion of historical discrete speech markers during the training phase. This prevents the pre-trained language generation network from completely relying on the replication of preceding speech when predicting the current position, and instead forces it to rely more on speaker information, dialect information, text information, and paralinguistic information to infer subsequent content. This is because historical discrete speech markers in multi-speaker continuous speech data often carry strong local acoustic repetition. Without control, the network can easily become overly reliant on preceding speech details, thus weakening its ability to comprehensively model semantic, identity, and dialect information. Random discarding does not delete the entire interleaved sequence, but rather partially masks the discrete speech markers preceding the current prediction position, ensuring that the current prediction position retains necessary context while avoiding a simple replication tendency of historical speech. The resulting target language generation network can more stably maintain content continuity and expression consistency when facing scenarios involving multi-turn dialogues, speaker switching, dialect changes, and paralinguistic insertion. Random dropout processing can be implemented using position-independent sampling, generating a retention or masking marker for each discrete speech identifier preceding the current prediction position; alternatively, it can be implemented using segment-based sampling, selecting consecutive segments from historical discrete speech identifiers for overall masking. Masked discrete positions can be replaced with empty markers, masking markers, or zero-vector placeholders, maintaining consistency with the input format of the embedding layer. Random dropout processing only applies to discrete speech identifiers, not speaker information, dialect information, text information, or paralinguistic information, thus ensuring that identity constraints, semantic constraints, dialect constraints, and paralinguistic constraints remain visible throughout the training process.
[0043] The pre-trained language generation network receives discrete identifier indices from the interleaved sequence as input. These indices are mapped to fixed-length vectors via an input embedding layer. The vector dimension can be set to 256, 512, or 768, with 512 being preferred. The position encoding layer superimposes the sequence position information onto the input vector. The context modeling layer is configured with a stacked structure of 8 to 16 layers, preferably 12 layers. Each layer contains 4 to 12 multi-head attention units and a feedforward transformation unit, preferably 8 multi-head attention units. The training segment length is set to 256, 512, or 1024 discrete positions, with 512 discrete positions being preferred. The output prediction layer outputs a discrete identifier probability distribution for each predicted position. During training, the true discrete results of the current position in the interleaved sequence are used as the supervision target, and cross-entropy loss is used as the basic error term. Random discarding is performed during each round of parameter updates. The random dropout rate is set to 10% to 40%, with the first-stage supervised fine-tuning using a random dropout rate of 10% to 20% and the second-stage supervised fine-tuning using a random dropout rate of 20% to 40%. The training batch size is set to 16, 32, or 64, preferably 32. The learning rate for the first-stage supervised fine-tuning is set to 5×10^-5 to 2×10^-4, preferably 1×10^-4, and the learning rate for the second-stage supervised fine-tuning is set to 1×10^-5 to 1×10^-4, preferably 5×10^-5. The warm-up steps account for 3% to 8% of the total training steps, preferably 5%. The learning rate decay method uses cosine decay or step decay. The first-stage supervised fine-tuning has 4 to 12 training epochs, preferably 8 epochs, and the second-stage supervised fine-tuning has 6 to 18 training epochs, preferably 12 epochs. The validation and evaluation cycle is set to be executed once every 500 to 2000 parameter updates, preferably once every 1000 parameter updates. The training termination condition is set to be met if at least one of the following conditions is met: the joint error function decreases by less than 0.1% in three consecutive validation evaluations, or the total number of training epochs reaches 20, or the learning rate decays to below 1×10^-6, or the joint error on the validation set no longer decreases in three consecutive evaluations. After the training termination condition is met, the target language generation network is output. The target language generation network's parameter distribution is adapted to sequence expression tasks involving multiple speakers, dialects, and paralinguistics.
[0044] In fintech business environments, interleaved sequences can consist of multiple rounds of voice interaction content, such as customer inquiries, customer service responses, risk disclosures, credit limit explanations, and repayment reminders. The input content includes not only the switching of different speakers but also financial keywords such as profit ranges, fee explanations, risk levels, and credit conditions. When using such interleaved sequences for supervised fine-tuning, randomly discarding some historical voice discrete identifiers can reduce the network's dependence on the surface acoustic morphology of a particular round of speech, allowing the network to utilize more textual content, speaker identity, and dialect attributes to understand subsequent outputs. The target language generation network trained in this way is more likely to maintain consistency in the expression of terms and conditions, product announcements, and risk warnings in financial voice interaction scenarios. In this business environment, the training input can be set as an interleaved sequence containing content such as financial product descriptions, return range announcements, fee terms announcements, risk disclosure announcements, credit approval communication, and repayment plan reminders. The training output is the discrete identifier prediction distribution and intermediate representation corresponding to the current position. The speaker consistency error term focuses on constraining the identity boundary between the customer role and the customer service role. The dialect classification error term focuses on constraining the regional pronunciation differences of customers in different regions in expressing financial keywords. The paralinguistic event classification error term focuses on constraining the expression position of communication states such as pauses, hesitations, and emphasis in the terms and risk warnings.
[0045] In a healthcare business environment, interleaved sequences can consist of multi-turn voice interactions such as health consultations, symptom descriptions, lifestyle suggestions, and rehabilitation reminders. The input content includes changes in the speaker, semantic continuity, and paralinguistic expressions. Random dropout processing reduces the network's reliance on local historical speech forms during training, thereby strengthening its overall grasp of health description content and interactive tone, making the trained target language generation network more suitable for complex health consultation expressions. In this business environment, the training input can be set as an interleaved sequence containing symptom descriptions, health suggestions, rehabilitation reminders, medication tips, and lifestyle suggestions. The training output is also the discrete identifier prediction distribution and intermediate representation corresponding to the current position. The speaker consistency error term focuses on constraining the distinction between the consultation and service provider's expression styles; the paralinguistic event classification error term focuses on constraining the expression positions of events such as sighs, pauses, inhalations, and coughs in symptom descriptions; the dialect classification error term focuses on constraining the pronunciation differences of users from different regions in expressing health vocabulary; and the reconstruction error term focuses on constraining the stability of health consultation statements in overall acoustic reconstruction.
[0046] In one implementation, the pre-trained language generation network employs a multi-layer self-attention structure. After embedding, the interleaved sequence enters a multi-layer context modeling layer. In each layer, the multi-head attention unit is responsible for establishing the dependency between the current position and previous positions, while the feedforward transformation unit is responsible for performing non-linear feature transformations. During training, the interleaved sequence is divided into multiple training segments of fixed length. Before each training segment is input into the network, it first retrieves historical speech discrete identifiers based on the current predicted position, and then performs random discarding according to a preset probability. The discarded training segments are input into the network, and the error between the output predicted distribution and the actual discrete results is calculated. Then, parameters are updated through backpropagation. The random discarding ratio can be dynamically adjusted according to the sequence length; the discarding ratio is increased for longer sequences and decreased for shorter sequences, allowing the network to maintain a relatively stable training intensity across different contexts.
[0047] In another implementation, the pre-trained language generation network employs a combination of gated temporal units and attention units. The interleaved input sequence is first processed by the gated units for local temporal modeling, and then by the attention units for long-range dependency modeling. This allows for the preservation of speaker turn-by-turn information while controlling computational complexity. During training, a candidate set of all discrete speech identifiers prior to the current prediction position is established. Then, based on the sampling ratio, a subset of discrete speech identifiers from the candidate set is selected for masking or empty label replacement, enabling the network to perceive missing historical speech during prediction. Parameter updates can utilize an optimizer with weight decay, where the learning rate gradually decreases with each turn. The training termination condition can be determined by the validation set error stabilizing, the number of training turns reaching a preset upper limit, or the parameter update magnitude falling below a threshold.
[0048] This embodiment uses interleaved sequences as input for supervised fine-tuning of a pre-trained language generation network. During training, it randomly discards discrete speech markers before the current prediction position, reducing the network's over-reliance on historical speech surface acoustics when updating parameters. This enhances the network's ability to comprehensively model speaker information, dialect information, text information, and paralinguistic information, thereby improving the content continuity, identity consistency, and expression stability of the trained target language generation network in multi-speaker continuous expression scenarios.
[0049] S50, receive generation condition prompts containing text to be generated and speaker prompts; when the generation condition prompts have dialect indication information, obtain dialect guidance statements according to the dialect indication information, and concatenate the dialect guidance statements to the beginning of the text to be generated. In this embodiment, the generation condition prompt represents a set of structured information input to the speech generation stage, which includes the text to be generated and speaker prompt information. The text to be generated is used to express the target semantic content, and the speaker prompt information is used to define the characteristics of the speaker, including identity, timbre attributes, or role type. The process of receiving the generation condition prompt corresponds to obtaining input data from an external interactive interface, business system, or application terminal, and parsing the format of the input data so that the text to be generated and the speaker prompt information can be distinguished and accessed independently in the internal representation structure.
[0050] The generated conditional prompts include dialect indication information, meaning that the input information contains a control field used to define the regional characteristics of pronunciation. This field can originate from user selection, business system configuration, or contextual inference results. The presence or absence of dialect indication information is determined through field detection or tagging. When dialect indication information is detected, the subsequent dialect guidance statement generation process is triggered. Dialect indication information can use discrete labels to represent different regional pronunciation categories, such as area codes, dialect category identifiers, or pronunciation pattern numbers.
[0051] Dialect guidance statements are obtained based on dialect indication information, which means mapping dialect indication information to specific textual expressions. These expressions guide the language generation network to form the corresponding dialect style during subsequent generation. The sources of dialect guidance statements can be typical expressions from a pre-set corpus, sentence fragments generated according to dialect feature rules, or short texts with dialect features generated by a language model. Dialect guidance statements typically contain representative pronunciation words, modal particles, or sentence structures, enabling the generation network to inherit these dialect features when processing subsequent text.
[0052] By appending dialect guidance statements to the beginning of the text to be generated, the input text is reconstructed at the semantic sequence level, placing the dialect guidance statements before the text to be generated, forming a new text sequence. This concatenation operation not only changes the text order but also affects the subsequent generation model's understanding of the context, ensuring that dialect features continue to play a role throughout the text generation process. After concatenation, the text to be generated structurally already contains dialect guidance information, thus eliminating the need for additional dialect control information in subsequent processing stages. This process maintains that speaker prompts do not participate in text concatenation, but only retain their original position within the generated conditional prompt structure, thereby ensuring the independence and consistency of the speaker's information and the text's semantic information.
[0053] In a fintech business environment, generated conditional prompts can include product name, description of return range, fee rate explanation, risk warning statements, and customer or customer service identification information. Dialect guidance information can be derived from user region settings or business rule configurations. After generating dialect guidance statements based on this guidance information, they are appended to the beginning of the text to be generated, ensuring that content related to amount descriptions, return descriptions, and risk warnings presents the corresponding regional pronunciation style during the speech generation process. In a healthcare business environment, generated conditional prompts can include symptom descriptions, health advice, rehabilitation reminders, and interaction role information. Dialect guidance information can be derived from the user's location or historical interaction records. The appended dialect guidance statements ensure regional pronunciation consistency in health-related expressions during the speech generation stage.
[0054] In one implementation, the generated condition prompts are stored in a structured field format. The text to be generated and the speaker prompt information field are stored in separate locations, while the dialect indication information field is represented by a Boolean identifier or category code. After receiving the generated condition prompts, each field is parsed and extracted. When dialect indication information is detected, the corresponding dialect expression template is retrieved from a preset corpus based on this information. The template is then used as a dialect guiding statement and appended to the beginning of the text to be generated to form a new text sequence. This new text sequence is then combined with the original speaker prompt information to form the updated generated condition prompts.
[0055] In another implementation, the dialect guidance statement does not rely on a fixed corpus set, but is generated based on dialect indication information through rules or a language generation model. For example, specific modal particles, sentence structures, or pronunciation features are selected according to the dialect category to form a short sentence, which is then appended to the beginning of the text to be generated as the dialect guidance statement. The appending method can be direct concatenation, or a separator can be added between the dialect guidance statement and the text to be generated to enhance the subsequent model's ability to distinguish different semantic segments.
[0056] This implementation introduces dialect indication information into the generation condition prompts and generates corresponding dialect guidance statements. The dialect guidance statements are then appended to the beginning of the text to be generated, so that the input text already carries clear dialect expression features before entering the language generation stage. This enables the unified expression of semantic content and regional pronunciation features in the subsequent generation process.
[0057] S60, the generation condition prompt is input into the target language generation network, the target text discrete identifier is output in an autoregressive manner, the corresponding target speech discrete identifier is output based on the target text discrete identifier, and the target speech discrete identifier is decoded into a target audio waveform.
[0058] In this embodiment, generating conditional prompts inputs into the target language generation network, meaning that the already organized conditional information is sent into the inference structure for speech generation. The generated conditional prompts carry text content, speaker information, and dialect guidance content. After entering the target language generation network, this information no longer exists as isolated data items but is uniformly mapped into conditional representations that can participate in sequence prediction. The target language generation network may include a conditional encoding unit, a context modeling unit, a text prediction unit, and a speech prediction unit. The conditional encoding unit receives the generated conditional prompts and completes the vector mapping of the discretized input. The context modeling unit performs temporal modeling on the input conditional representations. The text prediction unit outputs the target text discrete identifier based on the context state. The speech prediction unit outputs the target speech discrete identifier based on the text result and the context state. The units are connected in the order of conditional encoding, context update, text generation, and speech generation, ensuring a continuous transmission relationship between different types of information within the same round of inference.
[0059] The target text discrete identifiers are output in an autoregressive manner, meaning that when the target language generation network generates text results, the output at the current position depends on the results generated at the previous position and the current conditional context. The autoregressive approach aims to maintain the sequential consistency of the text generation process, allowing the target text discrete identifiers to accumulate position by position. The target text discrete identifiers are not static mappings of the complete text, but rather a progressively generated sequence of discrete representations. In implementation, the text prediction unit outputs a discrete text result at the current position, which is written into the current context representation and participates in the prediction of the next position. After continuously repeating this process, a complete sequence of target text discrete identifiers is formed. The reason for this approach is that text expressions in dialogue scenarios are context-dependent. If a one-time, complete output is used, it is easy to destroy the semantic constraints between the preceding and following parts. However, outputting in a position-by-position manner ensures that the text content remains consistent with the input speaker information, dialect information, and historical text content.
[0060] The output of the corresponding discrete target speech identifier based on the discrete target text identifier indicates that the text generation result is not directly used as the final output, but rather as a prerequisite for subsequent speech generation in the next stage. The discrete target text identifier carries the function of discrete expression of semantic content, while the discrete target speech identifier carries the function of discrete expression of reconstructable acoustic features of speech. They differ in function but are sequentially related. In implementation, the speech prediction unit receives the discrete target text identifier and, combined with the speaker information and dialect guidance information retained in the conditional coding unit, outputs the corresponding discrete target speech identifier position by position. This processing ensures that text content, speaker features, and dialect attributes work together to affect the speech expression, preventing a disconnect between semantic content and acoustic expression. For content such as fee explanations, profit announcements, risk warnings, and credit limit notifications in fintech businesses, this method helps maintain semantic consistency in numerical expressions, product names, and prompts; for content such as symptom descriptions, rehabilitation suggestions, medication reminders, and lifestyle recommendations in healthcare businesses, this method helps maintain semantic integrity and expression continuity.
[0061] Decoding the discrete target speech identifier into a target audio waveform means recovering the discrete acoustic representation into a playable continuous speech signal. The discrete target speech identifier itself is a discrete index sequence and cannot be directly used as audio output; therefore, it requires decoding. The decoding process can include two stages: spectrogram mapping and waveform synthesis. The spectrogram mapping stage converts the discrete target speech identifier into a continuous acoustic representation, such as spectral features, time-frequency representation, or other acoustic representations usable by the vocoder. The waveform synthesis stage recovers the time-domain waveform based on the acoustic representation to obtain the target audio waveform. In the implementation, a spectrogram mapping unit and a waveform synthesis unit can be set up. The spectrogram mapping unit receives the discrete target speech identifier and outputs the acoustic representation, while the waveform synthesis unit receives the acoustic representation and generates a continuous waveform. This processing establishes a clear division of labor between the semantic layer and the acoustic layer: the text discrete identifier is responsible for content expression, the speech discrete identifier is responsible for acoustic expression, and the target audio waveform is responsible for the final playback output.
[0062] The target language generation network can internally include a conditional encoding unit, a context modeling unit, a text prediction unit, and a speech prediction unit. The output of the conditional encoding unit is input to the context modeling unit. The state of the context modeling unit simultaneously connects to both the text prediction unit and the speech prediction unit. The discrete target text identifier output by the text prediction unit is fed back to the context modeling unit or directly input to the speech prediction unit. The speech prediction unit outputs the discrete target speech identifier, which then enters the spectrogram mapping unit and finally the waveform synthesis unit to generate the target audio waveform. In fintech applications, the input can be set to include generation conditional prompts containing product name, return range, fee structure, risk disclosure statements, and role information; the output is the target audio waveform broadcast to customers. In healthcare applications, the input can be set to include generation conditional prompts containing symptom descriptions, health advice, rehabilitation tips, and interaction role information; the output is the corresponding target audio waveform.
[0063] In one implementation, the conditional coding unit employs a combined embedding layer and positional coding layer structure to uniformly encode the text content, speaker information, and dialect guidance content in the generated conditional prompts. The context modeling unit uses a multi-layer attention structure to perform temporal modeling on the encoding results. The text prediction unit outputs the text category distribution at the current position, selects the discrete result with the highest score as the target text discrete identifier, and then sends this result to the speech prediction unit. The speech prediction unit combines the context state and the target text discrete identifier to output the corresponding target speech discrete identifier. The spectrogram mapping unit converts the target speech discrete identifier into a spectrogram representation, and the waveform synthesis unit recovers the target audio waveform based on the spectrogram representation. The number of discrete outputs, the number of context modeling layers, the embedding dimension, and the spectrogram representation resolution can be adjusted according to the sequence length and application environment.
[0064] In another implementation, the text prediction unit and the speech prediction unit adopt a shared context modeling structure. After the conditional prompts are generated and enter the shared modeling layer, the text prediction unit first outputs the discrete target text identifier, and then the speech prediction unit receives the discrete target text identifier and generates the discrete target speech identifier based on the shared context state. The spectrogram mapping unit can use a frame-by-frame mapping method to convert each set of discrete target speech identifiers into the corresponding acoustic representation; the waveform synthesis unit can use a parallel waveform synthesis method to generate the target audio waveform from all acoustic representations at once. This implementation is suitable for business environments with high real-time requirements.
[0065] This embodiment inputs the generation condition prompts into the target language generation network and gradually forms discrete identifiers for the target text in an autoregressive manner. Then, it generates corresponding discrete identifiers for the target speech based on the discrete identifiers for the target text and further decodes them into target audio waveforms. This allows the text content, speaker information, and dialect information to form a continuous transmission relationship in the same generation process, thereby improving the consistency between semantic expression and acoustic expression and enhancing the continuity and controllability of the final audio output.
[0066] In one embodiment, step S10 includes: S101, Obtain the original multi-speaker audio material, and perform audio enhancement and background noise suppression processing on the original multi-speaker audio material to obtain an enhanced audio segment; S102, the enhanced audio segment is processed by voice activity detection to obtain multiple voice segments, and the voice segments are segmented and merged according to the conversation boundary to obtain a short dialogue audio segment; S103, perform speaker separation processing on the short dialogue audio segment, and perform clustering processing based on the voiceprint embedding vector of the short dialogue audio segment to obtain speaker identifier; S104, Transcribe the short dialogue audio segment to obtain transcribed text; S105, perform sub-language event recognition on the short dialogue audio segment, and perform dialect retrieval and dialect recognition to obtain sub-language event annotation information and dialect annotation information; S106, construct a dialogue sample unit based on the short dialogue audio segment, the speaker identifier, the transcribed text, the sub-language event annotation information, and the dialect annotation information; S107, perform timestamp alignment processing on the dialogue sample units, and summarize the dialogue sample units that have completed timestamp alignment processing to obtain a multi-channel aligned metadata dataset.
[0067] In this embodiment, the original multi-speaker audio material represents a continuous set of speech data from a real interactive environment. This data includes speech content from multiple speakers within different time intervals, and may also include complex elements such as background music, environmental noise, crosstalk, double-tones, interruptions, and other intrusive elements. The acquisition stage does not require the input data to have a standardized structure; instead, it accepts various data sources, including audio files, streaming audio buffers, business call recordings, original podcast recordings, financial customer service recordings, and health consultation voice recordings. The purpose of performing audio enhancement and background removal processing is to compress irrelevant acoustic components such as environmental noise, reverberation, background music, and prompts from interfering with subsequent segmentation and recognition results. This can be achieved using a combination of spectral subtraction, masking enhancement, speech enhancement network inference, bandpass filtering, and noise reduction. Alternatively, a dual-branch structure can be used, with one branch suppressing steady-state noise and the other separating non-steady-state interference. The output is an enhanced audio segment. The enhanced audio segment does not require the complete removal of all background components, but rather the preservation of the temporal contour, articulation boundaries, and recognizability of the human voice, so that subsequent speech activity detection, speaker separation, transcription, and dialect recognition can be performed under unified acoustic conditions.
[0068] After the enhanced audio segment enters the speech activity detection process, the processing objective changes from continuous audio to multiple speech segments. Speech activity detection identifies speech start points, end points, and silence intervals, thus segmenting unstructured long audio into manageable short segments. Speech activity detection can employ energy and zero-crossing rate-based thresholding or a frame-level speech activity classification network. The input is the frame-level acoustic features of the enhanced audio segment, and the output is the speech presence marker for each frame. After obtaining multiple speech segments, segmentation and merging processing is performed in conjunction with conversation boundaries. Conversation boundaries are not simply pause boundaries, but a comprehensive result of semantic switching boundaries, turn-based switching boundaries, and topic-based switching boundaries. Segments with short intervals and semantic continuity are merged to form short dialogue audio segments; for positions with longer intervals, significant topic shifts, or changes in speaker turn-based switching, the boundaries are preserved. This results in short dialogue audio segments that control duration while retaining the semantic relationships within a complete interactive segment, facilitating subsequent multi-speaker sample organization.
[0069] After generating short dialogue audio segments, speaker separation processing is required, followed by clustering based on the speaker embedding vectors of the audio segments to form speaker identifiers. The purpose of speaker separation processing is to distinguish segments from different speakers within a short dialogue audio segment. This can be achieved by combining a speaker activity detection unit and a speaker embedding extraction unit. The speaker activity detection unit defines the speaker boundary for each time window, while the speaker embedding extraction unit extracts fixed-length identity vectors from the segmented segments. The speaker embedding vectors reflect the stable acoustic characteristics of the speaker. Clustering processing groups segments belonging to the same speaker into the same cluster based on the distance between vectors, thus outputting speaker identifiers. Clustering methods can include hierarchical clustering, density clustering, or centroid iterative clustering. The clustering input is a set of speaker embedding vectors, and the output is the correspondence between segments and speaker identities. After speaker separation and clustering processing, each valid segment in the short dialogue audio segment can be bound to a traceable speaker identifier, thus establishing a clear correspondence between speech content and speaker entities during subsequent metadata construction.
[0070] Transcription processing of short dialogue audio segments yields transcribed text. The input to transcription processing is the short dialogue audio segment and speaker boundary information; the output is a text sequence corresponding to the speech content. Transcription processing can employ an end-to-end speech recognition model or a combination of acoustic models, word graph decoding units, and language constraint units. The output text not only records semantic content but also serves as a semantic reference in subsequent multi-channel alignment. To improve the readability of business text and the recognition rate of specialized terms, a domain-specific vocabulary enhancement mechanism can be introduced. This constraint can be applied to texts such as credit limits, repayment plans, fee ranges, risk levels, and product names in fintech businesses, and to texts such as symptom names, health indicators, drug names, exercise suggestions, and rehabilitation movements in healthcare businesses. The resulting transcribed text not only expresses the speech content but also forms associative data with subsequent dialect information, paralinguistic information, and time stamps.
[0071] The short dialogue audio segment undergoes paralinguistic event recognition processing, followed by dialect retrieval and recognition processing, outputting paralinguistic event annotation information and dialect annotation information. Paralinguistic event recognition processing focuses on non-purely semantic vocalizations such as laughter, sighs, inhalations, coughs, prolongations, and obvious pauses. In implementation, short-time spectral features, prosodic features, and energy envelope features are extracted first, and then the paralinguistic category for each time interval is output through an event classification unit. Paralinguistic event annotation information includes the event category, occurrence time, and duration. Dialect retrieval and dialect recognition processing need to be distinguished into two levels. Dialect retrieval processing is used to filter out segments that may contain dialect features from the short dialogue audio segment, while dialect recognition processing is used to assign specific dialect categories to candidate segments. This two-level processing avoids performing high-complexity recognition directly on the entire audio. The dialect recognition results form dialect annotation information, which at least includes the dialect category and the corresponding time interval. For local customer service recordings in fintech businesses, there are significant differences in dialect expressions among customers from different regions when discussing loan inquiries, wealth management advice, and overdue payment communication. Similarly, for health service consultations in healthcare businesses, there are noticeable differences in pronunciation among users from different regions when describing symptoms and expressing themselves in daily life. The introduction of paralinguistic event annotation information and dialect annotation information means that short dialogue audio segments no longer only contain semantic and identity information, but also include tone of voice and regional pronunciation attributes.
[0072] Dialogue sample units are constructed based on short dialogue audio segments, speaker identifiers, transcribed text, paralinguistic event annotations, and dialect annotations. This represents organizing the aforementioned heterogeneous data into a unified record corresponding to the same interaction segment. Here, the dialogue sample unit is not a single-field result, but a multi-field, traceable, and alignable data object. The short dialogue audio segment provides the original acoustic carrier, the speaker identifier provides clues to the subject's identity, the transcribed text provides semantic content, the paralinguistic event annotations provide communication status information, and the dialect annotations provide regional pronunciation attributes. During the construction process, the short dialogue audio segment is used as the main index unit, and the remaining fields are bound according to time intervals and segment boundaries, thereby ensuring that all types of data within the same dialogue sample unit originate from the same interaction segment. After this organization, each dialogue sample unit can independently represent a complete or relatively complete multi-speaker interaction segment.
[0073] The dialogue sample units are timestamped and then aggregated to obtain a multi-channel aligned metadata dataset. The purpose of timestamping is to map audio time, text time, speaker activity time, paralinguistic event time, and dialect occurrence time onto the same timeline. This is achieved by using the frame index of enhanced audio segments or short dialogue audio segments as a unified time reference, performing forced alignment on the transcribed text to establish interval correspondences between text units and audio frames; speaker boundaries, paralinguistic event boundaries, and dialect segment boundaries are mapped to the same frame index system. After timestamping, speech content, speaker changes, paralinguistic events, and dialect features occurring within the same time interval can be structurally presented synchronously. Subsequently, multiple timestamped dialogue sample units are aggregated to form a multi-channel aligned metadata dataset. "Multi-channel" refers to parallel data dimensions such as audio channels, text channels, identity channels, paralinguistic channels, and dialect channels; the metadata dataset represents a collection of samples organized according to a unified index, format, and time reference. The resulting dataset not only preserves the original acoustic content, but also the semantics, identity, expression methods, and regional attributes surrounding the acoustic content, providing directly accessible data input for subsequent discretization, sequence construction, and network training.
[0074] Speech activity detection processing can employ a frame-level classification structure, with input being short-time acoustic features of the enhanced audio segment and output being a speech presence marker sequence. Speaker separation processing can utilize a speaker boundary detection unit and a speaker embedding extraction unit, with input being a short dialogue audio segment and output being speaker segment boundaries and speaker embedding vectors. Transcription processing can employ a speech recognition structure, with input being a short dialogue audio segment and output being the transcribed text. Paralinguistic event recognition processing can employ an event classification structure, with input being the time-frequency and prosodic representations of the short dialogue audio segment and output being paralinguistic event annotation information. Dialect retrieval and dialect recognition processing can employ a two-level classification structure, with input being a short dialogue audio segment or candidate segment and output being dialect annotation information. In fintech businesses, these processing units can take in recordings of customer consultations, product explanations, and risk warnings, and output transcribed text containing financial keywords such as amount, term, rate, and risk level, along with relevant identity and dialect annotations. In healthcare businesses, they can take in recordings of health consultations, rehabilitation guidance, and lifestyle advice, and output transcribed text containing symptom terms, indicator terms, and advice, along with relevant identity and dialect annotations.
[0075] This embodiment performs audio enhancement and background removal processing, speech activity detection, segmentation and merging, speaker separation, transcription, paralinguistic event recognition, dialect retrieval and dialect recognition on the original multi-speaker audio materials. Based on this, dialogue sample units are constructed, and then timestamp alignment processing is performed on the dialogue sample units and summarized to form a multi-channel aligned metadata dataset. This enables audio content, semantic content, speaker information, paralinguistic information and dialect information to establish a clear correspondence under the same time reference, thereby improving the consistency of multi-dimensional data organization, the completeness of sample expression and the data availability in subsequent processing stages.
[0076] In one embodiment, step S20 above includes: S201, obtain dialogue sample units from the multi-channel aligned metadata set, and extract short dialogue audio segments, speaker identifiers, transcribed text, paralinguistic event annotation information and dialect annotation information from the dialogue sample units; S202, perform residual vector quantization on the short dialogue audio segment to decompose the short dialogue audio segment into a multi-codebook acoustic tagging layer; S203, perform hierarchical discrete coding on the multi-codebook acoustic tag layer to generate discrete speech identifiers; S204, perform text tokenization on the transcribed text to generate discrete text identifiers; S205, mark and map the speaker identifier, the sub-language event annotation information and the dialect annotation information respectively, and generate speaker discrete identifier, sub-language discrete identifier and dialect discrete identifier accordingly; S206, combine the speech discrete identifier, the speaker discrete identifier, the text discrete identifier, the secondary language discrete identifier, and the dialect discrete identifier to obtain a discrete identifier set.
[0077] In this embodiment, after the dialogue sample units in the multi-channel aligned metadata set enter the discretization processing stage, the processing target is transformed from a heterogeneous combination of continuous audio, text, identity information, event information, and dialect information into a unified set of discrete representations. This transformation is not a simple numbering; rather, it compresses data of different modalities, structures, and temporal granularities into a set of discrete tags that are indexable, sortable, combinable, and parallelizable. Each dialogue sample unit contains at least a short dialogue audio segment, speaker identifier, transcribed text, paralinguistic event annotation information, and dialect annotation information. Extracting each field after acquiring the dialogue sample unit is to separate continuous acoustic information, symbolic semantic information, and category attribute information into mutually independent but temporally corresponding processing objects. The short dialogue audio segment carries the function of speech acoustic content, the speaker identifier distinguishes the speaker, the transcribed text expresses semantic content, the paralinguistic event annotation expresses the communication state, and the dialect annotation expresses regional pronunciation attributes. Only by extracting these objects separately can we adopt targeted discretization methods to avoid forcibly putting data that should use different processing logic into the same transformation process, which would cause a loss of expression.
[0078] When performing residual vector quantization on short dialogue audio segments, the input is a continuous acoustic waveform or a pre-processed continuous acoustic representation, and the output is a multi-codebook acoustic labeling layer. The role of residual vector quantization is to decompose continuously changing acoustic features into multiple levels of discrete indices. Each index represents a portion of acoustic information, and the multiple layers are superimposed to collectively represent the complete speech content. In implementation, acoustic coding units, residual decomposition units, and quantization unit groups can be set up. The acoustic coding unit is responsible for converting the short dialogue audio segment into a frame-level or segment-level continuous feature representation. The residual decomposition unit is responsible for calculating the residual error after each round of quantization. The quantization unit group indexes the current residual according to multiple codebooks. The first quantization layer records the main acoustic contours, and subsequent quantization layers gradually supplement finer-grained pronunciation details, prosodic variations, and timbre information. The resulting multi-codebook acoustic labeling layer can be understood as a set of discrete indices stored in parallel on multiple quantization layers, with each codebook index position corresponding to a portion of acoustic features. This processing no longer retains the continuous waveform form, but transforms it into a hierarchical discrete representation, enabling the speech content to be uniformly organized with discrete information such as text, identity, and dialect in subsequent temporal expressions. For voice messages in fintech businesses, such as revenue announcements, credit limit notifications, fee explanations, and risk disclosures, this quantitative expression helps to compress the rhythm of numerical announcements, the pronunciation of product names, and the outline of customer service tone into a discrete index system. For voice messages in healthcare businesses, such as symptom descriptions, medication reminders, rehabilitation suggestions, and health Q&A, this quantitative expression helps to retain short pauses, stressed syllables, and semantic emphasis in a multi-layered discrete structure.
[0079] After hierarchical discrete encoding, the multi-codebook acoustic tag layer generates discrete speech identifiers. The relationship between the multi-codebook acoustic tag layer and the discrete speech identifiers is not a simple renaming; rather, it transforms a hierarchical organizational structure into a unified and callable discrete speech representation. The purpose of hierarchical discrete encoding is to organize the indices from multiple codebook layers into a complete discrete speech identifier representation according to a predetermined order, allowing the quantization results from different levels to be called as a whole. In implementation, the indices of each layer can be arranged in a time-position priority manner or a codebook-level priority manner, and then the arrangement result is written into a fixed-format discrete speech identifier using a unified encoding rule. After the discrete speech identifiers are formed, continuous speech has departed from its waveform form and transformed into a discrete sequence with a sequential structure and encoding rules. This discrete sequence retains the recoverable acoustic information of the speech and is suitable for unified management with other types of discrete identifiers.
[0080] After text transcription is tokenized, discrete text identifiers are generated. The object of text tokenization is a sequence of text characters, words, or semantic fragments; the goal is to convert semantic content into a sequence of discrete symbols. Implementation can include text segmentation units and text mapping units. The text segmentation unit performs character-level, word-level, sub-word-level, or mixed segmentation on the transcribed text according to preset tokenization rules, transforming the semantic content into a set of sequentially arranged text units. The text mapping unit maps each text unit to a corresponding discrete index according to a preset tokenization table; multiple indices sequentially form the discrete text identifier. In fintech business environments, product names, return ranges, term expressions, fee expressions, and risk level expressions often contain a large number of business-specific terms. Text tokenization can maintain the integrity of these financial keywords during segmentation, reducing semantic bias caused by the disassembly of financial terms. In healthcare business environments, symptom names, health indicators, drug names, and lifestyle recommendations can also maintain terminology integrity through customized tokenization rules, enabling discrete text identifiers to more accurately carry semantic information.
[0081] Speaker identifiers, paralinguistic event annotations, and dialect annotations are mapped separately to generate discrete speaker identifiers, discrete paralinguistic identifiers, and discrete dialect identifiers, respectively. Although these three objects are all label-type information, they express different meanings, therefore the mapping process must be performed separately. The goal of speaker identifier mapping is to transform different speakers into stable identity indexes, ensuring consistency in the representation of the same speaker across different dialogue sample units. In implementation, an identity mapping unit can be set up to assign a unique index to each speaker identifier and output the discrete speaker identifier. The goal of paralinguistic event annotation mapping is to transform paralinguistic events such as laughter, sighs, inhalations, coughs, prolongations, and pauses into discrete markers that can participate in sequence organization. In implementation, an event mapping unit can be set up to create a discrete index table for each type of event and output the discrete paralinguistic identifier based on the event category and location. The goal of dialect annotation mapping is to transform pronunciation attributes from different regions into discrete category indexes, allowing regional pronunciation differences to participate in subsequent expressions symbolically. In implementation, a dialect mapping unit can be set up to output the discrete dialect identifier based on dialect category, regional affiliation, or pronunciation pattern. After this processing, the speaker, communication status, and regional pronunciation attributes are all transformed from the original descriptive form into a unified index form, which facilitates the formation of a unified set together with discrete speech identifiers and discrete text identifiers.
[0082] Combining discrete speech tags, speaker discrete tags, text discrete tags, paralinguistic discrete tags, and dialect discrete tags yields a discrete tag set, representing the organization of five different functional discrete results into a structured data object. This discrete tag set is not a random stack, but a unified representation structure with corresponding relationships. In implementation, set organization units can be set up to write the five types of discrete tags into the same data record and save their correspondence with the current dialogue sample unit. The set organization unit can adopt a field-based structure or a partitioned structure. The field-based structure writes the discrete speech tags, speaker discrete tags, text discrete tags, paralinguistic discrete tags, and dialect discrete tags into separate fields; the partitioned structure divides regions according to category to record different discrete results. Regardless of the organization method, the discrete tag set must simultaneously meet the requirements of traceability, indexability, and reusability. This discrete tag set allows the originally continuous and heterogeneous data to enter a unified representation space in form. Subsequent processing stages no longer need to deal with multiple data formats such as waveforms, text, and tags separately, but can directly deal with standardized discrete expression results.
[0083] This embodiment extracts short dialogue audio segments, speaker identifiers, transcribed text, paralinguistic event annotation information, and dialect annotation information from the dialogue sample unit within the multi-channel aligned metadata dataset. It then generates discrete speech identifiers, discrete speaker identifiers, discrete text identifiers, discrete paralinguistic identifiers, and discrete dialect identifiers, respectively. Finally, it organizes the five types of discrete results into a unified discrete identifier set, transforming the originally structurally distinct continuous speech, text semantics, identity information, communication status information, and regional pronunciation information into a unified and indexable expression form, thereby improving the consistency and organization of multidimensional data expression.
[0084] In one embodiment, step S30 above includes: S301, based on the time alignment information corresponding to the sub-language discrete identifier, locate the pronunciation insertion position within the text discrete identifier; S302, the sub-language discrete identifier is placed into the pronunciation insertion position to obtain the text discrete identifier after inserting the sub-language discrete identifier; S303, using the speaker discrete identifier as the starting unit, and concatenating the dialect discrete identifier to the end of the starting unit to generate a prefix sequence; S304, according to the pronunciation sequence, the text discrete identifier after inserting the sub-language discrete identifier is concatenated to the end of the prefix sequence to generate the transition sequence; S305, the discrete speech identifier is concatenated to the end of the transition sequence according to the pronunciation sequence to obtain an interleaved sequence.
[0085] In this embodiment, the time alignment information corresponding to the paralinguistic discrete identifier is used to describe the occurrence position and duration of the paralinguistic event on a unified time axis. The paralinguistic discrete identifier itself represents non-semantic vocal events such as laughter, sighs, inhalations, coughs, and prolonged pauses. The time alignment information indicates the actual time interval in which such events occur in the speech expression. Based on the time alignment information corresponding to the paralinguistic discrete identifier, the pronunciation insertion position is located within the text discrete identifier. This means that the paralinguistic discrete identifier is not arbitrarily inserted, but rather the specific writing position is determined in the text discrete identifier based on the synchronization relationship between the paralinguistic event and the text pronunciation content. The text discrete identifier originally carries the discrete expression of semantic content. When a paralinguistic event occurs before, after, or in the middle of a certain text position, it is necessary to find the position that best matches the paralinguistic event through the alignment relationship between the time interval and the text discrete units. In implementation, the text discrete identifier can be regarded as a discrete sequence arranged in the order of pronunciation. The time alignment information is mapped to the index space of the discrete sequence, and then the pronunciation insertion position is determined based on the start time, end time, and coverage range of the paralinguistic event. For paralinguistic events occurring at the beginning of a sentence, the insertion position will fall near the beginning of the text discrete marker; for paralinguistic events occurring in the middle of a sentence, the insertion position will fall between the corresponding text discrete units; and for paralinguistic events occurring at the end of a sentence, the insertion position will fall after the text discrete marker. The resulting pronunciation insertion positions are not abstract positional descriptions, but rather sequential positional descriptions consistent with the actual speech pronunciation time.
[0086] Inserting paralinguistic discrete identifiers into pronunciation insertion positions yields text discrete identifiers after the insertion. This means embedding paralinguistic discrete identifiers into the sequential structure of the original text discrete identifiers, transforming the text discrete identifiers from purely semantic expressions into a composite expression sequence with non-semantic vocalization information. The insertion action is not a simple addition of explanations, but rather a rewriting of the sequence relationship within the discrete sequence, creating a clear local correspondence between paralinguistic events and adjacent text content. Implementation can be achieved using a sequence insertion method, writing the paralinguistic discrete identifiers at specified indices; or using a dual-positional information encoding method, adding paralinguistic event positional encoding to the original text discrete identifier sequence before outputting the reorganized text discrete identifiers. After obtaining the text discrete identifiers with inserted paralinguistic discrete identifiers, the text content no longer only represents words or semantic fragments, but also carries communicative tone, breath state, or emotional expression state. For interactive statements in fintech businesses, such as explanations of returns, risk disclosures, repayment reminders, and anti-fraud verification, this inserted text discrete identifier can embed expressions such as pauses, emphasis, hesitation, and drawn-out tones into the vicinity of keywords such as rates, terms, amounts, and risk levels. For expressions in healthcare businesses, such as symptom descriptions, medication reminders, rehabilitation suggestions, and follow-up responses, this insertion method can embed communication traces such as sighs, inhalations, coughs, and pauses into the vicinity of symptom words, action words, and reminder words, thereby enabling the text expression to have a higher degree of communication fidelity at the sequence level.
[0087] Using the speaker discrete identifier as the starting unit and appending the dialect discrete identifier to the end of the starting unit, a prefix sequence is generated. This indicates that before the subsequent overall expression begins, a sequential expression structure of speaker subject information and regional pronunciation attribute information is established. The speaker discrete identifier serves as a speaker subject marker, indicating which speaker the current content corresponds to; the dialect discrete identifier serves as a regional pronunciation attribute marker, indicating the pronunciation style that the current content should follow. Using the speaker discrete identifier as the starting unit means that the entire sequence first provides speaker subject constraints at the beginning, and then adding the dialect discrete identifier means that pronunciation attribute constraints are further provided after the subject constraints. The resulting prefix sequence is not arbitrarily arranged, but rather fixes identity attributes and dialect attributes as preconditions, ensuring that subsequent content expression always revolves around this subject and pronunciation attribute. In implementation, the speaker discrete identifier can be placed at the beginning of the prefix sequence, followed immediately by the dialect discrete identifier, resulting in a prefix sequence longer than the original starting unit. After the prefix sequence is formed, the current expression unit already possesses basic constraint information about who is speaking and in which dialect.
[0088] Following the phonological sequence, the text discrete identifiers after inserting the paralinguistic discrete identifiers are concatenated to the end of the prefix sequence to generate a transition sequence. This transition sequence represents placing the previously rewritten semantic expression after identity and dialect constraints, ensuring that subject information, dialect information, text content, and paralinguistic information are arranged sequentially in temporal logic. The phonological sequence here does not represent an abstract ordering rule, but rather the sequential relationship in actual speech expression. When the current expression actually occurs, the speaker is determined first, followed by the pronunciation style, and then the specific content spoken and its paralinguistic modifications. By appending the text discrete identifiers after inserting the paralinguistic discrete identifiers to the end of the prefix sequence in this order, the resulting transition sequence can fully express the identity, dialect, semantic, and non-semantic communication information of a particular speaker unit. The transition sequence is called transitional because acoustic-level speech discrete identifiers have not yet been added; the sequence already possesses readable semantics and identity attributes, but it does not yet fully carry the information for recovering the acoustic expression. In implementation, a sequence concatenation method can be directly used, appending the entire text discrete identifier after inserting the sub-language discrete identifier to the end of the prefix sequence. Alternatively, it can be concatenated in segments according to the text length and pronunciation boundaries. However, the concatenated order must maintain the arrangement relationship of identity first, dialect in the middle, and semantics and sub-language content last.
[0089] By concatenating discrete speech markers to the end of a transition sequence according to the phonological sequence, an interleaved sequence is obtained. This represents the addition of acoustic representations corresponding to the existing sequence expression, which already contains information about identity, dialect, text, and paralinguistics. This extends the entire sequence from the semantic and attribute layers to the acoustic layer. Discrete speech markers serve to recover phonological details, timbre variations, and prosodic contours. Placing discrete speech markers at the end of the transition sequence is not a simple addition, but rather establishes a strict correspondence between the preceding speaker discrete markers, dialect discrete markers, and the text discrete markers after the insertion of paralinguistic discrete markers, and the subsequent acoustic expression. In this interleaved sequence, the first part provides the speaker and phonological mode, the middle part provides semantic content and non-linguistic events, and the last part provides specific acoustic discrete representations. The interleaved sequence thus becomes a temporal structure with a unified multi-layered information organization. Any expression in the interleaved sequence can be traced back to who uttered the speech, what dialect was used, what content was said, where the paralinguistic events were included, and the corresponding acoustic manifestations. In fintech operations, when customers inquire about returns, customer service explains fees, customers confirm redemption rules, and customer service provides supplementary risk warnings, the identity, dialect, text, paralinguistic information, and voice used in different rounds of expression can be organized into a continuous, interwoven sequence according to the order of pronunciation. This ensures that financial keywords such as return ranges, fee ranges, risk levels, and term descriptions are synchronized with the tone of the communication. In healthcare operations, when users describe symptoms, the server provides health advice, users supplement lifestyle information, and the server provides rehabilitation reminders, the identity information, dialect information, text information, paralinguistic information, and voice information used in different rounds can also be continuously organized within a unified temporal structure, ensuring consistency in the expression of symptoms, reminders, and emotional states.
[0090] This embodiment locates the pronunciation insertion position within the text discrete identifier based on the time alignment information corresponding to the secondary language discrete identifier, and places the secondary language discrete identifier into that position, thus establishing a clear positional correspondence between the secondary language information and the text semantic content. Then, the speaker discrete identifier, dialect discrete identifier, text discrete identifier after inserting the secondary language discrete identifier, and speech discrete identifier are sequentially concatenated according to the pronunciation time sequence to form a unified interleaved sequence. This allows identity information, regional pronunciation attributes, semantic content, non-verbal communication information, and acoustic expression to be organized synchronously in the same temporal structure, thereby improving the consistency and correspondence accuracy of multidimensional information expression.
[0091] In one embodiment, step S40 above includes: S401, construct a first-stage training sample containing single-speaker samples and dialogue samples based on the interleaved sequence, and construct a second-stage training sample based on multi-speaker dialogue samples containing dialect annotation information and sub-language event annotation information in the interleaved sequence. S402, using the first-stage training samples, perform first-stage supervised fine-tuning of the pre-trained language generation network to obtain the intermediate language generation network; S403, The intermediate language generation network is subjected to second-stage supervised fine-tuning using the second-stage training samples; S404, during the first stage of supervised fine-tuning and the second stage of supervised fine-tuning, a course learning strategy is adopted to provide training samples in the following order: from short single-speaker samples, to multi-speaker samples, to overlapping speech samples and long context samples, and finally to samples containing dialect annotation information and paralinguistic event annotation information. S405, In the second stage of supervised fine-tuning, the current prediction position of the intermediate language generation network on the interleaved sequence is located according to the autoregressive prediction order; S406, Identify the discrete speech identifiers before the current prediction position in the interleaved sequence, and select a portion of the discrete speech identifiers from the discrete speech identifiers before the current prediction position according to a preset probability value. S407, Randomly discard selected discrete speech identifiers and input the processed interleaved sequence into the intermediate language generation network; S408, obtain the output of the intermediate language generation network based on the processed interleaved sequence, and construct speech prediction cross-entropy error term, speaker consistency error term, paralinguistic event classification error term, dialect classification error term and reconstruction error term based on the output result respectively; S409, combine the speech prediction cross-entropy error term, the speaker consistency error term, the paralinguistic event classification error term, the dialect classification error term, and the reconstruction error term to construct a multi-task joint error function; S410, the intermediate language generation network is updated using the multi-task joint error function until the preset training termination condition is met, thus obtaining the target language generation network.
[0092] In this embodiment, interleaved sequences are input into the pre-trained language generation network for supervised fine-tuning. This means that interleaved sequences, which have already been organized with multidimensional discrete information, are used as training samples to readjust the parameters of the pre-trained language generation network for multi-speaker continuous speech expression tasks. The interleaved sequences simultaneously contain discrete speaker identifiers, dialect identifiers, text identifiers, paralinguistic identifiers, and speech identifiers. Different categories of discrete identifiers are arranged according to the pronunciation sequence, so that the same sequence expression simultaneously contains identity information, regional pronunciation information, semantic information, non-linguistic event information, and acoustic expression information. When the pre-trained language generation network receives interleaved sequences, it does not model a single modality, but rather models the conditional dependencies of multidimensional discrete information in a unified temporal sequence. The purpose of supervised fine-tuning is to transform the network from a general sequence prediction capability to a speech generation and expression capability that adapts to the joint participation of multiple speakers, dialects, and paralinguistics, so that the network parameters form a stable response to speaker switching, dialect changes, text content continuity, and paralinguistic insertion positions.
[0093] A pre-trained language generation network can include an input embedding unit, a positional encoding unit, a context modeling unit, and an output prediction unit. The input embedding unit maps discrete identifiers of different categories to a vector space of uniform dimension. The positional encoding unit overlays temporal positional information into the input representation, enabling the network to distinguish semantic differences of the same discrete identifier at different locations. The context modeling unit captures long-range dependencies and can be composed of multiple stacked self-attention layers and feedforward transform layers, or a combination of gated temporal layers and self-attention layers. Multiple self-attention layers establish dependencies between speaker switching, dialect continuation, text cohesion, and paralinguistic insertions globally across interleaved sequences. The feedforward transform layer performs nonlinear mapping on context-related representations. The output prediction unit outputs the category distribution of the discrete result at the current position based on the context representation preceding the current position. If a multi-layer stacked structure is used, the number of layers in the context modeling unit can range from several to dozens, the embedding dimension can be set to hundreds of dimensions, and the number of attention heads can be configured according to processing granularity and computational resources. Residual connections and normalization structures can maintain deep training stability between layers. The input to the pre-trained language generation network is the discrete identifier index in the interleaved sequence, and the output is the predicted distribution or intermediate representation of the corresponding position.
[0094] The first stage of training samples, consisting of single-speaker and dialogue samples, is constructed based on interleaved sequences. The second stage of training samples, based on multi-speaker dialogue samples containing dialect and paralinguistic event annotations within the interleaved sequences, represents a two-tiered organization of training data. The first stage training samples allow the network to learn single-turn expressions, basic semantic-to-acoustic correspondences, and relatively simple speaker transition rules. Single-speaker samples can be derived from segments in the interleaved sequences containing only one speaker, while dialogue samples can be derived from low-complexity interactive segments with turn transitions. The second stage training samples introduce more complex multi-speaker interactions and require explicit dialect and paralinguistic event annotations in the samples, enabling the network to further learn complex turn variations, dialect control, and paralinguistic control. The two-stage training samples are not simply collected in batches but are hierarchically organized based on sample complexity, number of speakers, dialect richness, paralinguistic event density, and context length. This approach prevents the network from directly encountering all complex interaction conditions in the early stages of training, allowing it to learn complex control relationships only after mastering basic expression rules, thus improving training convergence stability.
[0095] The pre-trained language generation network is fine-tuned using the first-stage training samples to obtain an intermediate language generation network. This means the pre-trained language generation network updates its parameters on a relatively simple data distribution. During the first-stage supervised fine-tuning, the interleaved sequence is segmented into training segments, which are input into the network according to a preset length. The real discrete result at each position serves as the supervision target, and the error between the predicted distribution and the real target is calculated. The network parameters are then updated through backpropagation. After the first-stage processing, the intermediate language generation network is obtained. This intermediate language generation network already possesses basic speaker condition modeling capabilities, basic text-to-speech discrete correspondence capabilities, and simple round-by-round continuous expression capabilities in terms of parameter states.
[0096] The intermediate language generation network is fine-tuned under second-stage supervision using second-stage training samples. This means retraining by introducing more complex dialogue samples based on the parameters from the first stage. The second-stage samples contain more complex speaker turn alternations, richer dialect categories and paralinguistic event categories, and a longer contextual range. During this second-stage supervision, the intermediate language generation network no longer only learns basic generative relations but also learns the joint constraints between discrete identifiers of different categories in complex dialogues. This allows the intermediate language generation network to transition from a basic modeling state in terms of parameter distribution to a target language generation network oriented towards complex, realistic interactive states.
[0097] In both the first and second phases of supervised fine-tuning, a curriculum learning strategy is employed, meaning that the input order of training samples is organized according to an increasing complexity principle. This curriculum learning strategy is not an additional category label, but rather a training scheduling rule. Training samples are provided progressively, starting with short single-speaker samples, then moving to multi-speaker samples, then to overlapping speech samples and long context samples, and finally to samples containing dialect annotations and paralinguistic event annotations. This allows the network to first learn basic conditional dependencies under short sequences, single-subject, and low-interference environments, and then gradually adapt to speaker switching, context expansion, overlapping speech, dialect control, and paralinguistic control. Short single-speaker samples are used to establish basic mapping relationships between discrete identifiers; multi-speaker samples are used to establish identity transformation relationships; overlapping speech samples and long context samples are used to enhance cross-round stability; and samples containing dialect annotations and paralinguistic event annotations are used to establish multi-dimensional control capabilities. The curriculum learning strategy does not change the content of individual samples but alters the order in which samples enter the training process, thereby controlling the rate of increase in training difficulty.
[0098] In the second stage of supervised fine-tuning, the current prediction position of the intermediate language generation network on the interleaved sequence is located according to the autoregressive prediction order. This means that conditional predictions are performed sequentially for each position to be predicted in the sequence. The current prediction position is the position of the discrete result being solved at the training time. All discrete identifiers before the current position can be considered as contextual conditions. The network can only utilize the valid inputs before the current position at the current position and does not utilize the discrete results after the current position. This ensures that the training objective remains consistent with the autoregressive conditions in the inference stage. After locating the current prediction position, the discrete speech identifiers before the current position can be identified from the interleaved sequence. The discrete speech identifiers in the interleaved sequence have clear category attributes, so the recognition process can be completed through category filtering, and the filtering output is a set of historical discrete speech identifiers.
[0099] Based on a preset probability value, a subset of discrete speech identifiers is selected from those preceding the current prediction position, and these selected identifiers are randomly discarded. This represents the degree to which the network relies on historical acoustic details during training. The preset probability value determines whether each historical discrete speech identifier should be retained. In implementation, Bernoulli sampling can be used to generate retention or discard markers for each historical discrete speech identifier, or the number of candidates can be determined first according to a preset ratio, and then a specific position can be randomly selected. The selected discrete speech identifiers for random discarding can be replaced by empty markers, masked markers, or directly removed from the visible context in the input. The random discarding process is limited to the discrete speech identifiers preceding the current prediction position and does not apply to speaker discrete identifiers, dialect discrete identifiers, text discrete identifiers, or paralinguistic discrete identifiers. The reason for this design is that historical discrete speech identifiers carry strong local acoustic repetition information. Without constraints, the network easily relies on the surface patterns of preceding speech to complete the current position prediction, weakening the comprehensive utilization of text content, identity conditions, dialect conditions, and paralinguistic conditions. By randomly dropping out data, the network is forced to continue predicting within an incomplete historical speech context, thereby enhancing its ability to jointly model multidimensional conditions.
[0100] After the processed interleaved sequence is input into the intermediate language generation network, the network outputs the prediction result for the current position. The output is not a single numerical value, but rather a category distribution and intermediate representation of the discrete results for the current prediction position. Based on the output, speech prediction cross-entropy error terms, speaker consistency error terms, paralinguistic event classification error terms, dialect classification error terms, and reconstruction error terms are constructed to represent the training objective as a multi-part objective. The speech prediction cross-entropy error term constrains the correctness of the speech discrete results prediction for the current position and is calculated using the cross-entropy between the predicted distribution and the true discrete index. The speaker consistency error term constrains the acoustic style of the prediction result to be consistent with the current speaker's identity; similarity constraints can be applied to the network output representation and the speaker reference representation during calculation. The paralinguistic event classification error term constrains the network output's discriminability of paralinguistic categories, ensuring that the paralinguistic insertion position and its type are stably maintained. The dialect classification error term constrains the regional pronunciation attributes in the output representation, ensuring that dialect conditions are continuously expressed in sequence prediction. The reconstruction error term is used to constrain the closeness between the output representation and the reference acoustic representation, and can employ time-frequency domain distance metrics or sequence reconstruction distance metrics. Five types of error terms constrain the same network output from different perspectives, ensuring that the output simultaneously satisfies acoustic correctness, identity consistency, paralinguistic correctness, dialect correctness, and overall reconstruction consistency.
[0101] A multi-task joint error function is constructed by combining the speech prediction cross-entropy error term, speaker consistency error term, paralinguistic event classification error term, dialect classification error term, and reconstruction error term. This function represents a unified optimization objective that can be used for parameter updates, unifying multiple training objectives into a single comprehensive optimization objective. The combination can be achieved using a weighted summation method or a dynamic weight adjustment method. The weights of each error term can be adjusted based on the training stage, sample complexity, or validation set performance. For example, in the first stage of supervised fine-tuning, the weights of the speech prediction cross-entropy error term and the reconstruction error term can be increased; in the second stage of supervised fine-tuning, the weights of the speaker consistency error term, paralinguistic event classification error term, and dialect classification error term can be increased to adapt to the learning focus at different stages. After the multi-task joint error function is formed, the gradient is propagated to the parameters of each layer of the intermediate language generation network through backpropagation.
[0102] The intermediate language generation network is updated using a multi-task joint error function until a preset training termination condition is met, resulting in the target language generation network. This indicates that the training process achieves network parameter convergence through iterative optimization. Parameter updates can employ adaptive moment estimation optimizers, momentum optimizers, or gradient descent optimizers. The learning rate can be set to a small order of magnitude and decrease with each training epoch. The batch size can be selected based on the length of the interleaved sequences and available memory resources. During training, each batch of interleaved sequence samples undergoes current prediction location localization, historical speech discrete identifier filtering, random discarding, output result calculation, error term construction, multi-task joint error function combination, and parameter updates. The preset training termination condition can be determined by validation error stability, reaching the upper limit of the training epochs, parameter change amplitude falling below a threshold, or the joint error decrease rate falling below a threshold. Upon reaching the training termination condition, the target language generation network is output. The target language generation network is now capable of simultaneously handling speaker switching, text continuity, dialect control, paralinguistic control, and acoustic expression consistency in terms of parameters.
[0103] In a fintech business environment, interleaved sequences can include multiple rounds of content such as product descriptions, return ranges, fee schedules, repayment reminders, and risk disclosures. Online input should focus on covering expressions of amount, term, risk, and speaker switching. Output should maintain clear financial keywords and consistency with the user's identity and style. In a healthcare business environment, interleaved sequences can include symptom descriptions, health advice, rehabilitation reminders, and lifestyle guidance. Online input should focus on covering health-related vocabulary, paralinguistic events, and dialect variations. Output should maintain stable and clear health-related expressions and consistency with the user's identity and attributes.
[0104] This embodiment constructs phased training samples based on interleaved sequences, introduces a course learning strategy during the two-stage supervised fine-tuning process, and randomly discards some discrete speech identifiers before the current prediction position. Then, it constructs a multi-task joint error function to update parameters by combining speech prediction cross-entropy error term, speaker consistency error term, paralinguistic event classification error term, dialect classification error term, and reconstruction error term. This reduces the pre-trained language generation network's over-reliance on historical local acoustic patterns during training and enhances its comprehensive modeling ability of speaker information, text information, dialect information, and paralinguistic information. As a result, it improves the post-trained target language generation network's ability to retain continuous expressions from multiple speakers, multiple dialect expressions, and paralinguistic expressions.
[0105] In one embodiment, step S50 above includes: S501, Receive generation condition prompts containing text to be generated and speaker prompt information; S502, parse the generation condition prompt, and extract the text to be generated and the speaker prompt information; S503, the dialect indication information is detected for the generated condition prompt; when the generated condition prompt has dialect indication information, the dialect indication information is extracted. S504, Obtain the corresponding dialect guidance statement based on the dialect indication information; S505, the dialect guidance statement is appended to the beginning of the text to be generated, and the original text to be generated in the generation condition prompt is replaced with the text to be generated after appending the dialect guidance statement, while keeping the position of the speaker prompt information in the generation condition prompt unchanged.
[0106] In this embodiment, the generation condition prompt represents the input data structure provided for the subsequent speech generation stage, which contains at least the text to be generated and speaker prompt information. The text to be generated is used to limit the semantic content of the subsequent output, and the speaker prompt information is used to limit the identity attribute, role attribute, or timbre attribute of the speaker. Receiving the generation condition prompt containing the text to be generated and the speaker prompt information indicates that structured input is obtained from an external input interface, business service interface, interactive terminal, or task scheduling module, and the text content and subject information therein are imported into a unified prompt representation space. This input is not simply natural language text, but composite condition data that simultaneously carries semantic constraints and speaker subject constraints. After receiving and processing, the field boundaries need to be kept clear so that the text to be generated and the speaker prompt information can be accessed separately later.
[0107] Content parsing is performed on the generated conditional prompts to extract the text to be generated and speaker prompt information, which involves splitting, identifying, and classifying the data fields within the generated conditional prompts. The purpose of content parsing is to separate the semantic part from the main body of the mixed input, so that different categories of information can perform different functions in subsequent processing. This can be implemented using field parsing to identify predefined field names; structural template matching to decode the position of input in a fixed order; or marker boundary recognition to segment content with separator markers. After extracting the text to be generated, the original text order, phrase structure, and semantic expression are preserved. After extracting the speaker prompt information, the role name, identity number, voice subject label, or timbre description information are preserved. The result of this processing is a clear distinction between the text content and main body attributes in the generated conditional prompts, avoiding field confusion when writing dialect guidance content later.
[0108] Dialect indication information is detected for the generated condition prompts. When dialect indication information is present, it is extracted to determine whether a regional pronunciation attribute control item exists in the prompt structure. Dialect indication information can originate from explicit user selection, business region configuration, historical interaction context, region service tags, or role preset attributes. The purpose of the detection process is to identify whether the current task requires the introduction of regional pronunciation priors. If no dialect indication information is present, the original structure of the text to be generated remains unchanged; if dialect indication information is present, the process proceeds to the dialect guidance statement construction stage. This can be implemented through field existence detection, category identifier detection, region code detection, or task attribute determination. After extracting the dialect indication information, the output object is a control item that can be directly used to map dialect expressions, such as regional category tags, dialect category codes, pronunciation attribute indexes, or dialect text prompt identifiers. This processing ensures that dialect control is not dependent on the original semantic content of the text to be generated, but exists as an independent condition, thereby avoiding structural mixing of semantic content and regional pronunciation attributes.
[0109] The process of obtaining corresponding dialect guidance statements based on dialect indication information transforms abstract dialect control items into concrete guidance content that can directly participate in text expression. Dialect guidance statements provide prior knowledge of regional pronunciation to subsequent generation processes, enabling subsequent semantic content to inherit corresponding dialect features during expression. The sources of dialect guidance statements can be typical short phrases from a pre-set text set, representative expressions formed according to dialect category rules, or fixed sentence patterns generated based on regional language characteristics. The acquisition process requires a one-to-one correspondence between the output content and the dialect indication information, without altering the original semantics of the text to be generated; it only requires providing sufficient prior information at the beginning of the text to trigger dialect expression tendencies. After the correspondence is established, different dialect indication information can be mapped to different dialect guidance statements, thus creating regional pronunciation differences at the input level. For example, in fintech businesses, when customers in different regions receive profit explanations, fee notifications, credit limit reminders, or risk disclosures, dialect guidance statements can be used to make the broadcast content sound closer to local language habits in terms of pronunciation; in healthcare businesses, when users in different regions receive health advice, rehabilitation reminders, or symptom feedback voice messages, dialect guidance statements can be used to make the voice expression more in line with local communication styles.
[0110] By appending dialect-guided statements to the beginning of the text to be generated, dialect priors are written into the text sequence, ensuring that dialect-related expressions are given priority in subsequent text processing. This appending is not a simple text concatenation process; rather, it involves reconstructing the order of the text to be generated, placing the dialect-guided statements at the beginning and the text to be generated at the end, together forming a new textual unit. After appending, the text to be generated is no longer a purely semantic segment, but an enhanced text sequence carrying regional pronunciation guidance. The reason for using a pre-positional appending rather than a mid-positional insertion or a post-positional appending is that the pre-position position more easily establishes stable preconditions in subsequent processing stages, ensuring that the phonetic expression of subsequent content is guided by dialect from the outset. This sequential arrangement helps ensure consistent control over homographs, dialectal intonations, and local expression habits in subsequent processing.
[0111] The process replaces the original text to be generated in the conditional prompt with the text to be generated after concatenating the dialect-guided statement. This indicates that the concatenated text result needs to be written back into the conditional prompt to form an updated conditional input structure. The significance of this replacement is that the externally generated dialect-guided statement is not suspended as an independent addendum outside the prompt structure, but is directly incorporated into the original text field to be generated. This ensures that all subsequent processing steps based on the conditional prompt read the dialect-enhanced text content. After the replacement, the semantic capacity of the text field within the conditional prompt changes. The semantic content of the original text to be generated is still retained, but regional pronunciation prior information is added at the beginning. This approach eliminates the disconnect between text reconstruction and the prompt structure, ensuring that only one text field is always retained within the conditional prompt, avoiding object conflicts caused by the coexistence of the original text to be generated and the enhanced text.
[0112] Maintaining the speaker prompt information's position within the generated conditional prompts means that the field position, index position, or logical position of the speaker prompt information is not changed during dialect guidance statement concatenation and text replacement. The speaker prompt information serves as a constraint on the speaker's identity, while the dialect guidance statement serves as a priori function for pronunciation attributes; their roles in the generated conditional prompts differ. Maintaining the speaker prompt information's position ensures that the subject attribute constraints do not drift due to text reconstruction and avoids identity resolution errors caused by changes in field order during subsequent readings. In implementation, after the text fields are replaced, only the text field content is updated; the storage order, calling order, or index identifier of the speaker prompt information fields within the generated conditional prompts is not adjusted. This approach ensures that dialect enhancement of the text to be generated and speaker identity preservation are simultaneously achieved, maintaining a unified structure and stable access method for the generated conditional prompts after dialect injection.
[0113] This embodiment performs content parsing and dialect indication information detection on the generated condition prompts, obtains the corresponding dialect guidance statements based on the dialect indication information, and then concatenates the dialect guidance statements at the beginning of the text to be generated, replacing the original text to be generated in the generated condition prompts. At the same time, the position of the speaker prompt information in the generated condition prompts remains unchanged. This allows the generated condition prompts to introduce prior information on regional pronunciation while maintaining the stability of the main constraints, thereby improving the organizational consistency between semantic content, main information and dialect attributes in the input conditions.
[0114] In one embodiment, step S60 above includes: S601, The generation condition prompt is input into the target language generation network to trigger the autoregressive prediction mechanism of the target language generation network; S602, Under the autoregressive prediction mechanism, the target language generation network outputs discrete identifiers of the target text sequentially based on the generation condition prompts; S603, the discrete identifier of the target text is concatenated to the end of the sequence of the generated condition prompt to construct a pre-sequence for speech generation containing prior information of the text; S604, the pre-sequence of speech generation is input into the acoustic prediction channel of the target language generation network, and the corresponding discrete target speech identifier is output based on the discrete identifier of the target text. S605, the target speech discrete identifier is mapped to a spectrogram representation, and a vocoder is used to perform waveform synthesis processing on the spectrogram representation to obtain the target audio waveform.
[0115] In this embodiment, generating conditional prompts is input into the target language generation network, meaning that the already organized prompt information is sent into the speech generation inference structure. The generated conditional prompts include the text to be generated, speaker prompts, and dialect guidance statements written after preprocessing. After entering the target language generation network, these contents are no longer processed as raw text fields, but are transformed into a unified conditional representation. The target language generation network may include a prompt encoding unit, an autoregressive prediction unit, a text generation unit, an acoustic prediction channel, a spectrogram mapping unit, and a waveform synthesis unit. The prompt encoding unit is responsible for receiving the generated conditional prompts and performing discrete symbol embedding, positional encoding overlay, and conditional compression; the autoregressive prediction unit is responsible for maintaining the dependency relationship between the current position and the preceding information in temporal order; the text generation unit is responsible for outputting the discrete identifier of the target text; the acoustic prediction channel is responsible for outputting the discrete identifier of the target speech under the prior conditions of the text; the spectrogram mapping unit is responsible for restoring the discrete identifier of the target speech into a continuous acoustic representation; and the waveform synthesis unit is responsible for restoring the continuous acoustic representation into a time-domain waveform. The connection relationship between the units is as follows: the output of the prompt encoding unit is connected to the autoregressive prediction unit; the state of the autoregressive prediction unit is connected to the text generation unit; the target text discrete identifier output by the text generation unit enters the speech generation pre-sequence construction position; the speech generation pre-sequence is then input into the acoustic prediction channel; the acoustic prediction channel outputs the target speech discrete identifier; the target speech discrete identifier enters the spectrogram mapping unit and the waveform synthesis unit in sequence; and finally, the target audio waveform is output.
[0116] Triggering the autoregressive prediction mechanism of the target language generation network signifies initiating a position-based, progressive prediction process within the network. The autoregressive prediction mechanism ensures that the output at the current position depends on previously determined content, rather than outputting the entire result independently at once. This ensures that both text and speech generation are constrained by the preceding context, preventing the generated content from being disconnected from previous expressions. In implementation, the autoregressive prediction unit can employ a multi-layered contextual modeling structure, sequentially reading the cue encoding results and historical output results to calculate the generation conditions at the current position. If an attention structure is used, the autoregressive constraint can be controlled by causal masking to ensure that the current position only accesses preceding information; if a temporal recursive structure is used, the autoregressive constraint can be achieved through hidden state propagation. Regardless of the implementation method, the autoregressive prediction mechanism ensures a sequential dependency in the generation process, thereby maintaining the continuity in the expression of the subsequently output discrete target text and discrete target speech identifiers.
[0117] Under the autoregressive prediction mechanism, the target language generation network outputs discrete identifiers of the target text sequentially based on generation condition prompts. This means that the text generation unit outputs a discrete text result at each prediction position and writes this result into the already generated text sequence. The discrete identifiers of the target text are not formed all at once, but rather progressively form a complete discrete text expression as the autoregressive prediction process unfolds. The significance of sequential output is to ensure that the current generated result is always influenced by the previous text content, speaker prompts, and dialect guidance statements. For example, in fintech, if the generation condition prompts include product names, profit ranges, fee descriptions, risk disclosure phrases, and customer service role information, the sequentially output discrete identifiers of the target text will continuously maintain consistency in the expression of amount, term, and role within the sequence. In healthcare, if the generation condition prompts include symptom descriptions, rehabilitation reminders, health advice, and service role information, the sequentially output discrete identifiers of the target text will continuously maintain consistency in the expression of health vocabulary and tone of communication within the sequence. The input to the text generation unit is the current state of the autoregressive prediction unit, and the output is the text category distribution at the current position. After selection, the category distribution yields the discrete identifier of the target text at the current position. The discrete identifier of the target text continues to be fed back to the autoregressive prediction unit to participate in the context update of the next position.
[0118] By concatenating the discrete target text identifier to the end of the sequence for generating conditional prompts, a pre-sequence for speech generation containing prior text information is constructed. This indicates that the generated text result is not directly used as the final audio output, but is incorporated into the conditional sequence for the next stage of speech prediction. The initial conditional prompts contain conditional text, speaker prompts, and dialect guidance information; this sequence controls the text generation stage of the target language generation network. After the discrete target text identifier is generated, concatenating it to the end of the sequence for generating conditional prompts means that the generated result continues to participate in subsequent acoustic prediction as a new conditional context. This allows speech prediction to no longer rely solely on the original prompts but on the already determined textual expression. The constructed pre-sequence for speech generation thus includes original semantic constraints, identity constraints, dialect constraints, and newly generated prior text information. This pre-sequence serves as a transition from semantic expression to acoustic expression. For business content involving continuous expressions from multiple speakers, using only the original prompts for speech generation can easily lead to inconsistencies between the acoustic output and the generated text. However, by adding the discrete target text identifier to the pre-sequence, speech prediction can directly revolve around the already confirmed textual result.
[0119] The acoustic prediction channel of the target language generation network inputs the pre-sequence of speech generation into the target language generation network. Based on the discrete identifier of the target text, it outputs the corresponding discrete identifier of the target speech. This means that within the target language generation network, a prediction branch specifically for acoustic representation receives the text prior conditions and generates discrete acoustic results consistent with the text content position by position. The acoustic prediction channel can be considered as an internal prediction structure parallel to but functionally different from the text generation unit. The text generation unit focuses on the discrete semantic results, while the acoustic prediction channel focuses on the discrete acoustic results of recoverable speech. After receiving the pre-sequence of speech generation, the acoustic prediction channel does not regenerate the text. Instead, it uses the discrete identifier of the target text as semantic constraints, the speaker's prompts as subject constraints, and the dialect guidance statements as pronunciation style constraints, outputting discrete identifiers of the target speech consistent with these conditions. In implementation, the acoustic prediction channel may include a pre-sequence encoding layer, a conditional fusion layer, and an acoustic output layer. The pre-sequence encoding layer maps the pre-sequence of speech generation to a contextual representation, the conditional fusion layer integrates the text prior, identity information, and dialect information, and the acoustic output layer predicts the corresponding discrete identifier of the target speech position. After this processing, the discrete identifier of the target speech remains consistent with the discrete identifier of the target text in terms of temporal order and semantic correspondence, while also remaining consistent with the generated conditional prompts in terms of the speaker and regional pronunciation attributes.
[0120] The process involves mapping the discrete target speech identifier to a spectrogram representation, then using a vocoder to perform waveform synthesis on this representation to obtain the target audio waveform. This process essentially restores the discrete acoustic representation to a continuous audio signal. Since the discrete target speech identifier itself is a discrete index sequence, it cannot be played directly; therefore, it requires continuous acoustic representation restoration and time-domain waveform reconstruction. The spectrogram mapping unit receives the discrete target speech identifier and converts the discrete index into a continuous spectrogram representation. The spectrogram representation can be time-arranged spectral features, a time-frequency joint representation, or other continuous acoustic representations available to the waveform synthesis unit. After receiving the spectrogram representation, the vocoder performs waveform synthesis to restore the continuous representation in the frequency domain or time-frequency domain to a time-domain audio waveform. The input-output relationship in this process is as follows: the spectrogram mapping unit inputs the discrete target speech identifier and outputs the spectrogram representation, while the waveform synthesis unit inputs the spectrogram representation and outputs the target audio waveform. The vocoder can employ an autoregressive waveform generation structure or a parallel waveform generation structure, both constrained by the spectrogram representation. The resulting target audio waveform not only semantically corresponds to the discrete identifier of the target text, but also inherits the control information generated in the conditional prompts in terms of speaker identity, dialect pronunciation, and speech style.
[0121] The generated conditional prompts are input data into the prompt encoding unit. The prompt encoding unit outputs a conditional representation and passes it to the autoregressive prediction unit. The autoregressive prediction unit controls the text generation unit to sequentially output discrete identifiers of the target text. The discrete identifiers of the target text are concatenated with the original prompt content to form a speech generation pre-sequence. The speech generation pre-sequence enters the acoustic prediction channel, which outputs a discrete identifier of the target speech. The discrete identifier of the target speech enters the spectrogram mapping unit, which outputs a spectrogram representation. The spectrogram representation enters the vocoder, which outputs the target audio waveform. In a fintech business environment, the input can be set to include the generated conditional prompts containing product name, return range, fee structure, risk disclosure statement, customer service identity information, and regional pronunciation information, and the output is a target audio waveform broadcast to the customer. In a healthcare business environment, the input can be set to include the generated conditional prompts containing symptom description, health advice, rehabilitation reminders, service role information, and regional pronunciation information, and the output is the corresponding target audio waveform.
[0122] This embodiment inputs the generation condition prompts into the target language generation network, uses an autoregressive prediction mechanism to successively form discrete identifiers of the target text, then concatenates the discrete identifiers of the target text with the generation condition prompts to form a speech generation pre-sequence, and outputs the discrete identifiers of the target speech based on the pre-sequence. Subsequently, the discrete identifiers of the target speech are mapped to a spectrogram representation and recovered into the target audio waveform by a vocoder. This ensures that semantic content, speaker information, and dialect information maintain a continuous transmission relationship between the text generation stage and the acoustic generation stage, thereby improving the consistency between the text expression and the final speech expression, and enhancing the controllability and coherence of the output audio.
[0123] In one embodiment, a speech generation apparatus based on interleaved sequences is provided, which corresponds one-to-one with the speech generation method based on interleaved sequences described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the speech generation device based on interleaved sequences of the present invention. The modules include a multi-channel alignment construction module 10, a discrete identifier generation module 20, an interleaved sequence construction module 30, a generation network training module 40, a conditional prompting construction module 50, and a speech generation and decoding module 60. Detailed descriptions of each functional module are as follows: The multi-channel alignment construction module 10 is used to acquire original multi-speaker audio material, extract dialogue sample units from the original multi-speaker audio material, and perform time stamp alignment on the dialogue sample units to construct a multi-channel alignment metadata dataset. The discrete identifier generation module 20 is used to convert the dialogue sample unit in the multi-channel aligned metadata set into a discrete identifier set, wherein the discrete identifier set includes speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, non-language discrete identifiers, and dialect discrete identifiers. The interleaved sequence construction module 30 is used to insert the sub-language discrete identifier into the text discrete identifier, and to concatenate the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier and the speech discrete identifier according to the pronunciation time sequence to obtain the interleaved sequence; The network training module 40 is used to input the interleaved sequence into the pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech identifiers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network. The condition prompt construction module 50 is used to receive a generation condition prompt containing text to be generated and speaker prompt information. When the generation condition prompt has dialect indication information, it obtains a dialect guidance statement according to the dialect indication information and concatenates the dialect guidance statement to the beginning of the text to be generated. The speech generation and decoding module 60 is used to input the generation condition prompts into the target language generation network, output the target text discrete identifier in an autoregressive manner, output the corresponding target speech discrete identifier based on the target text discrete identifier, and decode the target speech discrete identifier into a target audio waveform.
[0124] Specific limitations regarding the speech generation device based on interlaced sequences can be found in the aforementioned limitations on the speech generation method based on interlaced sequences, and will not be repeated here. Each module in the aforementioned speech generation device based on interlaced sequences can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0125] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a server-side speech generation method based on interleaved sequences.
[0126] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a speech generation method based on interleaved sequences.
[0127] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain original multi-speaker audio material, extract dialogue sample units from the original multi-speaker audio material, and align the dialogue sample units with timestamps to construct a multi-channel aligned metadata dataset; The dialogue sample units in the multi-channel aligned metadata set are converted into a discrete identifier set, which includes speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, paralinguistic discrete identifiers, and dialect discrete identifiers. The sub-language discrete identifier is inserted into the text discrete identifier, and the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier and the speech discrete identifier are concatenated according to the pronunciation time sequence to obtain an interleaved sequence; The interleaved sequence is input into the pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech identifiers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network. Receive a generation condition prompt containing text to be generated and speaker prompt information. When the generation condition prompt has dialect indication information, obtain a dialect guidance statement according to the dialect indication information and concatenate the dialect guidance statement to the beginning of the text to be generated. The generation condition prompts are input into the target language generation network, which outputs a discrete target text identifier in an autoregressive manner, outputs a corresponding discrete target speech identifier based on the discrete target text identifier, and decodes the discrete target speech identifier into a target audio waveform.
[0128] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon. When the computer program is executed by a processor, it performs the following steps: Obtain original multi-speaker audio material, extract dialogue sample units from the original multi-speaker audio material, and align the dialogue sample units with timestamps to construct a multi-channel aligned metadata dataset; The dialogue sample units in the multi-channel aligned metadata set are converted into a discrete identifier set, which includes speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, paralinguistic discrete identifiers, and dialect discrete identifiers. The sub-language discrete identifier is inserted into the text discrete identifier, and the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier and the speech discrete identifier are concatenated according to the pronunciation time sequence to obtain an interleaved sequence; The interleaved sequence is input into the pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech identifiers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network. Receive a generation condition prompt containing text to be generated and speaker prompt information. When the generation condition prompt has dialect indication information, obtain a dialect guidance statement according to the dialect indication information and concatenate the dialect guidance statement to the beginning of the text to be generated. The generation condition prompts are input into the target language generation network, which outputs a discrete target text identifier in an autoregressive manner, outputs a corresponding discrete target speech identifier based on the discrete target text identifier, and decodes the discrete target speech identifier into a target audio waveform.
[0129] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0130] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0131] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0132] It should be noted that any software tools or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0133] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A speech generation method based on interleaved sequences, characterized in that, Includes the following steps: Obtain original multi-speaker audio material, extract dialogue sample units from the original multi-speaker audio material, and align the dialogue sample units with timestamps to construct a multi-channel aligned metadata dataset; The dialogue sample units in the multi-channel aligned metadata set are converted into a discrete identifier set, which includes speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, paralinguistic discrete identifiers, and dialect discrete identifiers. The sub-language discrete identifier is inserted into the text discrete identifier, and the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier and the speech discrete identifier are concatenated according to the pronunciation time sequence to obtain an interleaved sequence; The interleaved sequence is input into the pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech identifiers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network. Receive a generation condition prompt containing text to be generated and speaker prompt information. When the generation condition prompt has dialect indication information, obtain a dialect guidance statement according to the dialect indication information and concatenate the dialect guidance statement to the beginning of the text to be generated. The generation condition prompts are input into the target language generation network, which outputs a discrete target text identifier in an autoregressive manner, outputs a corresponding discrete target speech identifier based on the discrete target text identifier, and decodes the discrete target speech identifier into a target audio waveform.
2. The speech generation method based on interleaved sequences as described in claim 1, characterized in that, Obtain original multi-speaker audio material, extract dialogue sample units from the original multi-speaker audio material, and perform timestamp alignment on the dialogue sample units to construct a multi-channel aligned metadata dataset, including: Obtain the original multi-speaker audio material, and perform audio enhancement and background noise suppression processing on the original multi-speaker audio material to obtain the enhanced audio segment; The enhanced audio segment is processed by voice activity detection to obtain multiple voice segments. The voice segments are then segmented and merged according to the conversation boundaries to obtain short dialogue audio segments. Speaker separation processing is performed on the short dialogue audio segments, and clustering processing is performed based on the voiceprint embedding vectors of the short dialogue audio segments to obtain speaker identifiers; Transcription processing is performed on the short dialogue audio segment to obtain transcribed text; The short dialogue audio segment is subjected to paralinguistic event recognition, and dialect retrieval and dialect recognition are performed to obtain paralinguistic event annotation information and dialect annotation information; Dialogue sample units are constructed based on the short dialogue audio segments, the speaker identifiers, the transcribed text, the sublingual event annotation information, and the dialect annotation information; The dialogue sample units are timestamped and then the timestamped dialogue sample units are aggregated to obtain a multi-channel aligned metadata dataset.
3. The speech generation method based on interleaved sequences as described in claim 1, characterized in that, The dialogue sample units in the multi-channel aligned metadata set are converted into a discrete identifier set, which includes speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, paralinguistic discrete identifiers, and dialect discrete identifiers, including: Dialogue sample units are obtained from the multi-channel aligned metadata set, and short dialogue audio segments, speaker identifiers, transcribed text, paralinguistic event annotation information, and dialect annotation information are extracted from the dialogue sample units. Residual vector quantization is performed on the short dialogue audio segment to decompose the short dialogue audio segment into a multi-codebook acoustic tagging layer; The multi-codebook acoustic tagging layer is subjected to hierarchical discrete encoding to generate discrete speech identifiers; The transcribed text is text-tagged to generate discrete text identifiers; The speaker identifier, the sub-language event annotation information, and the dialect annotation information are respectively marked and mapped to generate corresponding discrete speaker identifiers, discrete sub-language identifiers, and discrete dialect identifiers. By combining the discrete speech identifier, the discrete speaker identifier, the discrete text identifier, the discrete secondary language identifier, and the discrete dialect identifier, a discrete identifier set is obtained.
4. The speech generation method based on interleaved sequences as described in claim 1, characterized in that, The secondary language discrete identifier is inserted into the text discrete identifier, and the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier, and the speech discrete identifier are concatenated according to the pronunciation sequence to obtain an interleaved sequence, including: Based on the time alignment information corresponding to the sub-language discrete identifier, locate the pronunciation insertion position within the text discrete identifier; The sub-language discrete identifier is placed into the pronunciation insertion position to obtain the text discrete identifier after inserting the sub-language discrete identifier; Using the speaker discrete identifier as the starting unit, the dialect discrete identifier is concatenated to the end of the starting unit to generate a prefix sequence; According to the pronunciation sequence, the text discrete identifiers after inserting the secondary language discrete identifiers are concatenated to the end of the prefix sequence to generate a transition sequence; The discrete speech identifiers are concatenated to the end of the transition sequence according to the pronunciation sequence to obtain an interleaved sequence.
5. The speech generation method based on interleaved sequences as described in claim 1, characterized in that, The interleaved sequence is input into a pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech markers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network, which includes: A first-stage training sample consisting of single-speaker samples and dialogue samples is constructed based on the interleaved sequence, and a second-stage training sample consisting of multi-speaker dialogue samples containing dialect annotation information and paralinguistic event annotation information is constructed based on the interleaved sequence. The pre-trained language generation network is fine-tuned under first-stage supervision using the first-stage training samples to obtain the intermediate language generation network. The intermediate language generation network is then fine-tuned under second-stage supervision using the second-stage training samples. During the first and second phases of supervised fine-tuning, a course learning strategy is adopted, and training samples are provided step by step in the following order: from short single-speaker samples, to multi-speaker samples, to overlapping speech samples and long context samples, and finally to samples containing dialect annotation information and paralinguistic event annotation information. During the second stage of supervised fine-tuning, the current prediction position of the intermediate language generation network on the interleaved sequence is located according to the autoregressive prediction order. In the interleaved sequence, the speech discrete identifiers preceding the current prediction position are identified, and a portion of the speech discrete identifiers preceding the current prediction position are selected according to a preset probability value. Randomly discard selected discrete speech identifiers, and input the processed interleaved sequence into the intermediate language generation network; Obtain the output of the intermediate language generation network based on the processed interleaved sequence, and construct speech prediction cross-entropy error term, speaker consistency error term, paralinguistic event classification error term, dialect classification error term, and reconstruction error term based on the output. The speech prediction cross-entropy error term, the speaker consistency error term, the paralinguistic event classification error term, the dialect classification error term, and the reconstruction error term are combined to construct a multi-task joint error function; The intermediate language generation network is updated using the multi-task joint error function until the preset training termination condition is met, thus obtaining the target language generation network.
6. The speech generation method based on interleaved sequences as described in claim 1, characterized in that, Receive generation condition prompts containing text to be generated and speaker prompts. When the generation condition prompts contain dialect indication information, obtain a dialect guidance statement based on the dialect indication information, and append the dialect guidance statement to the beginning of the text to be generated, including: Receive generation condition prompts containing the text to be generated and speaker prompts; The generation condition prompts are analyzed to extract the text to be generated and the speaker prompt information; Dialect indication information is detected in the generated condition prompt. When the generated condition prompt has dialect indication information, the dialect indication information is extracted. Obtain the corresponding dialect guidance statement based on the dialect indication information; The dialect guidance statement is appended to the beginning of the text to be generated, and the text to be generated after appending the dialect guidance statement is used to replace the original text to be generated in the generation condition prompt, while keeping the speaker prompt information in the position of the generation condition prompt unchanged.
7. The speech generation method based on interleaved sequences as described in claim 1, characterized in that, The generation condition prompts are input into the target language generation network, which outputs discrete target text identifiers in an autoregressive manner. Based on the discrete target text identifiers, corresponding discrete target speech identifiers are output, and the discrete target speech identifiers are decoded into target audio waveforms, including: The generation condition prompts are input into the target language generation network to trigger the autoregressive prediction mechanism of the target language generation network; Under the autoregressive prediction mechanism, the target language generation network outputs discrete identifiers of the target text sequentially based on the generation condition prompts; The discrete identifier of the target text is concatenated to the end of the sequence of generated condition prompts to construct a pre-sequence for speech generation containing prior text information; The speech generation pre-sequence is input into the acoustic prediction channel of the target language generation network, and the corresponding target speech discrete identifier is output based on the target text discrete identifier. The target speech discrete identifier is mapped to a spectrogram representation, and a vocoder is used to perform waveform synthesis processing on the spectrogram representation to obtain the target audio waveform.
8. A speech generation device based on interleaved sequences, characterized in that, The speech generation device based on interleaved sequences includes: A multi-channel alignment construction module is used to acquire original multi-speaker audio materials, extract dialogue sample units from the original multi-speaker audio materials, and perform time stamp alignment on the dialogue sample units to construct a multi-channel alignment metadata dataset. The discrete identifier generation module is used to convert the dialogue sample units in the multi-channel aligned metadata set into a discrete identifier set, wherein the discrete identifier set includes speech discrete identifiers, speaker discrete identifiers, text discrete identifiers, non-language discrete identifiers, and dialect discrete identifiers. An interleaved sequence construction module is used to insert the sub-language discrete identifier into the text discrete identifier, and to concatenate the speaker discrete identifier, the dialect discrete identifier, the text discrete identifier and the speech discrete identifier according to the pronunciation time sequence to obtain an interleaved sequence; A network training module is used to input the interleaved sequence into a pre-trained language generation network for supervised fine-tuning. During the supervised fine-tuning process, the discrete speech identifiers located before the current prediction position in the interleaved sequence are randomly discarded to obtain the target language generation network. The condition prompt construction module is used to receive a generation condition prompt containing the text to be generated and speaker prompt information. When the generation condition prompt has dialect indication information, the module obtains a dialect guidance statement based on the dialect indication information and concatenates the dialect guidance statement to the beginning of the text to be generated. The speech generation and decoding module is used to input the generation condition prompts into the target language generation network, output the target text discrete identifier in an autoregressive manner, output the corresponding target speech discrete identifier based on the target text discrete identifier, and decode the target speech discrete identifier into a target audio waveform.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and an interleaved sequence-based speech generation program stored in the memory and executable on the processor, wherein the interleaved sequence-based speech generation program, when executed by the processor, implements the steps of the interleaved sequence-based speech generation method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a speech generation program based on interleaved sequences, which, when executed by a processor, implements the steps of the speech generation method based on interleaved sequences as described in any one of claims 1-7.