Time-varying emotion-guided speech generation method and device, equipment and medium

CN122676801APending Publication Date: 2026-09-01SHENZHEN PINGAN COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611083131.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0005]本发明的主要目的在于提供一种时变情感引导的语音生成方法、装置、设备及存储介质,旨在解决现有语音生成技术将情感信息作为全局静态条件输入,难以对语音生成过程中情感随时间变化的过程进行细粒度建模,导致生成语音的情感过渡不自然的技术问题

Benefits of technology

[0010] Beneficial Effects: This invention relates to the field of speech and semantic technology, and discloses a method, apparatus, device, and medium for time-varying emotion-guided speech generation. The method includes: acquiring training speech and corresponding training text; performing emotion encoding, emotion parameter determination, and time alignment on the training speech to obtain a time-varying emotion vector; constructing a hybrid expert speech segmentation structure and jointly training it with semantic features, the time-varying emotion vector, and identity features to obtain a speech token sequence; generating an emotion guidance token sequence based on the speech token sequence; combining the emotion guidance token sequence and the speech token sequence into an interleaved token sequence; and inputting the training text and the interleaved token sequence into a language model to be trained; and receiving task text and generating target speech using the trained language model. This invention can be applied to business scenarios such as fintech and healthcare. By converting the emotional changes in the training speech into a time-varying emotion vector and further forming an emotion guidance token sequence corresponding to the speech token sequence, the language model learns semantic content, emotional changes, and identity features simultaneously during training and generation. This reduces the problem of singular emotional expression caused by global static emotion control, resulting in more natural emotional transitions and more stable timbre in the generated speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676801A_ABST
    Figure CN122676801A_ABST
Patent Text Reader

Abstract

This invention relates to the field of speech and semantic technology, and discloses a method, apparatus, device, and medium for time-varying emotion-guided speech generation. The method includes: performing emotion encoding, emotion parameter determination, and time alignment on training speech to obtain a time-varying emotion vector; constructing a hybrid expert speech segmentation structure and jointly training it with semantic features, the time-varying emotion vector, and identity features to obtain a speech token sequence; generating an emotion guidance token sequence based on the speech token sequence; combining the emotion guidance token sequence and the speech token sequence into an interleaved token sequence; and training a language model using the interleaved token sequence; and generating target speech using the trained language model after receiving task text. This invention can be applied to business scenarios such as fintech and healthcare. By using time-varying emotion vectors and emotion guidance token sequences in modeling, it reduces the problem of emotion uniformity caused by static emotion control, and improves the naturalness of emotion and the stability of timbre in generated speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech semantics technology, and in particular to a method, apparatus, device and medium for time-varying emotion-guided speech generation. Background Technology

[0002] In existing speech generation technologies, emotion control typically employs static modeling, associating the entire speech segment with a single emotion tag or fixed emotion parameter. This approach struggles to express the emotional processes that evolve over time within the speech content. Furthermore, existing speech segmenters often prioritize semantic preservation, neglecting the joint modeling of emotional changes and speaker identity features. This results in generated speech prone to issues such as simplistic emotional expression, abrupt emotional transitions, and insufficient timbre consistency.

[0003] In the fintech business, scenarios such as intelligent customer service, robo-advisors, loan approval, transaction alerts, risk warnings, and anti-fraud verification typically involve continuous dialogues including informing, explaining, reminding, and reassuring. Existing voice generation methods often treat emotion as a global input condition, making it difficult to reflect phased emotional changes in voice content such as abnormal transaction alerts, overdue repayment reminders, investment risk disclosures, and account security verification. This can easily result in voice prompts lacking urgency, credibility, and identity stability.

[0004] In the healthcare sector, scenarios such as intelligent consultation, chronic disease follow-up, medication reminders, rehabilitation guidance, and psychological counseling require natural variations in voice between inquiries, reminders, explanations, and reassurances. Existing speech generation methods struggle to achieve fine-grained control over the emotional state of speech at the frame level, and also find it difficult to maintain consistency in semantics, emotion, and speaker identity simultaneously. This can easily result in mechanical expressions in the generated speech, affecting the naturalness and continuity of the interaction process. Summary of the Invention

[0005] The main objective of this invention is to provide a time-varying emotion-guided speech generation method, apparatus, device, and storage medium, aiming to solve the technical problem that existing speech generation technologies use emotional information as a global static condition input, making it difficult to perform fine-grained modeling of the process of emotion changing over time during speech generation, resulting in unnatural emotional transitions in the generated speech.

[0006] To achieve the above objectives, the present invention provides a time-varying emotion-guided speech generation method, comprising: Acquire training speech and training text corresponding to the training speech, perform emotion encoding on the training speech to obtain an emotion feature sequence, determine an emotion parameter sequence based on the emotion feature sequence, and perform time alignment on the emotion parameter sequence to obtain a time-varying emotion vector; A hybrid expert speech segmentation structure is constructed. Semantic features and identity features are extracted from the training speech. The hybrid expert speech segmentation structure is jointly trained based on the semantic features, time-varying sentiment vector, and identity features. The training speech is then segmented using the jointly trained hybrid expert speech segmentation structure to obtain a speech token sequence. Based on the number of tokens in the voice token sequence, the time-varying emotion vector is mapped and aggregated to obtain an emotion guidance token sequence corresponding to the voice token sequence; The emotion guidance token sequence and the voice token sequence are combined in an alternating order to form an alternating token sequence, and the training text and the alternating token sequence are input into the language model to be trained to obtain the prediction token sequence; The training loss is determined based on the predicted token sequence and the interleaved token sequence, and the language model to be trained is updated based on the training loss to obtain the trained language model. The task text is received and processed by the trained language model to obtain the target speech.

[0007] Furthermore, to achieve the above objectives, the present invention provides a time-varying emotion-guided speech generation device, comprising: The emotion feature extraction module is used to acquire training speech and training text corresponding to the training speech, perform emotion encoding on the training speech to obtain an emotion feature sequence, determine an emotion parameter sequence based on the emotion feature sequence, and perform time alignment on the emotion parameter sequence to obtain a time-varying emotion vector. The hybrid expert speech segmentation module is used to construct a hybrid expert speech segmentation structure, extract semantic features and identity features from the training speech, and jointly train the hybrid expert speech segmentation structure based on the semantic features, time-varying sentiment vector and identity features. The trained hybrid expert speech segmentation structure is then used to segment the training speech to obtain a speech token sequence. The emotion mapping and aggregation module is used to map and aggregate the time-varying emotion vector based on the number of tokens in the voice token sequence to obtain an emotion guidance token sequence corresponding to the voice token sequence. An interleaved token construction module is used to combine the emotion guidance token sequence and the voice token sequence in an interleaved order to form an interleaved token sequence, and input the training text and the interleaved token sequence into the language model to be trained to obtain a predicted token sequence; The language model training module is used to determine the training loss based on the predicted token sequence and the interleaved token sequence, and update the language model to be trained based on the training loss to obtain the trained language model. The speech generation module is used to receive task text, process the task text through the trained language model, and obtain the target speech.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a time-varying emotion-guided speech generation program stored in the memory and executable on the processor, wherein when the time-varying emotion-guided speech generation program is executed by the processor, it implements the steps of the time-varying emotion-guided speech generation method as described above.

[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a time-varying emotion-guided speech generation program, which, when executed by a processor, implements the steps of the time-varying emotion-guided speech generation method as described above.

[0010] Beneficial Effects: This invention relates to the field of speech and semantic technology, and discloses a method, apparatus, device, and medium for time-varying emotion-guided speech generation. The method includes: acquiring training speech and corresponding training text; performing emotion encoding, emotion parameter determination, and time alignment on the training speech to obtain a time-varying emotion vector; constructing a hybrid expert speech segmentation structure and jointly training it with semantic features, the time-varying emotion vector, and identity features to obtain a speech token sequence; generating an emotion guidance token sequence based on the speech token sequence; combining the emotion guidance token sequence and the speech token sequence into an interleaved token sequence; and inputting the training text and the interleaved token sequence into a language model to be trained; and receiving task text and generating target speech using the trained language model. This invention can be applied to business scenarios such as fintech and healthcare. By converting the emotional changes in the training speech into a time-varying emotion vector and further forming an emotion guidance token sequence corresponding to the speech token sequence, the language model learns semantic content, emotional changes, and identity features simultaneously during training and generation. This reduces the problem of singular emotional expression caused by global static emotion control, resulting in more natural emotional transitions and more stable timbre in the generated speech. Attached Figure Description

[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a time-varying emotion-guided speech generation method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the time-varying emotion-guided speech generation method of the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the time-varying emotion-guided speech generation device of the present invention. Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0013] The time-varying emotion-guided speech generation method provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain training speech and corresponding training text from the client, perform emotion encoding, emotion parameter determination, and time alignment on the training speech to obtain a time-varying emotion vector; construct a hybrid expert speech segmentation structure, and perform joint training combining semantic features, time-varying emotion vectors, and identity features to obtain a speech token sequence; generate an emotion guidance token sequence based on the speech token sequence, combine the emotion guidance token sequence and the speech token sequence into an interleaved token sequence, and input the training text and the interleaved token sequence into the language model to be trained for training; after receiving the task text, generate the target speech using the trained language model. This invention can be applied to business scenarios such as fintech and healthcare. By converting the emotional changes in the training speech into a time-varying emotion vector, and further forming an emotion guidance token sequence corresponding to the speech token sequence, the language model learns semantic content, emotional changes, and identity features simultaneously during training and generation, thereby reducing the problem of singular emotional expression caused by global static emotion control, and making the generated speech have more natural emotional transitions and more stable timbre performance. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the time-varying emotion-guided speech generation method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0015] like Figure 2 As shown, the time-varying emotion-guided speech generation method proposed in this invention includes the following steps: S10, acquire training speech and training text corresponding to the training speech, perform emotion encoding on the training speech to obtain an emotion feature sequence, determine an emotion parameter sequence based on the emotion feature sequence, and perform time alignment on the emotion parameter sequence to obtain a time-varying emotion vector; In this embodiment, the training speech can use speech data with a complete timeline, preserving pronunciation intervals, pause intervals, energy changes, fundamental frequency changes, and speech rate changes. During processing, the training speech undergoes sampling rate unification, volume normalization, silence boundary marking, and abnormal segment removal, ensuring that speech from different sources enters emotion encoding under the same acoustic scale. The training speech does not need to come from a single recording environment; it can come from customer service voice, interactive voice, prompt voice, or reading aloud, as long as the speech content can establish a correspondence with the training text.

[0016] Training text corresponding to the training speech is used to provide the boundaries of the speech content. The correspondence can be established through manual transcription, speech recognition-corrected text, or annotated text. During processing, the training text undergoes word segmentation, punctuation correction, digit pronunciation standardization, and pause mark adjustment to ensure that text units can match pronunciation segments in the training speech. The correspondence does not need to be accurate to each individual character; it can be accurate to phrases, short sentences, or semantic segments. This correspondence is used to reduce the misalignment between speech segments and text content during emotion encoding.

[0017] Emotion coding is used to convert acoustic variations in training speech into a processable emotional representation. In implementation, Mel spectrum, energy curve, fundamental frequency curve, speech rate features, and pause features are first extracted from the training speech and then input into the emotion encoder. The emotion encoder can include an acoustic input layer, a local feature extraction layer, a temporal modeling layer, and an emotion representation output layer. The local feature extraction layer extracts short-term energy, spectral variations, and pitch fluctuations; the temporal modeling layer processes continuous emotional changes between adjacent speech segments; and the emotion representation output layer outputs emotional features arranged in time.

[0018] The emotional feature sequence is formed by arranging emotional features at multiple time locations in chronological order. Each emotional feature can correspond to a group of acoustic frames, a speech segment, or a pronunciation interval corresponding to a text unit. In implementation, the training speech is segmented into continuous time segments, each time segment is processed by an emotional encoder to obtain an emotional feature, and then concatenated in chronological order to form the emotional feature sequence. Individual emotional features in the emotional feature sequence can contain information such as tone intensity, emotional tendency, energy changes, pitch changes, and pause states.

[0019] The emotional parameter sequence is obtained by parameter mapping from the emotional feature sequence. Emotional parameters can include dominance, arousal, and valence parameters, and can also include continuous parameters such as urgency, stability, or soothingness. In implementation, the emotional feature sequence is input into the parameter mapping layer, which can consist of a normalization layer, a temporal smoothing layer, and a parameter output layer. The normalization layer reduces differences in recorded volume, the temporal smoothing layer handles transitions between adjacent emotional features, and the parameter output layer outputs the emotional parameters corresponding to each time position. The emotional parameter sequence preserves temporal order, allowing the emotional differences between different segments of the speech to be read in subsequent processing.

[0020] Time alignment is used to adjust the sentiment parameter sequence to a uniform time scale. The sentiment parameter sequence may be generated according to speech segments, text units, or acoustic frame groups, and subsequent processing typically requires the application of sentiment information along a stable acoustic timeline. In implementation, the acoustic frame boundaries or text pronunciation boundaries of the training speech can be obtained, and the sentiment parameter sequence can be mapped to the target time position. For time positions where there are gaps between adjacent sentiment parameters, smoothing compensation can be performed; for cases where multiple sentiment parameters fall into the same time position, weighted merging can be used. After time alignment, each target time position has a corresponding sentiment parameter.

[0021] Time-varying emotion vectors consist of time-aligned emotion parameters. Each time position corresponds to a vector, which contains multiple emotion dimensions. In implementation, the aligned dominance, arousal, valence, and extended emotion parameters are combined in a fixed order to form a time-ordered vector sequence. Time-varying emotion vectors can preserve the change in emotion within a sentence, from calm to urgency, from reminder to reassurance, and from explanation to guidance, avoiding the situation where an entire speech segment corresponds to only a single emotion value.

[0022] In one implementation, the training speech is first converted into Mel spectrum, fundamental frequency curve, and energy curve, and the training text is processed through word segmentation and punctuation cleanup to obtain text units. The emotion encoder uses convolutional layers to extract short-term acoustic changes and a temporal modeling layer to obtain emotional changes across time segments. The emotion feature sequence is input to a parameter mapping layer to obtain a dominance parameter sequence, an arousal parameter sequence, and a valence parameter sequence. Acoustic frame boundaries are used to map the emotion parameter sequence to a unified frame-level time axis, and missing positions between frames are compensated for using adjacent parameters to obtain a time-varying emotion vector.

[0023] Alternatively, a text boundary-assisted implementation can be used. Text units in the training text are matched with pronunciation intervals in the training speech, and acoustic features are extracted independently for each pronunciation interval. The emotion encoder generates emotion features within each pronunciation interval and combines them with acoustic changes in adjacent pronunciation intervals to form an emotion feature sequence. The parameter mapping layer generates an emotion parameter sequence according to the pronunciation interval corresponding to the text unit, and then aligns the emotion parameter sequence to the acoustic frame boundary. This method is suitable for training data with clear text structure and stable pause boundaries.

[0024] Sensitive segments can also be processed with increased density. When the training text contains segments with risk warnings, monetary reminders, medication reminders, or follow-up reminders, shorter time segments are used to extract sentiment features from the corresponding pronunciation intervals. Non-sensitive segments are extracted using longer time segments. The parameter mapping layer outputs a sequence of sentiment parameters of the same dimension for different segments, and during time alignment, different segments are mapped to a unified acoustic frame boundary. This approach can preserve finer-grained sentiment changes near key content.

[0025] In fintech business scenarios, training speech can be used to alert users of abnormal account transactions, while training text can include text units such as account anomalies, identity verification, and financial risks. After establishing a correspondence between the training text and training speech, the emotion encoder extracts energy increases, pause shortenings, and pitch changes within the corresponding speech intervals to form an emotion feature sequence. The parameter mapping layer outputs higher arousal parameters in the risk warning interval and more stable emotion parameters in the operation guidance interval. The time-varying emotion vector formed after time alignment can preserve the emotion differences between anomaly alerts and verification guidance.

[0026] In healthcare scenarios, training speech can serve as reminders for chronic disease follow-ups, while training text can include text units such as blood pressure monitoring, medication adherence, and status feedback. An emotion encoder extracts speech rate, pauses, pitch, and energy variations from the training speech to generate an emotion feature sequence. A parameter mapping layer outputs neutral emotion parameters in the observation / explanation interval, relatively stable prompt parameters in the medication reminder interval, and gentler, reassuring parameters in the status feedback interval. After time alignment, the time-varying emotion vector can express the continuous changes between explanations, reminders, and reassurances.

[0027] This embodiment uses emotion encoding on training speech to form an emotion feature sequence, then uses this sequence to determine an emotion parameter sequence, and aligns the emotion parameter sequence to a unified time scale. This transforms a static emotion value into a time-varying emotion representation. The correspondence between training text and training speech ensures that emotion changes fit the boundaries of the speech content, and time alignment allows emotion parameters of different granularities to be included on the same timeline, thereby reducing the problems of monotonous emotion and abrupt transitions in the entire speech.

[0028] S20, construct a hybrid expert speech segmentation structure, extract semantic features and identity features from the training speech, and jointly train the hybrid expert speech segmentation structure based on the semantic features, time-varying sentiment vector and identity features, and use the jointly trained hybrid expert speech segmentation structure to perform speech segmentation on the training speech to obtain a speech token sequence; In this embodiment, a hybrid expert speech segmentation structure is used to convert continuous training speech into discrete speech token sequences, while preserving semantic content, emotional variations, and identity differences in the discrete results. The hybrid expert speech segmentation structure may include a shared speech coding layer, a semantic preservation expert branch, an emotional dynamics preservation expert branch, an identity preservation expert branch, an expert gating unit, and a token quantization unit. The shared speech coding layer receives the acoustic representation corresponding to the training speech and outputs an initial speech representation. The acoustic representation may consist of Mel spectrum, fundamental frequency curve, energy curve, speech rate variations, and pause boundaries. The shared speech coding layer may use convolutional layers to extract local articulation features, and then use a temporal modeling layer to preserve the variation relationships between continuous articulations, so that the initial speech representation simultaneously includes phoneme-level variations and segment-level acoustic states.

[0029] Semantic features can be extracted from training speech, reflecting the pronunciation content and semantic boundaries of the training speech. In implementation, the training speech can be input into a speech recognition model, and intermediate semantic representations can be obtained from the intermediate encoding layer of the model. Compared to the final recognized text, intermediate semantic representations can preserve the correspondence between pronunciation time position and semantic content, making them more suitable for speech segmentation training. Intermediate semantic representations can be aggregated according to the segmentation time scale of a hybrid expert speech segmentation structure, ensuring that the temporal granularity of the semantic features remains consistent with the initial speech representation. These semantic features are then used as supervision information for the semantic preservation expert branch, enabling the semantic branch representation output by the semantic preservation expert branch to retain the content information in the training speech.

[0030] Identity features can be extracted from training speech, reflecting the speaker's timbre, vocal tract features, vocal habits, and differences in tone quality. In implementation, the training speech is input into a speaker analysis model, and the raw identity representation is obtained from the model's identity representation layer. This raw identity representation is then normalized and projected, ensuring that identity features, semantic features, and time-varying sentiment vectors reside in a jointly processable representation space. The identity features are subsequently used as supervisory information for the identity preservation expert branch, enabling the identity branch representation output by the expert branch to retain the timbre differences and identity differentiation information in the training speech.

[0031] Time-varying sentiment vectors represent the emotional state of training speech at different time points. During joint training, the hybrid expert speech segmentation structure needs to adjust the time-varying sentiment vectors to the time scale of the initial speech representation. In implementation, this can be achieved by resampling, segment aggregation, or boundary alignment of the time-varying sentiment vectors according to the time boundaries of the initial speech representation, resulting in a supervised sentiment representation. This supervised sentiment representation then enters the supervision process of the emotion dynamics preservation expert branch, ensuring that the sentiment branch representation output by the emotion dynamics preservation expert branch retains the emotional intensity, intonation changes, and emotional transitions in the training speech.

[0032] Before participating in joint training, semantic features and identity features can be scaled uniformly according to the time scale of the initial speech representation. Semantic features can be aggregated, resampled, or boundary-aligned based on the time positions of the intermediate coding layers of the speech recognition model and the time positions of the initial speech representation to obtain a semantically supervised representation. Identity features can be projected to obtain an identity vector with the same dimension as the initial speech representation, and then extended to each time position of the initial speech representation to obtain an identity-supervised representation. Thus, semantic, sentiment, and identity-supervised representations can all correspond to semantic, sentiment, and identity branch representations on the same time scale.

[0033] Semantic features, time-varying sentiment vectors, and identity features form different supervision directions during joint training. Semantic supervision constrains the semantic preservation expert branch, sentiment supervision constrains the sentiment dynamics preservation expert branch, and identity supervision constrains the identity preservation expert branch. The expert gating unit determines the participation ratio of different expert branches based on the semantic, sentiment, and identity supervision representations. In implementation, the expert gating unit can receive the concatenation result or similarity result of the three types of supervision representations and output the gating weights for the corresponding semantic preservation expert branch, sentiment dynamics preservation expert branch, and identity preservation expert branch. The gating weights are used to fuse the semantic branch representation, sentiment branch representation, and identity branch representation to generate a fused speech representation.

[0034] Joint training can be performed based on multiple constraints. The matching degree between the semantic branch representation and the semantic supervision representation constrains the semantic preservation capability; the matching degree between the sentiment branch representation and the sentiment supervision representation constrains the sentiment dynamic preservation capability; and the matching degree between the identity branch representation and the identity supervision representation constrains the identity preservation capability. The fused speech representation is then processed by a token quantization unit to obtain candidate speech tokens. The token quantization unit can contain a trainable codebook, and the fused speech representation is mapped to a codebook vector. The index corresponding to the codebook vector is used to form discrete candidate speech tokens. The quantization difference between the codebook vector corresponding to the candidate speech token and the fused speech representation constrains the discretization process, ensuring that the discrete result retains necessary acoustic information while reducing the complexity of the continuous representation.

[0035] During training, the input acoustic representation of the shared speech coding layer can use an 80-dimensional memristor spectrum, and the fundamental frequency curve, energy curve, and pause boundary markers are input simultaneously. The hidden dimension of the shared speech coding layer can be set to 256, 512, or 768. The output dimensions of the semantic preservation expert branch, the emotion dynamics preservation expert branch, and the identity preservation expert branch are kept consistent and can be set to 256 or 512, so that the expert gating unit can perform weighted fusion of the representations of different branches. The gating temperature of the expert gating unit can be set to 0.5 to 2.0; when the gating temperature is lower, the difference in the participation ratio of different expert branches is more obvious, and when the gating temperature is higher, the participation ratio of different expert branches is smoother.

[0036] The number of codebooks in the token quantization unit can be set to 1 or 2, and the number of codebook vectors in each codebook can be set to 512, 1024, or 2048. The dimension of the codebook vectors is consistent with the dimension of the fused speech representation, and can be set to 256 or 512. Candidate speech tokens are obtained by matching the codebook vectors with the fused speech representation. During training, the codebook vectors are updated according to the quantization difference between the codebook vectors corresponding to the candidate speech tokens and the fused speech representation, so that the codebook vectors cover different pronunciation states, emotional states, and timbre states in the training speech.

[0037] The joint training loss can be obtained by weighting semantic matching error, sentiment matching error, identity matching error, and quantization error. The weights for semantic matching error can be set to 0.2 to 0.4; sentiment matching error to 0.3 to 0.5; identity matching error to 0.1 to 0.3; and quantization error to 0.05 to 0.2. In fintech applications, voice prompts for risk warnings, transaction anomalies, repayment reminders, and identity verification require high semantic accuracy and emotional urgency; therefore, the semantic matching error weight can be set to 0.35, and the sentiment matching error weight to 0.45. In healthcare applications, voice prompts for medication reminders, chronic disease follow-ups, and rehabilitation guidance require high stability of tone and consistency of timbre; therefore, the sentiment matching error weight can be set to 0.4, and the identity matching error weight to 0.3.

[0038] When updating parameters, the learning rate can be set to 0.0001 to 0.001; the batch size can be set to 16, 32, or 64; and the number of training epochs can be set to 50 to 200. The shared speech coding layer and the three expert branches can use the same learning rate. The learning rate of the expert gating unit can be set to 0.5 to 1 times the learning rate of the shared speech coding layer, and the codebook update rate of the token quantization unit can be set to 0.001 to 0.01. To reduce fluctuations between different supervision items, the changes in semantic matching error, sentiment matching error, identity matching error, and quantization error can be statistically analyzed after each training epoch, and the corresponding error weights can be adjusted according to the changes.

[0039] When training is stopped, both the magnitude of the change in joint training loss and the number of training epochs can be used for judgment. Parameter updates can be stopped when the decrease in joint training loss is less than 0.001 within 5 consecutive training epochs; parameter updates can also be stopped when the number of training epochs reaches 200. After stopping parameter updates, the parameters of the shared speech coding layer, semantic preservation expert branch, sentiment dynamics preservation expert branch, identity preservation expert branch, expert gating unit, and token quantization unit are saved to obtain the jointly trained hybrid expert speech segmentation structure.

[0040] A jointly trained hybrid expert speech segmentation structure is used to segment training speech. In implementation, the training speech is passed through a shared speech coding layer to obtain an initial speech representation. This initial representation is then passed through multiple expert branches to obtain branch representations. An expert gating unit determines the participation ratio, and the branch representations are weighted and fused to obtain a fused speech representation. A token quantization unit maps the fused speech representation to discrete speech tokens. Multiple discrete speech tokens are arranged in the temporal order of the training speech to form a speech token sequence. The speech token sequence preserves the semantic content, time-varying emotional state, and identity differences in the training speech, providing discrete speech representations for subsequent training.

[0041] In one implementation, training speech is converted into Mel spectrum and fundamental frequency curve. A shared speech coding layer uses a multi-scale convolutional structure to extract acoustic representations at different time spans. The intermediate semantic representation output by the speech recognition model is temporally aggregated according to the output frame rate of the shared speech coding layer to obtain semantic features. The speaker analysis model outputs the original identity representation, which is normalized and linearly projected to obtain identity features. The time-varying sentiment vector is resampled according to the temporal boundaries of the initial speech representation to obtain the sentiment supervision representation. The semantic features and identity features are respectively formed into semantic supervision representation and identity supervision representation according to the time scale of the initial speech representation. The semantic preservation expert branch uses a temporal coding layer, the sentiment dynamic preservation expert branch uses an attention layer, and the identity preservation expert branch uses a residual mapping layer. The expert gating unit outputs the participation ratio of the three branches, and the token quantization unit maps the fused speech representation into discrete indices to obtain a speech token sequence.

[0042] Alternatively, a segmented alignment approach can be used. The training speech is first divided into multiple speech segments based on pronunciation pauses and text boundaries. Each speech segment is input into a shared speech coding layer to obtain an initial segment-level speech representation. Semantic features are aggregated according to the speech segment boundaries, identity features are extracted across the entire training speech and copied to each speech segment, and time-varying sentiment vectors are aggregated according to the speech segment boundaries to form segment sentiment representations. During joint training, the semantic preservation expert branch primarily constrains segment semantics, the sentiment dynamics preservation expert branch primarily constrains emotional changes between segments, and the identity preservation expert branch primarily constrains timbre consistency between different segments. This approach is suitable for scenarios where the training speech has obvious pauses and clear speech segmentation.

[0043] Another implementation method is to enhance the representation of business-sensitive content. Based on the aforementioned training text, business-sensitive content such as amount, risk, verification, repayment, credit granting, review, medication, and observation is identified, and the corresponding pronunciation intervals are determined according to the correspondence between the training text and training speech. The shared speech coding layer outputs a denser initial speech representation in the pronunciation intervals corresponding to the business-sensitive content, and semantic features and time-varying sentiment vectors are scaled uniformly according to the same time granularity. The expert gating unit increases the participation ratio of the semantic preservation expert branch and the sentiment dynamic preservation expert branch in the pronunciation intervals corresponding to the business-sensitive content, and increases the participation ratio of the identity preservation expert branch in the stable explanatory intervals. This implementation method enables the voice token sequence to retain finer semantic and sentiment differences near key content.

[0044] In fintech business scenarios, training voice recordings can be used to deliver account risk alerts to intelligent customer service. The voice content includes financial keywords such as abnormal transactions, account verification, fund security, and repayment reminders. The speech recognition model extracts semantic features from the training voice, preserving the pronunciation of amounts, risk levels, and verification actions. A speaker analysis model extracts identity features, ensuring the same customer service representative's voice remains consistent across different prompt segments. A time-varying sentiment vector represents the urgency changes in risk warning segments and the stable tone in instructional segments. The resulting voice token sequence, after joint training, retains the content boundaries, emotional changes, and timbre characteristics of financial communication.

[0045] In healthcare scenarios, training voice messages can be used for chronic disease follow-up reminders, medication reminders, or rehabilitation guidance. The voice content includes information such as blood pressure monitoring, timely medication administration, check-up reminders, and feedback on physical condition. Semantic features are used to retain the health reminder content, identity features are used to maintain a stable tone for the same voice actor, and time-varying emotion vectors are used to express emotional changes between explanations, reminders, and reassurances. After joint training, a hybrid expert voice segmentation structure is used to segment the training voice messages to obtain a voice token sequence, ensuring that the voice token sequence balances content accuracy, emotional variation, and tone consistency.

[0046] This embodiment introduces semantic features, time-varying sentiment vectors, and identity features simultaneously during speech segmentation, and performs joint training using a hybrid expert speech segmentation structure. This allows the speech token sequence to simultaneously retain semantic content, sentiment variations, and speaker identity differences. Semantic supervision reduces semantic loss after discretization, sentiment supervision reduces the loss of emotional changes caused by static sentiment modeling, and identity supervision reduces the weakening of timbre information during speech segmentation. The token quantization process converts the fused speech representation into discrete speech tokens that can be used for subsequent processing, thereby improving the speech token sequence's ability to carry multi-dimensional information from the training speech.

[0047] S30, based on the number of tokens in the voice token sequence, the time-varying emotion vector is mapped and aggregated to obtain an emotion guidance token sequence corresponding to the voice token sequence; In this embodiment, the voice token sequence is formed from the front-end speech segmentation results. Each voice token in the sequence represents a discrete speech state of the training speech within a certain time range. The number of tokens is used to determine how many discrete control positions the emotional information needs to be divided into. In implementation, the length of the voice token sequence can be read, and the position of each voice token in the sequence can be recorded. The arrangement of the voice tokens preserves the order of pronunciation in the training speech, allowing continuous emotional changes to be assigned to corresponding discrete speech positions.

[0048] Time-varying emotion vectors (TVRs) consist of emotion vectors at multiple time points. Each emotion vector may include dimensions such as dominance, arousal, valence, tone intensity, stationarity, and urgency. TVRs are typically arranged according to acoustic frames, speech segments, or text pronunciation intervals, and their temporal granularity may differ from that of the voice token sequence. Processing requires reading the temporal boundaries, number of vectors, and emotion values ​​at each time point of the TVRs, and adjusting the TVRs to a discrete granularity acceptable to the voice token sequence.

[0049] Mapping is used to assign time-varying sentiment vectors over continuous time to various speech token positions in a speech token sequence. In implementation, the time range containing the time-varying sentiment vector can be divided into an equal number of sentiment intervals based on the number of tokens in the speech token sequence. Each sentiment interval corresponds to a speech token position. For training speech with relatively stable speech rate, the division can be uniform according to the time length. For training speech with significant sentiment changes, a non-uniform division can be performed based on the location of the sentiment change, allowing for finer-grained segmentation of segments with greater sentiment changes.

[0050] Aggregation is used to combine multiple sentiment vectors within the same sentiment interval into a single token-level sentiment representation. In implementation, sentiment vectors within the sentiment interval can be averaged, or weighted according to temporal distance, the magnitude of sentiment change, or the weight of vocal boundaries. If the sentiment change within the sentiment interval is relatively gradual, averaging yields a stable token-level sentiment representation. If there are significant sentiment jumps within the sentiment interval, weighting increases the participation proportion of key time positions, allowing the token-level sentiment representation to retain local sentiment changes.

[0051] The emotion guidance token sequence consists of multiple emotion guidance tokens arranged in the order of the speech token sequence. Each emotion guidance token can be obtained from a token-level emotion representation through vector mapping, embedding mapping, or discrete codebook matching. Vector mapping is suitable for scenarios where subsequent models receive continuous emotion representations, embedding mapping is suitable for scenarios where subsequent models receive emotion input with a uniform dimension, and discrete codebook matching is suitable for scenarios where subsequent models receive emotion control information with discrete indices. The number of emotion guidance tokens is consistent with the number of tokens in the speech token sequence, and each emotion guidance token forms a positional relationship with a speech token at the same position.

[0052] The relationship between voice token sequences is established through positional order. In implementation, the first emotion guidance token can be assigned to the first voice token, and subsequent emotion guidance tokens can be assigned to subsequent voice tokens in the same order. If the voice token sequence contains localized dense regions created by pauses, prolonged sounds, or short pronunciations, the corresponding emotion intervals can be adjusted according to the time range of the voice tokens, allowing the emotion guidance tokens within these dense regions to reflect finer-grained emotional changes.

[0053] In one implementation, the number of tokens in the voice token sequence is directly used as the number of segments, and the time-varying sentiment vector is uniformly divided into multiple sentiment intervals according to the complete time range. The sentiment vectors within each sentiment interval are averaged to obtain a token-level sentiment representation. The token-level sentiment representation is converted into a fixed-dimensional vector through a linear mapping layer, and then used to form a sentiment guidance token sequence according to the arrangement order of the voice token sequence. This method is suitable for training voice data with relatively stable speech pronunciation speed and continuous sentiment changes.

[0054] Alternatively, a change-aware implementation can be used. The emotional differences between adjacent time positions in the time-varying emotional vector are detected to obtain the emotional change locations. These emotional change locations are used to divide emotional intervals, resulting in shorter emotional intervals for periods of greater emotional change and longer emotional intervals for periods of lesser change. The emotional vectors within each emotional interval are weighted according to the magnitude of the emotional change, and then weighted and aggregated. The aggregation result is input into the emotional embedding layer to obtain an emotional guidance token sequence. This method is suitable for training speech in scenarios where there are tonal variations such as reminders, explanations, reassurances, and warnings.

[0055] Another approach is to use pronunciation boundary assistance. When the training speech already has text pronunciation boundaries or pause boundaries, these boundaries can be used as a reference for dividing emotional regions. The number of speech tokens controls the total number of emotional regions, while the pronunciation boundaries adjust the position of the region boundaries. During aggregation, emotional vectors near the pronunciation boundaries can be assigned boundary weights to preserve emotional changes near text segment transitions. This method is suitable for training speech with clear utterance structure and obvious pauses.

[0056] In fintech business scenarios, training voice messages can include account anomaly alerts, transaction risk notifications, repayment date reminders, and identity verification guidance. Time-varying sentiment vectors exhibit high urgency at the account anomaly and transaction risk locations, a stable reminder state at the repayment date location, and a calm explanation state at the identity verification guidance location. The number of tokens in the voice token sequence is used to divide sentiment intervals. Sentiment vectors within the risk notification interval aggregate into urgency-type sentiment guidance tokens, while sentiment vectors within the verification guidance interval aggregate into stable sentiment guidance tokens, ultimately forming a sentiment guidance token sequence corresponding to the position in the voice token sequence.

[0057] In healthcare scenarios, training voice can include reminders for blood pressure monitoring, timely medication administration, follow-up appointment arrangements, and guidance on physical condition feedback. The time-varying emotion vector presents a neutral explanatory state in the monitoring reminder section, a stable reminder state in the medication administration reminder section, and a gentle, reassuring state in the feedback guidance section. After mapping and aggregating the time-varying emotion vector based on the number of tokens in the voice token sequence, each voice token position receives a corresponding emotion guidance token, ensuring that the variations between explanations, reminders, and reassurances are preserved in the discrete speech representation.

[0058] This embodiment determines the discrete allocation granularity of time-varying emotion vectors by the number of tokens in the voice token sequence, and aggregates the emotion vectors within each emotion interval. This converts continuous emotion changes into emotion guidance token sequences corresponding to the voice token positions. Emotional information is transformed from the global state of the entire speech segment into local control information for each voice token position. This reduces the problems of monotonous emotion expression and abrupt transitions caused by static emotion input, and improves the ability of subsequent speech generation to utilize fine-grained emotion changes.

[0059] S40, the emotion guidance token sequence and the voice token sequence are combined in an alternating order to form an alternating token sequence, and the training text and the alternating token sequence are input into the language model to be trained to obtain the prediction token sequence; In this embodiment, the emotion guidance token sequence is formed by arranging multiple emotion guidance tokens in chronological order, with each emotion guidance token carrying the emotion state of the corresponding speech position. The speech token sequence is formed by arranging multiple speech tokens in chronological order, with each speech token carrying the discrete acoustic content of the corresponding speech position. The two types of tokens are corresponding in quantity and position. When combining them, it is necessary to keep the emotion guidance tokens and speech tokens at the same time position adjacent to each other, so that the language model can read the emotion state and speech content simultaneously during the prediction process.

[0060] The interleaving order can be achieved by arranging the emotion guidance token before its corresponding voice token. In implementation, voice tokens are read sequentially according to their order, and the corresponding emotion guidance token is placed before each voice token, resulting in an interleaved arrangement of emotion guidance tokens, voice tokens, emotion guidance tokens, and voice tokens. During the interleaving process, the positional relationship between each pair of emotion guidance tokens and voice tokens must be preserved to avoid mismatches between emotion guidance tokens and non-corresponding voice tokens.

[0061] Interleaved token sequences are used to carry the combined results of emotion guidance token sequences and speech token sequences. Adjacent pairs of emotion guidance tokens and speech tokens in an interleaved token sequence can be considered as combined units at the same speech location. In implementation, different type embeddings can be set for emotion guidance tokens and speech tokens, or the same group position encoding can be set for the same combined unit, enabling the language model to distinguish token types and identify the positional relationship between emotion guidance tokens and their corresponding speech tokens.

[0062] Training text provides the textual conditions for generating speech content. The training text can undergo word segmentation, text unit partitioning, pronunciation normalization, and business term standardization before being converted into a textual conditional representation. This textual conditional representation can be output by the text encoding layer or directly generated by the text embedding layer in the language model. The text units in the training text need to be consistent with the training speech content so that the language model learns the correspondence between the text content and the interleaved token sequence.

[0063] The language model to be trained is used to generate predicted token sequences based on training text and interleaved token sequences. The language model can employ an autoregressive decoding structure or a sequence prediction structure with textual conditional input. During input, the textual conditional representation corresponding to the training text can serve as a conditional prefix, and the interleaved token sequence can serve as the input sequence corresponding to the predicted object. Internally, the model processes the textual conditional representation and interleaved token sequence through token embedding layers, positional encoding layers, and attention layers, and generates the predicted token corresponding to each position at the output layer.

[0064] The predicted token sequence is output by the language model to be trained, and its arrangement corresponds to the arrangement of the interleaved token sequence. The predicted token sequence can simultaneously include both sentiment guidance token predictions and speech token predictions, or it can output the prediction distribution for each interleaved position separately. In implementation, after receiving the preceding interleaved information and the conditional representation of the training text at the current position, the model outputs the prediction result for the next or current position, enabling the predicted token sequence to be used for subsequent training loss determination.

[0065] In one implementation, the emotion guidance token sequence and the voice token sequence have the same length. The processing unit reads the two types of tokens sequentially according to their position indices, placing each emotion guidance token before its corresponding voice token to obtain an interleaved token sequence. The training text is processed through a text encoding layer to generate a text conditional representation, which serves as the input prefix for the language model. The interleaved token sequence is then processed through a token embedding layer and a position encoding layer before being input into the language model. The language model outputs a predicted token sequence with the same length as the interleaved token sequence. This approach is suitable for training data where the emotion guidance tokens and voice tokens have already been positionally aligned.

[0066] Alternatively, a combined unit embedding approach can be used. Each emotion guidance token and its corresponding speech token constitute a combined unit, with the emotion-first, speech-last arrangement preserved within the combined unit. Each combined unit is assigned a group position code, and the emotion guidance token and speech token are assigned type embeddings respectively. After the training text generates a text conditional representation, it is input into the language model to be trained along with the staggered token sequence obtained from the expanded combined units. This approach can enhance the model's ability to recognize the relationship between emotional state and acoustic content within the same speech location.

[0067] Text-based conditional fusion can also be used. After text encoding, the training text yields multiple text conditional vectors. The interleaved token sequence is then embedded to obtain an interleaved token representation. The language model simultaneously receives the text conditional vectors and interleaved token representations in the attention layer, enabling each interleaved token position to read the corresponding content from the training text. For text units such as amount, risk, verification, and repayment in financial business language, the participation weight of the corresponding text conditional vector at relevant interleaved token positions can be increased. For text units such as medication, observation, follow-up, and feedback in medical and health texts, the relevant interleaved token positions can read the corresponding text conditional vector.

[0068] This embodiment combines emotion guidance token sequences and speech token sequences in an interleaved order. This allows the language model to read emotional states and speech content in adjacent positions. The training text simultaneously provides semantic conditions, ensuring that the predicted token sequence is constrained by the text content, emotion guidance information, and discrete speech representation. The interleaved arrangement reduces the ambiguity of emotion location caused by using emotional information as a global conditional input, ensuring that each speech token has a corresponding emotional input nearby. This helps improve the matching degree between emotional changes and speech content during speech generation.

[0069] S50, determine the training loss based on the predicted token sequence and the interleaved token sequence, and update the language model to be trained based on the training loss to obtain the trained language model; In this embodiment, the predicted token sequence is output by the language model to be trained, and the staggered token sequence serves as the training target reference. When determining the training loss, it is necessary to establish a correspondence between each predicted position in the predicted token sequence and each target position in the staggered token sequence. If the language model to be trained uses the next position prediction method, the predicted token sequence can be aligned with the staggered token sequence by one bit offset, so that each predicted token corresponds to the next target token in the staggered token sequence. If the language model to be trained uses the current position prediction method, it can be directly aligned according to the same sequence position. After the position correspondence is determined, the prediction distribution, prediction vector, or prediction index in the predicted token sequence can be compared with the target tokens in the staggered token sequence to determine the difference.

[0070] The interleaved token sequence contains both sentiment guidance tokens and speech tokens. The training loss can be determined holistically based on the interleaving positions or separately for each token category. When determined holistically, all target tokens in the interleaved token sequence are used as supervision, enabling the language model to learn the alternating generation relationship between sentiment guidance tokens and speech tokens. When determined separately, the sentiment guidance token positions and speech token positions can be processed separately to obtain sentiment guidance token prediction errors and speech token prediction errors, which are then combined into the training loss. The sentiment guidance token prediction error is used to constrain sentiment state prediction, while the speech token prediction error is used to constrain discrete speech content prediction.

[0071] The training loss can consist of token prediction error, sequence continuity error, and interleaving consistency error. Token prediction error is determined based on the target difference between the predicted token sequence and the interleaved token sequence, and can be calculated using cross-entropy error, vector distance error, or codebook indexing error. Sequence continuity error is determined based on the magnitude of change between adjacent predicted tokens and is used to reduce aberrant jumps in emotion guidance tokens or speech tokens at adjacent positions. Interleaving consistency error is determined based on the degree of matching between the prediction results of adjacent emotion guidance tokens and speech tokens, and is used to ensure consistency between the emotional state and speech content near the same speech location.

[0072] The language model to be trained can include a text embedding layer, a token embedding layer, a positional encoding layer, an attention layer, a feedforward transform layer, and a prediction output layer. The text embedding layer receives the text conditional representation corresponding to the training text; the token embedding layer receives the interleaved token sequence; the positional encoding layer provides permutation information for different token positions; the attention layer fuses the text conditional and interleaved token contexts; the feedforward transform layer performs a non-linear transformation on the context representation output by the attention layer; and the prediction output layer generates the prediction token sequence. The prediction output layer can use a unified output space to jointly output the prediction distribution for both sentiment guidance tokens and speech tokens; alternatively, it can use separate sentiment prediction output layers and speech prediction output layers to output the sentiment guidance token prediction distribution and speech token prediction distribution, respectively. The output dimension of the sentiment prediction output layer can correspond to the sentiment guidance token space, and the output dimension of the speech prediction output layer can correspond to the number of speech token codebook entries.

[0073] The embedding dimension of the language model to be trained can be set to 512, 768, or 1024. The attention layers can be set to 6 to 24 layers. The number of attention heads in each attention layer can be set to 8, 12, or 16. The hidden dimension of the feedforward transform layer can be set to 2 to 4 times the embedding dimension. The maximum sequence length can be set to 1024, 2048, or 4096. If the interleaved token sequence length exceeds the maximum sequence length, training can be performed in segments according to the training text boundaries or speech pause boundaries, and overlapping tokens can be retained between adjacent segments.

[0074] Once the training loss is formed, it can be applied to the prediction output layer, feedforward transform layer, attention layer, positional encoding layer, token embedding layer, and text embedding layer, allowing the language model to be trained to gradually adjust its parameters. When updating the language model based on the training loss, the direction of parameter update can be determined according to the training loss, and the model parameters can be adjusted accordingly. Parameter update objects can include text embedding parameters, token embedding parameters, positional encoding parameters, attention parameters, feedforward transform parameters, and output layer parameters. The training process can be performed on batches of speech and text data, with each batch containing training text and interleaved token sequences. The prediction token sequence is output by the language model to be trained based on the training text and the interleaved token sequence. After each round of parameter updates, the prediction token sequence can be regenerated, and the training loss can be determined again.

[0075] The trained language model is formed by the updated model structure. The trained state is represented by text embedding parameters, token embedding parameters, positional encoding parameters, attention parameters, feedforward transformation parameters, and predicted output parameters. The trained language model retains the correspondence between the training text and the interleaved token sequence, and also retains the interleaved generation relationship between sentiment guidance tokens and speech tokens. Through this training, the trained language model can generate token outputs that match the text content and sentiment state when processing task texts in subsequent tasks.

[0076] In one implementation, the language model to be trained employs an autoregressive decoding structure. The interleaved token sequence, shifted one position to the right, serves as the token input, while the original interleaved token sequence serves as the prediction target. The training text is processed through a text embedding layer to obtain a text conditional representation, which is then input into the language model as a conditional prefix. The prediction output layer outputs the corresponding token probability distribution at each interleaving position. The training loss uses cross-entropy error, and the weights for the sentiment guidance token position and the speech token position can be set separately. The weight for the sentiment guidance token position can be set to 0.4 to 0.6; the weight for the speech token position can be set to 0.6 to 0.8. The learning rate can be set to 0.0001 to 0.001. The batch size can be set to 16, 32, or 64. The number of training epochs can be set to 50 to 200.

[0077] Conditional masking prediction can also be used. During training, a portion of the sentiment guidance token positions and a portion of the speech token positions in the interleaved token sequence are selected and masked. The language model to be trained predicts the masked positions based on the training text and the unmasked interleaved tokens. The training loss is determined based on the prediction results of the masked positions and the target tokens at the same positions in the interleaved token sequence. The masking ratio can be set to 0.1 to 0.4. The sentiment guidance token positions and speech token positions can use the same masking ratio, or a higher masking ratio can be set for the sentiment guidance token positions, for example, setting the masking ratio for the sentiment guidance token positions to 0.3 and the masking ratio for the speech token positions to 0.2, to improve the model's ability to predict sentiment states.

[0078] Another approach is to use a grouped error update method. The positions of the sentiment guidance tokens in the interleaved token sequence form the sentiment prediction group, and the positions of the speech tokens form the speech prediction group. The prediction token sequence is divided into sentiment prediction results and speech prediction results based on their positions. The sentiment prediction group uses the sentiment guidance token prediction error and adjacent sentiment continuity error, while the speech prediction group uses the speech token prediction error and adjacent speech continuity error. The training loss is formed by weighting the two sets of errors. In fintech businesses, content such as risk warnings, transaction anomalies, repayment reminders, and identity verification has high requirements for the synchronization of sentiment changes and semantic cues, so the weight of the sentiment prediction group can be increased. In healthcare businesses, content such as medication reminders, follow-up feedback, and rehabilitation guidance has high requirements for smooth tone transitions, so the weight of adjacent sentiment continuity error can be increased.

[0079] Training can be stopped based on both loss convergence and the maximum number of training epochs. Parameter updates can be stopped when the training loss decreases by less than 0.001 over five consecutive training epochs. Parameter updates can also be stopped when the number of training epochs reaches 200. After training stops, the parameters of the text embedding layer, token embedding layer, positional encoding layer, attention layer, feedforward transform layer, and prediction output layer are saved to obtain the trained language model. If the training data includes financial prompts and medical / health reminders, the training loss for each type of data can be calculated separately, and the proportion of the corresponding batch data can be adjusted when the loss for one type of data decreases slowly.

[0080] In fintech business scenarios, training input data can include transaction alert texts, account risk warning texts, loan repayment reminder texts, identity verification reminder texts, and corresponding interleaved token sequences. The sentiment guidance tokens in the interleaved token sequence can carry the sentiment states corresponding to risk warnings, anomaly alerts, repayment reminders, and verification guidance, while the voice tokens can carry the discrete acoustic content of the corresponding voice segments. After the language model to be trained outputs the predicted token sequence, the token prediction errors for the amounts, risk levels, transaction anomalies, and identity verification locations can be determined. The training loss can assign higher weights to the risk warning and verification guidance locations, enabling the trained language model to better learn the changing relationships between urgent reminders, stable explanations, and operational guidance in financial language.

[0081] In healthcare scenarios, training input data can include medication reminder texts, chronic disease follow-up texts, rehabilitation guidance texts, indicator observation reminder texts, and corresponding interleaved token sequences. The emotional guidance tokens in the interleaved token sequence can carry the emotional states corresponding to observation instructions, medication reminders, follow-up appointments, and status feedback, while the voice tokens can carry the discrete acoustic content of the corresponding voice segments. After the language model to be trained outputs the predicted token sequence, the token prediction errors for the observation instruction position, medication reminder position, follow-up appointment position, and status feedback position can be determined respectively. The training loss can increase the participation ratio of adjacent emotional continuity errors, enabling the trained language model to better learn the natural transitions between instructions, reminders, and reassurances in healthcare reminder content.

[0082] This embodiment determines the training loss by comparing the predicted token sequence with the staggered token sequence, and uses the training loss to update the language model to be trained. This allows the model to learn the correspondence between training text, sentiment guidance tokens, and speech tokens. The training loss simultaneously constrains the accuracy of token prediction, the continuity of adjacent positions, and the consistency of staggered positions. This reduces mismatches between sentiment guidance information and speech content, mitigates the problem of singular sentiment expression caused by global sentiment control, and improves the synchronization between sentiment changes and speech content during the generation process.

[0083] S60: Receive task text, process the task text using the trained language model, and obtain target speech.

[0084] In this embodiment, the task text can be text content to be converted into speech. This text content can come from manual input, prompts generated by the business system, text filled into a message template, or responses that need to be broadcast during the voice interaction process. After receiving the task text, character standardization, punctuation correction, number pronunciation conversion, abbreviation expansion, and pause marker addition can be performed on the text content to ensure that the input content can be stably read by the trained language model. In fintech businesses, the task text can contain terms such as account balance, transaction risk, repayment date, credit limit, identity verification, and abnormal login. In healthcare businesses, the task text can contain terms such as medication time, indicator observation, follow-up reminders, rehabilitation training, and physical status feedback. Amounts, dates, dosages, frequencies, and other content in the text can be converted into pronunciations suitable for speech output, reducing pronunciation ambiguity in subsequent speech generation.

[0085] The trained language model is used to generate outputs that express speech content based on task text. The trained language model may include a text embedding layer, a positional encoding layer, an attention layer, a feedforward transform layer, and an output layer. The text embedding layer converts characters, words, or text units in the task text into text representations. The positional encoding layer adds permutation information to the text units. The attention layer reads the contextual relationships within the task text and identifies content segments that need to be expressed with different tones. The feedforward transform layer transforms the representation output by the attention layer. The output layer generates an intermediate representation or speech token corresponding to the target speech. The model input is the task text or a text representation derived from the task text, and the model output can be a target speech token sequence, an acoustic representation, or a speech-generated representation for speech reconstruction.

[0086] When the task text is processed by the trained language model, it can first be divided into text units. Text units can be segments separated by characters, words, phrases, punctuation marks, or semantic segments. After segmentation, a sequence of task text units can be formed, and a positional representation is generated for each text unit. This sequence of task text units is then fed into the trained language model, where a text embedding layer generates a conditional representation of the task text. This conditional representation carries the semantic content, text order, and pause positions of the task text. For text units with clear expressive needs, such as those related to amounts, risks, reminders, or reassurance, the conditional representation can also carry business type markers or tone cue markers, enabling the model to distinguish between different content segments such as prompts, notifications, explanations, and guidance when generating target speech.

[0087] When processing task text, the trained language model can employ an autoregressive generation approach. The model generates a speech representation corresponding to the current speech output position based on the conditional representation of the task text, and then uses the generated speech representation as context input for subsequent generation. The generation process can continue until an output end marker is reached, the corresponding generation length of the text is reached, the maximum generation length is reached, or a pause termination condition is met. This approach is suitable for longer task texts, allowing for segment-by-segment generation of the target speech, ensuring continuity of text content, pause positions, and intonation changes during the output process.

[0088] The target speech can be playable audio or a speech waveform obtained after speech reconstruction. If the trained language model directly outputs an acoustic representation, this representation can be input into a vocoder or speech reconstruction network to obtain the target speech. If the trained language model outputs a target speech token sequence, this sequence can be input into a speech reconstruction network, which will then reconstruct the audio waveform. The speech reconstruction network can include a token embedding layer, an acoustic decoding layer, and a waveform synthesis layer. The token embedding layer converts the target speech tokens into a continuous acoustic representation, the acoustic decoding layer generates a Mel spectrum or other acoustic parameters, and the waveform synthesis layer outputs the target speech.

[0089] The generation of target speech can be adjusted by incorporating text pauses, intonation, and pronunciation duration. Punctuation marks, numbers, amounts, dates, medical indicators, and risk warnings in the task text can influence pronunciation rhythm. The trained language model can generate output with pauses and speech rhythm based on the conditional representation of the task text. When generating target speech, the speech reconstruction network can determine pitch, energy, duration, and pauses based on the acoustic representation output by the model, ensuring that the target speech corresponds to the task text content.

[0090] In one implementation, the task text, after text normalization, is input into the trained language model. Text normalization includes conversion of monetary amounts and dates, generation of punctuation and pause markers, and segmentation of business terms. The trained language model employs an autoregressive decoding structure. The text embedding layer outputs a conditional representation of the task text, and the attention layer reads the conditional representation and generates a target speech token position by position. The target speech token is input into the speech reconstruction network to generate a Mel spectrum, which is then synthesized by a vocoder to produce the target speech. The model embedding dimension can be set to 512, 768, or 1024; the maximum text length can be set to 256, 512, or 1024; the generated length can be estimated based on the task text length and average pronunciation duration.

[0091] Alternatively, a segmented generation approach can be used. The task text is divided into multiple text segments based on punctuation, business term boundaries, or semantic fragments, and each text segment generates its own segment speech representation. The trained language model generates a corresponding speech token for each text segment and pause markers between adjacent text segments. The speech reconstruction network concatenates the speech representations of each segment in sequence to generate the complete target speech. This method is suitable for longer texts containing multiple types of content, such as risk warnings followed by operation guidance, or health instructions followed by medication reminders.

[0092] Another approach is to use business-type tags to assist in generation. Before the task text is input into the trained language model, business type tags are generated based on the text content. In fintech, business type tags can be generated for text units such as abnormal transactions, account verification, repayment reminders, risk disclosure, and credit approval. In healthcare, business type tags can be generated for text units such as medication reminders, indicator monitoring, follow-up appointments, rehabilitation suggestions, and status feedback. The business type tags and task text units are input together into the trained language model, enabling the model to distinguish the pronunciation rhythm and tone of different content segments when generating target speech.

[0093] Regarding the generation parameter settings, the temperature parameter can be set to 0.7 to 1.2 during autoregressive generation; the candidate token truncation ratio can be set to 0.8 to 0.95; the maximum generation length can be estimated based on the number of characters in the task text, the number of text segments, and the average speech rate. If the task text contains sensitive content such as amounts, risk levels, follow-up examination times, or medication frequency, the generation temperature can be lowered to make the pronunciation more stable. If the task text contains reassuring, explanatory, or guiding content, the pause duration and tone smoothness can be appropriately increased to make the target speech more suitable for interactive scenarios.

[0094] Speech reconstruction networks can employ a combination of acoustic decoders and waveform synthesizers. The acoustic decoder receives the target speech token sequence and outputs acoustic parameters, which may include Mel-frequency spectra, fundamental frequency curves, energy curves, and duration information. The waveform synthesizer generates the target speech based on these acoustic parameters. The acoustic decoder can use convolutional, recurrent, or attention-based structures. The waveform synthesizer can employ a neural vocoder. After generation, loudness normalization, silence clipping, and sampling rate unification can be performed on the target speech to ensure it meets the requirements of the playback environment.

[0095] In fintech business scenarios, the task text could be something like: "An account has experienced abnormal transactions. Please complete identity verification within fifteen minutes to prevent further financial risk." Upon receiving the task text, the system can segment the text into units representing abnormal transactions, identity verification, and financial risk, and convert the fifteen minutes into a suitable audio reading format. The trained language model generates a target voice token based on the task text, and the speech reconstruction network converts the target voice token into target speech. The target speech can have clearer pauses and a stronger prompting tone in the risk warning section, and a more stable guiding tone in the identity verification section.

[0096] In healthcare scenarios, task text might include instructions to continue monitoring recent blood pressure fluctuations, take medication as prescribed, and provide feedback on your health status at the next follow-up visit. Upon receiving the task text, the system can segment the text into units representing blood pressure fluctuations, continued monitoring, medication adherence, and follow-up feedback. The trained language model generates corresponding speech representations based on the task text, and the speech reconstruction network outputs the target speech. The target speech should maintain a steady delivery in the observation instructions, provide clear prompts in the medication reminders, and employ a gentle tone in the feedback guidance.

[0097] This embodiment processes the task text using a trained language model and converts the model output into target speech, enabling the model to read the text content as a pronounceable speech representation during the generation process. After normalization and text unit segmentation, the task text can form stable inputs containing information such as amounts, dates, risk warnings, and medication reminders. The trained language model generates an intermediate representation or voice token corresponding to the target speech based on the conditional representation of the task text. Then, the speech reconstruction network generates the target speech, thereby reducing ambiguity in text pronunciation and awkward speech expression, and improving the matching degree between the target speech and the task text content.

[0098] In one embodiment, step S10 includes: S101, Obtain training speech and training text corresponding to the training speech, perform text-speech forced alignment on the training text and the training speech to obtain a text-speech alignment mark; S102, pause detection is performed on the training speech to obtain pause detection results, and based on the text speech alignment marker and the pause detection results, boundary-preserving segmentation is performed on the training speech to obtain an emotion coding segment sequence; S103, for each emotion coding segment in the emotion coding segment sequence, extract segment acoustic features, and extract contextual acoustic features from the emotion coding segments adjacent to each emotion coding segment; S104, the segment acoustic features and the context acoustic features are fused to obtain the emotional features corresponding to each emotional coding segment, and the emotional features corresponding to each emotional coding segment are arranged according to the segment order in the emotional coding segment sequence to obtain the emotional feature sequence; S105, Based on the text-speech alignment mark, a set of local emotion windows is constructed in the emotion feature sequence. Each local emotion window in the set of local emotion windows covers multiple emotion features, and the window boundary of each local emotion window matches the text unit boundary in the text-speech alignment mark. S106, based on the time distance between each emotion feature in each local emotion window and the text unit boundary corresponding to the emotion coding segment to which each emotion feature belongs, determine the attention weight of each emotion feature in each local emotion window, and perform weighted aggregation of each emotion feature in each local emotion window according to the attention weight to obtain a local emotion representation sequence; S107, Input the local emotion representation sequence into the emotion parameter mapping structure to obtain the emotion parameter sequence; S108, Obtain the acoustic frame boundary sequence corresponding to the training speech, wherein the acoustic frame boundary sequence includes the frame-level temporal boundary of the training speech during the acoustic feature extraction process; S109, using the text-speech alignment marker as the boundary anchor point, the emotion parameter sequence is time-aligned according to the acoustic frame boundary sequence, the emotion parameters in the emotion parameter sequence are mapped to the acoustic frames corresponding to the acoustic frame boundary sequence, and the emotion parameters corresponding to each acoustic frame between adjacent boundary anchor points are smoothed to obtain a time-varying emotion vector.

[0099] In this embodiment, a content consistency relationship needs to be established between the training speech and the training text. The training speech retains pronunciation time, pause positions, energy changes, fundamental frequency changes, and speech rate changes, while the training text retains text units consistent with the content of the training speech. Text units can be characters, words, phrases, or semantic segments. During processing, the training text is divided into text units, the training speech is acoustically time-axis-calibrated, and then forced text-speech alignment is used to map the text units to the pronunciation intervals in the training speech, resulting in text-speech alignment markers. These text-speech alignment markers contain the start and end positions of each text unit in the training speech, used for subsequent segmentation, window construction, and time alignment.

[0100] Forced text-to-speech alignment can be achieved through a combination of acoustic models and text pronunciation sequences. The training text is converted into a sequence of phonetic units through word-to-phoneme conversion, and the training speech is processed by acoustic feature extraction to obtain a frame-level acoustic representation. The acoustic model then matches the frame-level acoustic representation with the sequence of phonetic units, outputting the start and end positions for each text unit. If the training text contains numbers, amounts, dates, or abbreviations, pronunciation normalization can be performed first to ensure consistency between the text units and the pronunciation content in the training speech. After the text-to-speech alignment markers are generated, each text unit has a temporal range that can be used to locate emotional changes.

[0101] Pause detection results are used to supplement speech boundaries that cannot be fully represented by text-speech alignment markers. Pause detection can be obtained based on short-term energy, zero-crossing rate, fundamental frequency missing intervals, silence duration, and speech rate variations. Intervals in the training speech where energy is below a threshold and persists for a set duration can be marked as pause intervals, and the two ends of these intervals can serve as candidate segmentation points. Text-speech alignment markers reflect text content boundaries, while pause detection results reflect actual pronunciation intervals. Combining the two allows segmentation positions to simultaneously align with both text content and acoustic pauses.

[0102] Boundary-preserving segmentation is used to obtain a sequence of emotion-coded segments. During segmentation, the boundaries of text units in the text-speech alignment markers are used as the primary boundaries, and pause positions from pause detection results are used as auxiliary boundaries. If a pause position exists near the text unit boundary, the segmentation point can be adjusted to that pause position; if there is no obvious pause between consecutive text units, the text unit boundaries can be preserved or adjacent text units can be merged to form a longer emotion-coded segment. Each emotion-coded segment in the sequence has a clear start position, end position, and segment order, facilitating subsequent extraction of segment acoustic features and contextual acoustic features.

[0103] Segment acoustic features are used to represent the acoustic state within the current emotion-coded segment. Segment acoustic features can include Mel-spectral statistics, fundamental frequency mean and rate of change, energy mean and rate of change, speech rate features, pause adjacency features, and spectral envelope variations. During extraction, acoustic features are read frame-by-frame within each emotion-coded segment, and then pooling, temporal coding, or attention convergence are used to form a segment-level representation. Segment acoustic features preserve the emotional intensity, intonation, speech stability, and rhythmic variations within the current phonological segment.

[0104] Contextual acoustic features are used to represent the acoustic state of adjacent segments of the current sentiment encoding segment. During extraction, similar acoustic features can be read from the previous and next sentiment encoding segments to generate a contextual representation. If the current sentiment encoding segment is located at a sequence boundary, the contextual acoustic features can be supplemented using blank vectors, boundary marker vectors, or by copying adjacent segments. Contextual acoustic features can reflect the trend of sentiment transition, avoiding local abrupt changes caused by judging the sentiment state based on only a single segment.

[0105] Segment acoustic features and contextual acoustic features are fused to form sentiment features. Fusion methods can include post-concatenation linear mapping, gated fusion, attention fusion, or residual fusion. Post-concatenation linear mapping is suitable for uniform dimensionality, gated fusion is suitable for adjusting the participation ratio based on the differences between the current segment and adjacent segments, and attention fusion is suitable for highlighting the source of change when there are significant sentiment shifts between segments. Each sentiment-encoded segment generates corresponding sentiment features, which are then arranged according to the segment order in the sentiment-encoded segment sequence to form a sentiment feature sequence. The sentiment feature sequence retains the acoustic state of each segment itself as well as the sentiment transition information brought by adjacent segments.

[0106] Local sentiment windows are used to form localized sentiment observation regions within a sentiment feature sequence. During construction, window boundaries are determined based on text unit boundaries in the text-speech alignment markers, ensuring that each local sentiment window covers multiple consecutive sentiment features. Matching window boundaries to text unit boundaries ensures that the sentiment features within the window align with the text content segment. A local sentiment window can cover the sentiment features corresponding to a single text unit or multiple adjacent text units. The window length can be adjusted based on text unit duration, pause positions, and the number of sentiment-coded segments.

[0107] Attention weights are determined based on the temporal distance between each sentiment feature and the text unit boundary within the local sentiment window. Each sentiment feature belongs to a sentiment coding segment with a time interval, and the text unit boundaries are provided by text-speech alignment markers. During processing, the distance between the temporal position corresponding to the sentiment feature and the nearest text unit boundary can be determined, as can the distance between the temporal position corresponding to the sentiment feature and the start and end boundaries of its respective text unit. The closer the distance is to the key boundary, the higher the participation ratio of the sentiment feature in the local sentiment representation can be; when the distance is in the middle of the segment and changes smoothly, the participation ratio can be relatively smooth. Attention weights are used to weight and aggregate the sentiment features, enabling the local sentiment representation to retain sentiment variations near the text boundaries.

[0108] The local sentiment representation sequence is formed by arranging the weighted aggregate results of each local sentiment window in chronological order. Each local sentiment representation integrates multiple sentiment features within the window and highlights sentiment information related to text boundary changes through attention weights. Local sentiment representations can use fixed-dimensional vectors, with the dimension consistent with the input dimension of the subsequent sentiment parameter mapping structure. Compared to sentiment feature sequences, local sentiment representation sequences exhibit more stable local sentiment expression and can reduce the impact of acoustic fluctuations in individual segments on sentiment parameters.

[0109] The sentiment parameter mapping structure is used to convert a sequence of local sentiment representations into a sequence of sentiment parameters. The structure includes a normalization layer, a nonlinear transformation layer, a temporal smoothing layer, and a parameter output layer. The normalization layer reduces differences in recording volume, speaking habits, and segment length. The nonlinear transformation layer extracts the intensity and direction of emotion from local sentiment representations. The temporal smoothing layer handles the transition between adjacent local sentiment representations. The parameter output layer outputs a dominance parameter sequence, an arousal parameter sequence, and a valence parameter sequence. The dominance parameter sequence represents the strength of tone control, the arousal parameter sequence represents the degree of emotional activation, and the valence parameter sequence represents the emotional tendency. These three parameter sequences are arranged in the same temporal order and combined to form the sentiment parameter sequence.

[0110] Acoustic frame boundary sequences provide frame-level temporal boundaries for training speech during acoustic feature extraction. These sequences can be determined based on frame length, frame shift, and sampling rate, with each acoustic frame having a start and end time. Emotion parameter sequences may be generated according to local emotion windows, with a different temporal granularity than the acoustic frame boundary sequences; therefore, they require temporal alignment with the acoustic frame boundary sequences. The acoustic frame boundary sequences provide a finer temporal scale, enabling emotion parameters to be mapped to frame-level temporal locations.

[0111] Boundary anchors are provided by text-speech alignment markers. The start and end positions of each text unit can serve as temporal anchors for sentiment parameter mapping. During temporal alignment, sentiment parameters in the sentiment parameter sequence are mapped to the acoustic frames corresponding to the acoustic frame boundary sequence. If an acoustic frame is located near the text unit boundary, the sentiment parameters of the corresponding text unit can be used preferentially; if the acoustic frame is located between adjacent boundary anchors, smoothing compensation can be performed based on the distance between the acoustic frame and the adjacent boundary anchors. Smoothing compensation can employ linear interpolation, distance weighting, spline smoothing, or gated smoothing to ensure continuous changes between adjacent sentiment parameters.

[0112] The time-varying sentiment vector is formed by combining aligned sentiment parameters in the temporal order of acoustic frames. Each acoustic frame corresponds to a sentiment vector containing dimensions such as dominance, arousal, and valence. The time-varying sentiment vector preserves both the semantic segmentation caused by text unit boundaries and the temporal changes at the acoustic frame level, enabling it to express the changing state of sentiment in the training speech over time. Compared to a single sentiment value for the entire speech, the time-varying sentiment vector can provide different sentiment parameters at different acoustic frame positions.

[0113] This embodiment determines the start and end positions of text units in the training speech through forced text-to-speech alignment. Combined with pause detection results, boundary-preserving segmentation of the training speech is performed, making the emotion-encoded segments closer to the boundaries of real pronunciation and text content. The fusion of segment acoustic features and contextual acoustic features forms an emotion feature sequence, enabling the emotional state to simultaneously reflect the transition relationship between the current segment and adjacent segments. A set of local emotion windows is constructed according to the text-to-speech alignment markers, and attention weights are determined based on the temporal distance between the emotion features and the text unit boundaries, allowing local emotion representations to more effectively preserve emotional changes near the text boundaries. After the emotion parameter mapping structure outputs the dominance parameter sequence, arousal parameter sequence, and valence parameter sequence, it is time-aligned according to the acoustic frame boundary sequence and smoothed, resulting in frame-level continuous time-varying emotion vectors. This reduces the problems of monotonous emotion expression and abrupt transitions caused by static emotion modeling of the entire segment.

[0114] In one embodiment, step S20 above includes: S201, Construct a hybrid expert speech segmentation structure, which includes a shared speech coding layer, a semantic preservation expert branch, an emotion dynamics preservation expert branch, an identity preservation expert branch, an expert gating unit, and a token quantization unit; S202, the training speech is input into the automatic speech recognition model, the semantic feature sequence is extracted from the intermediate coding layer of the automatic speech recognition model, and the semantic feature sequence is time-aggregated according to the speech segmentation time scale corresponding to the hybrid expert speech segmentation structure to obtain semantic features that match the speech segmentation time scale. S203, input the training speech into the speaker analysis model, extract the original identity representation from the identity representation layer of the speaker analysis model, and normalize and project the original identity representation to obtain the identity features; S204, the training speech is input into the shared speech coding layer to obtain an initial speech representation, and the semantic features, the time-varying sentiment vector and the identity features are scaled according to the time scale of the initial speech representation to obtain a semantic supervision representation, a sentiment supervision representation and an identity supervision representation; S205, the initial speech representation is input into the semantic preservation expert branch, the emotional dynamics preservation expert branch and the identity preservation expert branch respectively to obtain the semantic branch representation, the emotional branch representation and the identity branch representation; S206, Based on the semantic supervision representation, the sentiment supervision representation, and the identity supervision representation, the participation ratio of the semantic preservation expert branch, the sentiment dynamic preservation expert branch, and the identity preservation expert branch is determined by the expert gating unit; S207, Based on the participation ratio, the semantic branch representation, the emotion branch representation, and the identity branch representation are weighted and fused to obtain a fused speech representation; S208, the fused speech representation is input into the token quantization unit, and the fused speech representation is quantized based on the codebook vector in the token quantization unit to obtain candidate speech tokens; S209, based on the matching degree between the semantic supervision representation and the semantic branch representation, the matching degree between the sentiment supervision representation and the sentiment branch representation, the matching degree between the identity supervision representation and the identity branch representation, and the quantization difference between the codebook vector corresponding to the candidate speech token and the fused speech representation, determine the joint training loss; S210, update the parameters of the shared speech coding layer, the semantic preservation expert branch, the emotion dynamics preservation expert branch, the identity preservation expert branch, the expert gating unit and the token quantization unit according to the joint training loss to obtain the jointly trained hybrid expert speech segmentation structure; S211, The training speech is segmented using the jointly trained hybrid expert speech segmentation structure to obtain a speech token sequence.

[0115] In this embodiment, a hybrid expert speech segmentation structure is used to convert continuous speech representations into discrete speech tokens, while preserving semantic, emotional, and identity information during the discretization process. A shared speech coding layer performs unified acoustic representation extraction; inputs can include Mel spectrum, fundamental frequency curve, energy curve, speech rate variation features, and pause boundary features, with the output being the initial speech representation. Semantic preservation expert branches, emotional dynamic preservation expert branches, and identity preservation expert branches receive the initial speech representation in parallel, generating semantic branch representations, emotional branch representations, and identity branch representations, respectively. An expert gating unit determines the participation ratio of each expert branch based on the semantic supervision representation, emotional supervision representation, and identity supervision representation. A token quantization unit receives the fused speech representation and maps it to a codebook vector to form candidate speech tokens.

[0116] The semantic feature sequence is output by the intermediate coding layer of the automatic speech recognition model. This intermediate coding layer, typically located between the acoustic input layer and the text decoding layer, preserves the correspondence between pronunciation temporal positions, phoneme transitions, and semantic content. Compared to the final recognized text, the output of the intermediate coding layer maintains strong temporal resolution, making it more suitable for matching with the speech segmentation time scale. During processing, after training the automatic speech recognition model with the input speech, the intermediate coding layer outputs a semantic feature sequence arranged by frame or segment, which is then temporally aggregated according to the speech segmentation time scale corresponding to the hybrid expert speech segmentation structure. Temporal aggregation can employ average aggregation, attention aggregation, boundary-weighted aggregation, or resampling to ensure that the semantic features have a consistent temporal granularity with the subsequent initial speech representation.

[0117] The speech segmentation timescale is used to determine the temporal granularity at which semantic features participate in speech segmentation training. The speech segmentation timescale can be determined based on the output frame rate of the shared speech coding layer, the output frequency of the token quantization unit, or the target token interval in the training speech. If the frame rate of the semantic feature sequence output by the automatic speech recognition model is higher than the speech segmentation timescale, multiple consecutive semantic features can be aggregated into a single semantic feature. If the frame rate of the semantic feature sequence output by the automatic speech recognition model is lower than the speech segmentation timescale, it can be extended to the corresponding time position through interpolation or duplication. The resulting semantic features can serve as a source of supervision for the semantic preservation expert branch.

[0118] Identity features are output by the speaker analysis model. The speaker analysis model can include an acoustic input layer, a speaker embedding layer, and an identity representation layer. After training the speaker analysis model on the speech input, the identity representation layer outputs the raw identity representation. The raw identity representation is typically a vector at the whole-speech level or segment level, containing timbre, vocal habits, vocal tract features, and acoustic stability information. Normalization is used to reduce amplitude shifts caused by differences in volume and recording conditions, and projection mapping is used to transform the raw identity representation into a feature space that can be processed together with semantic features and time-varying sentiment vectors. Projection mapping can be implemented by linear layers, nonlinear layers, or multi-layer perceptual structures, and the output is the identity feature.

[0119] The initial speech representation is generated by a shared speech coding layer. This shared speech coding layer can employ a multi-scale convolutional structure to extract short-term articulation changes and acoustic trends over longer time spans, or it can combine convolutional layers with self-attention layers to incorporate both local articulation states and cross-segment variations into the initial speech representation. After training the shared speech coding layer with the speech input, acoustic frames or segments are converted into initial speech representations of a unified dimension. The temporal scale of the initial speech representation becomes a reference for subsequent scale unification, enabling semantic features, time-varying sentiment vectors, and identity features to correspond with the outputs of the three expert branches.

[0120] Scale unification is used to adjust semantic features, time-varying sentiment vectors, and identity features to the time scale of the initial speech representation. Semantic features can be aggregated or resampled based on the correspondence between the time positions of the intermediate coding layers in the automatic speech recognition model and the time positions of the initial speech representation to obtain a semantic supervised representation. Time-varying sentiment vectors can be resampled, aligned to boundaries, or aggregated into segments based on the time boundaries of the initial speech representation to obtain a sentiment supervised representation. Identity features can be copied to various time positions, expanded at the segment level, or projected over time to form an identity supervised representation. After scale unification, the semantic supervised representation, sentiment supervised representation, and identity supervised representation correspond to the semantic branch representation, sentiment branch representation, and identity branch representation in terms of time position, respectively.

[0121] The semantic preservation expert branch receives the initial speech representation and outputs a semantic branch representation. This branch can employ temporal coding layers, convolutional coding layers, or attention coding layers to preserve phonological content and semantic boundaries. The emotion dynamics preservation expert branch receives the initial speech representation and outputs an emotion branch representation. This branch can employ time-sensitive attention layers, gated recurrent layers, or local change enhancement layers to preserve tone intensity, emotional changes, and emotional transitions between adjacent segments. The identity preservation expert branch receives the initial speech representation and outputs an identity branch representation. This branch can employ residual mapping layers, speaker conditional normalization layers, or identity projection layers to preserve timbre stability and identity differentiation information.

[0122] The expert gating unit generates participation ratios based on semantic, emotional, and identity-supervised representations. The expert gating unit can concatenate, calculate similarity, or perform attention transformations on the three types of supervisory representations to obtain the gating weights for the corresponding semantic-preserving, emotional-dynamic-preserving, and identity-preserving expert branches. Participation ratios can be time-position-varying weights or segment-level weights. Time-position-varying weights are suitable for speech with dense emotional changes, while segment-level weights are suitable for speech with clear pause boundaries. Participation ratios are used to adjust the proportion of the three branch representations in the fused speech representation.

[0123] Weighted fusion is used to synthesize semantic branch representations, sentiment branch representations, and identity branch representations into a fused speech representation. In implementation, the three branch representations can be weighted and summed position-by-position according to the participation ratio output by the expert gating unit, or the three branch representations can be concatenated first, and then a linear transformation can be performed based on the participation ratio. The fused speech representation simultaneously includes semantic content, time-varying sentiment state, and identity differences, serving as input to the token quantization unit. The dimension of the fused speech representation can be consistent with the codebook vector dimension, or it can be converted to a codebook matching dimension through a projection layer.

[0124] The token quantization unit is used to convert the fused speech representation into candidate speech tokens. Each token quantization unit can contain one or more trainable codebooks, each containing multiple codebook vectors. During processing, the fused speech representation is matched against the codebook vectors using distance or similarity matching. The codebook vector with the highest matching degree is selected, and its corresponding index is determined as the candidate speech token. Candidate speech tokens are discrete indices, while the corresponding codebook vectors are continuous vectors. The quantization difference is determined based on the difference between the codebook vector corresponding to the candidate speech token and the fused speech representation, and is used to constrain the information deviation between the continuous representation and the discrete index during discretization.

[0125] The joint training loss is determined by multiple matching relationships. The matching degree between the semantic supervision representation and the semantic branch representation is used to form the semantic matching error. The matching degree between the sentiment supervision representation and the sentiment branch representation is used to form the sentiment matching error. The matching degree between the identity supervision representation and the identity branch representation is used to form the identity matching error. The quantization difference between the codebook vector corresponding to the candidate speech token and the fused speech representation is used to form the quantization error. The joint training loss can be obtained by weighting the semantic matching error, sentiment matching error, identity matching error, and quantization error. The weights can be set according to the semantic complexity, sentiment variation amplitude, and identity stability requirements in the training speech.

[0126] When updating model parameters based on the joint training loss, the shared speech coding layer, semantic preservation expert branch, emotion dynamics preservation expert branch, identity preservation expert branch, expert gating unit, and token quantization unit jointly participate in parameter adjustment. The shared speech coding layer adjusts the acoustic representation extraction method through error signals, the three expert branches adjust their ability to preserve semantics, emotion, and identity respectively, the expert gating unit adjusts the participation ratio of different expert branches, and the token quantization unit adjusts the codebook vector distribution. After training, a jointly trained hybrid expert speech segmentation structure is obtained.

[0127] The training process for a hybrid expert speech segmentation structure can utilize training speech, semantic supervision representation, sentiment supervision representation, and identity supervision representation. The training speech undergoes acoustic feature extraction to obtain Mel spectrum, fundamental frequency curve, energy curve, and pause boundary markers, which serve as input to a shared speech coding layer. The semantic supervision representation is derived from the output of the intermediate coding layer of the automatic speech recognition model and resampled or aggregated according to the output time scale of the shared speech coding layer. The sentiment supervision representation is obtained by aligning the time-varying sentiment vector to the time scale of the initial speech representation. The identity supervision representation is obtained from the identity features output by the speaker analysis model through projection mapping and temporal expansion.

[0128] The shared speech coding layer can include an acoustic input layer, a multi-scale convolutional layer, a temporal coding layer, and a normalization layer. The acoustic input layer receives an 80-dimensional memristor spectrum and simultaneously receives the fundamental frequency curve, energy curve, and pause boundary markers. The multi-scale convolutional layer can be configured with three parallel convolutional branches, with kernel sizes of 3, 5, and 7, used to extract local articulation variations across different time ranges. The temporal coding layer can employ a bidirectional gated recurrent layer or a self-attention layer, with a hidden dimension of 256, 512, or 768. The normalization layer unifies the feature scale of the temporal coding layer output to obtain an initial speech representation.

[0129] The semantic preservation expert branch can include a semantic projection layer and a temporal preservation layer. The semantic projection layer maps the initial speech representation to the feature space of the semantic supervision representation, and the temporal preservation layer models the semantic changes between adjacent time positions, outputting the semantic branch representation. The emotion dynamics preservation expert branch can include an emotion projection layer, a local change enhancement layer, and a temporal smoothing layer. The emotion projection layer maps the initial speech representation to the feature space of the emotion supervision representation, the local change enhancement layer extracts the emotion changes between adjacent time positions, and the temporal smoothing layer reduces invalid short-term jumps, outputting the emotion branch representation. The identity preservation expert branch can include an identity projection layer and a residual preservation layer. The identity projection layer maps the initial speech representation to the feature space of the identity supervision representation, and the residual preservation layer preserves the timbre consistency across time positions, outputting the identity branch representation.

[0130] The expert gating unit can receive the concatenated results of semantic supervision representation, sentiment supervision representation, and identity supervision representation, as well as the similarity results between the semantic branch representation, sentiment branch representation, and identity branch representation and their corresponding supervision representations. The expert gating unit can include a gating input layer, a nonlinear transformation layer, and a normalized output layer. The gating input layer receives the three types of supervision information, the nonlinear transformation layer generates the gating intermediate representation, and the normalized output layer outputs the participation ratios of the semantic preservation expert branch, the sentiment dynamic preservation expert branch, and the identity preservation expert branch. The participation ratios can be output according to time position, so that different speech segments correspond to different expert branch weights.

[0131] The token quantization unit may include a fusion projection layer, a codebook matching layer, and an index output layer. The fusion projection layer maps the fused speech representation to a codebook vector space. The codebook matching layer searches the codebook for the codebook vector that is closest to or has the highest similarity to the fused speech representation. The index output layer indexes the matched codebook vector as candidate speech tokens. The number of codebooks can be set to one or two, and the number of codebook vectors in each codebook can be set to 512, 1024, or 2048. The dimension of the codebook vectors can be set to 256 or 512, consistent with the dimension of the fused speech representation.

[0132] The joint training process can be performed on batches of data. Each batch of data includes training speech, semantic supervision representation, sentiment supervision representation, and identity supervision representation. The training speech is input into a shared speech coding layer to obtain an initial speech representation. This initial speech representation is then input into the semantic preservation expert branch, the sentiment dynamic preservation expert branch, and the identity preservation expert branch, respectively, to obtain semantic branch representations, sentiment branch representations, and identity branch representations. The expert gating unit generates participation ratios based on the semantic supervision representation, sentiment supervision representation, and identity supervision representation, and performs weighted fusion of the three types of branch representations based on the participation ratios to obtain a fused speech representation. The token quantization unit performs codebook matching on the fused speech representation to obtain candidate speech tokens and their corresponding codebook vectors.

[0133] Semantic matching error can be determined based on the distance or similarity between the semantic supervision representation and the semantic branch representation. Sentiment matching error can be determined based on the multidimensional parameter differences between the sentiment supervision representation and the sentiment branch representation. Identity matching error can be determined based on the similarity between the identity supervision representation and the identity branch representation. Quantization error can be determined based on the distance between the codebook vector corresponding to the candidate speech token and the fused speech representation. The joint training loss is obtained by weighting the semantic matching error, sentiment matching error, identity matching error, and quantization error. The weights for semantic matching error can be set to 0.2 to 0.4; for sentiment matching error, 0.3 to 0.5; for identity matching error, 0.1 to 0.3; and for quantization error, 0.05 to 0.2.

[0134] When updating parameters, the learning rate can be set to 0.0001 to 0.001; the batch training size can be set to 16, 32, or 64; and the number of training epochs can be set to 50 to 200. The shared speech coding layer, semantic preservation expert branch, sentiment dynamics preservation expert branch, and identity preservation expert branch can all use the same learning rate. The learning rate of the expert gating unit can be set to 0.5 to 1 times the learning rate of the shared speech coding layer. The codebook update rate of the token quantization unit can be set to 0.001 to 0.01. Parameter updates can be stopped when the joint training loss decreases by less than 0.001 over five consecutive training epochs; parameter updates can also be stopped when the number of training epochs reaches 200. After stopping parameter updates, the parameters of the shared speech coding layer, semantic preservation expert branch, sentiment dynamics preservation expert branch, identity preservation expert branch, expert gating unit, and token quantization unit are saved to obtain the jointly trained hybrid expert speech segmentation structure.

[0135] When the jointly trained hybrid expert speech segmentation structure performs speech segmentation on the training speech, the training speech is passed through a shared speech coding layer to obtain an initial speech representation. This initial speech representation is then processed by three expert branches to form a semantic branch representation, an emotion branch representation, and an identity branch representation. An expert gating unit determines the participation ratio based on the semantic, emotion, and identity supervision representations, and performs weighted fusion of these representations according to the participation ratio to obtain a fused speech representation. A token quantization unit maps the fused speech representation into discrete speech tokens. Multiple discrete speech tokens are arranged in the chronological order of the training speech to obtain a speech token sequence.

[0136] This embodiment constructs a hybrid expert speech segmentation structure comprising a shared speech coding layer, multiple expert branches, an expert gating unit, and a token quantization unit. Before discretization, the training speech is represented by semantic, emotional, and identity branches. Semantic features, time-varying emotional vectors, and identity features, after scaling, constrain the three expert branches, ensuring that the speech segmentation process simultaneously preserves pronunciation content, emotional variations, and identity differences. The expert gating unit adjusts the participation ratio of different expert branches based on the three types of supervised representations, and the token quantization unit converts the fused speech representation into candidate speech tokens, constraining the discretization error through quantization differences. The resulting jointly trained hybrid expert speech segmentation structure outputs a speech token sequence containing semantic, emotional, and identity information, reducing the problems of insufficient emotional consistency and unstable timbre control caused by the excessive emphasis on semantic content in existing speech segmentation methods.

[0137] In one embodiment, step S30 above includes: S301, obtain the number of tokens in the voice token sequence, and construct a token position sequence according to the time order of each voice token in the voice token sequence, wherein each position index in the token position sequence corresponds to a voice token in the voice token sequence; S302, perform temporal emotion change detection on the time-varying emotion vector, determine the multidimensional emotion difference degree of the time-varying emotion vector between consecutive time frames, and mark the time frame where the multidimensional emotion difference degree meets the preset emotion change condition as a candidate emotion transition point; S303, perform boundary aggregation based on the candidate emotional transition points, merge candidate emotional transition points whose time intervals meet the preset aggregation conditions into aggregated emotional transition points, and use the aggregated emotional transition points as emotional change nodes; S304, determine the number of internal interval boundaries of the target based on the number of tokens, and select boundary emotion change nodes that meet the boundary selection conditions from the emotion change nodes based on the number of internal interval boundaries of the target; S305, construct a target interval boundary set based on the starting boundary of the time-varying sentiment vector, the ending boundary of the time-varying sentiment vector, and the boundary sentiment change nodes; S306, when the number of sentiment mapping intervals formed by the target interval boundary set is less than the number of tokens, the interval boundaries are supplemented along the time direction of the time-varying sentiment vector, and the supplemented interval boundaries are added to the target interval boundary set, wherein the number of sentiment mapping intervals formed by the target interval boundary set is consistent with the number of tokens; S307, Based on the target interval boundary set, the time-varying sentiment vector is dynamically segmented to obtain a sentiment mapping interval sequence, wherein the number of sentiment mapping intervals in the sentiment mapping interval sequence is consistent with the number of tokens; S308, establish a one-to-one mapping relationship between each emotional mapping interval in the emotional mapping interval sequence and each position index in the token position sequence to form an interval token correspondence table; S309, for each emotion mapping interval, determine the set of emotion carrying reference positions from the interval boundary of the corresponding emotion mapping interval and the emotion change nodes located in the vicinity of the corresponding emotion mapping interval, determine the normalized time distance between the time-varying emotion vector of each time frame in each emotion mapping interval and the nearest emotion carrying reference position in the set of emotion carrying reference positions, and determine the emotion carrying weight of each time frame based on the normalized time distance. S310, based on the emotional carrying weight, the time-varying emotional vectors of each time frame in each emotional mapping interval are weighted and aggregated to obtain the initial interval emotional representation corresponding to each emotional mapping interval; S311, Input the initial interval sentiment representation corresponding to each sentiment mapping interval into the trend enhancement network to obtain the sentiment change trend features corresponding to each sentiment mapping interval, and fuse the initial interval sentiment representation and sentiment change trend features corresponding to each sentiment mapping interval to obtain the interval sentiment representation corresponding to each sentiment mapping interval; S312, Arrange the interval sentiment representations corresponding to each sentiment mapping interval according to the order of the token position sequence to form an interval sentiment representation sequence; S313, based on the directional change relationship and amplitude change relationship between adjacent interval emotional representations in the interval emotional representation sequence, and by filling in the missing adjacent change relationships at the beginning and end positions of the interval emotional representation sequence, the emotional change direction feature sequence and emotional fluctuation intensity feature sequence corresponding to each interval emotional representation in the interval emotional representation sequence are determined; S314, input the interval sentiment representation sequence, the sentiment change direction feature sequence and the sentiment fluctuation intensity feature sequence into the sentiment token embedding network to obtain the sentiment guidance token corresponding to each sentiment mapping interval; S315, arrange the emotion guidance tokens corresponding to each emotion mapping interval according to the interval token correspondence table to obtain the emotion guidance token sequence corresponding to the voice token sequence.

[0138] In this embodiment, after obtaining the voice token sequence, the total number of voice tokens in the sequence can be counted, and a token position sequence can be generated according to the order of the voice tokens in the sequence. The position index in the token position sequence is used to record the arrangement position of each voice token in the discrete speech representation, so that subsequent emotion mapping can be performed according to the actual generation order of the voice tokens. The number of tokens is used to determine how many emotion mapping intervals the time-varying emotion vector needs to be divided into, avoiding inconsistencies in the number between continuous emotion representation and discrete voice tokens. The position index can be a continuous integer index or an index marker with a timestamp. If the voice token sequence is obtained from the speech segmentation results at a fixed frame rate, the position index can be generated directly according to the segmentation output order; if the voice token sequence contains pause tokens, extension tokens, or silence tokens, the position index can retain these special voice tokens, so that the emotion guidance tokens can cover pauses, extensions, and pronunciation segments.

[0139] When detecting changes in time-varying sentiment vectors across consecutive time frames, multidimensional sentiment values ​​can be read from adjacent time frames, and the differences between adjacent time frames can be determined. Multidimensional sentiment dissimilarity can be formed by changes in dimensions such as dominance, arousal, valence, tone intensity, stability, and urgency. During processing, the magnitude of change in each sentiment dimension between adjacent time frames can be determined separately, and then the results of changes in multiple dimensions can be synthesized into a multidimensional sentiment dissimilarity. Preset sentiment change conditions can include dissimilarity thresholds, continuous increase conditions, continuous decrease conditions, local peak conditions, or multidimensional joint change conditions. Time frames that meet the preset sentiment change conditions are marked as candidate sentiment transition points, which represent locations in the time-varying sentiment vector where significant sentiment changes may occur.

[0140] Candidate sentiment transition points may be affected by local noise, short-term speech jitter, or frame-level estimation errors, necessitating boundary aggregation. Boundary aggregation determines the distance between candidate sentiment transition points based on time intervals. When multiple candidate sentiment transition points are located within the same local time range and the time intervals meet preset aggregation conditions, they are merged into an aggregated sentiment transition point. Preset aggregation conditions can be determined based on acoustic frame length, speech token time span, or average text segment duration. Aggregation methods can include time-position averaging, maximum difference position preservation, or weighted center position determination. Aggregated sentiment transition points, as sentiment change nodes, can reduce over-segmentation caused by excessively dense candidate points, resulting in more stable subsequent sentiment mapping intervals.

[0141] The number of target internal interval boundaries is determined based on the number of tokens in the speech token sequence. The number of emotion mapping intervals needs to match the number of tokens; therefore, the number of internal interval boundaries, together with the start and end boundaries, needs to form a sufficient number of intervals. During processing, the number of internal boundaries to be selected can be estimated based on the number of tokens, and then boundary emotion change nodes that meet the boundary selection criteria can be selected from the emotion change nodes. Boundary selection criteria can include the intensity of emotion change, the evenness of temporal distribution, the minimum interval between adjacent boundaries, and the degree of matching between the boundary position and the change in speech rhythm. If the number of emotion change nodes exceeds the number of target internal interval boundaries, emotion change nodes with higher intensity and more even distribution can be retained first. If the number of emotion change nodes is insufficient, interval boundaries can be supplemented along the time direction in subsequent processing.

[0142] The target interval boundary set consists of the start boundary, end boundary, and boundary sentiment change nodes of the time-varying sentiment vector. The start and end boundaries cover the entire range of the time-varying sentiment vector, while the boundary sentiment change nodes form internal segmentation points at locations where sentiment changes are significant. When the number of sentiment mapping intervals formed by the target interval boundary set is less than the number of tokens, interval boundaries can be supplemented along the time direction of the time-varying sentiment vector. The supplemented interval boundaries can be determined according to proportional positions, the maximum blank area between adjacent boundaries, the time span of the voice tokens, or local sentiment stability intervals. After supplementation, the number of sentiment mapping intervals formed by the target interval boundary set is consistent with the number of tokens, ensuring that each voice token obtains an sentiment mapping interval.

[0143] Dynamic segmentation divides time-varying sentiment vectors based on a set of target interval boundaries. The boundaries in the target interval boundary set are ordered chronologically, forming a sentiment mapping interval between adjacent boundaries. Each sentiment mapping interval contains time-varying sentiment vectors from several time frames. The sequence of sentiment mapping intervals is arranged chronologically, with the number of intervals matching the number of tokens. Compared to fixed-length segmentation, dynamic segmentation can create boundaries at locations of significant sentiment changes and maintain a consistent number of tokens at locations of gradual sentiment changes by supplementing boundaries, thus balancing the location of sentiment changes with the number of voice tokens.

[0144] The interval token mapping table records the correspondence between emotion mapping intervals and token position sequences. During processing, the first emotion mapping interval can be assigned to the first position index in chronological order, and subsequent emotion mapping intervals can be assigned to subsequent position indices in turn. The interval token mapping table can contain an emotion mapping interval identifier, interval start time, interval end time, corresponding position index, and corresponding voice token identifier. This table is used to restore the arrangement relationship between emotion guidance tokens and voice tokens after the emotion guidance tokens are generated, preventing misalignment between emotion guidance tokens and voice tokens.

[0145] Within each sentiment mapping interval, a set of sentiment-carrying reference locations needs to be determined. This set can consist of the interval boundary and sentiment change nodes within its vicinity. The interval boundary provides the time range of the current interval, while the sentiment change nodes within its vicinity provide a reference for local sentiment changes. For each time frame within the interval, the normalized time distance between the time-varying sentiment vector of that time frame and the nearest sentiment-carrying reference location can be determined. The normalized time distance can be obtained based on the ratio between the actual distance from the time frame to the reference location and the interval length, making distances within intervals of different lengths comparable.

[0146] The emotional carrying weights are determined based on normalized temporal distance. Time frames closer to emotional change nodes or interval boundaries typically carry more significant emotional transition information and can be assigned higher weights; time frames in stable regions can be assigned smoother weights. Weight determination methods can include distance decay, gating functions, piecewise linear functions, or attention mapping. After obtaining the emotional carrying weights, the time-varying emotional vectors of each time frame within each emotional mapping interval are weighted and aggregated to obtain the initial interval emotional representation. The initial interval emotional representation retains the main emotional states within the current emotional mapping interval and reduces the impact of frame-level noise on the position of individual voice tokens.

[0147] Trend enhancement networks are used to extract the changing trends within sentiment mapping intervals. The initial interval sentiment representation reflects the overall sentiment state of the interval, and the trend enhancement network further processes the temporal variation information within the interval. Trend enhancement networks can employ small temporal convolutions, gated recurrent layers, attention-converging layers, or multilayer perceptron structures. Inputs can include the initial interval sentiment representation, the temporal sequence changes of sentiment vectors within the interval, sentiment differences at the start and end positions of the interval, and peak positions within the interval. The output is the sentiment change trend feature. Fusing the initial interval sentiment representation and the sentiment change trend feature yields the interval sentiment representation. Fusion methods can include post-concatenation projection, weighted summation, gated fusion, or residual fusion. The interval sentiment representation simultaneously contains both the overall sentiment state of the interval and the changing trends within the interval.

[0148] Interval sentiment representations are arranged in the order of the token position sequence to form an interval sentiment representation sequence. This sequence has the same number and order as the voice token sequence. The directional change relationship between adjacent interval sentiment representations can reflect whether the sentiment is enhanced, weakened, or reversed between adjacent voice token positions. The amplitude change relationship can reflect the intensity of the sentiment change between adjacent intervals. During processing, the directional feature of sentiment change can be determined based on the difference vector between adjacent interval sentiment representations, and the intensity feature of sentiment fluctuation can be determined based on the difference amplitude. If one adjacent interval is missing at the beginning and end positions, corresponding features can be generated by boundary padding. Boundary padding can use zero vectors, boundary replication, one-sided difference, or learned boundary vectors to ensure that the sequence of directional feature of sentiment change and the sequence of intensity feature of sentiment fluctuation correspond to the sentiment representations of each interval in the interval sentiment representation sequence.

[0149] The sentiment token embedding network receives interval sentiment representation sequences, sentiment change direction feature sequences, and sentiment fluctuation intensity feature sequences, and outputs sentiment guidance tokens corresponding to each sentiment mapping interval. The sentiment token embedding network can include an input concatenation layer, a normalization layer, a nonlinear mapping layer, and a token output layer. The input concatenation layer combines interval sentiment representations, sentiment change direction features, and sentiment fluctuation intensity features at the same position; the normalization layer reduces the numerical scale differences between different dimensions; the nonlinear mapping layer extracts interval sentiment states and change features; and the token output layer generates continuous sentiment guidance vectors or discrete sentiment guidance indices. If the subsequent model receives continuous representations, the sentiment guidance tokens can be fixed-dimensional vectors; if the subsequent model receives discrete indices, the sentiment guidance tokens can be formed through codebook matching or classification output.

[0150] After the emotion guidance tokens are generated, they are arranged according to the interval token mapping table. Each emotion guidance token is placed in the corresponding sequence position according to the correspondence between the emotion mapping interval and the position index. After the arrangement is completed, an emotion guidance token sequence is obtained that corresponds one-to-one with the voice token sequence. Each emotion guidance token in this sequence can be traced back to an emotion mapping interval and corresponds to a voice token in the voice token sequence. Continuous time-varying emotion information is thus converted into emotion control information at discrete positions.

[0151] This embodiment determines the number of emotion mapping intervals by the number of tokens and constructs a target interval boundary set by combining emotion change nodes. This allows the time-varying emotion vector to balance the number of voice tokens and the actual location of emotion changes during discretization. Candidate emotion transition points are aggregated through boundaries to form emotion change nodes, reducing over-segmentation caused by local noise. When the target interval boundary set is insufficient, interval boundaries are supplemented to ensure that the number of emotion mapping interval sequences is consistent with the number of voice token sequences. Within each emotion mapping interval, the emotion carrying weight is determined based on the emotion carrying reference position set, and the time-varying emotion vector is weighted and aggregated to retain representative emotion states within the interval. The trend enhancement network further extracts the change trend within the interval, and the emotion change direction feature sequence and emotion fluctuation intensity feature sequence supplement the change information between adjacent intervals. The emotion token embedding network generates emotion guidance tokens based on the above information and forms an emotion guidance token sequence that corresponds one-to-one with the voice token sequence according to the interval token correspondence table. This reduces the problem of insufficient emotion control granularity caused by global static emotion input, ensuring that each voice token position has corresponding emotion guidance information.

[0152] In one embodiment, step S40 above includes: S401, based on the correspondence between the emotion guidance token sequence and the voice token sequence, a token pairing unit sequence is constructed. Each token pairing unit in the token pairing unit sequence includes an emotion guidance token and a voice token, and the emotion guidance token in each token pairing unit is placed before the voice token. S402, according to the arrangement order of the token pairing unit sequence, each token pairing unit is unfolded in sequence to obtain an interleaved token sequence; S403, perform text encoding on the training text to obtain a conditional representation of the training text; S404, Generate a token type identifier sequence based on the interleaved token sequence, wherein the token type identifier sequence includes an emotion guidance token type identifier and a voice token type identifier; S405, an interleaved position encoding sequence is generated based on the token pairing unit sequence, and the interleaved position encoding sequence records the pairing position of each token pairing unit in the interleaved token sequence; S406, construct an interleaved attention constraint marker based on the token pairing unit sequence, wherein the interleaved attention constraint marker records the binding relationship between each emotion guidance token and the corresponding voice token, as well as the order between adjacent token pairing units; S407, the training text conditional representation, the interleaved token sequence, the token type identifier sequence, the interleaved position encoding sequence, and the interleaved attention constraint label are input into the language model to be trained. In the attention layer of the language model to be trained, the attention weights between the sentiment guidance tokens and the corresponding speech tokens in the interleaved token sequence are adjusted based on the interleaved attention constraint label to obtain the predicted token sequence.

[0153] In this embodiment, each emotion guidance token in the emotion guidance token sequence originates from the preceding emotion mapping result and carries the emotion state of the corresponding speech position. Each speech token in the speech token sequence originates from the speech segmentation result and carries the discrete acoustic content of the corresponding speech position. The two types of tokens need to maintain a correspondence in both quantity and position. Before combining, the sequence lengths of the emotion guidance token sequence and the speech token sequence can be read, and it can be checked whether the emotion guidance tokens and speech tokens at the same position can form a correspondence. If the two types of sequences are completely identical in length, they can be directly paired according to the same position; if there are end-filling tokens or pause tokens, the filling positions or pause positions can be retained so that the interleaved sequences still cover the complete speech time range.

[0154] The token pairing unit sequence carries the positional binding relationship between the emotion guidance token and the voice token. During construction, the emotion guidance token and the voice token are extracted sequentially according to their position in the sequence, and the emotion guidance token and the voice token at the same position are grouped into a token pairing unit. Within each token pairing unit, the emotion guidance token is fixed first, followed by the voice token, allowing the language model to read the emotional state at the same position before processing the voice token. The token pairing unit can record the pairing unit number, the emotion guidance token identifier, the voice token identifier, and the pairing position. The pairing unit number increments sequentially over time to maintain the pronunciation order within the training speech.

[0155] When unfolding each token pairing unit sequentially, the emotion guidance token and voice token in each token pairing unit can be written into the same sequence according to the order of the token pairing unit sequence. The resulting interleaved token sequence retains two layers of order relationship: one is the temporal order between token pairing units, and the other is the local order within each token pairing unit where the emotion guidance token precedes the voice token. The interleaved token sequence can be represented as alternating emotion guidance tokens and voice tokens, or it can retain the token type, pairing unit number, and position index in the internal storage structure, which facilitates the language model in recognizing the role of different tokens.

[0156] Training text is used to provide semantic conditions. The training text can be text-encoded before being input into the language model to generate a conditional representation of the training text. Text encoding can include text unit partitioning, pronunciation normalization, text embedding, positional encoding, and contextual encoding. Text units can be characters, words, phrases, or semantic fragments. Pronunciation normalization handles numbers, dates, abbreviations, and proper nouns, ensuring the text content matches the pronunciation in the training speech. Text embedding converts text units into vector representations, positional encoding supplements the arrangement of text units, and contextual encoding extracts the semantic relationships within the training text. The conditional representation of the training text can serve as conditional input to the language model, enabling the sentiment guidance tokens and speech tokens in the interleaved token sequence to establish a correspondence with the text content.

[0157] The token type identifier sequence is used to distinguish between sentiment guidance tokens and speech tokens in an interleaved token sequence. During generation, a corresponding type identifier can be written to each token position in the interleaved token sequence: sentiment guidance token type identifier for sentiment guidance token positions, and speech token type identifier for speech token positions. The token type identifier can be mapped to a type embedding and added to or concatenated with the token embeddings of the interleaved token sequence. Through the token type identifier sequence, the language model to be trained can distinguish between sentiment control information and speech content information within the same interleaved sequence, avoiding confusion between the two types of tokens in the embedding space.

[0158] The staggered position encoding sequence is used to record the pairing position of each token pairing unit in the staggered token sequence. During generation, the same pairing position encoding can be assigned to both the emotion guidance token and the speech token within the same token pairing unit, or an intra-unit relative position encoding can be added in addition to the same pairing position encoding. The pairing position encoding indicates which speech position the current token belongs to, while the intra-unit relative position encoding indicates whether the current token is an emotion guidance token or a speech token. The staggered position encoding sequence enables the language model to recognize the relationship between emotion guidance tokens and speech tokens within the same speech position, while preserving the sequential relationship between different speech positions.

[0159] Interleaved attention constraint markers are used to record the binding relationship between emotion guidance tokens and their corresponding voice tokens, as well as the sequential order of adjacent token pairs. During construction, an attention constraint matrix or an attention constraint index table can be generated. The row and column positions in the attention constraint matrix correspond to the token positions in the interleaved token sequence. Higher attention permission values ​​can be set between bound emotion guidance tokens and voice tokens, sequential attention permission values ​​can be set between adjacent token pairs, and lower attention permission values ​​can be set between non-corresponding and distant tokens. The attention constraint index table records the voice token position corresponding to each emotion guidance token, as well as the preceding and following positions of each token pair.

[0160] The training text conditional representation, interleaved token sequence, token type identifier sequence, interleaved position encoding sequence, and interleaved attention constraint label are all input into the language model to be trained. During input fusion, the interleaved token sequence can be converted into token embeddings, the token type identifier sequence into type embeddings, and the interleaved position encoding sequence into position embeddings, and then the token embeddings, type embeddings, and position embeddings are fused. The training text conditional representation can be input into the language model as a conditional prefix, or it can interact with the interleaved token representation through a cross-attention layer. After the interleaved attention constraint label enters the attention layer, it can adjust the attention weights between the sentiment guidance token and the corresponding speech token, so that the speech token at the corresponding position can more fully read the sentiment guidance token during modeling.

[0161] The language model to be trained can include a text conditional input layer, an interleaved token embedding layer, a type embedding layer, a positional encoding layer, an attention layer, a feedforward transform layer, and a prediction output layer. When calculating attention weights, the attention layer can convert the interleaved attention constraint labels into attention biases and incorporate the attention score between the sentiment guidance token and the speech token. For sentiment guidance tokens and speech tokens within the same token pairing unit, the attention bias can increase the attention intensity between them; for adjacent token pairing units, the attention bias can preserve the adjacent order dependency; and for non-corresponding tokens, it can reduce unnecessary cross-positional influence. Through these processes, the language model can simultaneously utilize the semantic conditions of the training text, token type information, interleaved positional relationships, and pairing attention relationships when predicting tokens.

[0162] The predicted token sequence is output by the language model to be trained. The prediction output layer can generate corresponding prediction results according to the positions of the interleaved token sequence. The sentiment guidance token position can output a sentiment guidance token prediction distribution or prediction vector, and the speech token position can output a speech token prediction distribution or prediction index. The length of the predicted token sequence can be the same as the interleaved token sequence, or it can correspond to the next position target of the interleaved token sequence in autoregressive training. The predicted token sequence retains the interleaved structure, allowing subsequent training stages to read the prediction results of the sentiment guidance token position and the speech token position separately.

[0163] This embodiment binds emotion guidance tokens and voice tokens at the same location through a token pairing unit sequence, ensuring that the emotion guidance token precedes the voice token in each token pairing unit. This allows the voice token to obtain adjacent emotion conditions during modeling. The interleaved token sequence preserves temporal and intra-unit order, the training text conditional representation provides semantic content, the token type identifier sequence distinguishes between emotion control information and voice content information, the interleaved position encoding sequence preserves the arrangement relationship between the same pairing unit and different pairing units, and the interleaved attention constraint label adjusts the attention weight between the corresponding emotion guidance token and voice token in the attention layer. The resulting predicted token sequence is simultaneously constrained by text semantics, emotion guidance, and voice token structure, reducing the problems of unclear position and voice content mismatch when emotion information is used as a global conditional input, and improving the matching degree between emotion state and corresponding voice position.

[0164] In one embodiment, step S50 above includes: S501, based on the token type identifier sequence and the interleaved position encoding sequence, separate the emotion guidance token prediction subsequence and the voice token prediction subsequence from the prediction token sequence; S502, based on the token type identifier sequence and the interleaved position encoding sequence, extract the emotion guidance token truth value subsequence and the voice token truth value subsequence from the interleaved token sequence; S503, determine the initial sentiment guidance token prediction loss based on the token prediction difference between the sentiment guidance token prediction subsequence and the sentiment guidance token truth value subsequence; S504, determine the emotional continuity loss based on the continuity difference between adjacent emotional guidance token prediction values ​​in the emotional guidance token prediction subsequence; S505, Based on the initial emotion guidance token prediction loss and the emotion continuity loss, determine the emotion guidance token prediction loss; S506, determine the initial voice token prediction loss based on the token prediction difference between the voice token prediction subsequence and the voice token truth subsequence; S507, the voice token prediction subsequence and the corresponding emotion guidance token truth subsequence are mapped to a common evaluation space, and the voice emotion consistency loss is determined according to the matching relationship in the common evaluation space; S508, Based on the initial voice token prediction loss and the voice emotion consistency loss, determine the voice token prediction loss; S509, determine the binding relationship between each emotion guidance token and the corresponding voice token based on the interleaved attention constraint marker, construct a binding token pair based on the binding relationship, and construct an unbound token pair from the emotion guidance token and voice token at non-corresponding positions, and determine the emotion-voice alignment constraint loss based on the representation difference between the binding token pair and the unbound token pair; S510, determine the training loss based on the emotion guidance token prediction loss, the voice token prediction loss, and the emotion-voice alignment constraint loss; S511, determine the parameter update direction of the language model to be trained according to the training loss, and update the language model to be trained according to the parameter update direction to obtain the updated language model; S512, the updated language model is used as the language model to be trained in the next round of training. The determination of training loss and the updating of the language model to be trained are repeated until the training loss corresponding to the updated language model meets the convergence condition. Then the updated language model is determined as the trained language model.

[0165] In this embodiment, the predicted token sequence includes predictions for both the emotion guidance token position and the voice token position. A token type identifier sequence identifies whether each predicted position belongs to an emotion guidance token or a voice token, and an interleaved position encoding sequence identifies the corresponding pairing position for each predicted position. During processing, emotion guidance token positions can be selected based on the token type identifier sequence, and then the original pairing order can be maintained according to the interleaved position encoding sequence to form an emotion guidance token prediction subsequence; simultaneously, voice token positions can be selected and paired in the same order to form a voice token prediction subsequence. The emotion guidance token prediction subsequence and the voice token prediction subsequence maintain a positional correspondence, enabling subsequent loss calculations to separately evaluate emotion prediction and voice prediction.

[0166] The staggered token sequence serves as the training objective. The token type identifier sequence and the staggered position encoding sequence are also used to extract truth subsequences from the staggered token sequence. The sentiment guidance token truth subsequence consists of sentiment guidance tokens from the staggered token sequence, and the voice token truth subsequence consists of voice tokens from the staggered token sequence. During extraction, the staggered position encoding sequence is used to ensure that the sentiment guidance token truth values ​​and voice token truth values ​​at the same paired position correspond to each other in subsequent calculations, avoiding positional confusion that would result from separating them solely by token type.

[0167] The initial sentiment guidance token prediction loss is determined based on the token prediction difference between the predicted subsequence and the true subsequence. If the sentiment guidance token is a discrete index, cross-entropy error can be used to determine the prediction difference; if the sentiment guidance token is a continuous vector, vector distance error, cosine difference, or multidimensional regression error can be used. Each predicted position in the predicted subsequence is compared with the corresponding position in the true subsequence to obtain the positional difference. These positional differences are then weighted and aggregated to form the initial sentiment guidance token prediction loss.

[0168] The emotional continuity loss is determined based on the continuity difference between adjacent emotional guidance token prediction values ​​in the emotional guidance token prediction subsequence. Adjacent emotional guidance token prediction values ​​correspond to the emotional state of adjacent speech positions. During processing, the magnitude, direction, and smoothness of change between adjacent prediction values ​​can be determined. If the change between adjacent positions exceeds a set range, a large continuity difference can be formed. The emotional continuity loss can be obtained using adjacent vector differences, adjacent rate of change differences, or local smoothing errors, ensuring that the predicted emotional guidance tokens maintain a continuous transition in the temporal direction.

[0169] The sentiment guidance token prediction loss is jointly determined by the initial sentiment guidance token prediction loss and the sentiment continuity loss. The initial sentiment guidance token prediction loss constrains the prediction accuracy of each sentiment guidance token position, while the sentiment continuity loss constrains the temporal variation between adjacent sentiment guidance tokens. These two losses can be weighted and combined, with the weights set according to the magnitude of sentiment variation in the training data. Data with gradual sentiment variation can increase the weight of the sentiment continuity loss, while data with dramatic sentiment variation can increase the weight of the initial sentiment guidance token prediction loss, ensuring that the prediction result closely approximates the true value while preserving reasonable variation.

[0170] The initial speech token prediction loss is determined based on the token prediction difference between the predicted speech token subsequence and the ground truth speech token subsequence. If the speech token is a codebook index, cross-entropy error can be used; if the speech token prediction result is a codebook vector or acoustic embedding, vector distance error or codebook distance error can be used. Each predicted position in the predicted speech token subsequence is compared with the corresponding position in the ground truth speech token subsequence to obtain the speech content prediction difference. This loss is used to constrain the language model's ability to predict discrete speech content.

[0171] A common evaluation space is used to compare the matching relationship between the predicted speech tokens and the corresponding ground truth values ​​of the sentiment guidance tokens. The predicted speech token subsequences and their corresponding ground truth value subsequences typically reside in different representation spaces and need to be transformed to the same evaluation space through a mapping layer. This mapping layer can consist of a linear projection layer, a nonlinear mapping layer, or a shared embedding layer. The speech token prediction result is passed through the speech projection layer to obtain the speech evaluation representation, and the sentiment guidance token ground truth value is passed through the sentiment projection layer to obtain the sentiment evaluation representation. Once the two evaluation representations are located in the same dimension, the speech-sentiment consistency loss can be determined using similarity, distance, or matching score.

[0172] The speech-emotion consistency loss is used to evaluate whether the predicted speech token matches the corresponding ground truth of the emotion guidance token. During processing, the ground truth of the emotion guidance token corresponding to each predicted speech token can be determined according to the interleaved position encoding sequence, and then the matching relationship is determined in the common evaluation space. The matching relationship can be obtained from vector similarity, distance difference, or contrastive learning objectives. Speech evaluation representations and emotion evaluation representations at corresponding positions should have a higher degree of matching, while speech evaluation representations and emotion evaluation representations at non-corresponding positions should have a lower degree of matching. This loss can reduce the decoupling between the predicted speech token and the emotion guidance information.

[0173] The voice token prediction loss is jointly determined by the initial voice token prediction loss and the voice emotion consistency loss. The initial voice token prediction loss constrains the accuracy of voice token prediction, while the voice emotion consistency loss constrains the degree of matching between the predicted voice token and the corresponding ground truth value of the emotion guidance token. The two can be weighted and synthesized. If the training objective is more inclined towards acoustic content accuracy, the weight of the initial voice token prediction loss can be increased; if the training objective is more inclined towards emotional expression consistency, the weight of the voice emotion consistency loss can be increased.

[0174] Interleaved attention constraint tags record the binding relationships between emotion guidance tokens and corresponding speech tokens. Based on these tags, bound token pairs can be constructed. Each bound token pair includes an emotion guidance token and a speech token at the same paired location. Unbound token pairs consist of emotion guidance tokens and speech tokens at non-corresponding locations, which can be selected from adjacent distant locations, different paired locations, or other training samples within the same batch. Bound and unbound token pairs are used together to evaluate the emotion-speech alignment constraints.

[0175] The emotion-speech alignment constraint loss is determined based on the representational differences between bound and unbound token pairs. Emotional and speech representations in bound token pairs should have a higher degree of matching in the evaluation space, while emotion and speech representations in unbound token pairs should have a lower degree of matching. Contrast error, interval constraint error, or ordering error can be used in the processing. The loss increases when the similarity of bound token pairs is lower than that of unbound token pairs; the loss decreases when the similarity of bound token pairs is higher than that of unbound token pairs and reaches a set interval. This loss is used to strengthen the emotion-speech binding relationships at corresponding positions in the interleaved structure.

[0176] The training loss is jointly determined by the emotion guidance token prediction loss, the speech token prediction loss, and the emotion-speech alignment constraint loss. These three types of losses constrain emotion prediction, speech prediction, and emotion-speech pairing relationships, respectively. During synthesis, either fixed weights or dynamic weights can be used. Fixed weights can be set to 0.3 to 0.5 for the emotion guidance token prediction loss, 0.4 to 0.6 for the speech token prediction loss, and 0.1 to 0.3 for the emotion-speech alignment constraint loss. Dynamic weights can be adjusted based on the rate of decrease of each loss during training rounds, maintaining a relative balance between different training objectives during training.

[0177] The language model to be trained may include a text embedding layer, a token embedding layer, a type embedding layer, a positional encoding layer, an attention layer, a feedforward transform layer, and a prediction output layer. The text embedding layer receives the conditional representation of the training text; the token embedding layer receives an interleaved token sequence; the type embedding layer receives a token type identifier sequence; the positional encoding layer receives an interleaved positional encoding sequence; the attention layer adjusts the attention intensity between corresponding tokens based on the interleaved attention constraint markers; the feedforward transform layer performs a non-linear transformation on the attention output; and the prediction output layer outputs the predicted token sequence. After the training loss is formed, the update direction of the parameters of each layer can be obtained through gradient backpropagation.

[0178] The parameter update direction is determined by the gradient of the training loss with respect to the model parameters. Updated parameters can include text embedding parameters, token embedding parameters, type embedding parameters, positional encoding parameters, attention parameters, feedforward transformation parameters, and predicted output parameters. The learning rate can be set to 0.0001 to 0.001; the batch size can be set to 16, 32, or 64; and the number of training epochs can be set to 50 to 200. The attention layer can be set to 6 to 24 layers; the number of attention heads can be set to 8, 12, or 16; and the embedding dimension can be set to 512, 768, or 1024.

[0179] The updated language model serves as the language model to be trained in the next round of training. In the next round of training, the training text, interleaved token sequence, token type identifier sequence, interleaved position encoding sequence, and interleaved attention constraint markers are re-inputted, and the predicted token sequence is re-output. The determination of the training loss and the updating of the language model to be trained are repeated. The determination of the training loss includes: separating the sentiment guidance token prediction subsequence and the speech token prediction subsequence from the predicted token sequence; extracting the sentiment guidance token ground truth subsequence and the speech token ground truth subsequence from the interleaved token sequence; determining the initial sentiment guidance token prediction loss, sentiment continuity loss, sentiment guidance token prediction loss, initial speech token prediction loss, speech-sentiment consistency loss, speech token prediction loss, and sentiment-speech alignment constraint loss; and determining the training loss based on the sentiment guidance token prediction loss, speech token prediction loss, and sentiment-speech alignment constraint loss. The updating of the language model to be trained includes: determining the parameter update direction of the language model to be trained based on the training loss; and updating the language model to be trained according to the parameter update direction to obtain the updated language model. During repeated training, the sentiment guidance token prediction loss, speech token prediction loss, and sentiment-speech alignment constraint loss can be recorded separately. The convergence condition can be set based on the magnitude of change in training loss, the magnitude of change in validation set loss, or the maximum number of training epochs. The training loss is considered to have met the convergence condition when the decrease in training loss is less than 0.001 over five consecutive training epochs; updates can also be stopped when the number of training epochs reaches 200. When the convergence condition is met, the updated language model is determined as the trained language model.

[0180] This embodiment separates the prediction token sequence and the interleaved token sequence using a token type identifier sequence and an interleaved position encoding sequence. These can form sentiment guidance token prediction subsequences, speech token prediction subsequences, and corresponding ground truth subsequences, allowing the training loss to evaluate sentiment prediction and speech prediction separately. The initial sentiment guidance token prediction loss, combined with a sentiment continuity loss, constrains the prediction accuracy and continuity of the sentiment guidance tokens in the temporal direction. The initial speech token prediction loss, combined with a speech sentiment consistency loss, constrains the matching relationship between the speech token prediction results and the corresponding sentiment guidance tokens. The sentiment-speech alignment constraint loss further strengthens the sentiment-speech relationship at corresponding positions by using the representation differences between bound and unbound token pairs. Updating the language model to be trained based on these three types of losses can reduce mismatches between sentiment guidance tokens and speech tokens, mitigating the problem of asynchronous emotional changes and speech content in generated speech.

[0181] In one embodiment, step S60 above includes: S601, Receive task text, divide the task text into text units to obtain a task text unit sequence, and generate a task text position sequence based on the task text unit sequence; S602, Input the task text unit sequence and the task text position sequence into the trained language model to obtain the task text conditional representation; S603, generate an initial emotion control state based on the task text condition representation, initialize the historical reasoning emotion guidance token sequence and the historical target voice token sequence, and determine the generation termination condition based on the task text condition representation; S604, At the current generation position, based on the task text conditional representation, the initial emotion control state, the historical inference emotion guidance token sequence, and the historical target speech token sequence, the current inference emotion guidance token is generated through the trained language model; S605, based on the task text conditional representation, the current inference sentiment guidance token, the historical inference sentiment guidance token sequence, and the historical target speech token sequence, the current target speech token is generated through the trained language model; S606, add the current reasoning emotion guidance token to the historical reasoning emotion guidance token sequence to obtain an updated historical reasoning emotion guidance token sequence, and add the current target voice token to the historical target voice token sequence to obtain an updated historical target voice token sequence; S607, the updated historical reasoning emotion guidance token sequence is used as the historical reasoning emotion guidance token sequence corresponding to the next generation position, and the updated historical target voice token sequence is used as the historical target voice token sequence corresponding to the next generation position. The current reasoning emotion guidance token and the current target voice token are repeatedly generated until the generation termination condition is met. S608, the historical reasoning emotion guidance token sequence obtained when the generation termination condition is met is determined as the reasoning emotion guidance token sequence, and the historical target speech token sequence obtained when the generation termination condition is met is determined as the target speech token sequence; S609, the target speech token sequence and the inference emotion guidance token sequence are input into the speech reconstruction network, and the speech reconstruction network performs speech reconstruction on the target speech token sequence based on the inference emotion guidance token sequence to obtain the target speech.

[0182] In this embodiment, after receiving the task text, character normalization, punctuation correction, number pronunciation conversion, abbreviation expansion, and pause mark supplementation can be performed to ensure the text content forms a stable input for speech generation. Text unit segmentation can be performed by separating segments or semantic segments based on characters, words, phrases, punctuation, etc. For task text containing numbers, dates, amounts, dosages, time durations, etc., it can be first converted into a text format suitable for pronunciation before text unit segmentation. The task text unit sequence is generated according to the arrangement order in the task text, with each task text unit retaining its text content and boundary position. The task text position sequence is generated based on the task text unit sequence and is used to record the arrangement position, pause position, and segment boundary of each task text unit in the task text, enabling the trained language model to read the sequential relationship between text units.

[0183] After receiving a sequence of task text units and a sequence of task text positions, the trained language model generates a conditional representation of the task text through a text embedding layer, a positional encoding layer, a context encoding layer, and a conditional representation output layer. The text embedding layer converts the task text units into vector representations, the positional encoding layer writes the task text position sequence into the text representation, the context encoding layer processes semantic and pause relationships between adjacent text units, and the conditional representation output layer generates a conditional representation of the task text that can be used for inference. The conditional representation of the task text can include text content, text order, pause positions, segment boundaries, and tone triggering information. The trained language model can employ an autoregressive decoding structure, which internally includes a text conditional input layer, a token embedding layer, a type embedding layer, a positional encoding layer, an attention layer, a feedforward transform layer, and a prediction output layer.

[0184] The initial sentiment control state is generated from the task text conditional representation. During processing, the task text conditional representation can be input into the sentiment state initialization layer. The sentiment state initialization layer can employ linear mapping, gating transformation, or attention convergence to extract the overall tone tendency, text length, pause distribution, and content strength variations from the task text conditional representation to generate the initial sentiment control state. The initial sentiment control state provides the initial sentiment conditions when no inference sentiment guidance token has yet been generated at the current generation position. The historical inference sentiment guidance token sequence and the historical target speech token sequence can be initialized as empty sequences or have a start marker written in them. The start marker can be an embedding of the start token used during the training phase, ensuring consistency between the input format of the generation process and the training process.

[0185] The generation termination condition can be determined based on the task text conditional representation. Termination conditions may include outputting an end marker, reaching the predicted generation length, reaching the maximum generation length, complete coverage of all text units in the task text, or reaching a set number of consecutive silence or pause tokens. The predicted generation length can be determined based on the number of task text units, the number of text pauses, the average speech rate parameter, and the text-to-speech token length relationship learned by the language model during training. For short texts, a smaller maximum generation length can be set; for long texts containing multiple segments, the maximum generation length can be increased based on the number of segments. The generation termination condition is used to limit the duration of the inference process and prevent the target speech token sequence from being too short or too long.

[0186] The current generation position identifies the location of the token pair to be generated. When generating the current inference sentiment guidance token, the trained language model reads the task text conditional representation, the initial sentiment control state, the historical inference sentiment guidance token sequence, and the historical target speech token sequence. The task text conditional representation provides the text content and text location, the initial sentiment control state provides the initial tone conditions, the historical inference sentiment guidance token sequence provides the generated sentiment changes, and the historical target speech token sequence provides the generated speech content. The trained language model can read the above inputs through the attention layer, and the sentiment prediction output layer generates the current inference sentiment guidance token. The current inference sentiment guidance token can be a continuous vector or a discrete sentiment token index, and its specific form can be consistent with the sentiment guidance token form during training.

[0187] The generation of the current target speech token occurs after the generation of the current inference sentiment guidance token. The trained language model reads the task text conditional representation, the current inference sentiment guidance token, the historical inference sentiment guidance token sequence, and the historical target speech token sequence, and generates the current target speech token through the speech prediction output layer. The current inference sentiment guidance token provides the sentiment condition for the current generation position, the historical inference sentiment guidance token sequence provides the preceding sentiment change trend, and the historical target speech token sequence provides the preceding speech context. The current target speech token can be a speech codebook index, a discrete acoustic token, or a speech representation readable by the speech reconstruction network. During generation, the attention layer can increase the attention intensity between the current inference sentiment guidance token and the current target speech token, making the current target speech token's pronunciation state influenced by the current sentiment condition.

[0188] The current inference sentiment guidance token is added to the historical inference sentiment guidance token sequence to form an updated historical inference sentiment guidance token sequence. Similarly, the current target speech token is added to the historical target speech token sequence to form an updated historical target speech token sequence. The update process preserves the generation order, ensuring that the sentiment guidance tokens and target speech tokens at each generation position are arranged chronologically. The updated historical inference sentiment guidance token sequence and the updated historical target speech token sequence serve as input for the next generation position, allowing for the re-execution of sentiment guidance token generation and target speech token generation at that position. Specifically, the trained language model generates the inference sentiment guidance token for the next generation position based on the task text conditional representation, the initial sentiment control state, the updated historical inference sentiment guidance token sequence, and the updated historical target speech token sequence; then, based on the task text conditional representation, the inference sentiment guidance token for the next generation position, the updated historical inference sentiment guidance token sequence, and the updated historical target speech token sequence, it generates the target speech token for the next generation position; and finally, the inference sentiment guidance token and target speech token for the next generation position are added to the corresponding historical sequences to form new historical inference sentiment guidance token sequences and new historical target speech token sequences. The above generation and sequence update process is executed sequentially according to the generation position until the generation termination condition is met. The trained language model continuously reads the task text conditional representation and generates new token pairs position by position based on the generated content.

[0189] When the generation termination condition is met, the historical inference emotion guidance token sequence is determined as the inference emotion guidance token sequence, and the historical target speech token sequence is determined as the target speech token sequence. The inference emotion guidance token sequence records the emotion state corresponding to each speech position during the complete generation process, while the target speech token sequence records the discrete speech content corresponding to each speech position during the complete generation process. The two types of sequences maintain a consistent generation order, enabling the subsequent speech reconstruction network to simultaneously read the speech content and emotion conditions.

[0190] After receiving the target speech token sequence and the inferred emotion guidance token sequence, the speech reconstruction network generates the target speech through a token embedding layer, an emotion modulation layer, an acoustic decoding layer, and a waveform synthesis layer. The token embedding layer converts the target speech token sequence into a continuous speech representation. The emotion modulation layer adjusts the energy, fundamental frequency, duration, and tone variations in the continuous speech representation based on the inferred emotion guidance token sequence. The acoustic decoding layer generates Mel spectrum, fundamental frequency curve, energy curve, or other acoustic parameters. The waveform synthesis layer converts the acoustic parameters into a speech waveform. If the target speech token sequence already contains acoustic parameters, the speech reconstruction network can directly perform waveform synthesis and loudness normalization. The target speech can be a playable audio file or speech waveform data used for subsequent playback interfaces.

[0191] The trained language model allows for adjustments to the generation temperature, candidate token truncation ratio, and maximum generation length during the inference phase. The generation temperature can be set between 0.7 and 1.2, the candidate token truncation ratio between 0.8 and 0.95, and the maximum generation length can be estimated based on the number of task text units and the average speech rate. When the task text involves amounts, dates, times, dosages, accounts, risks, medications, etc., the generation temperature can be lowered to improve the stability of the pronunciation. When the task text contains reassurance, explanation, or guidance, the emotional continuity constraint can be appropriately increased to make the change in the inference emotional guidance token sequence smoother between adjacent positions.

[0192] This embodiment divides the task text into text units and generates a task text position sequence. The trained language model can then obtain text content and text order information. Task text conditional representations are used to generate the initial emotion control state and generation termination conditions, providing the inference process with initial emotion input and termination constraints. The current inference emotion guidance token and the current target speech token are generated sequentially at the same generation position. The historical inference emotion guidance token sequence and the historical target speech token sequence are updated synchronously after each generation round, enabling subsequent generation to read previous emotion changes and previous speech content. The inference emotion guidance token sequence and the target speech token sequence obtained after satisfying the generation termination condition are jointly input into the speech reconstruction network, allowing the speech reconstruction process to simultaneously receive speech content and emotion state. This reduces the problem of insufficient emotion variation that occurs when generating speech directly from text alone, resulting in more continuous tone changes and more stable speech content in the target speech across different text segments.

[0193] In one embodiment, a time-varying emotion-guided speech generation device is provided, which corresponds one-to-one with the time-varying emotion-guided speech generation method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the time-varying emotion-guided speech generation device of the present invention. The modules include an emotion feature extraction module 10, a hybrid expert speech segmentation module 20, an emotion mapping and aggregation module 30, an interleaved token construction module 40, a language model training module 50, and a speech generation module 60. Detailed descriptions of each functional module are as follows: The emotion feature extraction module 10 is used to acquire training speech and training text corresponding to the training speech, perform emotion encoding on the training speech to obtain an emotion feature sequence, determine an emotion parameter sequence based on the emotion feature sequence, and perform time alignment on the emotion parameter sequence to obtain a time-varying emotion vector. The hybrid expert speech segmentation module 20 is used to construct a hybrid expert speech segmentation structure, extract semantic features and identity features from the training speech, and jointly train the hybrid expert speech segmentation structure based on the semantic features, time-varying sentiment vector and identity features, and use the jointly trained hybrid expert speech segmentation structure to perform speech segmentation on the training speech to obtain a speech token sequence. The emotion mapping and aggregation module 30 is used to map and aggregate the time-varying emotion vector based on the number of tokens in the voice token sequence to obtain an emotion guidance token sequence corresponding to the voice token sequence. The interleaved token construction module 40 is used to combine the emotion guidance token sequence and the voice token sequence in an interleaved order to form an interleaved token sequence, and input the training text and the interleaved token sequence into the language model to be trained to obtain the prediction token sequence; The language model training module 50 is used to determine the training loss based on the predicted token sequence and the interleaved token sequence, and update the language model to be trained based on the training loss to obtain the trained language model. The speech generation module 60 is used to receive task text, process the task text through the trained language model, and obtain target speech.

[0194] In one embodiment, the emotion feature extraction module 10 is specifically used for: Obtain training speech and training text corresponding to the training speech; perform text-speech forced alignment on the training text and the training speech to obtain a text-speech alignment mark. Pause detection is performed on the training speech to obtain pause detection results. Based on the text speech alignment markers and the pause detection results, boundary-preserving segmentation is performed on the training speech to obtain an emotion-coded segment sequence. For each emotion coding segment in the emotion coding segment sequence, segment acoustic features are extracted, and contextual acoustic features are extracted from the emotion coding segments adjacent to each emotion coding segment; The segment acoustic features and the context acoustic features are fused to obtain the emotional features corresponding to each emotional coding segment. The emotional features corresponding to each emotional coding segment are then arranged according to the segment order in the emotional coding segment sequence to obtain the emotional feature sequence. Based on the text-speech alignment marker, a set of local emotion windows is constructed in the emotion feature sequence. Each local emotion window in the set of local emotion windows covers multiple emotion features, and the window boundary of each local emotion window matches the boundary of the text unit in the text-speech alignment marker. Based on the time distance between each emotion feature within each local emotion window and the text unit boundary corresponding to the emotion coding segment to which each emotion feature belongs, the attention weight of each emotion feature within each local emotion window is determined, and the emotion features within each local emotion window are weighted and aggregated according to the attention weight to obtain the local emotion representation sequence. The local sentiment representation sequence is input into the sentiment parameter mapping structure to obtain the sentiment parameter sequence; Obtain the acoustic frame boundary sequence corresponding to the training speech, wherein the acoustic frame boundary sequence includes the frame-level temporal boundary of the training speech during the acoustic feature extraction process; Using the text-speech alignment marker as the boundary anchor point, the emotion parameter sequence is time-aligned according to the acoustic frame boundary sequence, the emotion parameters in the emotion parameter sequence are mapped to the acoustic frames corresponding to the acoustic frame boundary sequence, and the emotion parameters corresponding to each acoustic frame between adjacent boundary anchor points are smoothed to obtain the time-varying emotion vector.

[0195] In one embodiment, the hybrid expert speech segmentation module 20 is specifically used for: A hybrid expert speech segmentation structure is constructed, which includes a shared speech coding layer, a semantic preservation expert branch, an emotion dynamics preservation expert branch, an identity preservation expert branch, an expert gating unit, and a token quantization unit. The training speech is input into the automatic speech recognition model, and a semantic feature sequence is extracted from the intermediate coding layer of the automatic speech recognition model. The semantic feature sequence is then time-aggregated according to the speech segmentation time scale corresponding to the hybrid expert speech segmentation structure to obtain semantic features that match the speech segmentation time scale. The training speech is input into the speaker analysis model, the original identity representation is extracted from the identity representation layer of the speaker analysis model, and the original identity representation is normalized and projected to obtain the identity features. The training speech is input into the shared speech coding layer to obtain an initial speech representation. The semantic features, the time-varying sentiment vector, and the identity features are then scaled according to the time scale of the initial speech representation to obtain a semantic supervised representation, a sentiment supervised representation, and an identity supervised representation. The initial speech representation is input into the semantic preservation expert branch, the emotional dynamics preservation expert branch, and the identity preservation expert branch, respectively, to obtain the semantic branch representation, the emotional branch representation, and the identity branch representation; Based on the semantic supervision representation, the sentiment supervision representation, and the identity supervision representation, the participation ratio of the semantic preservation expert branch, the sentiment dynamic preservation expert branch, and the identity preservation expert branch is determined by the expert gating unit; Based on the participation ratio, the semantic branch representation, the emotion branch representation, and the identity branch representation are weighted and fused to obtain a fused speech representation; The fused speech representation is input into the token quantization unit, and the fused speech representation is quantized based on the codebook vector in the token quantization unit to obtain candidate speech tokens; The joint training loss is determined based on the matching degree between the semantic supervision representation and the semantic branch representation, the matching degree between the sentiment supervision representation and the sentiment branch representation, the matching degree between the identity supervision representation and the identity branch representation, and the quantization difference between the codebook vector corresponding to the candidate speech token and the fused speech representation. The parameters of the shared speech coding layer, the semantic preservation expert branch, the emotion dynamics preservation expert branch, the identity preservation expert branch, the expert gating unit, and the token quantization unit are updated according to the joint training loss to obtain the jointly trained hybrid expert speech segmentation structure; The training speech is segmented using the jointly trained hybrid expert speech segmentation structure to obtain a speech token sequence.

[0196] In one embodiment, the emotion mapping and aggregation module 30 is specifically used for: Obtain the number of tokens in the voice token sequence, and construct a token position sequence according to the time order of each voice token in the voice token sequence, wherein each position index in the token position sequence corresponds to a voice token in the voice token sequence; The time-varying emotion vector is subjected to temporal emotion change detection to determine the multidimensional emotion difference degree of the time-varying emotion vector between consecutive time frames, and the time frames in which the multidimensional emotion difference degree meets the preset emotion change conditions are marked as candidate emotion transition points. Based on the candidate emotional transition points, boundary aggregation is performed, and candidate emotional transition points whose time intervals meet the preset aggregation conditions are merged into aggregated emotional transition points, and the aggregated emotional transition points are used as emotional change nodes. The number of internal interval boundaries of the target is determined based on the number of tokens, and boundary emotion change nodes that meet the boundary selection conditions are selected from the emotion change nodes based on the number of internal interval boundaries of the target. Construct a target interval boundary set based on the starting boundary of the time-varying sentiment vector, the ending boundary of the time-varying sentiment vector, and the boundary sentiment change nodes; When the number of sentiment mapping intervals formed by the target interval boundary set is less than the number of tokens, interval boundaries are supplemented along the time direction of the time-varying sentiment vector, and the supplemented interval boundaries are added to the target interval boundary set. The number of sentiment mapping intervals formed by the target interval boundary set is consistent with the number of tokens. Based on the target interval boundary set, the time-varying sentiment vector is dynamically segmented to obtain a sentiment mapping interval sequence, wherein the number of sentiment mapping intervals in the sentiment mapping interval sequence is consistent with the number of tokens. Establish a one-to-one mapping relationship between each emotion mapping interval in the emotion mapping interval sequence and each position index in the token position sequence to form an interval token correspondence table; For each emotion mapping interval, an emotion carrying reference position set is determined from the interval boundary of the corresponding emotion mapping interval and the emotion change nodes located in the vicinity of the corresponding emotion mapping interval. The normalized time distance between the time-varying emotion vector of each time frame in each emotion mapping interval and the nearest emotion carrying reference position in the set of emotion carrying reference positions is determined, and the emotion carrying weight of each time frame is determined based on the normalized time distance. Based on the aforementioned emotional carrying weight, the time-varying emotional vectors of each time frame within each emotional mapping interval are weighted and aggregated to obtain the initial interval emotional representation corresponding to each emotional mapping interval; The initial interval sentiment representation corresponding to each sentiment mapping interval is input into the trend enhancement network to obtain the sentiment change trend features corresponding to each sentiment mapping interval. The initial interval sentiment representation and the sentiment change trend features corresponding to each sentiment mapping interval are then fused to obtain the interval sentiment representation corresponding to each sentiment mapping interval. Arrange the interval sentiment representations corresponding to each sentiment mapping interval in the order of the token position sequence to form an interval sentiment representation sequence; Based on the directional and amplitude change relationships between adjacent interval emotional representations in the interval emotional representation sequence, and by filling in the missing adjacent change relationships at the beginning and end positions of the interval emotional representation sequence, the emotional change direction feature sequence and emotional fluctuation intensity feature sequence corresponding to each interval emotional representation in the interval emotional representation sequence are determined. The interval sentiment representation sequence, the sentiment change direction feature sequence, and the sentiment fluctuation intensity feature sequence are input into the sentiment token embedding network to obtain the sentiment guidance token corresponding to each sentiment mapping interval. Arrange the emotion guidance tokens corresponding to each emotion mapping interval according to the interval token correspondence table to obtain the emotion guidance token sequence corresponding to the voice token sequence.

[0197] In one embodiment, the interleaved token construction module 40 is specifically used for: Based on the correspondence between the emotion guidance token sequence and the voice token sequence, a token pairing unit sequence is constructed. Each token pairing unit in the token pairing unit sequence includes an emotion guidance token and a voice token, and the emotion guidance token in each token pairing unit is placed before the voice token. According to the order of the token pairing unit sequence, each token pairing unit is unfolded sequentially to obtain an interleaved token sequence; The training text is text-encoded to obtain a conditional representation of the training text; A token type identifier sequence is generated based on the interleaved token sequence, the token type identifier sequence including an emotion guidance token type identifier and a voice token type identifier; An interleaved position encoding sequence is generated based on the token pairing unit sequence, and the interleaved position encoding sequence records the pairing position of each token pairing unit in the interleaved token sequence; An interleaved attention constraint marker is constructed based on the token pairing unit sequence. The interleaved attention constraint marker records the binding relationship between each emotion guidance token and the corresponding voice token, as well as the order between adjacent token pairing units. The training text conditional representation, the interleaved token sequence, the token type identifier sequence, the interleaved position encoding sequence, and the interleaved attention constraint label are input into the language model to be trained. In the attention layer of the language model to be trained, the attention weights between the sentiment guidance tokens and the corresponding speech tokens in the interleaved token sequence are adjusted based on the interleaved attention constraint label to obtain the predicted token sequence.

[0198] In one embodiment, the language model training module 50 is specifically used for: Based on the token type identifier sequence and the interleaved position encoding sequence, the emotion guidance token prediction subsequence and the voice token prediction subsequence are separated from the prediction token sequence; Based on the token type identifier sequence and the interleaved position encoding sequence, extract the emotion guidance token truth value subsequence and the voice token truth value subsequence from the interleaved token sequence; The initial sentiment guidance token prediction loss is determined based on the token prediction difference between the sentiment guidance token prediction subsequence and the sentiment guidance token truth value subsequence. The emotional continuity loss is determined based on the continuity difference between adjacent emotional guidance token prediction values ​​in the emotional guidance token prediction subsequence; Based on the initial sentiment guidance token prediction loss and the sentiment continuity loss, the sentiment guidance token prediction loss is determined; The initial voice token prediction loss is determined based on the token prediction difference between the voice token prediction subsequence and the voice token truth subsequence. The voice token prediction subsequence and the corresponding sentiment guidance token truth subsequence are mapped to a common evaluation space, and the voice sentiment consistency loss is determined based on the matching relationship in the common evaluation space. Based on the initial voice token prediction loss and the voice emotion consistency loss, the voice token prediction loss is determined; Based on the interleaved attention constraint marker, the binding relationship between each emotion guidance token and the corresponding voice token is determined. Based on the binding relationship, a binding token pair is constructed, and an unbinding token pair is constructed from the emotion guidance token and the voice token at non-corresponding positions. Based on the representation difference between the binding token pair and the unbinding token pair, the emotion-voice alignment constraint loss is determined. The training loss is determined based on the emotion-guided token prediction loss, the speech token prediction loss, and the emotion-speech alignment constraint loss. The parameter update direction of the language model to be trained is determined based on the training loss, and the language model to be trained is updated according to the parameter update direction to obtain the updated language model. The updated language model is used as the language model to be trained in the next round of training. The determination of training loss and the updating of the language model to be trained are repeated until the training loss corresponding to the updated language model meets the convergence condition. Then the updated language model is determined as the trained language model.

[0199] In one embodiment, the speech generation module 60 is specifically used for: Receive task text, divide the task text into text units to obtain a task text unit sequence, and generate a task text position sequence based on the task text unit sequence; The task text unit sequence and the task text position sequence are input into the trained language model to obtain the task text conditional representation; The starting emotion control state is generated based on the task text condition representation, and the historical reasoning emotion guidance token sequence and historical target voice token sequence are initialized, and the termination condition is determined based on the task text condition representation. At the current generation position, based on the task text conditional representation, the initial emotion control state, the historical inference emotion guidance token sequence, and the historical target speech token sequence, the current inference emotion guidance token is generated through the trained language model. Based on the task text conditional representation, the current inference sentiment guidance token, the historical inference sentiment guidance token sequence, and the historical target speech token sequence, the current target speech token is generated through the trained language model. The current reasoning emotion guidance token is added to the historical reasoning emotion guidance token sequence to obtain an updated historical reasoning emotion guidance token sequence, and the current target voice token is added to the historical target voice token sequence to obtain an updated historical target voice token sequence. The updated historical reasoning emotion guidance token sequence is used as the historical reasoning emotion guidance token sequence corresponding to the next generation position, and the updated historical target voice token sequence is used as the historical target voice token sequence corresponding to the next generation position. The current reasoning emotion guidance token and the current target voice token are repeatedly generated until the generation termination condition is met. The historical reasoning emotion guidance token sequence obtained when the generation termination condition is met is determined as the reasoning emotion guidance token sequence, and the historical target speech token sequence obtained when the generation termination condition is met is determined as the target speech token sequence. The target speech token sequence and the inference emotion guidance token sequence are input into the speech reconstruction network. The speech reconstruction network then performs speech reconstruction on the target speech token sequence based on the inference emotion guidance token sequence to obtain the target speech.

[0200] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a time-varying emotion-guided speech generation method on the server side.

[0201] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a time-varying emotion-guided speech generation method.

[0202] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire training speech and training text corresponding to the training speech, perform emotion encoding on the training speech to obtain an emotion feature sequence, determine an emotion parameter sequence based on the emotion feature sequence, and perform time alignment on the emotion parameter sequence to obtain a time-varying emotion vector; A hybrid expert speech segmentation structure is constructed. Semantic features and identity features are extracted from the training speech. The hybrid expert speech segmentation structure is jointly trained based on the semantic features, time-varying sentiment vector, and identity features. The training speech is then segmented using the jointly trained hybrid expert speech segmentation structure to obtain a speech token sequence. Based on the number of tokens in the voice token sequence, the time-varying emotion vector is mapped and aggregated to obtain an emotion guidance token sequence corresponding to the voice token sequence; The emotion guidance token sequence and the voice token sequence are combined in an alternating order to form an alternating token sequence, and the training text and the alternating token sequence are input into the language model to be trained to obtain the prediction token sequence; The training loss is determined based on the predicted token sequence and the interleaved token sequence, and the language model to be trained is updated based on the training loss to obtain the trained language model. The task text is received and processed by the trained language model to obtain the target speech.

[0203] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Acquire training speech and training text corresponding to the training speech, perform emotion encoding on the training speech to obtain an emotion feature sequence, determine an emotion parameter sequence based on the emotion feature sequence, and perform time alignment on the emotion parameter sequence to obtain a time-varying emotion vector; A hybrid expert speech segmentation structure is constructed. Semantic features and identity features are extracted from the training speech. The hybrid expert speech segmentation structure is jointly trained based on the semantic features, time-varying sentiment vector, and identity features. The training speech is then segmented using the jointly trained hybrid expert speech segmentation structure to obtain a speech token sequence. Based on the number of tokens in the voice token sequence, the time-varying emotion vector is mapped and aggregated to obtain an emotion guidance token sequence corresponding to the voice token sequence; The emotion guidance token sequence and the voice token sequence are combined in an alternating order to form an alternating token sequence, and the training text and the alternating token sequence are input into the language model to be trained to obtain the prediction token sequence; The training loss is determined based on the predicted token sequence and the interleaved token sequence, and the language model to be trained is updated based on the training loss to obtain the trained language model. The task text is received and processed by the trained language model to obtain the target speech.

[0204] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0205] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0206] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0207] It should be noted that any software tools or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0208] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A time-varying emotion-guided speech generation method, characterized in that, Includes the following steps: Acquire training speech and training text corresponding to the training speech, perform emotion encoding on the training speech to obtain an emotion feature sequence, determine an emotion parameter sequence based on the emotion feature sequence, and perform time alignment on the emotion parameter sequence to obtain a time-varying emotion vector; A hybrid expert speech segmentation structure is constructed. Semantic features and identity features are extracted from the training speech. The hybrid expert speech segmentation structure is jointly trained based on the semantic features, time-varying sentiment vector, and identity features. The training speech is then segmented using the jointly trained hybrid expert speech segmentation structure to obtain a speech token sequence. The process involves obtaining the number of tokens in the voice token sequence, constructing a token position sequence based on the temporal order of each voice token, detecting temporal emotional changes in the time-varying emotion vector, identifying candidate emotional transition points, performing boundary aggregation based on the candidate emotional transition points to obtain emotional change nodes, determining the number of target internal interval boundaries based on the number of tokens, selecting boundary emotional change nodes that meet the boundary selection conditions from the emotional change nodes based on the number of target internal interval boundaries, constructing a target interval boundary set based on the start boundary, end boundary, and boundary emotional change nodes of the time-varying emotion vector, and when the number of emotional mapping intervals formed by the target interval boundary set is less than the number of tokens, supplementing interval boundaries along the time direction of the time-varying emotion vector and adding the supplemented interval boundaries to the target interval boundary set, performing dynamic segmentation on the time-varying emotion vector based on the target interval boundary set to obtain an emotional mapping interval sequence, establishing a one-to-one mapping relationship between each emotional mapping interval in the emotional mapping interval sequence and each position index in the token position sequence, and aggregating the time-varying emotion vectors within each emotional mapping interval to obtain an emotional guidance token sequence corresponding to the voice token sequence. The emotion guidance token sequence and the voice token sequence are combined in an alternating order to form an alternating token sequence, and the training text and the alternating token sequence are input into the language model to be trained to obtain the prediction token sequence; The training loss is determined based on the predicted token sequence and the interleaved token sequence, and the language model to be trained is updated based on the training loss to obtain the trained language model. The task text is received and processed by the trained language model to obtain the target speech.

2. The time-varying emotion-guided speech generation method as described in claim 1, characterized in that, The process involves acquiring training speech and corresponding training text, performing emotion encoding on the training speech to obtain an emotion feature sequence, determining an emotion parameter sequence based on the emotion feature sequence, and performing time alignment on the emotion parameter sequence to obtain a time-varying emotion vector, including: Obtain training speech and training text corresponding to the training speech; perform text-speech forced alignment on the training text and the training speech to obtain a text-speech alignment mark. Pause detection is performed on the training speech to obtain pause detection results. Based on the text speech alignment markers and the pause detection results, boundary-preserving segmentation is performed on the training speech to obtain an emotion-coded segment sequence. For each emotion coding segment in the emotion coding segment sequence, segment acoustic features are extracted, and contextual acoustic features are extracted from the emotion coding segments adjacent to each emotion coding segment; The segment acoustic features and the context acoustic features are fused to obtain the emotional features corresponding to each emotional coding segment. The emotional features corresponding to each emotional coding segment are then arranged according to the segment order in the emotional coding segment sequence to obtain the emotional feature sequence. Based on the text-speech alignment marker, a set of local emotion windows is constructed in the emotion feature sequence. Each local emotion window in the set covers multiple emotion features, and the window boundary of each local emotion window matches the text unit boundary in the text-speech alignment marker. Based on the time distance between each emotion feature within each local emotion window and the text unit boundary corresponding to the emotion coding segment to which each emotion feature belongs, the attention weight of each emotion feature within each local emotion window is determined, and the emotion features within each local emotion window are weighted and aggregated according to the attention weight to obtain the local emotion representation sequence. The local sentiment representation sequence is input into the sentiment parameter mapping structure to obtain the sentiment parameter sequence; Obtain the acoustic frame boundary sequence corresponding to the training speech, wherein the acoustic frame boundary sequence includes the frame-level temporal boundary of the training speech during the acoustic feature extraction process; Using the text-speech alignment marker as the boundary anchor point, the emotion parameter sequence is time-aligned according to the acoustic frame boundary sequence, the emotion parameters in the emotion parameter sequence are mapped to the acoustic frames corresponding to the acoustic frame boundary sequence, and the emotion parameters corresponding to each acoustic frame between adjacent boundary anchor points are smoothed to obtain the time-varying emotion vector.

3. The time-varying emotion-guided speech generation method as described in claim 1, characterized in that, A hybrid expert speech segmentation structure is constructed. Semantic and identity features are extracted from the training speech. The hybrid expert speech segmentation structure is jointly trained based on the semantic features, time-varying sentiment vector, and identity features. The jointly trained hybrid expert speech segmentation structure is then used to segment the training speech to obtain a speech token sequence, including: A hybrid expert speech segmentation structure is constructed, which includes a shared speech coding layer, a semantic preservation expert branch, an emotion dynamics preservation expert branch, an identity preservation expert branch, an expert gating unit, and a token quantization unit. The training speech is input into the automatic speech recognition model, and a semantic feature sequence is extracted from the intermediate coding layer of the automatic speech recognition model. The semantic feature sequence is then time-aggregated according to the speech segmentation time scale corresponding to the hybrid expert speech segmentation structure to obtain semantic features that match the speech segmentation time scale. The training speech is input into the speaker analysis model, the original identity representation is extracted from the identity representation layer of the speaker analysis model, and the original identity representation is normalized and projected to obtain the identity features. The training speech is input into the shared speech coding layer to obtain an initial speech representation. The semantic features, the time-varying sentiment vector, and the identity features are then scaled according to the time scale of the initial speech representation to obtain a semantic supervised representation, a sentiment supervised representation, and an identity supervised representation. The initial speech representation is input into the semantic preservation expert branch, the emotional dynamics preservation expert branch, and the identity preservation expert branch, respectively, to obtain the semantic branch representation, the emotional branch representation, and the identity branch representation; Based on the semantic supervision representation, the sentiment supervision representation, and the identity supervision representation, the participation ratio of the semantic preservation expert branch, the sentiment dynamic preservation expert branch, and the identity preservation expert branch is determined by the expert gating unit; Based on the participation ratio, the semantic branch representation, the emotion branch representation, and the identity branch representation are weighted and fused to obtain a fused speech representation; The fused speech representation is input into the token quantization unit, and the fused speech representation is quantized based on the codebook vector in the token quantization unit to obtain candidate speech tokens; The joint training loss is determined based on the matching degree between the semantic supervision representation and the semantic branch representation, the matching degree between the sentiment supervision representation and the sentiment branch representation, the matching degree between the identity supervision representation and the identity branch representation, and the quantization difference between the codebook vector corresponding to the candidate speech token and the fused speech representation. The parameters of the shared speech coding layer, the semantic preservation expert branch, the emotion dynamics preservation expert branch, the identity preservation expert branch, the expert gating unit, and the token quantization unit are updated according to the joint training loss to obtain the jointly trained hybrid expert speech segmentation structure; The training speech is segmented using the jointly trained hybrid expert speech segmentation structure to obtain a speech token sequence.

4. The time-varying emotion-guided speech generation method as described in claim 1, characterized in that, The process involves obtaining the number of tokens in the voice token sequence, constructing a token position sequence based on the temporal order of each voice token, detecting temporal emotional changes in the time-varying emotional vector, identifying candidate emotional transition points, performing boundary aggregation based on these candidate emotional transition points to obtain emotional change nodes, determining the number of target internal interval boundaries based on the number of tokens, and selecting boundary emotional change nodes that meet the boundary selection criteria from the emotional change nodes based on the number of target internal interval boundaries, and constructing a target interval based on the start boundary, end boundary, and boundary emotional change nodes of the time-varying emotional vector. A boundary set is used to supplement interval boundaries along the time direction of the time-varying emotion vector when the number of emotion mapping intervals formed by the target interval boundary set is less than the number of tokens. The supplemented interval boundaries are then added to the target interval boundary set. Based on the target interval boundary set, the time-varying emotion vector is dynamically segmented to obtain an emotion mapping interval sequence. A one-to-one mapping relationship is established between each emotion mapping interval in the emotion mapping interval sequence and each position index in the token position sequence. The time-varying emotion vectors within each emotion mapping interval are aggregated to obtain an emotion guidance token sequence corresponding to the voice token sequence, including: Obtain the number of tokens in the voice token sequence, and construct a token position sequence according to the time order of each voice token in the voice token sequence, wherein each position index in the token position sequence corresponds to a voice token in the voice token sequence; The time-varying emotion vector is subjected to temporal emotion change detection to determine the multidimensional emotion difference degree of the time-varying emotion vector between consecutive time frames, and the time frames in which the multidimensional emotion difference degree meets the preset emotion change conditions are marked as candidate emotion transition points. Based on the candidate emotional transition points, boundary aggregation is performed, and candidate emotional transition points whose time intervals meet the preset aggregation conditions are merged into aggregated emotional transition points, and the aggregated emotional transition points are used as emotional change nodes. The number of internal interval boundaries of the target is determined based on the number of tokens, and boundary emotion change nodes that meet the boundary selection conditions are selected from the emotion change nodes based on the number of internal interval boundaries of the target. Construct a target interval boundary set based on the starting boundary of the time-varying sentiment vector, the ending boundary of the time-varying sentiment vector, and the boundary sentiment change nodes; When the number of sentiment mapping intervals formed by the target interval boundary set is less than the number of tokens, interval boundaries are supplemented along the time direction of the time-varying sentiment vector, and the supplemented interval boundaries are added to the target interval boundary set. The number of sentiment mapping intervals formed by the target interval boundary set is consistent with the number of tokens. Based on the target interval boundary set, the time-varying sentiment vector is dynamically segmented to obtain a sentiment mapping interval sequence, wherein the number of sentiment mapping intervals in the sentiment mapping interval sequence is consistent with the number of tokens. Establish a one-to-one mapping relationship between each emotion mapping interval in the emotion mapping interval sequence and each position index in the token position sequence to form an interval token correspondence table; For each emotion mapping interval, an emotion carrying reference position set is determined from the interval boundary of the corresponding emotion mapping interval and the emotion change nodes located in the vicinity of the corresponding emotion mapping interval. The normalized time distance between the time-varying emotion vector of each time frame in each emotion mapping interval and the nearest emotion carrying reference position in the set of emotion carrying reference positions is determined, and the emotion carrying weight of each time frame is determined based on the normalized time distance. Based on the aforementioned emotional carrying weight, the time-varying emotional vectors of each time frame within each emotional mapping interval are weighted and aggregated to obtain the initial interval emotional representation corresponding to each emotional mapping interval; The initial interval sentiment representation corresponding to each sentiment mapping interval is input into the trend enhancement network to obtain the sentiment change trend features corresponding to each sentiment mapping interval. The initial interval sentiment representation and the sentiment change trend features corresponding to each sentiment mapping interval are then fused to obtain the interval sentiment representation corresponding to each sentiment mapping interval. Arrange the interval sentiment representations corresponding to each sentiment mapping interval in the order of the token position sequence to form an interval sentiment representation sequence; Based on the directional and amplitude change relationships between adjacent interval emotional representations in the interval emotional representation sequence, and by filling in the missing adjacent change relationships at the beginning and end positions of the interval emotional representation sequence, the emotional change direction feature sequence and emotional fluctuation intensity feature sequence corresponding to each interval emotional representation in the interval emotional representation sequence are determined. The interval sentiment representation sequence, the sentiment change direction feature sequence, and the sentiment fluctuation intensity feature sequence are input into the sentiment token embedding network to obtain the sentiment guidance token corresponding to each sentiment mapping interval. Arrange the emotion guidance tokens corresponding to each emotion mapping interval according to the interval token correspondence table to obtain the emotion guidance token sequence corresponding to the voice token sequence.

5. The time-varying emotion-guided speech generation method as described in claim 1, characterized in that, The emotion guidance token sequence and the voice token sequence are combined in an interleaved order to form an interleaved token sequence. The training text and the interleaved token sequence are then input into the language model to be trained to obtain a predicted token sequence, including: Based on the correspondence between the emotion guidance token sequence and the voice token sequence, a token pairing unit sequence is constructed. Each token pairing unit in the token pairing unit sequence includes an emotion guidance token and a voice token, and the emotion guidance token in each token pairing unit is placed before the voice token. According to the order of the token pairing unit sequence, each token pairing unit is unfolded sequentially to obtain an interleaved token sequence; The training text is text-encoded to obtain a conditional representation of the training text; A token type identifier sequence is generated based on the interleaved token sequence, the token type identifier sequence including an emotion guidance token type identifier and a voice token type identifier; An interleaved position encoding sequence is generated based on the token pairing unit sequence, and the interleaved position encoding sequence records the pairing position of each token pairing unit in the interleaved token sequence; An interleaved attention constraint marker is constructed based on the token pairing unit sequence. The interleaved attention constraint marker records the binding relationship between each emotion guidance token and the corresponding voice token, as well as the order between adjacent token pairing units. The training text conditional representation, the interleaved token sequence, the token type identifier sequence, the interleaved position encoding sequence, and the interleaved attention constraint label are input into the language model to be trained. In the attention layer of the language model to be trained, the attention weights between the sentiment guidance tokens and the corresponding speech tokens in the interleaved token sequence are adjusted based on the interleaved attention constraint label to obtain the predicted token sequence.

6. The time-varying emotion-guided speech generation method as described in claim 5, characterized in that, The training loss is determined based on the predicted token sequence and the interleaved token sequence, and the language model to be trained is updated based on the training loss to obtain the trained language model, including: Based on the token type identifier sequence and the interleaved position encoding sequence, the emotion guidance token prediction subsequence and the voice token prediction subsequence are separated from the prediction token sequence; Based on the token type identifier sequence and the interleaved position encoding sequence, extract the emotion guidance token truth value subsequence and the voice token truth value subsequence from the interleaved token sequence; The initial sentiment guidance token prediction loss is determined based on the token prediction difference between the sentiment guidance token prediction subsequence and the sentiment guidance token truth value subsequence. The emotional continuity loss is determined based on the continuity difference between adjacent emotional guidance token prediction values ​​in the emotional guidance token prediction subsequence; Based on the initial sentiment guidance token prediction loss and the sentiment continuity loss, the sentiment guidance token prediction loss is determined; The initial voice token prediction loss is determined based on the token prediction difference between the voice token prediction subsequence and the voice token truth subsequence. The voice token prediction subsequence and the corresponding sentiment guidance token truth subsequence are mapped to a common evaluation space, and the voice sentiment consistency loss is determined based on the matching relationship in the common evaluation space. Based on the initial voice token prediction loss and the voice emotion consistency loss, the voice token prediction loss is determined; Based on the interleaved attention constraint marker, the binding relationship between each emotion guidance token and the corresponding voice token is determined. Based on the binding relationship, a binding token pair is constructed, and an unbinding token pair is constructed from the emotion guidance token and the voice token at non-corresponding positions. Based on the representation difference between the binding token pair and the unbinding token pair, the emotion-voice alignment constraint loss is determined. The training loss is determined based on the emotion-guided token prediction loss, the speech token prediction loss, and the emotion-speech alignment constraint loss. The parameter update direction of the language model to be trained is determined based on the training loss, and the language model to be trained is updated according to the parameter update direction to obtain the updated language model. The updated language model is used as the language model to be trained in the next round of training. The determination of training loss and the updating of the language model to be trained are repeated until the training loss corresponding to the updated language model meets the convergence condition. Then the updated language model is determined as the trained language model.

7. The time-varying emotion-guided speech generation method as described in claim 1, characterized in that, Receive task text, process the task text using the trained language model to obtain target speech, including: Receive task text, divide the task text into text units to obtain a task text unit sequence, and generate a task text position sequence based on the task text unit sequence; The task text unit sequence and the task text position sequence are input into the trained language model to obtain the task text conditional representation; The starting emotion control state is generated based on the task text condition representation, and the historical reasoning emotion guidance token sequence and historical target voice token sequence are initialized, and the termination condition is determined based on the task text condition representation. At the current generation position, based on the task text conditional representation, the initial emotion control state, the historical inference emotion guidance token sequence, and the historical target speech token sequence, the current inference emotion guidance token is generated through the trained language model. Based on the task text conditional representation, the current inference sentiment guidance token, the historical inference sentiment guidance token sequence, and the historical target speech token sequence, the current target speech token is generated through the trained language model. The current reasoning emotion guidance token is added to the historical reasoning emotion guidance token sequence to obtain an updated historical reasoning emotion guidance token sequence, and the current target voice token is added to the historical target voice token sequence to obtain an updated historical target voice token sequence. The updated historical reasoning emotion guidance token sequence is used as the historical reasoning emotion guidance token sequence corresponding to the next generation position, and the updated historical target voice token sequence is used as the historical target voice token sequence corresponding to the next generation position. The current reasoning emotion guidance token and the current target voice token are repeatedly generated until the generation termination condition is met. The historical reasoning emotion guidance token sequence obtained when the generation termination condition is met is determined as the reasoning emotion guidance token sequence, and the historical target speech token sequence obtained when the generation termination condition is met is determined as the target speech token sequence. The target speech token sequence and the inference emotion guidance token sequence are input into the speech reconstruction network. The speech reconstruction network then performs speech reconstruction on the target speech token sequence based on the inference emotion guidance token sequence to obtain the target speech.

8. A speech generation device with time-varying emotion guidance, characterized in that, The time-varying emotion-guided speech generation device includes: The emotion feature extraction module is used to acquire training speech and training text corresponding to the training speech, perform emotion encoding on the training speech to obtain an emotion feature sequence, determine an emotion parameter sequence based on the emotion feature sequence, and perform time alignment on the emotion parameter sequence to obtain a time-varying emotion vector. The hybrid expert speech segmentation module is used to construct a hybrid expert speech segmentation structure, extract semantic features and identity features from the training speech, and jointly train the hybrid expert speech segmentation structure based on the semantic features, time-varying sentiment vector and identity features. The trained hybrid expert speech segmentation structure is then used to segment the training speech to obtain a speech token sequence. The emotion mapping and aggregation module is used to obtain the number of tokens in the voice token sequence, construct a token position sequence according to the time order of each voice token in the voice token sequence, perform temporal emotion change detection on the time-varying emotion vector, determine candidate emotion transition points, perform boundary aggregation based on the candidate emotion transition points to obtain emotion change nodes, determine the number of target internal interval boundaries according to the number of target internal interval boundaries, and select boundary emotion change nodes that meet the boundary selection conditions from the emotion change nodes according to the starting boundary, ending boundary, and boundary emotion change of the time-varying emotion vector. The node constructs a target interval boundary set. When the number of emotion mapping intervals formed by the target interval boundary set is less than the number of tokens, the interval boundaries are supplemented along the time direction of the time-varying emotion vector, and the supplemented interval boundaries are added to the target interval boundary set. Based on the target interval boundary set, the time-varying emotion vector is dynamically segmented to obtain an emotion mapping interval sequence. A one-to-one mapping relationship is established between each emotion mapping interval in the emotion mapping interval sequence and each position index in the token position sequence. The time-varying emotion vectors in each emotion mapping interval are aggregated to obtain an emotion guidance token sequence corresponding to the voice token sequence. An interleaved token construction module is used to combine the emotion guidance token sequence and the voice token sequence in an interleaved order to form an interleaved token sequence, and input the training text and the interleaved token sequence into the language model to be trained to obtain a predicted token sequence; The language model training module is used to determine the training loss based on the predicted token sequence and the interleaved token sequence, and update the language model to be trained based on the training loss to obtain the trained language model. The speech generation module is used to receive task text, process the task text through the trained language model, and obtain the target speech.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a time-varying emotion-guided speech generation program stored in the memory and executable on the processor, wherein the time-varying emotion-guided speech generation program, when executed by the processor, implements the steps of the time-varying emotion-guided speech generation method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a time-varying emotion-guided speech generation program, which, when executed by a processor, implements the steps of the time-varying emotion-guided speech generation method as described in any one of claims 1-7.