A method and system for speech adaptive generation of traffic domain services based on chain-of-thought fine-tuning of large models

Through the traffic domain service voice adaptive generation method based on the thinking chain fine-tuning large model, the problem that the interactive voice response system cannot understand the user's natural language expression in the transportation field is solved, and high-fidelity emotional speech generation and robust synthesis in complex environments are realized, which improves the efficiency and accuracy of automated services.

CN120126484BActive Publication Date: 2025-07-22NANJING MICROVIDEO TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510604879.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-22
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In the transportation field, interactive voice response systems cannot directly understand the user's natural language expression demands, and lack dynamic response capabilities, resulting in interruption of the service chain, especially when facing complex services, and have limited automation processing capabilities.

Method used

The traffic domain service speech adaptive generation method based on the thinking chain fine-tuning large model is adopted, and the text output signal is generated using the speech encoder, text decoder and pinyin decoder. Combined with the accent type recognition decoder and multi-task speech recognition model, the big model is trained through the LoRA low-rank fine-tuning method and the mixed precision quantization method to realize high-fidelity emotional speech generation and robust synthesis in complex noise environments, supporting semantic-driven dynamic rhythm optimization and precise pronunciation of professional terms.

Benefits of technology

It realizes efficient processing of complex accents, many noises and multi-sounding words in the transportation field, improves logical reasoning capabilities and business intention recognition effects, reduces manual intervention, and improves the efficiency and accuracy of automated services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126484B_ABST
    Figure CN120126484B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for adaptively generating traffic domain service speech based on fine-tuning a large model with a chain of thought. First, a speech encoder is used to convert the input speech signal into a high-dimensional speech feature signal, and then a text decoder and a pinyin decoder are used to generate a text output signal according to the high-dimensional speech feature signal. The present invention realizes the functions of high-fidelity emotional speech generation and robust synthesis in complex noise environments by adopting a variational quantization joint data augmentation mechanism, a dual-channel multi-modal discriminator architecture, and a hierarchical composite loss function, and simultaneously supports semantic-driven dynamic prosody optimization and accurate pronunciation of professional terms. Moreover, in the application scenario of the traffic field, a multi-task speech recognition method can be adopted to achieve the efficient linkage of character-level recognition, audio-to-pinyin, and sentence-level accent classification modules, so as to effectively cope with the challenges of complex accents, many background noises, and polyphonic traffic terms, and is suitable for being widely promoted and used.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of service voice generation, and particularly to a traffic domain service voice adaptive generation method and system based on chain-of-thought fine-tuning of a large model. Background Art

[0002] An Interactive Voice Response (IVR) system is an automated telephone service technology that interacts with users through voice navigation and keypad input. It is widely used in fields such as enterprise customer service, banking, and telecommunications, helping users to complete operations such as inquiries, transfers, and appointments through voice menus self-service, while reducing labor costs and improving service efficiency.

[0003] Currently, the solutions for customer service and toll collection services in the transportation field mostly rely on the Interactive Voice Response system to complete. However, this system cannot directly understand the demands expressed by users in natural language, but relies on pre-designed fixed voice templates, guiding users to press keys to select through layer-by-layer announcement of options. This one-way tree-like interaction structure requires users to accurately remember and match the paths set by the system. For example, when querying traffic violations, users need to select hierarchical menus such as license plate jurisdiction and violation type in sequence. When users' demands involve cross-category or multi-condition combinations, they often need to experience multiple key presses and jumps, and even be forced to restart the process due to incorrect path selection. More critically, in the face of sudden consultations or personalized problems beyond the preset templates, the system lacks the ability to respond dynamically, resulting in the interruption of the service chain. In addition, there are obvious boundaries in the functional coverage of the IVR. Its automated processing capabilities are mostly limited to basic information query standardization services, such as balance inquiries and payment status confirmations. Once it comes to complex services that require logical judgment or data linkage, such as fine appeal acceptance, the system often cannot complete the retrieval and logical verification of multi-system data and must forcefully transfer to an artificial seat. Such service breakpoints not only cause users to repeat their demands, but also lead to a large amount of time-consuming information secondary verification by artificial seats. Especially during the peak period of transportation services, a large number of consultation requests that have not been automatically processed pour into the artificial channel, exacerbating the shortage of service resources and forming an efficiency paradox effect. Therefore, it is necessary to design a traffic domain service voice adaptive generation method and system based on chain-of-thought fine-tuning of a large model. Summary of the Invention

[0004] The object of the present invention is to overcome the deficiencies of the prior art and to provide a traffic domain service voice adaptive generation method and system based on a large model fine-tuned with a chain of thought, in order to better and effectively solve the problem that the current interactive voice response system generally cannot directly understand the demands expressed by users in natural language and lacks dynamic response capabilities, resulting in the interruption of the service chain. It realizes the functions of high-fidelity emotional speech generation and robust synthesis in complex noise environments by adopting a variational dequantization joint data augmentation mechanism, a dual-channel multi-modal discriminator architecture, and a hierarchical composite loss function, and simultaneously supports semantic-driven dynamic prosody optimization and accurate pronunciation of professional terms. Moreover, in the application scenarios of the traffic field, it can adopt a multi-task speech recognition method to achieve the efficient linkage of character-level recognition, audio-to-pinyin conversion, and sentence-level accent classification modules, thus effectively coping with challenges such as complex accents, a lot of background noise, and polyphonic traffic terms.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A traffic domain service voice adaptive generation method based on a large model fine-tuned with a chain of thought, comprising the following steps:

[0007] Step A: Use a speech encoder to convert the input speech signal into a high-dimensional speech feature signal, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional speech feature signal.

[0008] Step B: Use an accent type recognition decoder to recognize the input speech signal and output the probability of the audio accent type.

[0009] Step C: Based on the speech encoder, text decoder, pinyin decoder, and accent type recognition decoder, construct a multi-task speech recognition model, and then use the multi-task speech recognition model to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type.

[0010] Step D: Formulate a large model annotation specification, then construct a large model dataset and annotate the large model dataset according to the annotation specification to obtain an annotated large model dataset.

[0011] Step E: Use the LoRA low-rank fine-tuning method and the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a large model fine-tuned with a chain of thought.

[0012] Step F: Use the large model fine-tuned with a chain of thought to adaptively generate and output emotional traffic domain service voice according to the text representation, pinyin sequence, and audio accent type.

[0013] The aforementioned method for adaptively generating traffic domain service speech based on chain-of-thought fine-tuning of a large model, step A: Use a speech encoder to convert the input speech signal into a high-dimensional speech feature signal, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional speech feature signal. The specific steps are as follows:

[0014] Step A1: Use a speech encoder to convert the input speech signal into a high-dimensional speech feature signal. The speech encoder is specifically a Conformer encoding layer. The Conformer encoding layer is used to construct local audio feature dependencies and combine the global feature construction ability of the self-attention mechanism to improve the performance of the automatic speech recognition task. The Conformer encoding layer is specifically composed of a first feed-forward module, a self-attention module, a convolutional module, a second feed-forward module, and a first regularization module connected in sequence;

[0015] Step A2: Generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional speech feature signal. Both the text decoder and the pinyin decoder use a Transformer-ASR decoding layer. The Transformer-ASR decoding layer specifically uses a multi-head self-attention layer and uses a second regularization module, a linear output module, and a Softmax function module to generate the probability of the output text category according to the high-dimensional speech feature signal, thereby generating a text output signal.

[0016] The aforementioned method for adaptively generating traffic domain service speech based on chain-of-thought fine-tuning of a large model, step B: Use an accent type recognition decoder to recognize the input speech signal and output the audio accent type probability. The accent type recognition decoder is specifically an S-vector decoding layer. The S-vector decoding layer specifically uses a feed-forward neural network to map the audio features of the input speech signal to a high-dimensional space and uses a state pooling layer to perform statistical pooling on the audio features of the input speech signal, thereby forming a sentence-level feature output. Then, it is mapped to an accent type database through two feed-forward neural networks, and then passes through a linear layer and a Softmax layer to output the audio accent type probability, thereby outputting the audio accent type probability.

[0017] The aforementioned method for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model, step C: Construct a multi-task speech recognition model based on a speech encoder, a text decoder, a pinyin decoder, and an accent type recognition decoder, and then use the multi-task speech recognition model to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type. The multi-task speech recognition model is specifically constructed based on the speech encoder and in combination with the relevance of the speech recognition task, the audio-to-pinyin task, and the accent classification task. The multi-task speech recognition model specifically performs relationship embedding calculation on the outputs of the text decoder and the pinyin decoder, and after fusing the calculated cross-modal features with the original text representation, generates the final text probability through a linear layer and a Softmax layer. The multi-task speech recognition model also includes a gradient reversal mechanism and a pinyin-text relationship model. The gradient reversal mechanism is used to promote the encoder to extract deep acoustic features irrelevant to the accent through reverse optimization. The pinyin-text relationship model is used to correct semantic deviations by utilizing the acoustic alignment characteristics of the pinyin sequence and improve the adaptability to complex accent and noise scenarios.

[0018] The aforementioned method for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model, step D: Formulate a large model annotation specification, then construct a large model dataset and annotate the large model dataset according to the annotation specification to obtain an annotated large model dataset. The large model annotation specification includes mask annotation and data text description annotation, and the data text description annotation is formulated using the chain-of-thought method.

[0019] The aforementioned method for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model, step E: Use the LoRA low-rank fine-tuning method and the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a chain-of-thought fine-tuned large model. The specific steps are as follows.

[0020] Step E1: Introduce the LoRA low-rank fine-tuning method to adapt the fully connected layer of the annotated large model dataset. The LoRA low-rank fine-tuning method is specifically to add a side branch network in the large model dataset and update the weights of the fully connected layer using the product of two rank decomposition matrices, and then update the parameters in the side branch network.

[0021] Step E2: Use the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a chain-of-thought fine-tuned large model. The mixed-precision quantization method is specifically to use a low-precision data storage type and perform de-quantization processing during the calculation process, and then combine BFloat16 for high-precision calculation and implement the fine-tuning of the annotated large model dataset to obtain a chain-of-thought fine-tuned large model. The mixed-precision quantization method specifically includes the 4-bit NormalFloat quantization method and the double quantization method. The specific steps are as follows.

[0022] Step E21, model weight quantization and normalization processing, specifically, the quantile interval is calculated and the quantization mapping table is determined by using the property that the weights follow a zero-centered normal distribution. Specifically, when calculating the quantile interval, the weight tensor is normalized to the range [-1, 1], and then the discrete values are determined through the quantile quantization formula as shown in formula (1).

[0023] (1)

[0024] where, is the discrete value, is the i-th element value of the original weight, and are the weight mean and weight standard deviation respectively, and k is the quantization bit number;

[0025] Step E22, double quantization and parameter compression, specifically, a two-level quantization mechanism is introduced for the quantized weight matrix. Among them, the first level performs 4-bit NormalFloat quantization on the original weights, and the second level performs 8-bit floating-point requantization on the quantization constants and realizes the secondary compression of the quantization parameters, as shown in formula (2).

[0026] (2)

[0027] where, is the result after performing 4-bit NormalFloat quantization on the weight , is 4-bit NormalFloat quantization, is the quantization function, is the result after performing 8-bit floating-point quantization according to the intermediate calculation result , is 8-bit floating-point quantization;

[0028] Step E23, paging optimization and video memory management, specifically, a paging optimizer based on NVIDIA unified memory is constructed to dynamically manage the GPU-CPU memory exchange. If the video memory occupancy exceeds the threshold T during the calculation process, the gradient checkpoint data is automatically paged and transferred to the CPU memory, as shown in formula (3).

[0029] (3)

[0030] where, is the video memory, is the function to swap out the page in the memory to the disk and release the memory resources, is the data gradient;

[0031] Step E24, low-rank adaptation and parameter fusion. Specifically, a trainable low-rank adapter is inserted into the quantized base model to freeze the original quantized weights. The low-rank adapter consists of decomposition matrix A and decomposition matrix B, and the parameter fusion formula is as shown in formula (4).

[0032] (4)

[0033] where is the parameter fusion result, the quantized weight matrix, is the input vector, is the scaling coefficient, and are both additional weight matrices.

[0034] The aforementioned method for adaptively generating traffic domain service speech by thinking chain fine-tuning a large model, step F, uses the thinking chain to fine-tune the large model to adaptively generate and output emotional traffic domain service speech according to text representation, pinyin sequence, and audio accent type. The adaptive generation specifically uses a three-level architecture of acoustic generation model - prosody predictor - neural vocoder and integrates a polyphone disambiguation model and a context-aware dynamic adjustment mechanism to achieve semantic-driven prosody adaptive generation while ensuring accurate pronunciation of professional terms, so as to meet the high-quality speech synthesis requirements in real-time interaction and complex noise environments. The specific steps are as follows

[0035] Step F1, construct a prosody predictor. The prosody predictor specifically constructs a continuous phoneme duration prediction model using a joint mechanism of variational dequantization and data augmentation, and then uses the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. The specific steps are as follows

[0036] Step F11, construct a continuous phoneme duration prediction model using a joint mechanism of variational dequantization and data augmentation. The continuous phoneme duration prediction model specifically constructs a continuous latent space through two groups of variables with the same dimension as the phoneme sequence;

[0037] Step F12, use the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. Specifically, based on the variational lower bound optimization objective, posterior distribution sampling training is used and the gradient backpropagation of the stochastic duration predictor is blocked synchronously to maintain the independence of the module. Then, the reversible flow model generates continuous duration values from noise and outputs a phoneme alignment sequence that conforms to the rhythm of human speech after integer conversion;

[0038] Step F2: Build an acoustic generation model. The acoustic generation model specifically uses the feature enhancement layer of a dual-channel multi-modal discriminator to extract an emotional acoustic feature matrix, and then establishes a dual supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. The specific steps are as follows:

[0039] Step F21: Use the feature enhancement layer of a dual-channel multi-modal discriminator to extract an emotional acoustic feature matrix. Specifically, the dual-channel multi-modal discriminator uses a dual-channel architecture including the DiscriminatorS network and the DiscriminatorP network to re-extract the Mel Frequency Cepstral Coefficient (MFCC) parameters of the generated spectrum and the real spectrum, and then analyzes the emotional information labels of the generated spectrum through an acoustic feature analysis module based on a Bidirectional Long Short-Term Memory (Bi-LSTM) network. Then, the emotional information labels and the original emotional annotations are jointly used to form an auxiliary feature matrix, and the auxiliary feature matrix is injected into the convolutional processing channel of the dual-channel multi-modal discriminator in parallel to obtain the emotional acoustic feature matrix;

[0040] Step F22: Establish a dual supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. Specifically, the visual features of the emotional acoustic feature matrix are extracted through a conventional convolutional layer, and the implicit emotional associations in the acoustic feature space are analyzed synchronously. Then, the basic spectral feature map and the enhanced emotional feature map are maintained, and the consistency of emotional expression is judged through cross-modal feature cross-verification;

[0041] Step F3: Build a neural vocoder. The neural vocoder specifically uses a hierarchical composite loss function system to fuse spectral reconstruction, feature space constraint, and dual emotion supervision, and combines an environmental noise suppression method to enhance the robustness of the vocoder. The specific steps are as follows:

[0042] Step F31: Build a hierarchical composite loss function system. The hierarchical composite loss function system uses a multi-variable composite loss function to perform multi-objective optimization on speech synthesis. The multi-variable composite loss function includes a basic loss layer and a feature enhancement layer. The basic loss layer is specifically composed of a reconstruction loss, a KL divergence loss for controlling the stability of adversarial training, and a Wasserstein distance optimization adversarial loss. The feature enhancement layer is specifically composed of a Mel spectrum perception loss and a feature map comparison loss. The Mel spectrum perception loss specifically uses a Mel-scale frequency domain conversion to establish an acoustic physical constraint, and the feature map comparison loss specifically uses the multi-level feature maps output by the discriminator for cross-network layer feature matching;

[0043] Step F32: Build a dual supervision mechanism for emotional vectors. The dual supervision mechanism for emotional vectors specifically uses the distance loss in the emotional label embedding space and the emotional feature component loss in the discriminator feature map to implement dual emotional constraints.

[0044] A traffic domain service voice adaptive generation system based on chain of thought fine-tuning of large models, including a text output signal generation module, an audio accent type probability recognition module, a multi-task speech recognition module, a large model annotation specification formulation module, a chain of thought fine-tuning large model establishment module, and an emotional traffic domain service voice generation module. The text output signal generation module is used to convert the input voice signal into a high-dimensional voice feature signal by using a voice encoder, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional voice feature signal. The audio accent type probability recognition module is used to identify the input voice signal by using an accent type recognition decoder and output the audio accent type probability. The multi-task speech recognition module is used to construct a multi-task speech recognition model based on the voice encoder, the text decoder, the pinyin decoder, and the accent type recognition decoder, and then further process the text output signal and the audio accent type probability by using the multi-task speech recognition model to obtain a text representation, a pinyin sequence, and an audio accent type. The large model annotation specification formulation module is used to formulate a large model annotation specification, then construct a large model data set and annotate the large model data set according to the annotation specification to obtain an annotated large model data set. The chain of thought fine-tuning large model establishment module is used to perform lightweight training on the annotated large model data set by using the LoRA low-rank fine-tuning method and the mixed precision quantization method to obtain a chain of thought fine-tuning large model. The emotional traffic domain service voice generation module is used to adaptively generate and output an emotional traffic domain service voice by using the chain of thought fine-tuning large model according to the text representation, the pinyin sequence, and the audio accent type.

[0045] The beneficial effects of the present invention are as follows: A method and system for adaptively generating traffic domain service speech based on chain-of-thought fine-tuning of a large model. First, a voice encoder is used to convert the input voice signal into a high-dimensional voice feature signal. Then, a text decoder and a pinyin decoder generate a text output signal based on the high-dimensional voice feature signal. Next, an accent type recognition decoder is used to identify the input voice signal and output the probability of the audio accent type. Subsequently, a multi-task speech recognition model is constructed based on the voice encoder, text decoder, pinyin decoder, and accent type recognition decoder. Then, the multi-task speech recognition model is used to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type. Then, a large model annotation specification is formulated, and a large model dataset is constructed and annotated according to the annotation specification to obtain an annotated large model dataset. Finally, the LoRA low-rank fine-tuning method and the mixed-precision quantization method are used to perform lightweight training on the annotated large model dataset to obtain a chain-of-thought fine-tuned large model. Then, the chain-of-thought fine-tuned large model is used to adaptively generate and output emotional traffic domain service speech according to the text representation, pinyin sequence, and audio accent type; effectively realizing that the method and system for adaptively generating traffic domain service speech based on chain-of-thought fine-tuning of a large model have the functions of using a variational quantization joint data augmentation mechanism, a dual-channel multi-modal discriminator architecture, and a hierarchical composite loss function for high-fidelity emotional speech generation and robust synthesis in a complex noise environment, and simultaneously supporting semantic-driven dynamic prosody optimization and accurate pronunciation of professional terms. Moreover, in the application scenario of the traffic field, the multi-task speech recognition method can be used to achieve the efficient linkage of character-level recognition, audio-to-pinyin conversion, and sentence-level accent classification modules, so as to effectively cope with the challenges of complex accents, a large number of background noises, and multiple pronunciations of traffic terms. At the same time, by introducing a chain-of-thought module on the basis of the original large model to simulate the human thinking process, the close connection between the conclusion description and the reasoning process can be achieved, thereby stimulating logical reasoning ability and improving the reasoning accuracy. And by combining the LoRA low-rank fine-tuning method and the mixed-precision quantization method, high accuracy and generalization ability can be maintained while quickly training with limited computing power, thereby improving the business intention recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is the overall flowchart of a method for adaptively generating traffic domain service speech based on chain-of-thought fine-tuning of a large model according to the present invention;

[0047] Figure 2 is the adaptive generation principle diagram of a system for adaptively generating traffic domain service speech based on chain-of-thought fine-tuning of a large model according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The present invention will be further described below in conjunction with the accompanying drawings of the specification.

[0049] As Figure 1As shown in the figure, a traffic domain service voice adaptive generation method based on thinking chain fine-tuning of a large model according to the present invention includes the following steps:

[0050] Step A: Use a voice encoder to convert the input voice signal into a high-dimensional voice feature signal, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional voice feature signal. The specific steps are as follows:

[0051] Step A1: Use a voice encoder to convert the input voice signal into a high-dimensional voice feature signal. The voice encoder is specifically a Conformer encoding layer. The Conformer encoding layer is used to construct local audio feature dependencies and combine the global feature construction ability of the self-attention mechanism to improve the performance of the automatic speech recognition task. The Conformer encoding layer is specifically composed of a first feed-forward module, a self-attention module, a convolutional module, a second feed-forward module, and a first layer of regularization module connected in sequence;

[0052] Step A2: Generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional voice feature signal. Both the text decoder and the pinyin decoder use a Transformer-ASR decoding layer. The Transformer-ASR decoding layer specifically uses a multi-head self-attention layer and uses a second layer of regularization module, a linear output module, and a Softmax function module to generate the probability of the output text category according to the high-dimensional voice feature signal, thereby generating a text output signal.

[0053] Step B: Use an accent type recognition decoder to recognize the input voice signal and output the audio accent type probability. The accent type recognition decoder is specifically an S-vector decoding layer. The S-vector decoding layer specifically uses a feed-forward neural network to map the audio features of the input voice signal to a high-dimensional space and uses a state pooling layer to perform statistical pooling on the audio features of the input voice signal, thereby forming a sentence-level feature output. Then, it is mapped to an accent type database through two feed-forward neural networks, and then passes through a linear layer and a Softmax layer to output the audio accent type probability, thereby outputting the audio accent type probability.

[0054] Step C: Build a multi-task speech recognition model based on a speech encoder, a text decoder, a pinyin decoder, and an accent type recognition decoder. Then, use the multi-task speech recognition model to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type. Specifically, the multi-task speech recognition model is built based on the speech encoder and in combination with the relevance of the speech recognition task, the audio-to-pinyin task, and the accent classification task. The multi-task speech recognition model specifically performs a relationship embedding calculation on the outputs of the text decoder and the pinyin decoder, and fuses the calculated cross-modal features with the original text representation, and then generates the final text probability through a linear layer and a Softmax layer. The multi-task speech recognition model also includes a gradient reversal mechanism and a pinyin-text relationship model. The gradient reversal mechanism is used to promote the encoder to extract deep acoustic features irrelevant to the accent through reverse optimization. The pinyin-text relationship model is used to correct semantic biases and improve the adaptability to complex accents and noise scenarios by using the acoustic alignment characteristics of the pinyin sequence.

[0055] Step D: Formulate a large model annotation specification, then build a large model dataset and annotate the large model dataset according to the annotation specification to obtain an annotated large model dataset. The large model annotation specification includes masked annotation and data text description annotation. The data text description annotation is formulated using the chain of thought method.

[0056] Step E: Use the LoRA low-rank fine-tuning method and the mixed precision quantization method to perform lightweight training on the annotated large model dataset to obtain a chain of thought fine-tuned large model. The specific steps are as follows:

[0057] Step E1: Introduce the LoRA low-rank fine-tuning method to adapt the fully connected layer of the annotated large model dataset. Specifically, the LoRA low-rank fine-tuning method adds a side network to the large model dataset and updates the weights of the fully connected layer using the product of two rank decomposition matrices, and then updates the parameters in the side network.

[0058] Step E2: Use the mixed precision quantization method to perform lightweight training on the annotated large model dataset to obtain a chain of thought fine-tuned large model. Specifically, the mixed precision quantization method uses a low-precision data storage type and performs de-quantization processing during the calculation process, and then combines BFloat16 for high-precision calculation and realizes the fine-tuning of the annotated large model dataset to obtain a chain of thought fine-tuned large model. The mixed precision quantization method specifically includes the 4-bit NormalFloat quantization method and the double quantization method. The specific steps are as follows:

[0059] Step E21: Model weight quantization and normalization. Specifically, the quantile interval is calculated and the quantization mapping table is determined by using the property that the weights follow a zero-centered normal distribution. Specifically, when calculating the quantile interval, the weight tensor is normalized to the range [-1, 1], and then the discrete values are determined through the quantile quantization formula as shown in Equation (1).

[0060] (1)

[0061] where, is the discrete value, is the i-th element value of the original weight, and are the weight mean and weight standard deviation respectively, and k is the quantization bit number;

[0062] Step E22: Dual quantization and parameter compression. Specifically, a two-level quantization mechanism is introduced for the quantized weight matrix. Among them, the first level performs 4-bit NormalFloat quantization on the original weights, and the second level performs 8-bit floating-point re-quantization on the quantization constants and realizes the secondary compression of the quantization parameters, as shown in Equation (2).

[0063] (2)

[0064] where, is the result after performing 4-bit NormalFloat quantization on the weight , is 4-bit NormalFloat quantization, is the quantization function, is the result after performing 8-bit floating-point quantization according to the intermediate calculation result , is 8-bit floating-point quantization;

[0065] Step E23: Paging optimization and video memory management. Specifically, a paging optimizer based on NVIDIA unified memory is constructed to dynamically manage the GPU-CPU memory exchange. If the video memory occupancy exceeds the threshold T during the calculation process, the gradient checkpoint data is automatically paged and transferred to the CPU memory, as shown in Equation (3).

[0066] (3)

[0067] where, is the video memory, is the function to swap out the page in the memory to the disk and release the memory resources, is the data gradient;

[0068] Step E24, low-rank adaptation and parameter fusion. Specifically, a trainable low-rank adapter is inserted into the quantized base model to freeze the original quantized weights. The low-rank adapter consists of decomposition matrix A and decomposition matrix B, and the parameter fusion formula is shown in Formula (4).

[0069] (4)

[0070] where, is the parameter fusion result, the quantized weight matrix, is the input vector, is the scaling coefficient, and are both additional weight matrices.

[0071] Step F, use chain of thought to fine-tune the large model to adaptively generate and output emotional traffic domain service speech according to text representation, pinyin sequence and audio accent type. The adaptive generation specifically adopts a three-level architecture of acoustic generation model - prosody predictor - neural vocoder and integrates a polyphone disambiguation model and a context-aware dynamic adjustment mechanism to achieve semantic-driven prosody adaptive generation while ensuring the accurate pronunciation of professional terms, so as to meet the high-quality speech synthesis requirements in real-time interaction and complex noise environments. The specific steps are as follows

[0072] Step F1, construct a prosody predictor. The prosody predictor specifically constructs a continuous phoneme duration prediction model by using a joint mechanism of variational dequantization and data augmentation, and then uses the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. The specific steps are as follows

[0073] Step F11, construct a continuous phoneme duration prediction model by using a joint mechanism of variational dequantization and data augmentation. The continuous phoneme duration prediction model specifically constructs a continuous latent space through two groups of variables with the same dimension as the phoneme sequence;

[0074] Step F12, use the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. Specifically, based on the variational lower bound optimization objective, posterior distribution sampling training is adopted and the gradient backpropagation of the stochastic duration predictor is blocked synchronously to maintain the independence of the module. Then, the reversible flow model generates continuous duration values from noise and outputs a phoneme alignment sequence that conforms to the rhythm of human speech after integer conversion;

[0075] Step F2, construct an acoustic generation model. The acoustic generation model specifically extracts an emotional acoustic feature matrix by using the feature enhancement layer of a dual-channel multi-modal discriminator, and then establishes a dual supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the emotional expression consistency. The specific steps are as follows

[0076] Step F21: Use the feature enhancement layer of the dual-channel multi-modal discriminator to extract the emotional acoustic feature matrix. Specifically, the dual-channel multi-modal discriminator uses a dual-channel architecture including the DiscriminatorS network and the DiscriminatorP network to re-extract the Mel Frequency Cepstral Coefficient (MFCC) parameters of the generated spectrum and the real spectrum. Then, the emotional information label of the generated spectrum is parsed through the acoustic feature analysis module based on the Bidirectional Long Short-Term Memory Network (Bi-LSTM). Next, the emotional information label and the original emotional annotation are jointly used to form an auxiliary feature matrix. Then, the auxiliary feature matrix is injected into the convolutional processing channel of the dual-channel multi-modal discriminator in a parallel connection manner to obtain the emotional acoustic feature matrix;

[0077] Step F22: Establish a dual supervision mechanism of spectrum structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. Specifically, extract the visual features of the emotional acoustic feature matrix through the conventional convolutional layer and synchronously analyze the implicit emotional association in the acoustic feature space. Then, maintain the basic spectrum feature map and the enhanced emotional feature map, and judge the consistency of emotional expression through cross-modal feature cross-verification;

[0078] Step F3: Construct a neural vocoder. The neural vocoder specifically uses a hierarchical composite loss function system to fuse spectrum reconstruction, feature space constraint, and dual emotion supervision, and combines the environmental noise suppression method to enhance the robustness of the vocoder. The specific steps are as follows:

[0079] Step F31: Construct a hierarchical composite loss function system. The hierarchical composite loss function system uses a multi-source composite loss function to perform multi-objective optimization on speech synthesis. The multi-source composite loss function includes a basic loss layer and a feature enhancement layer. The basic loss layer is specifically composed of a reconstruction loss, a KL divergence loss for controlling the stability of adversarial training, and a Wasserstein distance optimization adversarial loss. The feature enhancement layer is specifically composed of a Mel spectrum perception loss and a feature map contrast loss. The Mel spectrum perception loss specifically uses the Mel-scale frequency domain conversion to establish acoustic physical constraints. The feature map contrast loss specifically uses the multi-level feature maps output by the discriminator to perform cross-network layer feature matching;

[0080] Step F32: Construct a dual supervision mechanism for emotional vectors. The dual supervision mechanism for emotional vectors specifically uses the distance loss in the emotional label embedding space and the emotional feature component loss in the discriminator feature map to implement dual emotional constraints.

[0081] such as Figure 2As shown in the figure, a traffic domain service voice adaptive generation system based on chain-of-thought fine-tuning of a large model includes a text output signal generation module, an audio accent type probability recognition module, a multi-task speech recognition module, a large model annotation specification formulation module, a chain-of-thought fine-tuning large model establishment module, and an emotional traffic domain service voice generation module. The text output signal generation module is used to convert the input voice signal into a high-dimensional voice feature signal by using a voice encoder, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional voice feature signal; the audio accent type probability recognition module is used to recognize the input voice signal by using an accent type recognition decoder and output the audio accent type probability; the multi-task speech recognition module is used to construct a multi-task speech recognition model based on the voice encoder, the text decoder, the pinyin decoder, and the accent type recognition decoder, and then further process the text output signal and the audio accent type probability by using the multi-task speech recognition model to obtain a text representation, a pinyin sequence, and an audio accent type; the large model annotation specification formulation module is used to formulate a large model annotation specification, then construct a large model data set and annotate the large model data set according to the annotation specification to obtain an annotated large model data set; the chain-of-thought fine-tuning large model establishment module is used to perform lightweight training on the annotated large model data set by using the LoRA low-rank fine-tuning method and the mixed precision quantization method to obtain a chain-of-thought fine-tuning large model; the emotional traffic domain service voice generation module is used to adaptively generate and output emotional traffic domain service voice by using the chain-of-thought fine-tuning large model according to the text representation, the pinyin sequence, and the audio accent type.

[0082] In summary, for a method and system for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model according to the present invention, first, a voice encoder is used to convert an input voice signal into a high-dimensional voice feature signal, then a text decoder and a pinyin decoder are used to generate a text output signal according to the high-dimensional voice feature signal. Next, an accent type recognition decoder is used to recognize the input voice signal and output the probability of the audio accent type. Subsequently, a multi-task voice recognition model is constructed based on the voice encoder, the text decoder, the pinyin decoder, and the accent type recognition decoder. Then, the multi-task voice recognition model is used to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type. Then, a large model annotation specification is formulated, and a large model data set is constructed and annotated according to the annotation specification to obtain an annotated large model data set. Finally, the LoRA low-rank fine-tuning method and the mixed precision quantization method are used to perform lightweight training on the annotated large model data set to obtain a chain-of-thought fine-tuned large model, and then the chain-of-thought fine-tuned large model is used to adaptively generate and output emotional traffic domain service voice according to the text representation, the pinyin sequence, and the audio accent type; effectively realizing that the method and system for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model have the functions of using a variational quantization joint data augmentation mechanism, a dual-channel multi-modal discriminator architecture, and a hierarchical composite loss function for high-fidelity emotional speech generation and robust synthesis in a complex noise environment, and simultaneously supporting semantic-driven dynamic prosody optimization and accurate pronunciation of professional terms. Moreover, in the application scenario of the traffic field, the multi-task voice recognition method can be used to achieve the efficient linkage of character-level recognition, audio-to-pinyin, and sentence-level accent classification modules, so as to effectively cope with the challenges of complex accents, a large number of background noises, and polyphonic traffic terms. At the same time, by introducing a chain-of-thought module on the basis of the original large model to simulate the human thinking process, the close connection between the conclusion description and the reasoning process can be achieved, thereby stimulating logical reasoning ability and improving the reasoning accuracy. And by combining the LoRA low-rank fine-tuning method and the mixed precision quantization method, high accuracy and generalization ability can be maintained while quickly training under limited computing power, thereby improving the business intention recognition effect.

[0083] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only used to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all of these changes and improvements fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A traffic domain service voice adaptive generation method based on chain-of-thought fine-tuning of large models, characterized in that: It includes the following steps: Step A: Use a speech encoder to convert the input speech signal into a high-dimensional speech feature signal, and then generate a text output signal through a text decoder and a pinyin decoder based on the high-dimensional speech feature signal; Step B: Use an accent type recognition decoder to recognize the input speech signal and output the audio accent type probability; Step C: Build a multi-task speech recognition model based on the speech encoder, text decoder, pinyin decoder, and accent type recognition decoder, and then use the multi-task speech recognition model to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type; Step D: Formulate a large model annotation specification, then build a large model dataset and annotate the large model dataset according to the annotation specification to obtain an annotated large model dataset; Step E: Use the LoRA low-rank fine-tuning method and the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a chain-of-thought fine-tuned large model. The specific steps are as follows: Step E1: Introduce the LoRA low-rank fine-tuning method to adapt the fully connected layer of the annotated large model dataset. Specifically, the LoRA low-rank fine-tuning method is to add a bypass network in the large model dataset and update the weights of the fully connected layer using the product of two rank decomposition matrices, and then update the parameters in the bypass network; Step E2: Use the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a chain-of-thought fine-tuned large model. Specifically, the mixed-precision quantization method is to use a low-precision data storage type and perform de-quantization processing during the calculation process, and then combine BFloat16 for high-precision calculation and implement fine-tuning of the annotated large model dataset to obtain a chain-of-thought fine-tuned large model. The mixed-precision quantization method specifically includes the 4-bit NormalFloat quantization method and the double quantization method. The specific steps are as follows: Step E21: Model weight quantization and normalization processing. Specifically, use the characteristic that the weights follow a zero-centered normal distribution to calculate the quantile interval and determine the quantization mapping table. Specifically, calculating the quantile interval is to normalize the weight tensor to the range [-1, 1], and then determine the discrete value through the quantile quantization formula as shown in formula (1); Among them, q i is a discrete value, w i is the i-th element value of the original weight, μ and σ are the weight mean and the weight standard deviation respectively, and k is the quantization bit number; Step E22: Double quantization and parameter compression. Specifically, introduce a two-level quantization mechanism for the quantized weight matrix. The first level performs 4-bit NormalFloat quantization on the original weights, and the second level performs 8-bit floating-point re-quantization on the quantization constants and realizes secondary compression of the quantization parameters, as shown in formula (2); Q1(w) = Quant(w,NF4),Q2(s) = Quant(s,FP8) (2) where Q1(w) is the result after performing 4-bit NormalFloat quantization on the weight w, NF4 is 4-bit NormalFloat quantization, Quant is the quantization function, Q2(s) is the result after performing 8-bit floating-point quantization on the intermediate calculation result s, and FP8 is 8-bit floating-point quantization; Step E23, Paging Optimization and Video Memory Management, specifically, build a paging optimizer based on NVIDIA unified memory to dynamically manage GPU-CPU memory swapping. When the video memory occupancy exceeds the threshold T during the calculation process, the gradient checkpoint data is automatically paged and transferred to the CPU memory, as shown in formula (3). Where Mused is the video memory, PageOut is the function to swap out the page in the memory to the disk and release the memory resources, and Dgradient is the data gradient. Step E24, Low-Rank Adaptation and Parameter Fusion, specifically, insert a trainable low-rank adapter on the quantized base model to freeze the original quantized weights. The low-rank adapter consists of decomposition matrix A and decomposition matrix B, and the parameter fusion formula is as shown in formula (4). h out = W quant ·x + α·(B·A·x)(4) Among them, h out is the parameter fusion result, W quant is the quantized weight matrix, x is the input vector, α is the scaling factor, and both A and B are additional weight matrices; Step F, Use chain of thought to fine-tune the large model to adaptively generate and output emotional traffic domain service speech according to text representation, pinyin sequence, and audio accent type. The adaptive generation specifically adopts a three-level architecture of acoustic generation model - prosody predictor - neural vocoder and integrates a polyphone disambiguation model and a context-aware dynamic adjustment mechanism to achieve semantic-driven prosody adaptive generation while ensuring the accurate pronunciation of professional terms, so as to meet the high-quality speech synthesis requirements in real-time interaction and complex noise environments. The specific steps are as follows. Step F1, Build a prosody predictor. The prosody predictor specifically uses a joint mechanism of variational dequantization and data augmentation to build a continuous phoneme duration prediction model, and then uses the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. Step F2, Build an acoustic generation model. The acoustic generation model specifically uses the feature enhancement layer of a dual-channel multi-modal discriminator to extract the emotional acoustic feature matrix, and then establishes a dual supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. Step F3, Build a neural vocoder. The neural vocoder specifically uses a hierarchical composite loss function system to fuse spectrum reconstruction, feature space constraint, and emotion dual supervision, and combines the environmental noise suppression method to enhance the robustness of the vocoder.

2. The traffic domain service voice adaptive generation method based on thought chain fine-tuning of a large model according to claim 1, wherein: Step A, Use a speech encoder to convert the input speech signal into a high-dimensional speech feature signal, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional speech feature signal. The specific steps are as follows. Step A1, Use a speech encoder to convert the input speech signal into a high-dimensional speech feature signal. The speech encoder is specifically a Conformer encoding layer. The Conformer encoding layer is used to build local audio feature dependencies and combine the global feature construction ability of the self-attention mechanism to improve the performance of the automatic speech recognition task. The Conformer encoding layer is specifically composed of a first feed-forward module, a self-attention module, a convolutional module, a second feed-forward module, and a first regularization module connected in sequence. Step A2: Generate a text output signal using a character decoder and a pinyin decoder based on the high-dimensional speech feature signal. Both the character decoder and the pinyin decoder adopt a Transformer-ASR decoding layer. The Transformer-ASR decoding layer specifically uses a multi-head self-attention layer and uses a second-layer regularization module, a linear output module, and a Softmax function module to generate the probability of the output character category based on the high-dimensional speech feature signal, thereby generating a text output signal.

3. The method for adaptively generating traffic domain service speech based on thought chain fine-tuning of a large model according to claim 2, wherein: Step B: Use an accent type recognition decoder to recognize the input speech signal and output the audio accent type probability. The accent type recognition decoder is specifically an S-vector decoding layer. The S-vector decoding layer specifically uses a feed-forward neural network to map the audio features of the input speech signal to a high-dimensional space and uses a state pooling layer to perform statistical pooling on the audio features of the input speech signal, thereby forming a sentence-level feature output. Then, it is mapped to an accent type database through two feed-forward neural networks, and then passes through a linear layer and a Softmax layer to output the audio accent type probability, thereby outputting the audio accent type probability.

4. The method for adaptively generating traffic domain service voice based on thought chain fine-tuning of a large model according to claim 3, wherein: Step C: Build a multi-task speech recognition model based on a speech encoder, a character decoder, a pinyin decoder, and an accent type recognition decoder. Then, use the multi-task speech recognition model to further process the text output signal and the audio accent type probability and obtain a text representation, a pinyin sequence, and an audio accent type. The multi-task speech recognition model is specifically built based on the speech encoder and combines the relevance of the speech recognition task, the audio-to-pinyin task, and the accent classification task. The multi-task speech recognition model specifically performs a relationship embedding calculation on the outputs of the character decoder and the pinyin decoder and fuses the calculated cross-modal features with the original text representation, and then generates the final text probability through a linear layer and a Softmax layer. The multi-task speech recognition model also includes a gradient reversal mechanism and a pinyin-character relationship model. The gradient reversal mechanism is used to promote the encoder to extract deep acoustic features independent of the accent by using reverse optimization. The pinyin-character relationship model is used to correct semantic biases by using the acoustic alignment characteristics of the pinyin sequence and improve the adaptability to complex accents and noise scenarios.

5. A method for adaptively generating traffic domain service voice based on thought chain fine-tuning of a large model according to claim 4, characterized in that: Step D: Formulate a large model annotation specification, then build a large model dataset and annotate the large model dataset according to the annotation specification to obtain an annotated large model dataset. The large model annotation specification includes masked annotation and data text description annotation. The data text description annotation is formulated using the chain of thought method.

6. The method for adaptively generating traffic domain service speech based on fine-tuning a large model with a chain of thought according to claim 5, characterized in that: The specific steps of Step F1 are as follows: Step F11: Build a continuous phoneme duration prediction model using a variational dequantization and data augmentation joint mechanism. The continuous phoneme duration prediction model specifically constructs a continuous latent space through two groups of variables with the same dimension as the phoneme sequence. Step F12: Use the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. Specifically, based on the variational lower bound optimization objective, sample from the posterior distribution for training and synchronously block the gradient backpropagation of the random duration predictor to maintain the independence of the module. Then, generate continuous duration values from noise through the reversible flow model and output a phoneme alignment sequence that conforms to the rhythm of human speech after integer conversion; The specific steps of Step F2 are as follows: Step F21: Extract the emotional acoustic feature matrix using the feature enhancement layer of the dual-channel multi-modal discriminator. Specifically, the dual-channel multi-modal discriminator uses a dual-channel architecture including the DiscriminatorS network and the DiscriminatorP network to re-extract the Mel Frequency Cepstral Coefficient (MFCC) parameters of the generated spectrum and the real spectrum. Then, analyze the emotional information label of the generated spectrum through the acoustic feature analysis module based on the Bidirectional Long Short-Term Memory (Bi-LSTM) network. Next, jointly construct an auxiliary feature matrix with the emotional information label and the original emotional annotation, and then inject the auxiliary feature matrix into the convolutional processing channel of the dual-channel multi-modal discriminator in parallel to obtain the emotional acoustic feature matrix; Step F22: Establish a dual supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. Specifically, extract the visual features of the emotional acoustic feature matrix through a conventional convolutional layer and synchronously analyze the implicit emotional correlation in the acoustic feature space. Then, maintain the basic spectral feature map and the enhanced emotional feature map and discriminate the consistency of emotional expression through cross-modal feature cross-validation; The specific steps of Step F3 are as follows: Step F31: Construct a hierarchical composite loss function system. The hierarchical composite loss function system uses a multi-variate composite loss function to perform multi-objective optimization on speech synthesis. The multi-variate composite loss function includes a basic loss layer and a feature enhancement layer. The basic loss layer is specifically composed of a reconstruction loss, an adversarial training stability control KL divergence loss, and a Wasserstein distance optimization adversarial loss. The feature enhancement layer is specifically composed of a Mel spectrogram perception loss and a feature map contrast loss. The Mel spectrogram perception loss specifically uses Mel-scale frequency domain conversion to establish acoustic physical constraints. The feature map contrast loss specifically performs cross-network layer feature matching using the multi-level feature maps output by the discriminator; Step F32: Construct a dual supervision mechanism for emotional vectors. The dual supervision mechanism for emotional vectors specifically implements dual emotional constraints using the distance loss in the emotional label embedding space and the emotional feature component loss in the discriminator feature map.

7. A traffic domain service voice adaptive generation system based on chain-of-thought fine-tuning of a large model, wherein the specific generation process of the traffic domain service voice adaptive generation system is based on the traffic domain service voice adaptive generation method according to any one of claims 1-6, and is characterized in that: It includes a text output signal generation module, an audio accent type probability recognition module, a multi-task speech recognition module, a large model annotation specification formulation module, a thought chain fine-tuned large model establishment module, and an emotional traffic domain service speech generation module. The text output signal generation module is used to convert the input speech signal into a high-dimensional speech feature signal by using a speech encoder, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional speech feature signal; The audio accent type probability recognition module is used to recognize the input speech signal by using an accent type recognition decoder and output the audio accent type probability; The multi-task speech recognition module is used to construct a multi-task speech recognition model based on a speech encoder, a text decoder, a pinyin decoder, and an accent type recognition decoder, and then further process the text output signal and the audio accent type probability by using the multi-task speech recognition model to obtain a text representation, a pinyin sequence, and an audio accent type; The large model annotation specification formulation module is used to formulate large model annotation specifications, then construct a large model data set and annotate the large model data set according to the annotation specifications to obtain an annotated large model data set; The thought chain fine-tuned large model establishment module is used to perform lightweight training on the annotated large model data set by using the LoRA low-rank fine-tuning method and the mixed precision quantization method to obtain a thought chain fine-tuned large model; The emotional traffic domain service speech generation module is used to adaptively generate and output emotional traffic domain service speech by using the thought chain fine-tuned large model according to the text representation, the pinyin sequence, and the audio accent type.

Citation Information

Patent Citations

  • Semi-structured interview system based on fine-tuning large language model

    CN118886503A

  • System for speech recognition text enhancement fusing multi-modal semantic invariance

    US11488586B1