Adaptive traffic domain service voice generation method and system based on thinking chain fine-tuning large model

By adopting multi-task speech recognition technology based on the thinking chain fine-tuning large model in the interactive voice response system, the problem that the system cannot directly understand the user's natural language expression and lacks dynamic response capabilities is solved, and efficient and continuous voice adaptive generation of traffic domain services is achieved.

CN120126484AActive Publication Date: 2025-06-10NANJING MICROVIDEO TECH +1

Patent Information

Application Number
CN202510604879.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-10
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In the transportation field, interactive voice response systems cannot directly understand the user's natural language expression demands, and lack dynamic response capabilities, resulting in interruption of service chains and inefficiency.

Method used

The traffic domain service voice adaptive generation method based on the thinking chain fine-tuning large model is adopted. A multi-task speech recognition model is constructed through a speech encoder, a text decoder, a pinyin decoder and accent type recognition decoder. Lightweight training is carried out in combination with LoRA low-rank fine-tuning method and mixed precision quantization method to achieve high-fidelity emotional speech generation and robust synthesis in complex noise environments.

Benefits of technology

It realizes the direct understanding and dynamic response of the voice system to the user's natural language, improves the continuity and efficiency of the service chain, can effectively respond to the challenges of complex accents and noise in the transportation field, and improves the accuracy of business intention recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126484A_ABST
    Figure CN120126484A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic domain service voice adaptive generation method and system based on a thinking chain fine-tuning large model, and the method comprises the steps: converting an input voice signal into a high-dimensional voice feature signal through a voice encoder, and generating a text output signal according to the high-dimensional voice feature signal through a text decoder and a pinyin decoder; according to the invention, a variational de-quantization joint data enhancement mechanism, a dual-channel multi-mode discriminator architecture and a layered composite loss function are adopted to carry out high-fidelity emotional speech generation and robustness synthesis in a complex noise environment, and semantic-driven dynamic rhythm optimization and professional term precise pronunciation are synchronously supported; in an application scene in the traffic field, a multi-task speech recognition method can be adopted to realize efficient linkage of character-level recognition, audio-to-pinyin conversion and sentence-level accent classification modules, so that complex accent, more noise and traffic term polyphone challenges are effectively handled, and the method is suitable for being widely popularized and used.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of service voice generation, and particularly to a traffic domain service voice adaptive generation method and system based on chain-of-thought fine-tuning of a large model. Background Art

[0002] An Interactive Voice Response (IVR) system is an automated telephone service technology that interacts with users through voice navigation and keypad input. It is widely used in fields such as enterprise customer service, banking, and telecommunications, helping users to complete operations such as inquiries, transfers, and appointments through voice menus self-service, while reducing labor costs and improving service efficiency.

[0003] Currently, the solutions for customer service and toll collection services in the transportation field mostly rely on the Interactive Voice Response system to complete. However, this system cannot directly understand the demands expressed in natural language by users, but relies on pre-designed fixed voice templates, guiding users to press keys to select through layer-by-layer announcement of options. This one-way tree-like interaction structure requires users to accurately remember and match the paths set by the system. For example, when querying traffic violations, users need to sequentially select hierarchical menus such as license plate jurisdiction and violation type. When users' demands involve cross-category or multi-condition combinations, they often need to experience multiple key presses to jump, and even be forced to restart the process due to incorrect path selection. More critically, in the face of sudden consultations or personalized problems beyond the preset templates, the system lacks dynamic response capabilities, resulting in the interruption of the service chain. In addition, there are obvious boundaries in the functional coverage of the IVR. Its automated processing capabilities are mostly limited to basic information query standardization services, such as balance inquiries and payment status confirmations. Once it involves complex services that require logical judgment or data linkage, such as fine appeal acceptance, the system often cannot complete multi-system data retrieval and logical verification and must forcefully transfer to an artificial agent. Such service breakpoints not only cause users to repeat their demands but also lead to artificial agents spending a large amount of time on secondary information verification. Especially during peak traffic service periods, a large number of consultation requests that have not been automatically processed flood into the artificial channel, further exacerbating the shortage of service resources and forming an efficiency paradox effect. Therefore, it is necessary to design a traffic domain service voice adaptive generation method and system based on chain-of-thought fine-tuning of a large model. Summary of the Invention

[0004] The object of the present invention is to overcome the deficiencies of the prior art and provide a traffic domain service voice adaptive generation method and system based on a large model fine-tuned with a chain of thought to better and effectively solve the problem that the current interactive voice response system generally fails to directly understand the demands expressed by users in natural language and lacks dynamic response capabilities, resulting in the interruption of the service chain. It realizes the functions of high-fidelity emotional speech generation and robust synthesis in complex noise environments by adopting a variational dequantization joint data augmentation mechanism, a dual-channel multimodal discriminator architecture, and a hierarchical composite loss function, and simultaneously supports semantic-driven dynamic prosody optimization and accurate pronunciation of professional terms. In the application scenario of the traffic field, it can adopt a multi-task speech recognition method to achieve the efficient linkage of character-level recognition, audio-to-pinyin conversion, and sentence-level accent classification modules, thus effectively coping with challenges such as complex accents, a lot of background noise, and polyphonic traffic terms.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows: A traffic domain service voice adaptive generation method based on a large model fine-tuned with a chain of thought, comprising the following steps: Step A: Use a voice encoder to convert the input voice signal into a high-dimensional voice feature signal, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional voice feature signal; Step B: Use an accent type recognition decoder to recognize the input voice signal and output the probability of the audio accent type; Step C: Based on the voice encoder, text decoder, pinyin decoder, and accent type recognition decoder, construct a multi-task speech recognition model, and then use the multi-task speech recognition model to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type; Step D: Formulate a large model annotation specification, then construct a large model data set and annotate the large model data set according to the annotation specification to obtain an annotated large model data set; Step E: Use the LoRA low-rank fine-tuning method and the mixed-precision quantization method to perform lightweight training on the annotated large model data set to obtain a large model fine-tuned with a chain of thought; Step F: Use the large model fine-tuned with a chain of thought to adaptively generate and output emotional traffic domain service voices according to the text representation, pinyin sequence, and audio accent type.

[0006] For the foregoing traffic domain service voice adaptive generation method based on a large model fine-tuned with a chain of thought, in Step A, using a voice encoder to convert the input voice signal into a high-dimensional voice feature signal, and then generating a text output signal through a text decoder and a pinyin decoder according to the high-dimensional voice feature signal, the specific steps are as follows: Step A1, use a voice encoder to convert the input voice signal into a high-dimensional voice feature signal. The voice encoder is specifically a Conformer encoding layer. The Conformer encoding layer is used to construct local audio feature dependencies and combine the global feature construction ability of the self-attention mechanism to improve the performance of the automatic speech recognition task. The Conformer encoding layer is specifically composed of a first feed-forward module, a self-attention module, a convolutional module, a second feed-forward module, and a first regularization module connected in sequence; Step A2, generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional voice feature signal. Both the text decoder and the pinyin decoder adopt a Transformer-ASR decoding layer. The Transformer-ASR decoding layer specifically uses a multi-head self-attention layer and uses a second regularization module, a linear output module, and a Softmax function module to generate the output text category probability according to the high-dimensional voice feature signal, thereby generating a text output signal.

[0007] For the aforementioned traffic domain service voice adaptive generation method based on chain-of-thought fine-tuning of a large model, in step B, an accent type recognition decoder is used to recognize the input voice signal and output the audio accent type probability. The accent type recognition decoder is specifically an S-vector decoding layer. The S-vector decoding layer specifically uses a feed-forward neural network to map the audio features of the input voice signal to a high-dimensional space and uses a state pooling layer to perform statistical pooling on the audio features of the input voice signal, thereby forming a sentence-level feature output. Then, it is mapped to an accent type database through two feed-forward neural networks, and then passes through a linear layer and a Softmax layer to output the audio accent type probability, thereby outputting the audio accent type probability.

[0008] The foregoing method for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model, step C: construct a multi-task speech recognition model based on a speech encoder, a text decoder, a pinyin decoder, and an accent type recognition decoder, and then use the multi-task speech recognition model to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type. The multi-task speech recognition model is specifically constructed based on the speech encoder and in combination with the relevance of the speech recognition task, the audio-to-pinyin task, and the accent classification task. The multi-task speech recognition model specifically performs relational embedding calculation on the outputs of the text decoder and the pinyin decoder, fuses the calculated cross-modal features with the original text representation, and then generates the final text probability through a linear layer and a Softmax layer. The multi-task speech recognition model also includes a gradient reversal mechanism and a pinyin-text relationship model. The gradient reversal mechanism is used to promote the encoder to extract deep acoustic features irrelevant to the accent through reverse optimization. The pinyin-text relationship model is used to correct semantic deviations by utilizing the acoustic alignment characteristics of the pinyin sequence and improve the adaptability to complex accent and noise scenarios.

[0009] The foregoing method for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model, step D: formulate a large model annotation specification, then construct a large model dataset and annotate the large model dataset according to the annotation specification to obtain an annotated large model dataset. The large model annotation specification includes masked annotation and data text description annotation, and the data text description annotation is formulated using the chain-of-thought method.

[0010] The foregoing method for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model, step E: use the LoRA low-rank fine-tuning method and the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a chain-of-thought fine-tuned large model. The specific steps are as follows. Step E1: Introduce the LoRA low-rank fine-tuning method to adapt the fully connected layer of the annotated large model dataset. The LoRA low-rank fine-tuning method specifically adds a side network to the large model dataset and updates the weights of the fully connected layer using the product of two rank decomposition matrices, and then updates the parameters in the side network. Step E2: Use the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a chain-of-thought fine-tuned large model. The mixed-precision quantization method specifically uses a low-precision data storage type and performs de-quantization processing during the calculation process, and then combines BFloat16 for high-precision calculation and realizes the fine-tuning of the annotated large model dataset to obtain a chain-of-thought fine-tuned large model. The mixed-precision quantization method specifically includes the 4-bit NormalFloat quantization method and the double quantization method. The specific steps are as follows. Step E21, model weight quantization and normalization processing. Specifically, the quantile interval is calculated and the quantization mapping table is determined by using the characteristic that the weights follow a zero-centered normal distribution. Specifically, when calculating the quantile interval, the weight tensor is normalized to the range of [-1, 1], and then the discrete values are determined through the quantile quantization formula as shown in formula (1). (1) where, is the discrete value, is the i-th element value of the original weight, and are the weight mean and weight standard deviation respectively, and k is the quantization bit number; Step E22, double quantization and parameter compression. Specifically, a two-level quantization mechanism is introduced for the quantized weight matrix. Among them, the first level performs 4-bit NormalFloat quantization on the original weights, and the second level performs 8-bit floating-point requantization on the quantization constants and realizes secondary compression of the quantization parameters, as shown in formula (2). (2) where, is the result after performing 4-bit NormalFloat quantization on the weight , is 4-bit NormalFloat quantization, is the quantization function, is the result after performing 8-bit floating-point quantization according to the intermediate calculation result , is 8-bit floating-point quantization; Step E23, paging optimization and video memory management. Specifically, a paging optimizer based on NVIDIA unified memory is constructed to dynamically manage the GPU-CPU memory exchange. If the video memory occupancy exceeds the threshold T during the calculation process, the gradient checkpoint data is automatically paged and transferred to the CPU memory, as shown in formula (3). (3) where, is the video memory, is the function to swap out the page in the memory to the disk and release the memory resources, is the data gradient; Step E24, low-rank adaptation and parameter fusion. Specifically, a trainable low-rank adapter is inserted into the quantized base model to freeze the original quantized weights. The low-rank adapter consists of the decomposition matrix A and the decomposition matrix B, and the parameter fusion formula is as shown in formula (4). (4) where, is the parameter fusion result, The quantized weight matrix, is the input vector, is the scaling factor, and are both additional weight matrices.

[0011] The aforementioned method for adaptively generating emotional traffic domain service speech by fine-tuning a large model based on the chain of thought, step F, uses the chain of thought to fine-tune the large model to adaptively generate and output emotional traffic domain service speech according to text representation, pinyin sequence, and audio accent type. The adaptive generation specifically adopts a three-level architecture of an acoustic generation model - prosody predictor - neural vocoder and integrates a polyphone disambiguation model and a context-aware dynamic adjustment mechanism to achieve semantic-driven prosody adaptive generation while ensuring the accurate pronunciation of professional terms, thereby meeting the high-quality speech synthesis requirements in real-time interaction and complex noise environments. The specific steps are as follows. Step F1, construct a prosody predictor. The prosody predictor specifically uses a joint mechanism of variational dequantization and data augmentation to construct a continuous phoneme duration prediction model, and then uses the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. The specific steps are as follows. Step F11, use a joint mechanism of variational dequantization and data augmentation to construct a continuous phoneme duration prediction model. The continuous phoneme duration prediction model specifically constructs a continuous latent space through two groups of variables with the same dimension as the phoneme sequence. Step F12, use the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. Specifically, based on the variational lower bound optimization objective, posterior distribution sampling training is used and the gradient backpropagation of the stochastic duration predictor is blocked synchronously to maintain the independence of the module. Then, the reversible flow model is used to generate continuous duration values from noise and after integer conversion, an aligned phoneme sequence that conforms to the rhythm of human speech is output. Step F2, construct an acoustic generation model. The acoustic generation model specifically uses the feature enhancement layer of a dual-channel multi-modal discriminator to extract an emotional acoustic feature matrix, and then establishes a dual supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. The specific steps are as follows. Step F21: Use the feature enhancement layer of the dual-channel multi-modal discriminator to extract the emotional acoustic feature matrix. Specifically, the dual-channel multi-modal discriminator uses a dual-channel architecture including the DiscriminatorS network and the DiscriminatorP network to re-extract the Mel Frequency Cepstral Coefficient (MFCC) parameters of the generated spectrum and the real spectrum. Then, the emotional information label of the generated spectrum is parsed through the acoustic feature analysis module based on the Bidirectional Long Short-Term Memory (Bi-LSTM) network. Next, the emotional information label and the original emotional annotation together form an auxiliary feature matrix, and the auxiliary feature matrix is injected into the convolutional processing channel of the dual-channel multi-modal discriminator in parallel to obtain the emotional acoustic feature matrix; Step F22: Establish a dual supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. Specifically, extract the visual features of the emotional acoustic feature matrix through the conventional convolutional layer and synchronously analyze the implicit emotional correlation in the acoustic feature space. Then, maintain the basic spectral feature map and the enhanced emotional feature map, and judge the consistency of emotional expression through cross-modal feature cross-verification; Step F3: Construct a neural vocoder. The neural vocoder specifically uses a hierarchical composite loss function system to fuse spectrum reconstruction, feature space constraint, and dual emotion supervision, and combines the environmental noise suppression method to enhance the robustness of the vocoder. The specific steps are as follows: Step F31: Construct a hierarchical composite loss function system. The hierarchical composite loss function system uses a multi-source composite loss function to perform multi-objective optimization on speech synthesis. The multi-source composite loss function includes a basic loss layer and a feature enhancement layer. The basic loss layer is specifically composed of a reconstruction loss, an adversarial training stability control KL divergence loss, and a Wasserstein distance optimization adversarial loss. The feature enhancement layer is specifically composed of a Mel spectrum perception loss and a feature map contrast loss. The Mel spectrum perception loss specifically uses Mel-scale frequency domain conversion to establish acoustic physical constraints. The feature map contrast loss specifically uses the multi-level feature maps output by the discriminator for cross-network layer feature matching; Step F32: Construct a dual supervision mechanism for emotional vectors. The dual supervision mechanism for emotional vectors specifically uses the distance loss in the emotional label embedding space and the emotional feature component loss in the discriminator feature map to implement dual emotional constraints.

[0012] A traffic domain service voice adaptive generation system based on chain-of-thought fine-tuning of large models, including a text output signal generation module, an audio accent type probability recognition module, a multi-task speech recognition module, a large model annotation specification formulation module, a chain-of-thought fine-tuning large model establishment module, and an emotional traffic domain service voice generation module. The text output signal generation module is used to convert the input voice signal into a high-dimensional voice feature signal by using a voice encoder, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional voice feature signal. The audio accent type probability recognition module is used to recognize the input voice signal by using an accent type recognition decoder and output the audio accent type probability. The multi-task speech recognition module is used to construct a multi-task speech recognition model based on the voice encoder, the text decoder, the pinyin decoder, and the accent type recognition decoder, and then further process the text output signal and the audio accent type probability by using the multi-task speech recognition model to obtain a text representation, a pinyin sequence, and an audio accent type. The large model annotation specification formulation module is used to formulate large model annotation specifications, then construct a large model data set and annotate the large model data set according to the annotation specifications to obtain an annotated large model data set. The chain-of-thought fine-tuning large model establishment module is used to perform lightweight training on the annotated large model data set by using the LoRA low-rank fine-tuning method and the mixed precision quantization method to obtain a chain-of-thought fine-tuning large model. The emotional traffic domain service voice generation module is used to adaptively generate and output emotional traffic domain service voices by using the chain-of-thought fine-tuning large model according to the text representation, the pinyin sequence, and the audio accent type.

[0013] The beneficial effects of the present invention are as follows: A traffic domain service voice adaptive generation method and system based on thought chain fine-tuning of a large model according to the present invention first uses a voice encoder to convert an input voice signal into a high-dimensional voice feature signal, and then uses a text decoder and a pinyin decoder to generate a text output signal based on the high-dimensional voice feature signal. Then, an accent type recognition decoder is used to identify the input voice signal and output the probability of the audio accent type. Subsequently, a multi-task voice recognition model is constructed based on the voice encoder, text decoder, pinyin decoder, and accent type recognition decoder. Then, the multi-task voice recognition model is used to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type. Then, a large model annotation specification is formulated, and a large model data set is constructed and annotated according to the annotation specification to obtain an annotated large model data set. Finally, the LoRA low-rank fine-tuning method and the mixed precision quantization method are used to perform lightweight training on the annotated large model data set to obtain a thought chain fine-tuned large model, and then the thought chain fine-tuned large model is used to adaptively generate and output emotional traffic domain service voices according to the text representation, pinyin sequence, and audio accent type; effectively realizing that the traffic domain service voice adaptive generation method and system based on thought chain fine-tuning of a large model have the functions of using a variational quantization joint data augmentation mechanism, a dual-channel multi-modal discriminator architecture, and a hierarchical composite loss function for high-fidelity emotional speech generation and robust synthesis in a complex noise environment, and simultaneously supporting semantic-driven dynamic prosody optimization and accurate pronunciation of professional terms. Moreover, in the application scenario of the traffic field, the multi-task voice recognition method can be used to efficiently link the character-level recognition, audio-to-pinyin, and sentence-level accent classification modules to effectively cope with the challenges of complex accents, many background noises, and polyphonic traffic terms. At the same time, by introducing a thought chain module on the basis of the original large model to simulate the human thinking process, the close connection between the conclusion description and the reasoning process can be achieved, thereby stimulating the logical reasoning ability and improving the reasoning accuracy. And by combining the LoRA low-rank fine-tuning method and the mixed precision quantization method, high accuracy and generalization ability can be maintained while quickly training with limited computing power, thereby improving the business intention recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is the overall flowchart of a traffic domain service voice adaptive generation method based on thought chain fine-tuning of a large model according to the present invention; Figure 2 is the adaptive generation principle diagram of a traffic domain service voice adaptive generation system based on thought chain fine-tuning of a large model according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0015] The present invention will be further described below in conjunction with the accompanying drawings of the specification.

[0016] Such as Figure 1As shown, a method for adaptively generating traffic domain service speech based on fine-tuning a large model with a chain of thought includes the following steps: Step A: Use a speech encoder to convert the input speech signal into a high-dimensional speech feature signal, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional speech feature signal. The specific steps are as follows: Step A1: Use a speech encoder to convert the input speech signal into a high-dimensional speech feature signal. The speech encoder is specifically a Conformer encoding layer. The Conformer encoding layer is used to construct local audio feature dependencies and combine the global feature construction ability of the self-attention mechanism to improve the performance of the automatic speech recognition task. The Conformer encoding layer is specifically composed of a first feed-forward module, a self-attention module, a convolutional module, a second feed-forward module, and a first regularization module connected in sequence. Step A2: Generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional speech feature signal. Both the text decoder and the pinyin decoder use a Transformer-ASR decoding layer. The Transformer-ASR decoding layer specifically uses a multi-head self-attention layer and uses a second regularization module, a linear output module, and a Softmax function module to generate the probability of the output text category according to the high-dimensional speech feature signal, thereby generating a text output signal.

[0017] Step B: Use an accent type recognition decoder to recognize the input speech signal and output the audio accent type probability. The accent type recognition decoder is specifically an S-vector decoding layer. The S-vector decoding layer specifically uses a feed-forward neural network to map the audio features of the input speech signal to a high-dimensional space and uses a state pooling layer to perform statistical pooling on the audio features of the input speech signal, thereby forming a sentence-level feature output. Then, it is mapped to an accent type database through two feed-forward neural networks, and then passes through a linear layer and a Softmax layer to output the audio accent type probability, thereby outputting the audio accent type probability.

[0018] Step C: Build a multi-task speech recognition model based on a speech encoder, a text decoder, a pinyin decoder, and an accent type recognition decoder, and then use the multi-task speech recognition model to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type. Specifically, the multi-task speech recognition model is built based on the speech encoder and in combination with the relevance of the speech recognition task, the audio-to-pinyin task, and the accent classification task. The multi-task speech recognition model specifically performs a relationship embedding calculation on the outputs of the text decoder and the pinyin decoder, fuses the calculated cross-modal features with the original text representation, and then generates the final text probability through a linear layer and a Softmax layer. The multi-task speech recognition model also includes a gradient reversal mechanism and a pinyin-text relationship model. The gradient reversal mechanism is used to promote the encoder to extract deep acoustic features independent of the accent through reverse optimization. The pinyin-text relationship model is used to correct semantic deviations using the acoustic alignment characteristics of the pinyin sequence and improve the adaptability to complex accent and noise scenarios.

[0019] Step D: Formulate a large model annotation specification, then build a large model dataset and annotate the large model dataset according to the annotation specification to obtain an annotated large model dataset. The large model annotation specification includes masked annotation and data text description annotation. The data text description annotation is formulated using the chain of thought method.

[0020] Step E: Use the LoRA low-rank fine-tuning method and the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a chain of thought fine-tuned large model. The specific steps are as follows: Step E1: Introduce the LoRA low-rank fine-tuning method to adapt the fully connected layer of the annotated large model dataset. Specifically, the LoRA low-rank fine-tuning method adds a side network to the large model dataset and updates the weights of the fully connected layer using the product of two rank decomposition matrices, and then updates the parameters in the side network. Step E2: Use the mixed-precision quantization method to perform lightweight training on the annotated large model dataset and obtain a chain of thought fine-tuned large model. Specifically, the mixed-precision quantization method uses a low-precision data storage type and performs de-quantization processing during the calculation process, and then combines BFloat16 for high-precision calculation and realizes the fine-tuning of the annotated large model dataset to obtain a chain of thought fine-tuned large model. The mixed-precision quantization method specifically includes the 4-bit NormalFloat quantization method and the double quantization method. The specific steps are as follows: Step E21: Model weight quantization and normalization processing. Specifically, the quantile interval is calculated using the characteristic that the weights follow a zero-centered normal distribution, and the quantization mapping table is determined. Specifically, when calculating the quantile interval, the weight tensor is normalized to the range [-1, 1], and then the discrete value is determined through the quantile quantization formula as shown in formula (1). (1) Among them, is a discrete value, is the i-th element value of the original weight, and are the weight mean and weight standard deviation respectively, and k is the quantization bit number; Step E22, double quantization and parameter compression, specifically introducing a two-level quantization mechanism for the quantized weight matrix. Among them, the first level performs 4-bit NormalFloat quantization on the original weight, and the second level performs 8-bit floating-point re-quantization on the quantization constant and realizes the secondary compression of quantization parameters, as shown in formula (2), (2) Among them, is the result after performing 4-bit NormalFloat quantization on the weight , is 4-bit NormalFloat quantization, is the quantization function, is the result after performing 8-bit floating-point quantization according to the intermediate calculation result , is 8-bit floating-point quantization; Step E23, paging optimization and video memory management, specifically constructing a paging optimizer based on NVIDIA unified memory to dynamically manage the GPU-CPU memory exchange. If the video memory occupancy exceeds the threshold T during the calculation process, the gradient checkpoint data will be automatically paged and transferred to the CPU memory, as shown in formula (3), (3) Among them, is the video memory, is the function to swap out the page in the memory to the disk and release the memory resources, is the data gradient; Step E24, low-rank adaptation and parameter fusion, specifically inserting a trainable low-rank adapter on the quantized base model to freeze the original quantized weight. The low-rank adapter consists of the decomposition matrix A and the decomposition matrix B, and the parameter fusion formula is as shown in formula (4), (4) Among them, is the parameter fusion result, is the quantized weight matrix, is the input vector, is the scaling factor, and are both additional weight matrices.

[0021] Step F: Use the chain-of-thought to fine-tune the large model to adaptively generate and output emotional voice for traffic domain services based on text representation, pinyin sequence, and audio accent type. The adaptive generation specifically adopts a three-level architecture of acoustic generation model - prosody predictor - neural vocoder and integrates a polyphone disambiguation model and a context-aware dynamic adjustment mechanism to achieve semantic-driven prosody adaptive generation while ensuring the accurate pronunciation of professional terms, thus meeting the high-quality speech synthesis requirements in real-time interaction and complex noise environments. The specific steps are as follows: Step F1: Construct a prosody predictor. The prosody predictor specifically constructs a continuous phoneme duration prediction model using a joint mechanism of variational dequantization and data augmentation, and then uses the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. The specific steps are as follows: Step F11: Construct a continuous phoneme duration prediction model using a joint mechanism of variational dequantization and data augmentation. The continuous phoneme duration prediction model specifically constructs a continuous latent space through two groups of variables with the same dimension as the phoneme sequence. Step F12: Use the continuous phoneme duration prediction model to deeply co-optimize prosody and emotion by combining the reversible flow generation method and the dynamic alignment method. Specifically, based on the variational lower bound optimization objective, posterior distribution sampling is used for training, and the gradient backpropagation of the random duration predictor is blocked synchronously to maintain the independence of the module. Then, the reversible flow model generates continuous duration values from noise and outputs a phoneme alignment sequence that conforms to the rhythm of human speech after integer conversion. Step F2: Construct an acoustic generation model. The acoustic generation model specifically uses the feature enhancement layer of a dual-channel multi-modal discriminator to extract an emotional acoustic feature matrix, and then establishes a dual supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. The specific steps are as follows: Step F21: Use the feature enhancement layer of a dual-channel multi-modal discriminator to extract an emotional acoustic feature matrix. The dual-channel multi-modal discriminator specifically adopts a dual-channel architecture including the DiscriminatorS network and the DiscriminatorP network to re-extract the Mel Frequency Cepstral Coefficients (MFCC) parameters of the generated spectrum and the real spectrum, and then analyzes the emotional information label of the generated spectrum through an acoustic feature analysis module based on a Bidirectional Long Short-Term Memory (Bi-LSTM) network. Then, the emotional information label and the original emotional annotation jointly form an auxiliary feature matrix, and the auxiliary feature matrix is injected into the convolutional processing channel of the dual-channel multi-modal discriminator in parallel to obtain the emotional acoustic feature matrix. Step F22: Establish a dual-supervision mechanism of spectral structure and semantic emotion based on the emotional acoustic feature matrix to improve the consistency of emotional expression. Specifically, extract the visual features of the emotional acoustic feature matrix through a conventional convolutional layer and synchronously analyze the implicit emotional associations in the acoustic feature space. Then, maintain the basic spectral feature map and the enhanced emotional feature map, and judge the consistency of emotional expression through cross-modal feature cross-verification; Step F3: Construct a neural vocoder. The neural vocoder specifically uses a hierarchical composite loss function system to fuse spectral reconstruction, feature space constraint, and dual emotional supervision, and combines the environmental noise suppression method to enhance the robustness of the vocoder. The specific steps are as follows: Step F31: Construct a hierarchical composite loss function system. The hierarchical composite loss function system uses a multi-component composite loss function to perform multi-objective optimization on speech synthesis. The multi-component composite loss function includes a basic loss layer and a feature enhancement layer. The basic loss layer is specifically composed of a reconstruction loss, an adversarial training stability control KL divergence loss, and a Wasserstein distance optimization adversarial loss. The feature enhancement layer is specifically composed of a Mel-spectrum perceptual loss and a feature map contrast loss. The Mel-spectrum perceptual loss specifically uses Mel-scale frequency domain conversion to establish acoustic physical constraints. The feature map contrast loss specifically uses the multi-level feature maps output by the discriminator to perform cross-network layer feature matching; Step F32: Construct a dual-supervision mechanism for emotional vectors. The dual-supervision mechanism for emotional vectors specifically implements dual emotional constraints using the distance loss in the emotional label embedding space and the emotional feature component loss in the discriminator feature map.

[0022] Such as Figure 2As shown in the figure, a traffic domain service voice adaptive generation system based on chain-of-thought fine-tuning of a large model includes a text output signal generation module, an audio accent type probability recognition module, a multi-task speech recognition module, a large model annotation specification formulation module, a chain-of-thought fine-tuning large model establishment module, and an emotional traffic domain service voice generation module. The text output signal generation module is used to convert the input speech signal into a high-dimensional speech feature signal using a speech encoder, and then generate a text output signal through a text decoder and a pinyin decoder according to the high-dimensional speech feature signal. The audio accent type probability recognition module is used to recognize the input speech signal using an accent type recognition decoder and output the audio accent type probability. The multi-task speech recognition module is used to construct a multi-task speech recognition model based on the speech encoder, text decoder, pinyin decoder, and accent type recognition decoder, and then further process the text output signal and the audio accent type probability using the multi-task speech recognition model to obtain a text representation, a pinyin sequence, and an audio accent type. The large model annotation specification formulation module is used to formulate large model annotation specifications, then construct a large model data set and annotate the large model data set according to the annotation specifications to obtain an annotated large model data set. The chain-of-thought fine-tuning large model establishment module is used to perform lightweight training on the annotated large model data set using the LoRA low-rank fine-tuning method and the mixed-precision quantization method to obtain a chain-of-thought fine-tuning large model. The emotional traffic domain service voice generation module is used to adaptively generate and output emotional traffic domain service voices using the chain-of-thought fine-tuning large model according to the text representation, pinyin sequence, and audio accent type.

[0023] In summary, for the method and system for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model of the present invention, first, a voice encoder is used to convert the input voice signal into a high-dimensional voice feature signal, then a text decoder and a pinyin decoder are used to generate a text output signal according to the high-dimensional voice feature signal. Next, an accent type recognition decoder is used to identify the input voice signal and output the probability of the audio accent type. Subsequently, a multi-task voice recognition model is constructed based on the voice encoder, text decoder, pinyin decoder, and accent type recognition decoder. Then, the multi-task voice recognition model is used to further process the text output signal and the audio accent type probability to obtain a text representation, a pinyin sequence, and an audio accent type. Then, a large model annotation specification is formulated, and a large model dataset is constructed and annotated according to the annotation specification to obtain an annotated large model dataset. Finally, the LoRA low-rank fine-tuning method and the mixed-precision quantization method are used to perform lightweight training on the annotated large model dataset to obtain a chain-of-thought fine-tuned large model, and the chain-of-thought fine-tuned large model is used to adaptively generate and output emotional traffic domain service voice according to the text representation, pinyin sequence, and audio accent type; effectively realizing that the method and system for adaptively generating traffic domain service voice based on chain-of-thought fine-tuning of a large model have the functions of using a variational dequantization joint data augmentation mechanism, a dual-channel multi-modal discriminator architecture, and a hierarchical composite loss function for high-fidelity emotional voice generation and robust synthesis in a complex noise environment, and simultaneously supporting semantic-driven dynamic prosody optimization and accurate pronunciation of professional terms. Moreover, in the application scenario of the traffic field, the multi-task voice recognition method can be used to achieve the efficient linkage of character-level recognition, audio-to-pinyin, and sentence-level accent classification modules, so as to effectively cope with the challenges of complex accents, a large number of background noises, and polyphonic traffic terms. At the same time, by introducing a chain-of-thought module on the basis of the original large model to simulate the human thinking process, the conclusion description and the reasoning process can be closely connected, thereby stimulating logical reasoning ability and improving the reasoning accuracy. And by combining the LoRA low-rank fine-tuning method and the mixed-precision quantization method, high accuracy and generalization ability can be maintained while quickly training with limited computing power, thereby improving the business intention recognition effect.

[0024] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for adaptively generating service speech in the traffic domain based on a thought chain fine-tuning large model, characterized by: The following steps are included: Step A, using a speech encoder to convert an input speech signal into a high-dimensional speech feature signal, and then using a text decoder and a pinyin decoder to generate a text output signal based on the high-dimensional speech feature signal; Step B, using an accent type recognition decoder to recognize the input speech signal and output the audio accent type probability; Step C, constructing a multi-task speech recognition model based on the speech encoder, the text decoder, the pinyin decoder and the accent type recognition decoder, and then using the multi-task speech recognition model to further process the text output signal and the audio accent type probability and obtain the text representation, the pinyin sequence and the audio accent type; Step D, formulate a large model annotation specification, then construct a large model data set and annotate the large model data set according to the annotation specification to obtain an annotated large model data set; Step E, using the LoRA low-rank fine-tuning method and the mixed precision quantization method to perform lightweight training on the labeled large model data set and obtain the thinking chain fine-tuning large model; Step F, use the thinking chain to fine-tune the large model to adaptively generate and output emotional traffic domain service speech based on text representation, pinyin sequence and audio accent type.

2. According to claim 1, a method for adaptively generating traffic domain service speech based on a thought chain fine-tuning large model is characterized by: Step A, using a speech encoder to convert the input speech signal into a high-dimensional speech feature signal, and then using a text decoder and a pinyin decoder to generate a text output signal according to the high-dimensional speech feature signal. The specific steps are as follows: Step A1, using a speech encoder to convert an input speech signal into a high-dimensional speech feature signal, wherein the speech encoder is specifically a Conformer encoding layer, the Conformer encoding layer is used to construct local audio feature dependencies and combine the global feature construction capability of the self-attention mechanism to improve the performance of the automatic speech recognition task, and the Conformer encoding layer is specifically composed of a first feedforward module, a self-attention module, a convolution module, a second feedforward module and a first layer regularization module connected in sequence; Step A2, generating a text output signal according to the high-dimensional speech feature signal through a text decoder and a pinyin decoder, wherein the text decoder and the pinyin decoder both adopt a Transformer-ASR decoding layer, and the Transformer-ASR decoding layer specifically adopts a multi-head self-attention layer and uses a second-layer regularization module, a linear output module and a Softmax function module to generate an output text category probability according to the high-dimensional speech feature signal, thereby generating a text output signal.

3. According to claim 2, a method for adaptively generating traffic domain service speech based on a thought chain fine-tuning large model is characterized by: Step B, using an accent type recognition decoder to recognize the input speech signal and output the audio accent type probability, wherein the accent type recognition decoder is specifically an S-vector decoding layer, and the S-vector decoding layer specifically uses a feedforward neural network to map the audio features of the input speech signal to a high-dimensional space and uses a state pooling layer to perform statistical pooling on the audio features of the input speech signal, thereby forming a sentence-level feature output, which is then mapped to an accent type database through two feedforward neural networks, and then outputs the audio accent type probability through a linear layer and a Softmax layer, thereby outputting the audio accent type probability.

4. According to claim 3, a method for adaptively generating traffic domain service speech based on a thought chain fine-tuning large model is characterized by: Step C, constructing a multi-task speech recognition model based on the speech encoder, text decoder, pinyin decoder and accent type recognition decoder, and then using the multi-task speech recognition model to further process the text output signal and the audio accent type probability and obtain text representation, pinyin sequence and audio accent type, wherein the multi-task speech recognition model is specifically constructed based on the speech encoder and in combination with the correlation of the speech recognition task, the audio to pinyin task and the accent classification task, the multi-task speech recognition model specifically performs relational embedding calculation on the outputs of the text decoder and the pinyin decoder and fuses the calculated cross-modal features with the original text representation to generate the final text probability through a linear layer and a Softmax layer, the multi-task speech recognition model also includes a gradient reversal mechanism and a pinyin-text relationship model, the gradient reversal mechanism is used to use reverse optimization to promote the encoder to extract deep acoustic features that are not related to accents, and the pinyin-text relationship model is used to use the acoustic alignment characteristics of the pinyin sequence to correct semantic deviations and improve the adaptability to complex accents and noise scenes.

5. According to claim 4, a method for adaptively generating traffic domain service speech based on a thought chain fine-tuning large model is characterized by: Step D, formulate a large model annotation specification, then construct a large model data set and annotate the large model data set according to the annotation specification to obtain an annotated large model data set, wherein the large model annotation specification includes mask annotation and data text description annotation, and the data text description annotation is formulated using a thinking chain method.

6. The method for adaptively generating traffic domain service speech based on the thought chain fine-tuning large model according to claim 5 is characterized by: Step E: Use the LoRA low-rank fine-tuning method and mixed precision quantization method to perform lightweight training on the labeled large model data set and obtain the thinking chain fine-tuning large model. The specific steps are as follows: Step E1, introducing the LoRA low-rank fine-tuning method to adapt the fully connected layer of the labeled large model data set, wherein the LoRA low-rank fine-tuning method specifically adds a side network to the large model data set and uses the product of two rank decomposition matrices to update the weight of the fully connected layer, and then updates the parameters in the side network; Step E2, use the mixed precision quantization method to perform lightweight training on the labeled large model data set and obtain the thinking chain fine-tuning large model, wherein the mixed precision quantization method specifically uses low-precision storage data types and performs dequantization processing during the calculation process, and then combines BFloat16 for high-precision calculation and realizes fine-tuning of the labeled large model data set to obtain the thinking chain fine-tuning large model. The mixed precision quantization method specifically includes 4-bit NormalFloat quantization method and double quantization method. The specific steps are as follows: Step E21, model weight quantization and normalization processing, specifically, using the characteristic that the weight obeys the zero-centered normal distribution to calculate the quantile interval and determine the quantization mapping table, wherein the calculation of the quantile interval is specifically to normalize the weight tensor to the range of [-1, 1], and then determine the discrete value through the quantile quantization formula as shown in formula (1), (1) in, is a discrete value, is the i-th element value of the original weight, and are weight mean and weight standard deviation respectively, and k is the number of quantization bits; Step E22, double quantization and parameter compression, specifically, a two-level quantization mechanism is introduced into the quantized weight matrix, wherein the first level performs 4-bit NormalFloat quantization on the original weight, and the second level uses 8-bit floating point to re-quantize the quantization constant and realizes secondary compression of the quantization parameter, as shown in formula (2). (2) in, For weight The result after 4-bit NormalFloat quantization is: Quantized to 4-bitNormalFloat, is the quantization function, Based on the intermediate calculation results The result after 8-bit floating point quantization is: 8-bit floating point quantization; Step E23, paging optimization and video memory management, specifically, building a paging optimizer based on NVIDIA unified memory to dynamically manage GPU-CPU memory exchange. If the video memory usage exceeds the threshold T during the calculation process, the gradient checkpoint data is automatically paged and transferred to the CPU memory, as shown in formula (3). (3) in, For video memory, A function that swaps out pages in memory to disk and releases memory resources. is the data gradient; Step E24, low-rank adaptation and parameter fusion, specifically inserting a trainable low-rank adapter into the quantized base model to freeze the original quantized weights. The low-rank adapter is composed of a decomposition matrix A and a decomposition matrix B, and the parameter fusion formula is shown in formula (4). (4) in, is the parameter fusion result, The quantized weight matrix, is the input vector, is the scaling factor, and are all additional weight matrices.

7. The method for adaptively generating traffic domain service speech based on the thought chain fine-tuning large model according to claim 6 is characterized by: Step F, use the thinking chain to fine-tune the large model to adaptively generate and output emotional traffic domain service speech according to text representation, pinyin sequence and audio accent type. The adaptive generation specifically adopts the three-level architecture of acoustic generation model-prosody predictor-neural vocoder and integrates the polyphone disambiguation model and context-aware dynamic adjustment mechanism to achieve semantic-driven prosody adaptive generation while ensuring the accurate pronunciation of professional terms, thereby meeting the needs of high-quality speech synthesis in real-time interaction and complex noise environments. The specific steps are as follows: Step F1, construct a prosody predictor, which specifically adopts a combination of variational dequantization and data enhancement mechanism to construct a continuous phoneme duration prediction model, and then uses the continuous phoneme duration prediction model combined with a reversible flow generation method and a dynamic alignment method to perform deep collaborative optimization of prosody and emotion. The specific steps are as follows: Step F11, using a combination of variational dequantization and data enhancement mechanism to construct a continuous phoneme duration prediction model, wherein the continuous phoneme duration prediction model specifically constructs a continuous latent space through two groups of variables with the same dimensions as the phoneme sequence; Step F12, using the continuous phoneme duration prediction model combined with the reversible flow generation method and the dynamic alignment method to perform deep collaborative optimization of rhythm and emotion, specifically, based on the variational lower bound optimization target, adopting posterior distribution sampling training and synchronously blocking the gradient back propagation of the random duration predictor to maintain the independence of the module, and then generating continuous duration values ​​from the noise through the reversible flow model and outputting a phoneme alignment sequence that conforms to the rhythm of human speech after integer conversion; Step F2, constructing an acoustic synthetic model, the acoustic synthetic model specifically uses the feature enhancement layer of a dual-channel multimodal discriminator to extract the emotional acoustic feature matrix, and then establishes a dual supervision mechanism of spectral structure and semantic emotion according to the emotional acoustic feature matrix to improve the consistency of emotional expression. The specific steps are as follows: Step F21, using the feature enhancement layer of the dual-channel multimodal discriminator to extract the emotional acoustic feature matrix, wherein the dual-channel multimodal discriminator specifically uses a dual-channel architecture including a DiscriminatorS network and a DiscriminatorP network to re-extract Mel-frequency cepstral coefficient MFCC parameters from the generated spectrum and the real spectrum, and then parses the emotional information label of the generated spectrum through an acoustic feature analysis module based on a bidirectional long short-term memory network Bi-LSTM, and then the emotional information label and the original emotional annotation are jointly formed into an auxiliary feature matrix, and then the auxiliary feature matrix is ​​injected into the convolution processing channel of the dual-channel multimodal discriminator in parallel to obtain the emotional acoustic feature matrix; Step F22, based on the emotional acoustic feature matrix, a dual supervision mechanism of spectral structure and semantic emotion is established to improve the consistency of emotional expression. Specifically, the visual features of the emotional acoustic feature matrix are extracted through a conventional convolutional layer and the implicit emotional association in the acoustic feature space is simultaneously analyzed. Then, the basic spectral feature map and the enhanced emotional feature map are maintained and the emotional expression consistency is judged through cross-modal feature cross-validation. Step F3, constructing a neural vocoder, the neural vocoder specifically adopts a hierarchical composite loss function system to fuse spectrum reconstruction, feature space constraints and emotional dual supervision and combines environmental noise suppression method to enhance the robustness of the vocoder. The specific steps are as follows: Step F31, constructing a hierarchical composite loss function system, wherein the hierarchical composite loss function system uses a multivariate composite loss function to perform multi-objective optimization on speech synthesis, wherein the multivariate composite loss function includes a basic loss layer and a feature enhancement layer, wherein the basic loss layer is specifically composed of a reconstruction loss, a KL divergence loss for adversarial training stability control, and a Wasserstein distance optimization adversarial loss, wherein the feature enhancement layer is specifically composed of a Mel spectrum perception loss and a feature spectrum contrast loss, wherein the Mel spectrum perception loss specifically uses a Mel scale frequency domain conversion to establish acoustic physical constraints, and the feature spectrum contrast loss specifically uses a multi-level feature map output by a discriminator to perform cross-network layer feature matching; Step F32, constructing a dual supervision mechanism for emotion vectors, wherein the dual supervision mechanism for emotion vectors specifically implements dual emotion constraints by using the distance loss in the emotion tag embedding space and the emotion feature component loss in the discriminator feature map.

8. A traffic domain service speech adaptive generation system based on a thought chain fine-tuning large model, wherein the specific generation process of the traffic domain service speech adaptive generation system is based on the traffic domain service speech adaptive generation method according to any one of claims 1 to 7, characterized in that: It includes a text output signal generation module, an audio accent type probability recognition module, a multi-task speech recognition module, a large model annotation specification formulation module, a thinking chain fine-tuning large model establishment module and an emotional traffic domain service speech generation module. The text output signal generation module is used to convert the input speech signal into a high-dimensional speech feature signal by using a speech encoder, and then generate a text output signal according to the high-dimensional speech feature signal through a text decoder and a pinyin decoder; The audio accent type probability recognition module is used to recognize the input speech signal using an accent type recognition decoder and output the audio accent type probability; The multi-task speech recognition module is used to construct a multi-task speech recognition model based on a speech encoder, a text decoder, a pinyin decoder and an accent type recognition decoder, and then use the multi-task speech recognition model to further process the text output signal and the audio accent type probability and obtain a text representation, a pinyin sequence and an audio accent type; The large model annotation specification formulation module is used to formulate the large model annotation specification, then construct the large model data set and annotate the large model data set according to the annotation specification to obtain the annotated large model data set; The thinking chain fine-tuning large model building module is used to use the LoRA low-rank fine-tuning method and the mixed precision quantization method to perform lightweight training on the labeled large model data set and obtain the thinking chain fine-tuning large model; The emotional traffic domain service speech generation module is used to adopt the thought chain fine-tuning large model to adaptively generate and output emotional traffic domain service speech according to text representation, pinyin sequence and audio accent type.

Citation Information

Patent Citations

  • Semi-structured interview system based on fine-tuning large language model

    CN118886503A

  • Speech synthesis method and device based on hierarchical emotion distribution, equipment and medium

    CN119207372A

  • End-cloud collaborative vehicle-mounted interaction method for multiple task scenes

    CN119207408A

  • Speech recognition model training method and device and speech recognition method and device

    CN119851658A

  • Low-resource language adaptive speech recognition method based on AdLoRA Plus

    CN119943052A

Cited By

  • Aviation speech generation method and system based on generative adversarial network

    CN120913539A

  • An aviation speech generation method and system based on a generative adversarial network

    CN120913539B

  • Thinking chain and thinking mode auxiliary voice generation method and device, equipment and medium

    CN120977289A

  • Thought chain and thought mode auxiliary speech generation method and device, equipment and medium

    CN120977289B

  • High-precision voice recognition and safety monitoring system and method for electric power operation

    CN121331111A