Voice interaction method and device based on thinking chain, computer equipment and medium

By generating an initial text thought chain using a large language model and allowing users to edit it, the problems of accumulated errors and semantic biases in traditional voice dialogue systems are solved, achieving high-quality, personalized voice responses.

CN121884792APending Publication Date: 2026-04-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-14
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional voice dialogue systems are prone to introducing cumulative errors during information transmission, leading to semantic deviations and the loss of paralinguistic information such as speaker identity, emotion, and tone. The generated response voices lack naturalness and personalization.

Method used

The system directly receives speech feature vectors from a large language model to generate an initial text thought chain. The user-editable thought chain is then corrected to generate the target text thought chain, and finally, the response speech is synthesized. This avoids the cumulative errors of traditional serial architectures and enables personalized customization.

Benefits of technology

It improves the semantic accuracy and naturalness of the response voice, and allows users to modify the content and style during the voice generation process to achieve personalized customization and improve the quality of the response voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884792A_ABST
    Figure CN121884792A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice generation, and particularly discloses a voice interaction method and device based on a thinking chain, computer equipment and a medium. The voice feature vectors are directly received through the large language model to generate the thinking chain, accumulative errors caused by a traditional series framework are avoided, semantic accuracy is improved, reply voice quality is improved, the initial text thinking chain capable of being edited by the user is generated through the large language model, and user experience is improved. And the initial text thinking chain is corrected according to the user input content to obtain the target text thinking chain, so that the user can modify the voice in the voice generation process, personalized customization is realized, and the reply voice quality is further improved. The method is applied to a voice interaction system of financial or medical services such as voice product explanation, remote medical consultation, medical education and training and the like, the text content, the voice emotion and the style modified by the user can be obtained through the text thinking chain, voice editing is achieved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech generation technology, and in particular to a speech interaction method, device, computer equipment and medium based on thought chain. Background Technology

[0002] With the development of artificial intelligence technology, voice interaction systems are widely used in financial and medical services such as voice product explanation, remote medical consultation, and medical education and training. Traditional voice dialogue systems employ a serial architecture of automatic speech recognition, natural language understanding, text generation, and text-to-speech. Information transmission between modules is prone to introducing cumulative errors, leading to semantic deviations. Furthermore, this architecture loses paralinguistic information such as speaker identity, emotion, tone, and rhythm during the conversion process, resulting in unnatural and personalized responses. Therefore, how to avoid the loss of linguistic information, reduce semantic deviations, and improve the quality of response speech has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides a voice interaction method, device, computer equipment, and medium based on thought chain, in order to avoid the loss of language information, reduce semantic deviation, and improve the quality of response voice.

[0004] Firstly, this application provides a voice interaction method based on thought chain, the method comprising: Upon receiving input speech, the input speech is processed to obtain the target speech feature vector; Autoregressive processing is performed on the target speech feature vector based on a large language model to generate an initial text thought chain. Receive user input based on the initial text thought chain, and update the initial text thought chain based on the input to obtain the target text thought chain; Speech synthesis is performed based on the target text's thought chain to generate a response speech corresponding to the input speech.

[0005] Secondly, this application also provides a voice interaction device based on a thought chain, the device comprising: The feature vector acquisition module is used to process the input speech upon receiving it to obtain the target speech feature vector. The thought chain generation module is used to perform autoregressive processing on the target speech feature vector based on a large language model to generate an initial text thought chain. The thought chain correction module is used to receive input content from the user based on the initial text thought chain, and update the initial text thought chain based on the input content to obtain the target text thought chain. The response speech generation module is used to perform speech synthesis based on the target text thought chain to generate a response speech corresponding to the input speech.

[0006] Thirdly, this application also provides a computer device, the computer device including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the thought chain-based voice interaction method as described above.

[0007] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the thought chain-based voice interaction method described above.

[0008] This application discloses a voice interaction method, apparatus, computer device, and medium based on a thought chain. Upon receiving input voice, the method processes the input voice to obtain a target voice feature vector; performs autoregressive processing on the target voice feature vector based on a large language model to generate an initial text thought chain; receives user input based on the initial text thought chain and updates the initial text thought chain based on the input to obtain a target text thought chain; and performs speech synthesis based on the target text thought chain to generate a response voice corresponding to the input voice. This application directly receives voice feature vectors from a large language model to generate a thought chain, avoiding the cumulative errors caused by traditional concatenated architectures, improving semantic accuracy, and enhancing the quality of the response voice. Furthermore, by generating an editable initial text thought chain through a large language model, and correcting the initial text thought chain based on user input to obtain the target text thought chain, and generating a response voice based on the target text thought chain, the method allows users to modify the voice during the voice generation process, achieving personalized customization and further improving the quality of the response voice. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a first schematic flowchart of a voice interaction method based on thought chain provided by an embodiment of this application; Figure 2 This is a second schematic flowchart of a voice interaction method based on thought chain provided by an embodiment of this application; Figure 3This is a third schematic flowchart of a voice interaction method based on thought chain provided by an embodiment of this application; Figure 4 A schematic block diagram of a thought chain-based voice interaction device provided for embodiments of this application; Figure 5 A schematic block diagram of the structure of a computer device provided for an embodiment of this application. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0013] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0014] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0015] This application provides a thought chain-based voice interaction method, apparatus, computer device, and medium. The thought chain-based voice interaction method can be applied to voice interaction systems and servers. It directly receives speech feature vectors through a large language model to generate thought chains, avoiding the cumulative errors caused by traditional serial architectures, improving semantic accuracy, and enhancing the quality of response speech. Furthermore, it generates an initial text thought chain that can be edited by the user through the large language model, corrects the initial text thought chain based on user input to obtain a target text thought chain, and generates response speech based on the target text thought chain. This allows users to modify the speech during the speech generation process, achieving personalized customization and further improving the quality of response speech. The server can be a standalone server or a server cluster. The voice interaction system includes: a voice acquisition module for acquiring input voice; a voice encoder module for encoding the input voice to obtain an initial voice feature vector; a linear projection module for mapping the initial voice feature vector according to a preset feature dimension to obtain a target voice feature vector; a large language model for processing the target voice feature vector to generate an initial text thought chain; an interactive editing interface for receiving user modification instructions for the initial text thought chain to update the initial text thought chain and generate a target text thought chain; an acoustic projection network module for generating a voice generation condition vector based on the target text thought chain and generating a Mel spectrogram based on the voice generation condition vector; and a vocoder for generating a response voice corresponding to the input voice based on the Mel spectrogram.

[0016] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0017] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating a thought-chain-based voice interaction method provided in an embodiment of this application. This thought-chain-based voice interaction method can be applied to voice interaction systems and servers. It generates thought chains by directly receiving speech feature vectors through a large language model, avoiding the cumulative errors caused by traditional concatenated architectures, improving semantic accuracy, and enhancing the quality of the response speech. Furthermore, it generates an initial text thought chain that can be edited by the user through the large language model, corrects the initial text thought chain based on user input to obtain the target text thought chain, and generates the response speech based on the target text thought chain. This allows users to modify the speech during the speech generation process, achieving personalized customization and further improving the quality of the response speech.

[0018] like Figure 1 As shown, the voice interaction method based on thought chain specifically includes steps S101 to S104.

[0019] S101. Upon receiving input speech, the input speech is processed to obtain the target speech feature vector; In one embodiment, for a voice interaction system used in services such as voice product explanation and telemedicine, the system receives user input voice through a voice acquisition module built into the system or connected to the system, converts the input voice into a spectrogram, then performs high-level feature extraction on the spectrogram through a voice encoder to obtain a voice feature vector containing semantic and acoustic features, and finally maps the voice feature vector to a preset spatial dimension through a projection layer to obtain the target voice feature vector.

[0020] Further, the step of processing the input speech to obtain the target speech feature vector includes: performing a spectrum transformation on the input speech to obtain a second Mel spectrogram corresponding to the input speech; transforming the second Mel spectrogram based on the speech encoder to obtain an initial speech feature vector; and mapping the initial speech feature vector according to a preset feature dimension to obtain the target speech feature vector.

[0021] In one embodiment, the input speech is first converted into a second Mel spectrogram through a short-time Fourier transform and Mel filtering process. Specifically, the speech waveform corresponding to the input speech is divided into frames of a preset length, a Hamming window is used to reduce spectral leakage, a short-time Fourier transform is performed on each frame of speech to obtain a linear spectrum, the amplitude value is taken to obtain the power spectrum, and the linear power spectrum is mapped to the Mel frequency domain through a Mel filter bank (composed of multiple triangular filters) to generate the second Mel spectrogram.

[0022] In one embodiment, the speech encoder can be a Conformer encoder, which combines the self-attention mechanism of the Transformer with the local feature extraction capabilities of a CNN (Convolutional Neural Network). Specifically, a multi-head self-attention layer is used to capture long-range dependencies in speech, and a depthwise separable convolution module is used to extract local acoustic features (such as phonemes and prosody). The features are then nonlinearly transformed through a feedforward network layer to obtain an initial speech feature vector.

[0023] In one embodiment, an initial speech feature vector is mapped to a preset feature dimension using a linear projection layer, achieving alignment between speech features and text features. The linear projection layer is a fully connected layer, and the mapping formula is:

[0024] Where P is the linear projection layer (linear transformation matrix). For the target speech feature vector, This is the initial speech feature vector. This is the second Mel spectrum.

[0025] In this embodiment, the preset feature dimension is the feature dimension of the large language model in subsequent processing.

[0026] Furthermore, before converting the second Mel spectrogram based on the speech encoder to obtain the initial speech feature vector, the method further includes: acquiring a training dataset and a joint loss function, wherein the training dataset includes the spectrogram corresponding to the training audio and the text corresponding to the training audio, and the training audio consists of a pair of question and answer speech; and jointly training the pre-trained speech encoder and the pre-trained large language model based on the training dataset and the joint loss function to obtain the speech encoder and the large language model.

[0027] In one embodiment, the training audio is a pair of consecutive question and answer voices, such as a user asking "How is the weather today?" and the answer being "Sunny turning cloudy today".

[0028] The training audio-corresponding text refers to the transcribed text of the question and the transcribed text of the answer.

[0029] The spectrogram corresponding to the training audio refers to the Mel spectrogram generated by the training audio after short-time Fourier transform and Mel filtering.

[0030] In one embodiment, the spectrogram of the training audio is randomly divided into prompt sections at a time point s (typically near the boundary between the question and answer) by randomly selecting a time point s. : The spectrogram corresponding to the question's audio; Continued writing section The spectrogram of the answer speech. Similarly, the text corresponding to the training speech is segmented into transcribed texts of the question speech. and the transcribed text of the response speech .

[0031] In one embodiment, the joint loss function incorporates cross-entropy loss. and speech reconstruction loss :

[0032] in, These are weight parameters used to avoid overly favoring text generation and neglecting speech quality during training. They can be set according to actual needs and adjusted based on the training status.

[0033] In one embodiment, cross-entropy loss includes supervised speech recognition loss. With response text loss Cross-loss function .in, For large language models based on prompts The transcribed text obtained from the corresponding feature vector analysis, The response text obtained from reasoning using a large language model.

[0034] In one embodiment, the reconstruction loss includes the basic loss:

[0035] Characteristic derivative loss:

[0036] Time derivative loss:

[0037] Reconstruction loss function .

[0038] In one embodiment, the large language model computes all outputs via forward propagation:

[0039] in, LM represents the response speech feature vector obtained from inference using a large language model.

[0040] During training, control whether the labels are automatically extracted from the training data (e.g., when using a sentiment classification model) or used as additional labeled input.

[0041] Specifically, the pre-trained speech encoder analyzes the cue spectrogram. Encode to generate initial audio feature vectors This is then mapped to the embedding space of a large language model to obtain the target audio feature vector. Large language model reception As a prefix, autoregressive generation of suggestion spectrum is performed. Corresponding transcribed text It then performs contextual reasoning based on the target speech feature vector and generates autoregressive response text content. , will reply text The text is converted into a text embedding vector, and then the text embedding vector is mapped to the speech feature space to obtain the response speech feature vector. The loss is calculated based on the joint loss function. When the loss value exceeds a preset loss threshold, the total loss is backpropagated to the speech encoder and the large language model. The optimizer then synchronously updates the parameters of the pre-trained speech encoder and the pre-trained large language model. The next round of iterative training is then performed based on the updated speech encoder and the large language model until the total loss value is less than or equal to the preset loss threshold. The preset loss threshold can be set according to actual needs.

[0042] In the above embodiments, the speech encoder and the large language model are deeply bound together through joint training, which solves the feature loss problem of traditional cascaded systems, improves the performance of the speech encoder and the large language model, and thus improves the quality of the generated response speech.

[0043] S102. Based on the large language model, perform autoregressive processing on the target speech feature vector to generate an initial text thought chain. In one embodiment, the large language model receives the target speech feature vector. As a prefix, autoregression generates a structured initial textual thought chain.

[0044] The initial text thought chain includes the descriptive text content corresponding to the input speech, the response text content obtained by reasoning based on the input speech, and the acoustic control tags (such as emotion tags, speech rate tags, stress tags, etc.) for generating synthesized speech.

[0045] In a specific embodiment, the large language model analyzes the target speech feature vector, generates a descriptive text corresponding to the input speech, and performs inference based on the target speech feature vector and the descriptive text to obtain the response text of the input speech. Then, it performs control label prediction based on the descriptive text and the response text to obtain acoustic control labels.

[0046] In one embodiment, when the large language model analyzes and infers the target speech feature vector and descriptive text, it can obtain inference tasks, such as continuation writing or question-and-answer. For continuation writing tasks, the generated response text is the continuation content; for question-and-answer tasks, the generated response text is the answer to the user's question.

[0047] S103. Receive the user's input content based on the initial text thinking chain, and update the initial text thinking chain based on the input content to obtain the target text thinking chain. In one embodiment, the voice interaction system further includes an interactive editing interface connected to the user device. After generating an initial text thought chain, the initial text thought chain is displayed on the interactive editing interface of the user device, and the user can edit the initial text thought chain (such as modifying some continuation content or adjusting emotion tags) or confirm it.

[0048] The edited text thought chain is used as the target text thought chain. If the user does not need to edit (i.e., the user confirms the instruction), the initial text thought chain is used as the target text thought chain.

[0049] Furthermore, the step of receiving user input content based on the initial text thought chain and updating the initial text thought chain based on the input content to obtain the target text thought chain includes: receiving the input content based on an interactive editing interface, parsing the input content to obtain the tag field to be modified and the tag modification value; and correcting the acoustic control tag in the initial text thought chain based on the tag field to be modified and the tag modification value to obtain the target text thought chain.

[0050] In one embodiment, a user selects a tag field to be modified via a menu / button in an interactive editing interface, and selects or enters the tag value corresponding to that tag field. After the user submits the operation, the interactive editing interface receives the user's input, parses the input, and determines the tag field to be modified and the modified value. Based on the tag field to be modified and the modified value, the initial text thought chain is corrected, including: retaining the unmodified tag content and replacing the user-modified tag value.

[0051] In another embodiment, if the reply text is a continuation text, it can also receive text modifications made by the user to the reply text in the text thought chain, and correct the text based on these modifications to obtain the target text thought chain. It is understood that when the user edits the continuation text, they can also edit the acoustic control tags.

[0052] For example, when a doctor generates medical orders for a child patient through a remote consultation platform, it is necessary to ensure that the dosage is accurate, the tone of voice, and the content of the orders are suitable for the child's family to understand. The doctor's voice input is "Patient Li Si (child), amoxicillin 0.25g, three times a day, after meals, pay attention to allergic reactions." The doctor's voice input is encoded by an encoder and mapped to a speech feature vector in the spatial dimension of a large language model. The large language model processes the speech feature vector to generate an initial text thought chain. The doctor can modify the initial text thought chain through an interactive editing interface, such as changing the emotion label to gentle, the speech rate label to slow, and can also modify the content of the medical order text generated by the large language model to make the subsequently generated voice medical orders suitable for the patient's family to understand.

[0053] In the above embodiments, the editable and interactive text thought chain allows users to correct and adjust during the speech generation process, achieving fine control over the speech content, emotion, and style. This provides speech editing solutions for fields such as speech content creation, personalized digital humans, and intelligent dubbing, overcoming the problem of rigid speech generated by traditional speech generation models.

[0054] S104. Based on the target text thought chain, perform speech synthesis to generate a response speech corresponding to the input speech.

[0055] In one embodiment, the voice interaction system further includes an acoustic projection network module for generating a Mel spectrogram based on the target text thought chain. Specifically, the acoustic projection network module encodes acoustic control tags in the target text thought chain to obtain tag embedding vectors, then fuses these tag embedding vectors with the target speech feature vectors, and drives acoustic generation based on the fused condition vector to obtain the Mel spectrogram.

[0056] In one embodiment, the response speech is obtained by processing the Mel spectrogram using a vocoder.

[0057] In one embodiment, after obtaining the target text thought chain, the Mel spectrogram and response speech are generated immediately in a streaming manner, without waiting for the entire Mel spectrogram to be generated, thus reducing response latency.

[0058] In the above embodiments, the large language model directly receives speech feature vectors to generate thought chains, avoiding the cumulative errors caused by traditional concatenated architectures, improving semantic accuracy, and enhancing the quality of response speech. Furthermore, the large language model generates an initial text thought chain that can be edited by the user. This initial text thought chain is then corrected based on user input to obtain the target text thought chain. The response speech is then generated based on the target text thought chain, allowing users to modify the speech during the speech generation process for personalized customization, further improving the quality of the response speech. Applying this method to voice interaction systems in financial or medical businesses such as voice product explanations, remote medical consultations, and medical education and training, it can obtain speech emotions and styles that meet user needs through text thought chains, generating response speech that meets user preferences and improving user experience. This method can also be applied to systems such as real-time voice assistants, interactive audio content generation, game character dialogue, and voice restoration and creation, improving the quality and richness of audio content, game character dialogue, and voice restoration and creation.

[0059] Please see Figure 2 , Figure 2 This is a schematic flowchart illustrating a thought-chain-based voice interaction method provided in an embodiment of this application. This thought-chain-based voice interaction method can be applied to voice interaction systems and servers. It generates thought chains by directly receiving speech feature vectors through a large language model, avoiding the cumulative errors caused by traditional serial architectures, improving semantic accuracy, and enhancing the quality of response speech. Through interpretable and interactive thought chains, users can edit the generated content and style, achieving transparency and precise control over the speech generation process, significantly improving the quality and personalization of response speech in voice interaction systems.

[0060] like Figure 2 As shown, step S102 of the voice interaction method based on thought chain specifically includes steps S201 to S203.

[0061] S201. Analyze the target speech feature vector based on the large language model to obtain the description text corresponding to the input speech, and perform autoregressive generation processing based on the target speech feature vector to obtain the response text corresponding to the input speech. In one embodiment, the target speech feature vector contains all the acoustic and semantic information of the input speech, and the target speech feature vector is used as a prefix cue and input into the large language model.

[0062] When the large language model receives a target speech feature vector, it analyzes the feature vector, interprets the text content within it, and generates descriptive text. It then performs contextual reasoning based on the target speech feature vector, using autoregression to generate text content as the response text.

[0063] S202. Based on the description text and the response text, predict the control label to obtain the acoustic control label; In one embodiment, the descriptive text and the response text are concatenated into a complete dialogue context, and the most matching acoustic attribute, i.e., the acoustic control tag, is predicted based on the text speech of the dialogue context.

[0064] Acoustic control tags include sentiment tags: obtained by analyzing the emotional tone of the response text; speech rate tags: the output speech rate can be determined based on the type (such as medication guidance, emergency reminder) and content of the response text; and tags for predicting stress and intonation can also be generated. For example, stress tags may be generated to emphasize medication precautions.

[0065] Furthermore, the step of predicting control tags based on the description text and the response text to obtain acoustic control tags includes: concatenating the description text and the response text to obtain a dialogue context; and predicting tags on the dialogue context based on a preset control tag prediction model to obtain the acoustic control tags.

[0066] In one embodiment, the descriptive text and the response text are concatenated in the logical order of the dialogue to form a complete context sequence, providing a semantic basis for controlling label prediction.

[0067] The splicing method can be to use special separators (such as...) <sep>Separate the two to ensure the model can distinguish between user input and system response. For example, if the description text is "How's the weather today?" and the response text is "Sunny turning cloudy today", then the concatenated dialogue context is "How's the weather today?". <sep>Today will be sunny turning cloudy.

[0068] In one embodiment, the control label prediction model is a submodule of LLM (Large Language Model) (or an independent lightweight classification model) that takes the concatenated dialogue context as input and outputs acoustic control labels such as emotion, speech rate, and stress.

[0069] Specifically, the text embedding layer of LLM is used to convert the dialogue context into semantic embedding vectors. The semantic embedding vectors are then input into the control label prediction model. The multi-task classification head of the control label prediction model processes the semantic embedding vectors and outputs the label values ​​and probabilities of various control labels. For each classification task, the value with the highest probability is taken as the final control label value.

[0070] S203. Based on the description text, the response text, and the acoustic control tag, generate the initial text thought chain.

[0071] In one embodiment, descriptive text, response text, and acoustic control tags are assembled in a predefined, machine-parseable structured format to generate an initial text thought chain.

[0072] In the above embodiments, the large language model directly receives speech feature vectors to generate thought chains, avoiding the cumulative errors caused by traditional serial architecture, improving semantic accuracy, and enhancing the quality of response speech. Through interpretable and interactive thought chains, users can edit the generated content and style, achieving transparency and precise control over the speech generation process, and significantly improving the quality and personalization of response speech in the voice interaction system.

[0073] Please see Figure 3 , Figure 3 This is a schematic flowchart illustrating a thought-chain-based voice interaction method provided in an embodiment of this application. This thought-chain-based voice interaction method can be applied to voice interaction systems and servers. It is used to deeply fuse control tags in the thought chain of the target text with the speech features of the input speech through an acoustic projection network module, ensuring that the control tags can be effectively mapped to acoustic features, generating high-fidelity, emotional, and expressive speech.

[0074] like Figure 3 As shown, step S104 of the voice interaction method based on thought chain specifically includes steps S301 to S304.

[0075] S301. Process the reply text in the target text thought chain to obtain the reply voice feature vector; S302. Encode the acoustic control tags in the target text thought chain to obtain tag embedding vectors; S303. The tag embedding vector and the response speech feature vector are fused to generate a speech generation condition vector; S304. Based on the speech generation condition vector, generate a first Mel spectrogram, and perform speech synthesis on the first Mel spectrogram based on the vocoder to generate the response speech.

[0076] In one embodiment, the acoustic projection network module parses the target text thought chain, extracts the acoustic control tags and response text, uses a text embedding layer to convert the response text into a text embedding vector, and then uses a projection layer to map the text embedding vector to the speech feature space to obtain the response speech feature vector.

[0077] In one embodiment, the label value of each acoustic control tag is treated as an independent word. For example, the emotion tag has the value "happy," and the speech rate tag has the value "fast." The system stores a pre-trained embedding matrix, which functions similarly to a dictionary. For instance, in this matrix, the label "happy" corresponds to a 256-dimensional vector [0.12, -0.45, ..., 0.88], and "sad" corresponds to another different 256-dimensional vector. Each control tag is converted into a fixed, dense numerical vector, and the embedding vectors of each tag are concatenated into a unified label embedding vector.

[0078] In another embodiment, acoustic control tags in the target text's thought chain can be categorized according to tag type, with each type corresponding to an independent preset embedding table. The tag values ​​are mapped to fixed-dimensional vectors through an embedding layer. For example, sentiment tags use a sentiment embedding table, and speech rate tags use a speech rate embedding table. The embedding vectors of each tag type are concatenated into a unified tag embedding vector (e.g., concatenated sentiment + speech rate + accent vectors).

[0079] In one embodiment, the label embedding vector is combined with the response speech feature vector. The data is fused to generate a speech generation conditional vector. , ,in, This indicates a feature concatenation operation; Embed indicates that the label will be controlled. Convert to an embedding vector.

[0080] In another embodiment, if the speech feature vector and the label embedding vector have the same dimension, the fusion operation can also be to directly add them element by element.

[0081] In one embodiment, the acoustic projection network module further includes a post-processing network layer (Post-net), which generates a first Mel spectrogram based on the speech generation conditional vector. , Then, a neural vocoder is used to convert the first Mel spectrogram into the response speech.

[0082] In the above embodiments, the control tags in the target text's thought chain are deeply fused with the speech features of the input speech through the acoustic projection network module, ensuring that the control tags can be effectively mapped to acoustic features, generating speech with high fidelity and rich emotion and expressiveness.

[0083] Please see Figure 4 , Figure 4 This application provides a schematic block diagram of a thought-chain-based voice interaction device, which is used to execute the aforementioned thought-chain-based voice interaction method. The thought-chain-based voice interaction device can be configured on a server.

[0084] like Figure 4 As shown, the thought chain-based voice interaction device 400 includes: The feature vector acquisition module 401 is used to process the input speech upon receiving the input speech to obtain the target speech feature vector; The thought chain generation module 402 is used to perform autoregressive processing on the target speech feature vector based on a large language model to generate an initial text thought chain. The thought chain correction module 403 is used to receive input content from the user based on the initial text thought chain, and update the initial text thought chain based on the input content to obtain the target text thought chain. The response speech generation module 404 is used to perform speech synthesis based on the target text thought chain to generate a response speech corresponding to the input speech.

[0085] Furthermore, the thought chain generation module 402 includes: The text generation unit is used to analyze the target speech feature vector based on the large language model to obtain the descriptive text corresponding to the input speech, and to perform autoregressive generation processing based on the target speech feature vector to obtain the response text corresponding to the input speech. A tag prediction unit is used to predict control tags based on the description text and the response text to obtain acoustic control tags; The thought chain generation unit is used to generate the initial text thought chain based on the description text, the response text, and the acoustic control tag.

[0086] Furthermore, the label prediction unit includes: The text splicing subunit is used to splice the description text and the response text to obtain the dialogue context; The tag prediction subunit is used to predict tags for the dialogue context based on a preset control tag prediction model to obtain the acoustic control tags.

[0087] Furthermore, the response voice generation module 404 includes: The response speech feature vector acquisition unit is used to process the response text in the target text thought chain to obtain the response speech feature vector; The tag embedding vector acquisition unit is used to encode the acoustic control tags in the target text thought chain to obtain the tag embedding vector; The speech generation condition vector generation unit is used to fuse the label embedding vector and the response speech feature vector to generate a speech generation condition vector; The response speech generation unit is used to generate a first Mel spectrogram based on the speech generation condition vector, and to perform speech synthesis on the first Mel spectrogram based on a vocoder to generate the response speech.

[0088] Furthermore, the feature vector acquisition module 401 includes: The Mel spectrogram acquisition unit is used to perform spectrum conversion on the input speech to obtain a second Mel spectrogram corresponding to the input speech; The initial speech feature vector acquisition unit is used to transform the second Mel spectrogram based on the speech encoder to obtain the initial speech feature vector; The target speech feature vector acquisition unit is used to map the initial speech feature vector according to a preset feature dimension to obtain the target speech feature vector.

[0089] Furthermore, the thought-chain-based voice interaction device 400 also includes a training module, which includes: The data acquisition unit is used to acquire a training dataset and a joint loss function, wherein the training dataset includes a spectrogram corresponding to the training audio and the text corresponding to the training audio, and the training audio consists of a pair of question speech and answer speech; A joint training unit is used to jointly train a pre-trained speech encoder and a pre-trained large language model based on a training dataset and the joint loss function, so as to obtain the speech encoder and the large language model.

[0090] Furthermore, the thought chain correction module 403 includes: The input content parsing unit is used to receive the input content based on the interactive editing interface, parse the input content, and obtain the tag field to be modified and the tag modification value; The thought chain correction unit is used to correct the acoustic control tags in the initial text thought chain based on the tag field to be modified and the tag modification value, so as to obtain the target text thought chain.

[0091] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0092] The aforementioned device can be implemented as a computer program, which can be used in, for example... Figure 5 It runs on the computer device shown.

[0093] Please see Figure 5 , Figure 5 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a voice interaction system or a server.

[0094] See Figure 5 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0095] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any thought-chain-based voice interaction method.

[0096] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0097] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When these computer programs are executed by the processor, the processor can execute any thought chain-based voice interaction method.

[0098] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0099] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0100] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: Upon receiving input speech, the input speech is processed to obtain the target speech feature vector; Autoregressive processing is performed on the target speech feature vector based on a large language model to generate an initial text thought chain. Receive user input based on the initial text thought chain, and update the initial text thought chain based on the input to obtain the target text thought chain; Speech synthesis is performed based on the target text's thought chain to generate a response speech corresponding to the input speech.

[0101] In one embodiment, when the processor performs autoregressive processing on the target speech feature vector based on a large language model to generate an initial text thought chain, it is used to: Based on the large language model, the target speech feature vector is analyzed to obtain the descriptive text corresponding to the input speech, and autoregressive generation processing is performed based on the target speech feature vector to obtain the response text corresponding to the input speech. Based on the description text and the response text, control label prediction is performed to obtain acoustic control labels; The initial text thought chain is generated based on the description text, the response text, and the acoustic control tag.

[0102] In one embodiment, when the processor performs control tag prediction based on the description text and the response text to obtain acoustic control tags, it is configured to: The description text and the response text are concatenated to obtain the dialogue context; Based on a preset control label prediction model, the dialog context is used to predict labels to obtain the acoustic control labels.

[0103] In one embodiment, when the processor performs speech synthesis based on the target text thought chain to generate a response speech corresponding to the input speech, it is configured to: The response text in the target text thought chain is processed to obtain the response voice feature vector; Encode the acoustic control tags in the target text thought chain to obtain tag embedding vectors; The label embedding vector and the response speech feature vector are fused to generate a speech generation condition vector; Based on the speech generation condition vector, a first Mel spectrogram is generated, and speech synthesis is performed on the first Mel spectrogram based on the vocoder to generate the response speech.

[0104] In one embodiment, when the processor processes the input speech to obtain the target speech feature vector, it is configured to: The input speech is subjected to spectrum conversion to obtain a second Mel spectrogram corresponding to the input speech; The second Mel spectrogram is transformed based on the speech encoder to obtain the initial speech feature vector; The initial speech feature vector is mapped according to a preset feature dimension to obtain the target speech feature vector.

[0105] In one embodiment, before the processor performs the transformation of the second Mel spectrogram based on the speech encoder to obtain the initial speech feature vector, it is also configured to perform: Obtain a training dataset and a joint loss function, wherein the training dataset includes a spectrogram corresponding to the training audio and the text corresponding to the training audio, and the training audio consists of a pair of question speech and answer speech; Based on the training dataset and the joint loss function, the pre-trained speech encoder and the pre-trained large language model are jointly trained to obtain the speech encoder and the large language model.

[0106] In one embodiment, when the processor receives user input based on the initial text thought chain and updates the initial text thought chain based on the input to obtain the target text thought chain, it is configured to: The input content is received based on the interactive editing interface, and the input content is parsed to obtain the tag field to be modified and the tag modification value; Based on the label field to be modified and the label modification value, the acoustic control label in the initial text thinking chain is corrected to obtain the target text thinking chain.

[0107] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the thought chain-based voice interaction methods provided in the embodiments of this application.

[0108] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0109] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.< / sep> < / sep>

Claims

1. A voice interaction method based on thought chain, characterized in that, include: Upon receiving input speech, the input speech is processed to obtain the target speech feature vector; Autoregressive processing is performed on the target speech feature vector based on a large language model to generate an initial text thought chain. Receive user input based on the initial text thought chain, and update the initial text thought chain based on the input to obtain the target text thought chain; Speech synthesis is performed based on the target text's thought chain to generate a response speech corresponding to the input speech.

2. The voice interaction method based on thought chain according to claim 1, characterized in that, The process of performing autoregressive processing on the target speech feature vector based on a large language model to generate an initial text thought chain includes: Based on the large language model, the target speech feature vector is analyzed to obtain the descriptive text corresponding to the input speech, and autoregressive generation processing is performed based on the target speech feature vector to obtain the response text corresponding to the input speech. Based on the description text and the response text, control label prediction is performed to obtain acoustic control labels; The initial text thought chain is generated based on the description text, the response text, and the acoustic control tag.

3. The voice interaction method based on thought chain according to claim 2, characterized in that, The step of predicting control labels based on the description text and the response text to obtain acoustic control labels includes: The description text and the response text are concatenated to obtain the dialogue context; Based on a preset control label prediction model, the dialog context is used to predict labels to obtain the acoustic control labels.

4. The voice interaction method based on thought chain according to claim 1, characterized in that, The step of synthesizing speech based on the target text's thought chain to generate a response speech corresponding to the input speech includes: The response text in the target text thought chain is processed to obtain the response voice feature vector; Encode the acoustic control tags in the target text thought chain to obtain tag embedding vectors; The label embedding vector and the response speech feature vector are fused to generate a speech generation condition vector; Based on the speech generation condition vector, a first Mel spectrogram is generated, and speech synthesis is performed on the first Mel spectrogram based on the vocoder to generate the response speech.

5. The voice interaction method based on thought chain according to claim 1, characterized in that, The process of processing the input speech to obtain the target speech feature vector includes: The input speech is subjected to spectrum conversion to obtain a second Mel spectrogram corresponding to the input speech; The second Mel spectrogram is transformed based on the speech encoder to obtain the initial speech feature vector; The initial speech feature vector is mapped according to a preset feature dimension to obtain the target speech feature vector.

6. The voice interaction method based on thought chain according to claim 5, characterized in that, Before converting the second Mel spectrogram based on the speech encoder to obtain the initial speech feature vector, the method further includes: Obtain the training dataset and joint loss function, wherein the training dataset includes the spectrogram corresponding to the training audio and the text corresponding to the training audio, and the training audio consists of a pair of question speech and answer speech; Based on the training dataset and the joint loss function, the pre-trained speech encoder and the pre-trained large language model are jointly trained to obtain the speech encoder and the large language model.

7. The voice interaction method based on thought chain according to any one of claims 1 to 6, characterized in that, The step of receiving user input based on the initial text thought chain and updating the initial text thought chain based on the input to obtain the target text thought chain includes: The input content is received based on the interactive editing interface, and the input content is parsed to obtain the tag field to be modified and the tag modification value; Based on the label field to be modified and the label modification value, the acoustic control label in the initial text thinking chain is corrected to obtain the target text thinking chain.

8. A voice interaction device based on thought chain, characterized in that, include: The feature vector acquisition module is used to process the input speech upon receiving it to obtain the target speech feature vector. The thought chain generation module is used to perform autoregressive processing on the target speech feature vector based on a large language model to generate an initial text thought chain. The thought chain correction module is used to receive input content from the user based on the initial text thought chain, and update the initial text thought chain based on the input content to obtain the target text thought chain. The response speech generation module is used to perform speech synthesis based on the target text thought chain to generate a response speech corresponding to the input speech.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, implement the thought chain-based voice interaction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement the thought chain-based voice interaction method as described in any one of claims 1 to 7.