Voice conversation method and device based on artificial intelligence, equipment and storage medium
By combining discrete speech coding and dialogue language models with dialogue context information to generate dialogue response speech, the problems of insufficient recognition of paralinguistic information and insufficient adjustment of turn intervals in existing technologies are solved, and more accurate and vivid spoken dialogue is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-17
AI Technical Summary
Existing spoken dialogue systems (ASR-LLM-TTS cascade system and end-to-end speech generation model) cannot effectively recognize paralinguistic information and dynamically adjust turn intervals, resulting in generated speech that lacks human touch and fluency.
Discrete speech unit sequences of the question speech are obtained through discrete speech coding, and prediction is performed using a dialogue language model. The dialogue response speech is generated by combining dialogue environment information, and the volume and speech rate are adjusted to improve the accuracy and vividness of the dialogue.
It enables accurate responses to dialogue questions and dynamic adjustments to the dialogue environment, thereby improving the accuracy and vividness of spoken dialogue.
Smart Images

Figure CN121884807A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a voice dialogue method, apparatus, device and storage medium based on artificial intelligence. Background Technology
[0002] Spoken language is the most direct and frequently used form of human communication, and oral expression skills are crucial in both work and daily life. Good oral expression skills enable the effective and accurate transmission of information, thereby improving communication efficiency. For example, in the financial industry, when clients inquire about financial products, a serious atmosphere and sincere tone are needed in the conversation to build trust and acceptance. Similarly, in the medical field, when patients inquire about medication methods, a gentle tone and a warm atmosphere are needed to provide more care and support, facilitating a faster recovery.
[0003] Currently, spoken dialogue systems (SDS) mainly include cascaded systems of ASR-LLM-TTS and end-to-end speech generation models. Although LLM and TTS have developed rapidly in recent years, they still have significant shortcomings and are difficult to meet the needs of natural human-computer interaction. For cascaded systems of ASR-LLM-TTS, firstly, non-verbal signals are completely lost. The ASR module only focuses on speech-to-text conversion and cannot recognize paralinguistic information such as laughter, sighs, and feedback words (such as "hmm" and "yeah"), resulting in speech generated by TTS lacking human warmth and being disconnected from the rich paralinguistic expressions in human dialogue. Secondly, turn-taking is mechanical and rigid. Text LM is based on text training for turn-taking segmentation, and its generation logic is limited to the "alternating speaking" mode. It cannot produce the reasonable overlaps common in human dialogue, nor can it accurately control the duration of turn-taking intervals. In real-world scenarios, the intervals in human dialogue are usually between several hundred milliseconds and one second, and they are dynamically adjusted according to the rhythm of the dialogue. However, cascaded systems often have excessively long silences (more than 2 seconds) or no gaps between turns, disrupting the fluency of the dialogue.
[0004] Therefore, how to improve the accuracy and vividness of spoken dialogue is an urgent problem to be solved. Summary of the Invention
[0005] The main purpose of this application is to provide an artificial intelligence-based voice dialogue method, device, equipment, and storage medium, which aims to improve the accuracy and vividness of spoken dialogue.
[0006] Firstly, this application provides an artificial intelligence-based voice dialogue method, which includes the following steps: Acquire the dialogue question speech and perform discrete speech coding processing on the dialogue question speech to obtain a discrete unit sequence of question speech; Obtain a dialogue language model, wherein the dialogue language model is obtained by pre-training based on sample data; The question speech discrete unit sequence is used to perform dialogue prediction on the dialogue language model to obtain the target answer speech discrete unit sequence. Based on the spoken dialogue questions, determine the dialogue environment information; Based on the target response speech discrete unit sequence and the dialogue environment information, spoken dialogue prediction is performed to generate dialogue response speech.
[0007] Secondly, this application also provides a voice dialogue device, which includes an acquisition module, a generation module, and a determination module, wherein: The acquisition module is used to acquire the voice of the dialogue questions; The generation module is used to perform discrete speech encoding processing on the dialogue question speech to obtain a discrete unit sequence of question speech; The acquisition module is also used to acquire a dialogue language model, which is a dialogue language model that has been trained in advance based on sample data. The generation module is further configured to perform dialogue prediction on the question speech discrete unit sequence using a dialogue language model to obtain the target answer speech discrete unit sequence. The determining module is used to determine the dialogue environment information based on the voice of the dialogue question; The generation module is further configured to perform spoken dialogue prediction based on the target response speech discrete unit sequence and the dialogue environment information, and generate dialogue response speech.
[0008] Thirdly, this application also provides a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the artificial intelligence-based voice dialogue method described above.
[0009] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the artificial intelligence-based voice dialogue method described above.
[0010] This application provides a voice dialogue method, apparatus, device, and storage medium based on artificial intelligence. The application acquires dialogue question speech and performs discrete speech encoding processing on the dialogue question speech to obtain a discrete unit sequence of question speech; acquires a dialogue language model; performs dialogue prediction on the discrete unit sequence of question speech using the dialogue language model to obtain a target response speech discrete unit sequence; determines dialogue environment information based on the dialogue question speech; and performs spoken dialogue prediction based on the target response speech discrete unit sequence and dialogue environment information, thereby accurately generating dialogue response speech. This application can directly perform question-and-answer spoken dialogue on the dialogue question speech, accurately improving the accuracy of spoken dialogue. Furthermore, by adjusting the output spoken dialogue based on dialogue environment information, it can effectively improve the vividness of spoken dialogue, thereby improving its accuracy. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating an artificial intelligence-based voice dialogue method provided in this application embodiment; Figure 2 A flowchart illustrating another AI-based voice dialogue method provided in this application embodiment; Figure 3 for Figure 1 A flowchart illustrating the sub-steps of an AI-based voice dialogue method. Figure 4 A schematic block diagram of a voice dialogue device provided in an embodiment of this application; Figure 5 for Figure 4 A schematic block diagram of a sub-module of the voice dialogue device in the diagram; Figure 6 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0013] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the described order. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0016] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0017] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0018] Spoken language is the most direct and frequently used form of human communication, and oral expression skills are crucial in both work and daily life. Good oral expression skills enable the effective and accurate transmission of information, thereby improving communication efficiency. For example, in the financial industry, when clients inquire about financial products, a serious atmosphere and sincere tone are needed in the conversation to build trust and acceptance. Similarly, in the medical field, when patients inquire about medication methods, a gentle tone and a warm atmosphere are needed to provide more care and support, facilitating a faster recovery.
[0019] Currently, spoken dialogue systems (SDS) mainly include cascaded systems of ASR-LLM-TTS and end-to-end speech generation models. Although LLM and TTS have developed rapidly in recent years, they still have significant shortcomings and are difficult to meet the needs of natural human-computer interaction. For cascaded systems of ASR-LLM-TTS, firstly, non-verbal signals are completely lost. The ASR module only focuses on speech-to-text conversion and cannot recognize paralinguistic information such as laughter, sighs, and feedback words (such as "hmm" and "yeah"), resulting in speech generated by TTS lacking human warmth and being disconnected from the rich paralinguistic expressions in human dialogue. Secondly, turn-taking is mechanical and rigid. Text LM is based on text training for turn-taking segmentation, and its generation logic is limited to the "alternating speaking" mode. It cannot produce the reasonable overlaps common in human dialogue, nor can it accurately control the duration of turn-taking intervals. In real-world scenarios, the intervals in human dialogue are usually between several hundred milliseconds and one second, and they are dynamically adjusted according to the rhythm of the dialogue. However, cascaded systems often have excessively long silences (more than 2 seconds) or no gaps between turns, disrupting the fluency of the dialogue.
[0020] To address the aforementioned problems, embodiments of this application provide an artificial intelligence-based voice dialogue method, apparatus, device, and storage medium. The artificial intelligence-based voice dialogue method includes: acquiring dialogue question speech; performing discrete speech encoding processing on the dialogue question speech to obtain a sequence of discrete units for the question speech; acquiring a dialogue language model, wherein the dialogue language model is pre-trained based on sample data; performing dialogue prediction on the sequence of discrete units for the question speech using the dialogue language model to obtain a target response speech sequence of discrete units; determining dialogue environment information based on the dialogue question speech; and performing spoken dialogue prediction based on the target response speech sequence of discrete units and the dialogue environment information to generate a dialogue response speech.
[0021] This AI-based voice dialogue method can be applied to computer devices, such as mobile phones, tablets, laptops, desktop computers, personal digital assistants, and wearable devices.
[0022] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0023] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an artificial intelligence-based voice dialogue method provided as an embodiment of this application.
[0024] like Figure 1 As shown, the AI-based voice dialogue method includes steps S101 to S105.
[0025] Step S101: Obtain the dialogue question speech and perform discrete speech encoding processing on the dialogue question speech to obtain a discrete unit sequence of question speech.
[0026] The dialogue question is a spoken dialogue question. For example, in the financial management industry, the dialogue question is a product consultation question; in the medical field, the dialogue question is a question about how to use medicine.
[0027] In some embodiments, as shown in FIG2, the AI-based voice dialogue method further includes steps S201 to S206.
[0028] Step S201: Obtain a sample dataset, which includes multiple sample data, including sample dialogue voice, which includes the dialogue voice of both parties in the dialogue.
[0029] The sample dataset includes multiple sample data, which includes sample dialogue voice recordings. These voice recordings consist of the voice recordings of two parties in a continuous spoken dialogue.
[0030] In some embodiments, a large amount of spoken dialogue voice data is acquired, and the large amount of spoken dialogue voice data is segmented and filtered to obtain multiple voice dialogue segments. Each voice dialogue segment is used as a sample data. The method of voice segmentation and filtering of spoken dialogue voice data can be selected according to the actual situation, and this application embodiment does not specifically limit it. For example, voice segmentation can be performed at a granularity of 5 minutes of dialogue, and the filtering method can be to select spoken dialogue segments of 4-6 minutes as sample data.
[0031] Step S202: Select a sample data from the sample dataset as the target sample data, wherein the target sample data includes first-party dialogue voice and second-party dialogue voice.
[0032] A sample data is randomly selected from the sample dataset as the target sample data, which includes first-party dialogue voice and second-party dialogue voice.
[0033] Step S203: Obtain a preset discrete speech coding model, and perform discrete speech coding processing on the first party dialogue speech and the second party dialogue speech according to the preset discrete speech coding model to obtain the first party speech discrete unit sequence and the second party speech discrete unit sequence.
[0034] The preset discrete speech coding model includes a speech feature extraction layer and a speech feature clustering layer. The speech feature extraction layer can be a HuBERT model, and the speech feature clustering layer includes the k-means clustering algorithm.
[0035] In some embodiments, a speech feature extraction layer is used to extract speech features and perform mask prediction on the first-party dialogue speech to obtain multiple first-party dialogue speech features; a speech feature clustering layer is used to perform discrete unit mapping on the multiple first-party dialogue speech features to obtain a first-party speech discrete unit sequence.
[0036] In some embodiments, the method of obtaining multiple first-party dialogue speech features by performing speech feature extraction and mask prediction on the first-party dialogue speech through the speech feature extraction layer can be as follows: performing feature extraction on the first-party dialogue speech to obtain multiple first-party speech characteristic features and multiple first-party speech dependency features; performing feature quantization processing and mask prediction processing on the multiple first-party speech characteristic features and multiple first-party speech dependency features to obtain multiple first-party dialogue speech features.
[0037] In some embodiments, the method of obtaining a first-party speech discrete unit sequence by mapping multiple first-party dialogue speech features to discrete units through a speech feature clustering layer can be as follows: the first-party speech discrete unit sequence is obtained by clustering multiple first-party dialogue speech features to discrete units through a preset k-means algorithm.
[0038] In some embodiments, a speech feature extraction layer is used to extract speech features and perform mask prediction on the second-party dialogue speech to obtain multiple second-party dialogue speech features; a speech feature clustering layer is used to perform discrete unit mapping on the multiple second-party dialogue speech features to obtain a second-party speech discrete unit sequence.
[0039] In some embodiments, the method of obtaining multiple second-party dialogue speech features by performing speech feature extraction and mask prediction on the second-party dialogue speech through the speech feature extraction layer can be as follows: performing feature extraction on the second-party dialogue speech to obtain multiple second-party speech characteristic features and multiple second-party speech dependency features; performing feature quantization processing and mask prediction processing on the multiple second-party speech characteristic features and multiple second-party speech dependency features to obtain multiple second-party dialogue speech features.
[0040] In some embodiments, the method of obtaining a second-party speech discrete unit sequence by mapping multiple second-party dialogue speech features to discrete units through a speech feature clustering layer can be as follows: the second-party speech discrete unit sequence is obtained by clustering multiple second-party dialogue speech features to discrete units through a preset k-means algorithm.
[0041] Step S204: Obtain a preset dialogue language model, and use the preset dialogue language model to perform dialogue prediction on the first party's speech discrete unit sequence to obtain the predicted response speech discrete unit sequence.
[0042] The preset dialogue language model can be a Dialogue Language Model.
[0043] In some embodiments, a dialogue prediction is performed on the first-party speech discrete unit sequence using a preset dialogue language model to obtain a predicted response speech discrete unit sequence. By performing dialogue prediction on the first-party speech discrete unit sequence using this preset dialogue language model, the predicted response speech discrete unit sequence can be accurately obtained.
[0044] Step S205: Determine the dialogue environment information based on the first party's dialogue voice, and perform spoken dialogue prediction based on the dialogue environment information and the predicted response voice discrete unit sequence to generate the predicted dialogue response voice.
[0045] The environmental information refers to the dialogue environment in which the speakers are located, such as a quiet indoor environment, a noisy street environment, and a quiet meeting room environment.
[0046] In some embodiments, the ambient sounds of the first-party dialogue speech are identified and classified to obtain the dialogue environment. The discrete unit sequence of the predicted response speech is decoded to generate the predicted initial dialogue response speech; an environment label is determined based on the dialogue environment information, and dialogue adjustment parameters are determined based on the environment label; the volume and speech rate parameters of the predicted initial dialogue response speech are adjusted according to the dialogue adjustment parameters to obtain the predicted dialogue response speech.
[0047] In some embodiments, determining the dialogue adjustment parameters based on the environment tags can be achieved by: obtaining a preset mapping table between dialogue adjustment parameters and environment tags; and querying the dialogue adjustment parameters corresponding to the environment tags from the mapping table. This mapping table is pre-established based on the dialogue adjustment parameters and environment tags, and can be established according to actual circumstances; this application embodiment does not specifically limit its implementation.
[0048] Step S206: Based on the predicted dialogue response speech and the second-party dialogue speech, determine whether the preset discrete speech coding model and the preset dialogue language model have converged. If at least one of the preset discrete speech coding model and the preset dialogue language model has not converged, adjust the model parameters of at least one of the preset discrete speech coding model and the preset dialogue language model, and continue to execute the step of selecting a sample data from the sample dataset as the target sample data until the model converges.
[0049] Based on the predicted dialogue response speech and the second-party dialogue speech, determine the model loss values of the preset discrete speech coding model and the preset dialogue language model. If the model loss value is greater than or equal to the preset model loss value, it is determined that at least one of the preset discrete speech coding model and the preset dialogue language model has not converged. Then, adjust the model parameters of at least one of the preset discrete speech coding model and the preset dialogue language model, and continue to select a sample data from the sample dataset as the target sample data. Perform discrete speech coding processing on the first-party dialogue speech and the second-party dialogue speech according to the preset discrete speech coding model to obtain the first-party speech discrete unit sequence and the second-party speech discrete unit sequence. Perform dialogue prediction on the first-party speech discrete unit sequence through the preset dialogue language model to obtain the predicted response speech discrete unit sequence. Based on the first-party dialogue speech, determine the dialogue environment information, and perform spoken dialogue prediction based on the dialogue environment information and the predicted response speech discrete unit sequence to generate the predicted dialogue response speech. Based on the predicted dialogue response speech and the second-party dialogue speech, determine whether the preset discrete speech coding model and the preset dialogue language model have converged, until the model converges. The preset model loss value can be set according to the actual situation. This application embodiment does not make specific limitations on this. For example, the preset loss value can be set to 0.02.
[0050] In some embodiments, the method for determining the model loss values of the preset discrete speech coding model and the preset dialogue language model based on the predicted dialogue response speech and the second-party dialogue speech can be as follows: calculate the similarity between the predicted dialogue response speech and the second-party dialogue speech to obtain the current similarity; obtain the historical similarity, which is the mean of the current similarities of each sample data that has been trained; calculate the mean of the current similarity and the historical similarity to obtain the target similarity; subtract the target similarity from the unit value to determine the model loss value. The method for calculating the similarity can be selected according to the actual situation, and this embodiment does not specifically limit it. For example, the cosine similarity between the predicted dialogue response speech and the second-party dialogue speech can be calculated.
[0051] In some embodiments, such as Figure 3 As shown, step S101 includes sub-steps S1011 to S1013.
[0052] Sub-step S1011: Obtain a discrete speech coding model, which includes a speech feature extraction layer and a speech feature clustering layer.
[0053] The discrete speech coding model includes a speech feature extraction layer and a speech feature clustering layer. The speech feature extraction layer is used to extract speech features and predict masks for dialogue question speech, and the speech feature clustering layer is used to map multiple speech features into discrete units.
[0054] Sub-step S1012: Extract speech features and perform mask prediction on the speech of the dialogue question through the speech feature extraction layer to obtain multiple speech features.
[0055] Feature extraction is performed on the speech of dialogue questions to obtain multiple speech characteristic features and multiple speech dependency features. These features are then subjected to feature quantization and mask prediction to obtain multiple speech features. By extracting dependency features and performing mask prediction on the speech of dialogue questions, speech features with dependencies and semantic associations can be obtained, thereby effectively improving the efficiency and accuracy of spoken dialogue prediction.
[0056] Sub-step S1013: The multiple speech features are mapped to discrete units through the speech feature clustering layer to obtain the discrete unit sequence of the problem speech.
[0057] By using a pre-defined k-means algorithm to cluster and map multiple speech features into discrete units, a sequence of discrete units representing the speech problem is obtained. This pre-defined k-means algorithm accurately yields the sequence of discrete units representing the speech problem, thereby effectively improving the efficiency and accuracy of spoken dialogue prediction.
[0058] Step S102: Obtain the dialogue language model, which is a dialogue language model that has been trained in advance based on sample data.
[0059] The dialogue language model is used to predict dialogues from discrete unit sequences of question speech.
[0060] In some embodiments, obtaining a dialogue language model can effectively improve the efficiency and accuracy of dialogue question responses.
[0061] Step S103: Perform dialogue prediction on the question speech discrete unit sequence using a dialogue language model to obtain the target answer speech discrete unit sequence.
[0062] By performing temporal dependency and dialogue interaction prediction processing on the discrete unit sequence of question speech using a dialogue language model, the target response discrete unit sequence can be obtained. This dialogue language model enables accurate acquisition of the target response discrete unit sequence, significantly improving the efficiency and accuracy of spoken dialogue.
[0063] Step S104: Determine the dialogue environment information based on the spoken dialogue question.
[0064] Background noise is extracted from the spoken dialogue to obtain dialogue environment information. This allows for accurate acquisition of dialogue environment information from the spoken dialogue.
[0065] For example, human voices are removed from the spoken dialogue, and the remaining sound is used as the dialogue context information. Using the sound retained after removing human voices as the dialogue context information can preserve the dialogue context information to the greatest extent.
[0066] Step S105: Perform spoken dialogue prediction based on the target response speech discrete unit sequence and the dialogue environment information to generate dialogue response speech.
[0067] In some embodiments, the discrete unit sequence of the target response speech is decoded to generate an initial dialogue response speech; environmental labels are determined based on dialogue environment information, and dialogue adjustment parameters are determined based on the environmental labels; the volume and speech rate parameters of the initial dialogue response speech are adjusted according to the dialogue adjustment parameters to obtain the final dialogue response speech. Adjusting the volume and speech rate parameters of the initial dialogue response speech using dialogue environment information can yield a more accurate and vivid dialogue response speech.
[0068] In some embodiments, determining the environment label based on the dialogue environment information can be achieved by: obtaining a preset environmental sound classification model, performing background sound analysis on the dialogue environment information using the preset environmental sound classification model, and obtaining the environment label. The preset environmental sound classification model can be set according to actual conditions, and this application embodiment does not specifically limit it. For example, the preset environmental sound classification model can be an audio event classification model. Performing background sound analysis on the dialogue environment information using the preset environmental sound classification model can accurately obtain the environment label.
[0069] In some embodiments, determining the dialogue adjustment parameters based on the environment tags can be achieved by: obtaining a preset mapping table between dialogue adjustment parameters and environment tags; and querying the dialogue adjustment parameters corresponding to the environment tags from the mapping table. The mapping table can be set according to actual conditions, and this embodiment does not impose specific limitations on it.
[0070] For example, in the financial management industry, when a customer inquires about financial products, the system obtains the customer's spoken product consultation voice, performs discrete speech encoding on the product consultation voice, and obtains a discrete unit sequence of product consultation voice. It then obtains a dialogue language model and uses this model to predict the dialogue within the discrete unit sequence of product consultation voice, resulting in a discrete unit sequence of product introduction voice. Based on the product consultation voice, it determines the dialogue environment information. Finally, it performs spoken dialogue prediction based on the discrete unit sequence of product introduction voice and the product consultation voice, generating the product introduction voice required by the customer.
[0071] For example, in the medical field, when a patient inquires about medication usage, the system acquires the patient's voice recording of the medication usage consultation; performs discrete speech encoding on the voice recording to obtain a discrete unit sequence of the voice recording; obtains a dialogue language model; uses the dialogue language model to predict the dialogue within the discrete unit sequence of the voice recording, resulting in a discrete unit sequence of the voice recording; determines the dialogue environment information based on the voice recording; and performs spoken dialogue prediction based on the discrete unit sequence of the voice recording and the dialogue environment information to generate the voice recording of the medication usage.
[0072] The AI-based voice dialogue method provided in the above embodiments acquires the dialogue question speech and performs discrete speech encoding processing on the dialogue question speech to obtain a discrete unit sequence of question speech; acquires a dialogue language model; performs dialogue prediction on the discrete unit sequence of question speech using the dialogue language model to obtain a discrete unit sequence of target response speech; determines dialogue environment information based on the dialogue question speech; and performs spoken dialogue prediction based on the discrete unit sequence of target response speech and dialogue environment information, which can accurately generate dialogue response speech. This application can directly perform question-and-answer spoken dialogue on the dialogue question speech, which can accurately improve the accuracy of spoken dialogue, and adjust the output spoken dialogue based on dialogue environment information, which can effectively improve the vividness of spoken dialogue, thereby improving the accuracy of spoken dialogue.
[0073] Please see Figure 4 , Figure 4 This is a schematic block diagram of a voice dialogue device provided in an embodiment of this application.
[0074] like Figure 4 As shown, the voice dialogue device 300 includes an acquisition module 310, a generation module 320, and a determination module 330, wherein: The acquisition module 310 is used to acquire the voice of the dialogue question; The generation module 320 is used to perform discrete speech encoding processing on the dialogue question speech to obtain a discrete unit sequence of question speech; The acquisition module 310 is also used to acquire a dialogue language model, wherein the dialogue language model is obtained by pre-training based on sample data. The generation module 320 is further configured to perform dialogue prediction on the question speech discrete unit sequence using a dialogue language model to obtain the target answer speech discrete unit sequence. The determining module 330 is used to determine dialogue environment information based on the voice of the dialogue question; The generation module 320 is further configured to perform spoken dialogue prediction based on the target response speech discrete unit sequence and the dialogue environment information, and generate dialogue response speech.
[0075] In some embodiments, such as Figure 5 As shown, the generation module 320 includes an acquisition submodule 321 and a generation submodule 322, wherein: The acquisition submodule 321 is used to acquire a discrete speech coding model, which includes a speech feature extraction layer and a speech feature clustering layer. The generation submodule 322 is used to extract speech features and predict masking for the speech of the dialogue question through the speech feature extraction layer to obtain multiple speech features; The generation submodule 322 is further configured to perform discrete unit mapping on the multiple speech features through the speech feature clustering layer to obtain the discrete unit sequence of the problem speech.
[0076] In some embodiments, the generation submodule 322 is further configured to: Feature extraction is performed on the spoken dialogue to obtain multiple speech characteristic features and multiple speech dependency features; The multiple speech feature features and the multiple speech dependency features are subjected to feature quantization and mask prediction processing to obtain the multiple speech features.
[0077] In some embodiments, the generation submodule 322 is further configured to: The multiple speech features are clustered and mapped to discrete units using a preset k-means algorithm to obtain the discrete unit sequence of the problem speech.
[0078] In some embodiments, the generation submodule 322 is further configured to: The question speech discrete unit sequence is processed by the dialogue language model to perform temporal dependence and dialogue interaction prediction, thereby obtaining the target answer speech discrete unit sequence.
[0079] In some embodiments, the generation submodule 322 is further configured to: The target response speech discrete unit sequence is decoded to generate the initial dialogue response speech; Based on the dialogue environment information, determine the environment label, and based on the environment label, determine the dialogue adjustment parameters; The volume and speech rate parameters of the initial dialogue response speech are adjusted according to the dialogue adjustment parameters to obtain the dialogue response speech.
[0080] In some embodiments, the generation submodule 322 is further configured to: Obtain the mapping table between preset dialogue adjustment parameters and environment labels; Retrieve the dialogue adjustment parameters corresponding to the environment label from the mapping table.
[0081] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the aforementioned voice dialogue device can be referred to the corresponding process in the aforementioned embodiment of the AI-based voice dialogue method, and will not be repeated here.
[0082] Please see Figure 6 , Figure 6 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0083] like Figure 6 As shown, the computer device 400 includes a processor 402 and a memory 406 connected via a system bus 401, wherein the memory may include a storage medium and internal memory.
[0084] The storage medium may store a computer program. This computer program includes program instructions that, when executed, cause the processor to perform any artificial intelligence-based voice dialogue method.
[0085] Processor 402 provides computing and control capabilities to support the operation of the entire computer device.
[0086] Internal memory provides an environment for the execution of computer programs stored in the storage medium. When the computer program is executed by the processor, it enables the processor to execute any artificial intelligence-based voice dialogue method.
[0087] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0088] It should be understood that processor 402 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0089] In one embodiment, the processor 402 is configured to run a computer program stored in a memory to perform the following steps: Acquire the dialogue question speech and perform discrete speech coding processing on the dialogue question speech to obtain a discrete unit sequence of question speech; Obtain a dialogue language model, wherein the dialogue language model is obtained by pre-training based on sample data; The question speech discrete unit sequence is used to perform dialogue prediction on the dialogue language model to obtain the target answer speech discrete unit sequence. Based on the spoken dialogue questions, determine the dialogue environment information; Based on the target response speech discrete unit sequence and the dialogue environment information, spoken dialogue prediction is performed to generate dialogue response speech.
[0090] In one embodiment, when the processor 402 performs discrete speech encoding processing on the dialogue question speech to obtain a discrete unit sequence of question speech, it is configured to: Obtain a discrete speech coding model, wherein the discrete speech coding model includes a speech feature extraction layer and a speech feature clustering layer; The speech feature extraction layer extracts speech features and performs mask prediction on the speech of the dialogue question to obtain multiple speech features. The multiple speech features are mapped to discrete units through the speech feature clustering layer to obtain the discrete unit sequence of the problem speech.
[0091] In one embodiment, when the processor 402 performs speech feature extraction and mask prediction on the dialogue question speech through the speech feature extraction layer to obtain multiple speech features, it is used to implement: Feature extraction is performed on the spoken dialogue to obtain multiple speech characteristic features and multiple speech dependency features; The multiple speech feature features and the multiple speech dependency features are subjected to feature quantization and mask prediction processing to obtain the multiple speech features.
[0092] In one embodiment, when the processor 402 performs discrete unit mapping on the plurality of speech features through the speech feature clustering layer to obtain the discrete unit sequence of the problem speech, it is configured to: The multiple speech features are clustered and mapped to discrete units using a preset k-means algorithm to obtain the discrete unit sequence of the problem speech.
[0093] In one embodiment, when the processor 402 performs dialogue prediction on the question speech discrete unit sequence using a dialogue language model to obtain the target answer speech discrete unit sequence, it is configured to: The question speech discrete unit sequence is processed by the dialogue language model to perform temporal dependence and dialogue interaction prediction, thereby obtaining the target answer speech discrete unit sequence.
[0094] In one embodiment, when the processor 402 performs spoken dialogue prediction based on the target response speech discrete unit sequence and the dialogue environment information to generate dialogue response speech, it is configured to: The target response speech discrete unit sequence is decoded to generate the initial dialogue response speech; Based on the dialogue environment information, determine the environment label, and based on the environment label, determine the dialogue adjustment parameters; The volume and speech rate parameters of the initial dialogue response speech are adjusted according to the dialogue adjustment parameters to obtain the dialogue response speech.
[0095] In one embodiment, when implementing the process of determining the dialogue adjustment parameters based on the environment label, the processor 402 is configured to: Obtain the mapping table between preset dialogue adjustment parameters and environment labels; Retrieve the dialogue adjustment parameters corresponding to the environment label from the mapping table.
[0096] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the computer device described above can be referred to the corresponding process in the aforementioned embodiment of the AI-based voice dialogue method, and will not be repeated here.
[0097] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, and the method implemented when the program instructions are executed can refer to the various embodiments of the artificial intelligence-based voice dialogue method of this application.
[0098] The computer-readable storage medium can be an internal storage unit of the computer device described in the foregoing embodiments, such as a hard disk or memory of the computer device. The computer-readable storage medium can be non-volatile or volatile. Alternatively, the computer-readable storage medium can be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0099] Furthermore, the computer-readable storage medium may primarily include a program storage area and a data storage area, wherein the program storage area may store the operating system, at least one application required for a function, etc.; and the data storage area may store data created based on the use of blockchain nodes, etc.
[0100] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0101] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” are intended to include the plural forms.
[0102] It should also be understood that the term "and / or" as used in this specification refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0103] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above descriptions are merely specific implementations of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. An artificial intelligence-based voice dialogue method, characterized by, include: Acquire the dialogue question speech and perform discrete speech coding processing on the dialogue question speech to obtain a discrete unit sequence of question speech; Obtain a dialogue language model, wherein the dialogue language model is obtained by pre-training based on sample data; The question speech discrete unit sequence is used to perform dialogue prediction on the dialogue language model to obtain the target answer speech discrete unit sequence. Based on the spoken dialogue questions, determine the dialogue environment information; Based on the target response speech discrete unit sequence and the dialogue environment information, spoken dialogue prediction is performed to generate dialogue response speech. 2.The artificial intelligence-based voice dialogue method of claim 1, wherein, The step of performing discrete speech coding on the dialogue question speech to obtain a sequence of discrete question speech units includes: Obtain a discrete speech coding model, wherein the discrete speech coding model includes a speech feature extraction layer and a speech feature clustering layer; The speech feature extraction layer extracts speech features and performs mask prediction on the speech of the dialogue question to obtain multiple speech features. The multiple speech features are mapped to discrete units through the speech feature clustering layer to obtain the discrete unit sequence of the problem speech. 3.The artificial intelligence-based voice dialogue method of claim 2, wherein, The process involves extracting speech features and predicting a mask on the dialogue question speech through the speech feature extraction layer, resulting in multiple speech features, including: Feature extraction is performed on the spoken dialogue to obtain multiple speech characteristic features and multiple speech dependency features; The multiple speech feature features and the multiple speech dependency features are subjected to feature quantization and mask prediction processing to obtain the multiple speech features. 4.The artificial intelligence-based voice dialogue method of claim 2, wherein, The step of mapping the multiple speech features to discrete units through the speech feature clustering layer to obtain the sequence of discrete units for the problem speech includes: The multiple speech features are clustered and mapped to discrete units using a preset k-means algorithm to obtain the discrete unit sequence of the problem speech. 5.The artificial intelligence-based voice dialogue method of claim 1, wherein, The step of performing dialogue prediction on the question speech discrete unit sequence using a dialogue language model to obtain the target answer speech discrete unit sequence includes: The question speech discrete unit sequence is processed by the dialogue language model to perform temporal dependence and dialogue interaction prediction to obtain the target answer speech discrete unit sequence. 6.The artificial intelligence-based voice dialogue method of claim 1, wherein, The step of predicting spoken dialogue based on the target response speech discrete unit sequence and the dialogue environment information, and generating the dialogue response speech, includes: The target response speech discrete unit sequence is decoded to generate the initial dialogue response speech; Based on the dialogue environment information, determine the environment label, and based on the environment label, determine the dialogue adjustment parameters; The volume and speech rate parameters of the initial dialogue response speech are adjusted according to the dialogue adjustment parameters to obtain the dialogue response speech. 7.The artificial intelligence-based voice dialogue method of claim 6, wherein, The step of determining the dialogue adjustment parameters based on the environment tags includes: Obtain the mapping table between preset dialogue adjustment parameters and environment labels; Retrieve the dialogue adjustment parameters corresponding to the environment label from the mapping table.
8. A voice dialog apparatus, characterized by The voice dialogue device includes an acquisition module, a generation module, and a determination module, wherein: The acquisition module is used to acquire the voice of the dialogue questions; The generation module is used to perform discrete speech encoding processing on the dialogue question speech to obtain a discrete unit sequence of question speech; The acquisition module is also used to acquire a dialogue language model, which is a dialogue language model that has been trained in advance based on sample data. The generation module is further configured to perform dialogue prediction on the question speech discrete unit sequence using a dialogue language model to obtain the target answer speech discrete unit sequence. The determining module is used to determine the dialogue environment information based on the voice of the dialogue question; The generation module is further configured to perform spoken dialogue prediction based on the target response speech discrete unit sequence and the dialogue environment information, and generate dialogue response speech.
9. A computer device, comprising: The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the artificial intelligence-based voice dialogue method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the artificial intelligence-based voice dialogue method as described in any one of claims 1 to 7.