Online human-computer conversation method, system, device, electronic equipment, storage medium and program product

By integrating automatic speech recognition and text-based dialogue models into a two-stage voice dialogue architecture, the problems of accuracy and response time in medical diagnosis of intelligent chatbots are solved, achieving more efficient voice dialogue interaction and improving user experience.

CN120723892BActive Publication Date: 2026-01-02ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511140707.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2026-01-02
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

Existing intelligent chatbots have low accuracy and reliability in voice dialogue interaction in the field of medical diagnosis, and their response time is long, resulting in a poor user experience.

Method used

A two-stage online human-computer voice dialogue interaction architecture is adopted, which integrates the automatic speech recognition module with the text dialogue model to form a voice dialogue model. The model retains the sound information in the user's voice and extracts speech features and text features through a preset model to generate response text. The model is then jointly trained with a multimodal dialogue model.

Benefits of technology

It improves the accuracy and reliability of responses, shortens response time, and enhances user experience, especially in healthcare scenarios, providing medical advice that is more tailored to the user's actual situation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723892B_ABST
    Figure CN120723892B_ABST
Patent Text Reader

Abstract

The embodiment of the present specification provides an online man-machine conversation method, system, device, electronic equipment, storage medium and program product. The scheme provided by the embodiment, after obtaining the conversation voice input by the user in the online man-machine voice conversation, the conversation voice is input to the preset model, so that the preset model performs: extracting the voice feature of the conversation voice, and performing voice recognition on the conversation voice to obtain the corresponding conversation text, and generating the reply text based on the association information (including the conversation text corresponding to the conversation voice of the current round voice conversation user) of the online man-machine voice conversation and the voice feature of the conversation voice. Further, the generated reply text is also output to the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of artificial intelligence, and in particular, to an online human-computer conversation method, system, device, electronic equipment, storage medium and program product. BACKGROUND

[0002] Intelligent conversation robots (also commonly referred to as intelligent conversation assistants) can interact with users in a conversation, and their mission is to help users solve problems or chat, which has been widely used in various fields such as medical health. At present, intelligent conversation robots mainly use text conversation as the main interaction mode. With the development of voice technology, more and more intelligent conversation robots have added the ability of voice conversation interaction. Voice conversation interaction is more in line with the conversation habits of users, especially for users who are not used to typing, which can bring a better experience. However, the current voice conversation interaction only relies on the text corresponding to the user's voice to reply to the user, and the accuracy and reliability of the reply are relatively low, especially in the field of medical diagnosis.

[0003] Therefore, there is an urgent need to provide an improved solution for voice conversation interaction of intelligent conversation robots. SUMMARY

[0004] The embodiments in the present specification provide an online human-computer conversation method, system, device, electronic equipment, storage medium and program product, which realizes voice conversation interaction of intelligent conversation robots, can reply to the user based on the voice features in the user's voice, and is beneficial to improve the accuracy and reliability of the reply. Among them,

[0005] In a first embodiment, an online human-computer conversation method is provided in the present specification. The method comprises:

[0006] Obtaining a conversation voice input by a user in an online human-computer voice conversation;

[0007] Inputting the conversation voice into a preset model, and performing the following by the preset model: performing feature extraction on the conversation voice to obtain voice features of the conversation voice; generating a reply text based on associated information of the online human-computer voice conversation and the voice features; wherein the associated information includes a conversation text obtained by the preset model performing voice recognition on the conversation voice;

[0008] Outputting the reply text to the user.

[0009] In a second embodiment, an online human-computer conversation method is also provided in the present specification. The method comprises:

[0010] Collecting a conversation voice input by a user in an online human-computer voice conversation;

[0011] send the dialogue voice to a server to trigger the server to input the dialogue voice into a preset model, and perform the following operations by the preset model: extracting features of the dialogue voice to obtain voice features of the dialogue voice; converting the dialogue voice into dialogue text; and generating the reply text based on the dialogue text and the voice features;

[0012] receive reply voice returned by the server; the reply voice is generated by performing voice synthesis on the reply text.

[0013] play the reply voice to the user.

[0014] In a third embodiment, the present specification also provides an online human-computer dialogue method. The method comprises:

[0015] displaying an interface for health services;

[0016] starting an online human-computer voice dialogue in response to a voice dialogue starting operation triggered through the interface;

[0017] in the online human-computer voice dialogue, collecting dialogue voice input by the user;

[0018] inputting the dialogue voice into a preset model, and performing the following operations by the preset model: extracting features of the dialogue voice to obtain voice features of the dialogue voice; and generating the reply text based on associated information of the online human-computer voice dialogue and the voice features; the associated information includes interaction generated in at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interaction includes dialogue text obtained by performing voice recognition on dialogue voice of the user in the current round of voice dialogue by the preset model;

[0019] outputting the reply text to the user.

[0020] In a fourth embodiment, the present specification also provides an intelligent dialogue robot training method. The intelligent dialogue robot comprises a preset model and a voice synthesis model; the preset model comprises a voice recognition module, a word segmenter, and a trained multi-modal dialogue model. The method comprises:

[0021] performing parameter freezing on the trained multi-modal dialogue model;

[0022] training the word segmenter based on a first training sample set to obtain the word segmenter after the first training; the first training sample set comprises a plurality of first sample voices and text content corresponding to the first sample voices;

[0023] Based on a second training sample set, the word segmenter after the first training is continuously trained to obtain the word segmenter after the second training; the second training sample set contains a plurality of second sample voices and speech features corresponding to the second sample voices;

[0024] Based on a third training sample set, the speech synthesis model is trained to obtain the trained speech synthesis model; the third training sample set contains a plurality of sample texts and speech corresponding to the sample texts;

[0025] Based on a fourth training sample set, the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model are jointly fine-tuned to obtain the trained intelligent dialogue robot;

[0026] Among them, the fourth training sample set includes a plurality of third sample voices and fourth sample voices for responding to the third sample voices; in the process of joint fine-tuning, the speech recognition module is used for converting the third sample voice into corresponding text content, the word segmenter is used for extracting the speech features of the third sample voice, the multi-modal dialogue model is used for generating a dialogue text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is used for converting the dialogue text output by the multi-modal dialogue model into a dialogue voice; according to the loss value of the dialogue voice and the corresponding fourth sample voice, the parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model are fine-tuned.

[0027] In the fifth embodiment, an online human-computer dialogue system is provided in the specification. The system includes:

[0028] The client is used for collecting dialogue voice input by a user in an online human-computer voice dialogue, and sending the dialogue voice to the server;

[0029] The server is deployed with a preset model, and is used for inputting the dialogue voice into the preset model, and performing the following by the preset model: extracting features of the dialogue voice to obtain speech features of the dialogue voice; generating the dialogue text based on associated information of the online human-computer voice dialogue and the speech features; wherein the associated information contains dialogue text obtained by the preset model performing speech recognition on the dialogue voice;

[0030] The server is also used for converting the dialogue text into a dialogue voice, and sending the dialogue voice to the client;

[0031] The client is also used for playing the dialogue voice to the user.

[0032] In a sixth embodiment, an online human-computer dialogue device is provided in the specification. The device comprises:

[0033] An acquisition module is configured to acquire dialogue speech input by a user in an online human-computer voice dialogue;

[0034] An execution module is configured to input the dialogue speech into a preset model, and the preset model is configured to perform the following operations: extracting features of the dialogue speech to obtain speech features of the dialogue speech; and generating the reply text based on associated information of the online human-computer voice dialogue and the speech features, wherein the associated information comprises dialogue text obtained by the preset model through speech recognition of the dialogue speech;

[0035] An output module is configured to output the reply text to the user.

[0036] In a seventh embodiment, an online human-computer dialogue device is provided in the specification. The device comprises:

[0037] An acquisition module is configured to acquire dialogue speech input by a user in an online human-computer voice dialogue;

[0038] A sending trigger module is configured to send the dialogue speech to a server to trigger the server to input the dialogue speech into a preset model, and the preset model is configured to perform the following operations: extracting features of the dialogue speech to obtain speech features of the dialogue speech; converting the dialogue speech into dialogue text; and generating the reply text based on the dialogue text and the speech features;

[0039] A receiving module is configured to receive reply speech returned by the server, wherein the reply speech is generated by performing speech synthesis on the reply text;

[0040] A broadcast module is configured to broadcast the reply speech to the user.

[0041] In an eighth embodiment, an online human-computer dialogue device is provided in the specification. The device comprises:

[0042] A display module is configured to display an interface for health services;

[0043] A starting module is configured to start an online human-computer voice dialogue in response to a voice dialogue starting operation triggered through the interface;

[0044] An acquisition module is configured to acquire dialogue speech input by a user in the online human-computer voice dialogue;

[0045] The execution module is configured to input the dialogue voice into a preset model, and perform the following operations by using the preset model: performing feature extraction on the dialogue voice to obtain voice features of the dialogue voice; and generating the reply text based on associated information of the online human-machine voice dialogue and the voice features, wherein the associated information includes interactions generated by at least one round of voice dialogue between a user and the intelligent dialogue robot, and the interactions include dialogue text obtained by performing voice recognition on dialogue voice of the user in the current round of voice dialogue by using the preset model.

[0046] The output module is configured to output the reply text to the user.

[0047] In a ninth embodiment, an intelligent dialogue robot training device is also provided in the specification. The intelligent dialogue robot includes a preset model and a voice synthesis model; the preset model includes a voice recognition module, a word segmenter, and a trained multi-modal dialogue model. The device includes:

[0048] The freezing module is configured to freeze parameters of the trained multi-modal dialogue model.

[0049] The training module is configured to: train the word segmenter based on a first training sample set to obtain the word segmenter after first training; the first training sample set includes a plurality of first sample voices and text content corresponding to the first sample voices; continue training the word segmenter after the first training based on a second training sample set to obtain the word segmenter after the second training; the second training sample set includes a plurality of second sample voices and voice features corresponding to the second sample voices; train the voice synthesis model based on a third training sample set to obtain the trained voice synthesis model; the third training sample set includes a plurality of sample texts and voices corresponding to the sample texts; and perform joint fine-tuning training on the voice recognition module, the word segmenter, the multi-modal dialogue model, and the voice synthesis model based on a fourth training sample set to obtain a trained intelligent dialogue robot; wherein the fourth training sample set includes a plurality of third sample voices and fourth sample voices for responding to the third sample voices; in the joint fine-tuning training process, the voice recognition module is configured to convert the third sample voices into corresponding text content, the word segmenter is configured to extract voice features of the third sample voices, the multi-modal dialogue model is configured to generate reply text based on the voice features output by the word segmenter and the text content output by the voice recognition module, and the voice synthesis model is configured to convert the reply text output by the multi-modal dialogue model into reply voice; and parameters of the voice recognition module, the word segmenter, the multi-modal dialogue model, and the voice synthesis model are fine-tuned based on a loss value of the reply voice and the corresponding fourth sample voice.

[0050] In the tenth embodiment, the electronic device is provided in the specification, comprising a memory and a processor, wherein the memory stores executable program instructions, and the processor executes the program instructions to implement the method provided in the first to fourth embodiments.

[0051] In the eleventh embodiment, the computer-readable storage medium is provided in the specification, and the computer program is stored on the computer-readable storage medium, wherein when the computer program is executed in the computer, the computer executes the method provided in the first to fourth embodiments.

[0052] In the twelfth embodiment, the computer program product is also provided in the specification, comprising computer programs / instructions, which are executed by the processor to implement the method provided in the first to fourth embodiments.

[0053] The above-mentioned embodiments of the present specification provide the following solutions: after obtaining the dialogue voice input by the user in the online human-computer voice dialogue, the dialogue voice is input to the preset model, so that the preset model performs the following operations: extracting the voice features of the dialogue voice, performing voice recognition on the dialogue voice to obtain the corresponding dialogue text, and generating the reply text based on the association information of the online human-computer voice dialogue (containing the dialogue text corresponding to the dialogue voice of the user in the current round of voice dialogue) and the voice features of the dialogue voice. It can be seen that when the corresponding reply text is generated for the dialogue voice input by the user, not only the text content of the dialogue voice is combined, but also the voice features of the dialogue voice are combined, which can effectively improve the reply quality. In particular, in the medical health scene, the voice features of the user are an important diagnostic basis, and the combination of the voice features of the user can improve the accuracy and reliability of medical diagnosis, so as to effectively ensure that the medical advice given in the reply is more suitable for the actual situation of the user. The above-mentioned dialogue text corresponding to the dialogue voice, the voice features, and the generation of the reply text are realized through a preset model, the preset model combines a voice recognition module and a multi-modal dialogue model (used for reply text generation), and the design of the preset model makes the entire online human-computer voice dialogue interaction realize two-stage type, which can reduce processing delay, improve response speed, and enhance user experience. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in the specification, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only a part of the embodiments disclosed in the specification, and other drawings can also be obtained by those skilled in the art without creative labor. In the drawings:

[0055] Figure 1A technical architecture diagram of a conventional online human-machine voice dialogue provided for an exemplary embodiment;

[0056] Figure 2A and Figure 2B A technical architecture diagram on which various method implementations in the present specification are based, provided for an exemplary embodiment;

[0057] Figure 3 A structural diagram of an online human-machine dialogue system provided for an exemplary embodiment in the present specification;

[0058] Figure 4 , Figure 5 and Figure 6 A flow diagram of an online human-machine dialogue method provided for an exemplary embodiment in the present specification;

[0059] Figure 7 A flow diagram of a smart dialogue robot training method provided for an exemplary embodiment in the present specification;

[0060] Figure 8 , Figure 9 and Figure 10 A structural diagram of an online human-machine dialogue device provided for an exemplary embodiment in the present specification;

[0061] Figure 11 A flow diagram of a smart dialogue robot training device provided for an exemplary embodiment in the present specification;

[0062] Figure 12 A structural diagram of an electronic device provided for an exemplary embodiment in the present specification. DETAILED DESCRIPTION

[0063] With the rapid development of artificial intelligence and mobile internet technology, intelligent dialogue robots have gradually become an important tool for providing human-computer interaction in various applications. Intelligent dialogue robots can interact with users in dialogue, and their mission is to help users solve problems or chat, which has been widely used in various fields such as medical health. Taking a health intelligent dialogue robot (such as a health manager provided on some applications) as an example, it is committed to helping users solve various problems before, during and after medical treatment, covering health consultation, disease screening, medical advice, medication reminders, rehabilitation guidance and other scenarios, which significantly improves the user's health management efficiency and life health level. At present, health intelligent dialogue robots or other intelligent dialogue robots (such as customer service robots, chat robots) mainly support text dialogue as the interaction mode. For example, users describe their health status by inputting text, and health intelligent dialogue robots can use their language models to generate corresponding medical diagnoses. With the development of voice technology, in recent years, more and more intelligent dialogue robots have added the ability of voice dialogue interaction, supporting users to interact with intelligent dialogue robots through voice. Voice dialogue interaction is more in line with the user's dialogue habits, especially for users who are not used to typing, which can bring them a better experience. Then, in the current voice dialogue interaction, the user's reply is mainly based on the text content of the user's voice dialogue.

[0064] Figure 1 The existing technical architecture for implementing online human-voice dialogue interaction is exemplarily shown in the following figure. As shown in the figure, Figure 1 The existing online human-voice dialogue interaction is mainly implemented based on a three-stage pipeline. Specifically, when the user inputs through voice, automatic speech recognition (ASR) is first performed on the user's input voice to obtain the text representation of the user's voice. Then, based on the text representation of the user's voice, text dialogue is performed to obtain the reply of the intelligent dialogue robot. Finally, the text content of the reply of the intelligent dialogue robot is synthesized into corresponding reply voice through text-to-speech (TTS), and the reply voice is played to the user.

[0065] The above Figure 1The online human-computer voice dialogue interaction mode shown may have relatively good interaction effects in some application scenarios, but the user experience is not very good in the medical scenario, which mainly reflects that: first, the user's voice is converted into a text statement after automatic speech recognition, and the sound information other than the text content is lost, such as hoarse voice of the user and excited emotions of the user, and these sound information is an important basis for diagnosis; second, automatic speech recognition (ASR), intelligent text dialogue (for text reply generation), and text-to-speech synthesis (TTS) need to be executed in series, and in the case of larger model parameter quantity (such as the use of a speech model (LLM) in the intelligent dialogue part), the execution time-consuming time will be relatively long (such as more than 3 seconds), which leads to a long response time, so that the user has to wait for a long time, and the experience effect is relatively poor.

[0066] To solve the above problems, the embodiment of the present specification provides a solution, in which a two-stage online human-computer voice dialogue interaction is proposed. In this two-stage online human-computer voice dialogue interaction, an automatic speech recognition (ASR) module and a text dialogue model (such as a language model (LLM)) are fused to form a new voice dialogue model. This solution not only retains the sound information (such as tone and emotion) in the user's voice, but also reduces processing delay, shortens response time, and improves user experience.

[0067] In order for those skilled in the art to better understand the technical solutions in the present specification, the technical solutions in the embodiments of the present specification will be described clearly and completely in conjunction with the accompanying drawings in the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present specification, not all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present specification.

[0068] It should be noted that only parts related to the technical solutions are shown in the drawings for the convenience of description. The embodiments in the specification and the features in the embodiments can be combined with each other without conflict. In addition, the terms "first", "second", "third" and the like in the embodiments of the specification are only used for information differentiation, and do not have any limiting effect. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, product or equipment. Without more limitation, it does not exclude the presence of other same or equivalent elements in the process, method, product or equipment including the elements. In addition, in the specification, unless explicitly stated, "reception and transmission of data" is not necessarily direct reception and transmission, but can be indirect reception and transmission. For example, A receives data sent by B, which can be understood as A directly receiving data sent by B, or A indirectly receiving data sent by B through C and other subjects; similarly, B sends data to A, which can be understood as B directly sending data to A, or B indirectly sending data to A through C and other subjects. Here, C can be one subject, or two or more subjects.

[0069] In addition, it should be noted that the specification uses specific words to describe the embodiments of the specification. As "one embodiment", "an embodiment", and / or "some embodiments" means a certain feature, structure or characteristic related to at least one embodiment of the specification. Therefore, it should be emphasized and noted that the "an embodiment" or "one embodiment" or "one alternative embodiment" mentioned in different places in the specification does not necessarily refer to the same embodiment. In addition, the skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples without conflict. Although one or more embodiments of the specification provide method steps as described in the embodiments or flowcharts, it can be understood that the order of steps listed in the embodiments or flowcharts is only one of the many execution orders, and does not represent the only execution order. Therefore, when the method steps are involved in the claims, the variation adjustment of the order of the steps or the parallel between the steps is also within the scope of protection of the claims.

[0070] Furthermore, it should be noted that the user data obtained by the specification is authorized by the user and does not involve user privacy.

[0071] The embodiments provided by the specification are described below in conjunction with the drawings.

[0072] First, the glossary involved in the embodiments of the present specification is explained. It can be understood that the explanation is for a clearer understanding of the embodiments of the present specification, and does not necessarily constitute a limitation on the embodiments of the present specification.

[0073] Automatic Speech Recognition (ASR): used to recognize human speech as text, i.e., convert human speech signals into text, which involves multiple processes including speech signal acquisition, feature extraction, acoustic model, language model, and decoding, etc.

[0074] Text-to-Speech (TTS): used to convert text information into natural flow field voice output.

[0075] Pre-set model, which is a well-trained artificial intelligence model. In the specification, the pre-set model is a voice dialogue model, which integrates the functions of automatic speech recognition (ASR) and text dialogue. The text dialogue function can be realized through a language model (LLM). That is, in an example, the voice dialogue model is mainly realized by combining automatic speech recognition (ASR) and language model (LLM). The language model (LLM) is an artificial intelligence model, which is a key component in natural language processing (NLP) technology. LLM is based on the Transformer architecture, which can understand and generate high-quality human language text, and is used for dialogue content generation in dialogue interaction systems. It is worth noting here that the embodiments of the present specification do not limit the number of parameters supported by the pre-set model, aiming to meet the actual application requirements.

[0076] Linguistic Tokenizer: used to extract structured high-level representations in speech, specifically, to extract structured information at the linguistic level in speech, such as phonemes, syllables, tones, etc.

[0077] Semantic Tokenizer: which is designed to encode the semantics and coarse-grained acoustic features in speech.

[0078] Speech Decoder: in text-to-speech synthesis (TTS), it is used to convert the encoding of text (such as the features output by the language model) into the corresponding speech waveform. For example: the word "hello"—> encoded into speech features—> speech decoder generates speech waveform.

[0079] The technical solutions provided by the embodiments of the present specification are based on Figure 2A andFigure 2B The technical architecture is implemented in the system as shown in the following figure. Figure 2A and Figure 2B As shown in the figure, the technical architecture is a two-stage online human-machine voice dialogue architecture, one stage of which is to process user voice by using a voice dialogue model to generate corresponding response text, and the other stage is to convert the response text into response voice by a TTS module and play it to the user.

[0080] Specifically, the voice dialogue model is obtained by fusing an ASR module and a text dialogue model. The input of the voice dialogue model is the dialogue voice of the user, and the output is the response text for responding to the dialogue voice of the user.

[0081] Among them, for the fused voice dialogue model, since the text dialogue model in it needs to combine the speech features of the user's input dialogue voice to output the text dialogue result (specifically the response text), in the following description, the text dialogue model is referred to as a Speech To Text (SST) model. The input of the SST model has multiple modal data. In specific implementation, the input of the SST model at least includes the following two modal data: speech modal data and text modal data. The speech modal data refers to the speech features of the user's dialogue voice, and the text modal data refers to the text content of the user's dialogue voice. Therefore, the SST model is actually a multi-modal dialogue model. The specific input of the SST model is described in detail below, and will not be described in detail here.

[0082] In this scheme, in addition to including the ASR module and the SST model (constructed based on the LLM model) in the voice dialogue model, two Tokenizers are also included. The two Tokenizers include a Linguistic Tokenizer and a Semantic Tokenizer. The Linguistic Tokenizer is used to extract structured linguistic features (high-level representations) in the voice, and the Semantic Tokenizer is used to encode semantic and coarse-grained acoustic features in the voice. For details of the two Tokenizers, please refer to the related content in the aforementioned term explanation section.

[0083] As mentioned in Figure 2B , the input of the STT model in the voice dialogue model includes the following two parts:

[0084] 1) Text features of the interactive text, which are input in the form of text. The interactive text includes the dialogue text corresponding to the current round of voice dialogue user input dialogue voice, and the historical interactive text generated by the historical round of voice dialogue (including the dialogue text corresponding to the user's historical dialogue voice, and can also include the historical reply text given by the intelligent dialogue robot). The dialogue voice input by the user can be converted into corresponding dialogue text by the ASR module.

[0085] 2) Speech features. Speech features are obtained by feature extraction of the user's dialogue voice by a language tokenizer and a semantic analyzer.

[0086] Further, the above interactive text (specifically the text features of the interactive text) and speech features are combined and input into the STT model, and the execution of the STT model generates encoded features corresponding to the reply text required to answer the user's dialogue voice. Then, the encoded features corresponding to the reply text output by the SST model are input into the TTS model, and the corresponding reply voice can be generated by the TTS model. In specific implementation, the TTS model can call the speech decoder therein to process the encoded features of the reply text, thereby generating the corresponding reply voice. The speech decoder described herein is a key component in the TTS model, which can convert the encoded features of the text (such as the features output by the language model) into corresponding speech waveforms.

[0087] The overall training process of the technical architecture can be, for example: first, based on a trained language model (LLM), the tokenizer required for speech feature extraction is trained using speech-text alignment training based on freezing the parameters of the language model (LLM); then, the tokenizer and the speech decoder are further trained using the data of the ASR / TTS task, for example, the tokenizer can be trained using the data of the ASR task to enable the tokenizer to learn the coding ability (i.e., feature extraction ability) of the speech, and the speech decoder can be trained using the data of the TTS task to enable the speech decoder to learn the decoding ability (i.e., the ability to convert the coding of the text into speech waveforms) of the text; finally, the full link is fine-tuned using voice dialogue data (such as user voice input + system voice reply).

[0088] In the technical architecture, a voice dialogue model is realized by combining the ASR module with the multi-modal dialogue model, so that the voice dialogue model can learn the reply conditions for different user voice, emotion, environmental sound, etc., and can reflect the recognition and empathy of the model to the user details, which is beneficial to improve the user experience.

[0089] Although, for example, Figure 1The problem of missing a large amount of audio details in the three-stage online human-machine voice dialogue interaction shown can also be solved in other ways, such as by customizing a voice feature extraction module to alleviate, such as emotion recognition, environment recognition, etc., but this way on the one hand will introduce more models and have accuracy loss problems, on the other hand, it needs additional models or rules to reflect the influence of the extracted features on the subsequent text generation, which is difficult to data-driven and long-term optimization effect is more difficult. In addition to the above way, a fully end-to-end speech model can also be used to implement online human-machine voice dialogue interaction. The fully end-to-end speech model can couple ASR, LLM and TTS modules to realize speech from input speech to dialogue answer. This speech model coupled with ASR, LLM and TTS is theoretically feasible, but it needs to rely on a large amount of supervised training, and the parameter quantity of LLM will also slow down the speed of TTS, which is easy to cause the use to have a high response delay, and it is difficult to be well applied.

[0090] Therefore, in the technical architecture provided in the present specification, a voice dialogue model is implemented by fusing an ASR module and a multi-modal dialogue model. This technical architecture has good scalability and optimizability, and can continuously improve the architecture performance and effect through data training and model iteration, and has a wide range of applications and strong competitive advantage in the market.

[0091] The technical architecture mentioned above is implemented based on a server and a client. As shown in Figure 1 The voice dialogue model and the TTS function in the technical architecture are implemented on the server. The playback function of the reply voice and the collection function of the user voice can be implemented on the client. The server can be a server, a server cluster, a virtual server, or a cloud, etc. The client can be, but is not limited to, a smart phone, a smart wearable device, a tablet computer, a notebook computer, a desktop computer, etc. The server provides corresponding function services for the client, such as intelligent dialogue, etc. The user can initiate an online human-machine voice dialogue through a browser, an application (APP), a web application H5 (HyperText Markup Language 5, the fifth generation of HTML, HyperText Markup Language), a light application (also known as a small program, a lightweight application program), or a cloud application on the client. The online human-machine voice dialogue will establish an online human-machine voice dialogue interaction between the client and the server after accessing the intelligent dialogue robot on the server.

[0092] Therefore, Figure 3 It is also shown that the online human-machine dialogue system (also referred to as a service system) provided in an embodiment of the present specification includes a client 200 and a server 100. Wherein,

[0093] The client 200 is used for collecting dialogue voice input by a user in an online man-machine voice dialogue and sending the dialogue voice to the server;

[0094] The server 100 is deployed with a preset model (i.e., the voice dialogue model mentioned above) and is used for inputting the dialogue voice into the preset model and executing the following by the preset model: performing feature extraction on the dialogue voice to obtain voice features of the dialogue voice; generating the reply text based on associated information of the online man-machine voice dialogue and the voice features; wherein the associated information comprises dialogue text obtained by the preset model performing voice recognition on the dialogue voice.

[0095] The server 100 is further used for converting the reply text into reply voice and sending the reply voice to the client.

[0096] The client 200 is further used for playing the reply voice to the user.

[0097] The online man-machine voice dialogue interaction function mentioned above can be integrated in an intelligent dialogue robot in any field. The intelligent dialogue robot is deployed on the server, the user can enter a service interface for man-machine interaction provided by the intelligent dialogue robot through the client, and the online man-machine voice dialogue interaction is started through the service interface. The preset model (i.e., the voice dialogue model) and the TTS model mentioned above are main parts of the architecture of the intelligent dialogue robot.

[0098] The functions of the server 100 and the client 200 and the starting of the online man-machine voice dialogue interaction will be described in detail in the following method embodiment, and will not be described in detail here.

[0099] The technical solutions provided in the specification will be described below in the form of a method embodiment.

[0100] Figure 4 A flowchart of an online man-machine dialogue method provided by one embodiment of the specification is shown. The execution subject of the method is the server in the system mentioned above. Referring to Figure 4 The online man-machine dialogue method comprises the following steps:

[0101] 102, obtaining dialogue voice input by a user in an online man-machine voice dialogue;

[0102] 104, inputting the dialogue voice into a preset model and executing the following by the preset model: performing feature extraction on the dialogue voice to obtain voice features of the dialogue voice; generating the reply text based on associated information of the online man-machine voice dialogue and the voice features; wherein the associated information comprises dialogue text obtained by the preset model performing voice recognition on the dialogue voice.

[0103] 106、outputting the reply text to the user.

[0104] In the embodiment, the online human-computer voice conversation is triggered and started by the user through a service interface for human-computer interaction displayed on the client, and the triggering and starting manner can be but is not limited to triggering through operating a corresponding control, inputting voice, inputting text, etc. The service interface is provided by a corresponding intelligent conversation robot, and specifically, the service interface can be a human-computer interaction interface supporting traditional text conversation or graphic-text conversation. Thus, in the specification, the intelligent conversation robot is to support online human-computer voice conversation interaction, and the intelligent conversation robot can be an intelligent conversation assistant (also referred to as an intelligent conversation system) in any field, such as a health service (including medical health service) field, an e-commerce customer service field, etc.

[0105] For example, taking an intelligent health manager (an intelligent conversation assistant for providing health services) as an example, the user enters a service interface for human-computer interaction provided by the intelligent health manager through the client 200, and in general, the service interface by default supports an interaction manner of text conversation or graphic-text conversation. As shown in Figure 3 , the user can trigger and start online human-computer voice conversation interaction by operating a "call" control 210 on the service interface 21, or can trigger and start online human-computer voice conversation interaction by inputting a voice conversation starting instruction (such as please start voice conversation) in an input box 212 provided by the service interface 21. Of course, in other examples, the voice conversation starting instruction can also be input in other manners, such as voice input, for example, operating a "voice input" control 213 to input the voice of "please start voice conversation". The interface for online human-computer voice conversation interaction is as shown in Figure 3 , for example. After starting online human-computer voice conversation interaction, the user and the intelligent health manager will perform online human-computer conversation interaction in the manner of voice call, and in the interaction process, the client 200 will collect the conversation voice input by the user through a sound pickup device thereon and send the conversation voice to the server 100 for processing.

[0106] Based on the above content, the step 102 "obtaining the conversation voice input by the user in the online human-computer voice conversation" comprises:

[0107] 1021, receiving the conversation voice sent by the client;

[0108] The conversation voice is collected by the client through a sound pickup device, and the sound pickup device can be but is not limited to a microphone.

[0109] After the server receives the dialogue voice initiated by the client, a preset model deployed thereon is called to process the dialogue voice. The preset model has functions of speech recognition, speech feature extraction and text generation. The speech feature extraction function is mainly used to extract corresponding language features and acoustic features from the received dialogue voice. The speech recognition function is used to perform speech recognition on the received dialogue voice to obtain corresponding dialogue text content. The text generation function is used to generate corresponding reply text for the received dialogue voice.

[0110] Based on the above, in a specific implementation, the "the preset model performs feature extraction on the dialogue voice to obtain speech features of the dialogue voice" in 104 can specifically include:

[0111] 1042, extract language features from the dialogue voice, the language features reflecting pronunciation structure and language rules of the dialogue voice;

[0112] 1044, extract acoustic features from the dialogue voice, the acoustic features reflecting auditory perception attributes of the dialogue voice.

[0113] In implementation, the preset model includes a first segmenter and a second segmenter. The first segmenter is, for example, a language segmenter shown in Figure 2B , and the second segmenter is, for example, a semantic segmenter shown in Figure 2B . The first segmenter is used to extract corresponding language features from the dialogue voice. The second segmenter is used to extract corresponding acoustic features from the dialogue voice.

[0114] That is, the implementation of "extracting language features from the dialogue voice" in 1042 can include:

[0115] inputting the dialogue voice into the first segmenter, and executing the first segmenter to output the language features.

[0116] The implementation of "extracting acoustic features from the dialogue voice" in 1044 can include:

[0117] inputting the dialogue voice into the second segmenter, and executing the second segmenter to output the acoustic features.

[0118] In the above, the language features include phoneme-level features and linguistic features. A phoneme is the smallest unit of sound in a language, and thus in this specification, phonemic features directly reflect the pronunciation structure of the dialogue speech, such as / p / , / b / , / t / , and the like. The linguistic features include syntax, part of speech, stress, word boundary, word pronunciation duration, sentence pause duration, and the like, which directly reflect the language rules of the dialogue speech.

[0119] In addition, the acoustic features include features that can reflect the perceptual attributes of the dialogue speech, such as emotion, timbre (e.g., hoarseness), prosody, intonation, speech rate, and tone of voice. In addition, the acoustic features can also include features that can reflect the physical attributes of the dialogue speech, such as the spectral and energy characteristics of the dialogue speech, such as fundamental frequency, amplitude, short-time energy, zero-crossing rate, spectrum, mel spectrum, MFCC, and the like.

[0120] According to the speech features of the dialogue speech and the corresponding dialogue text, the preset model can generate corresponding reply text for responding to the user. Of course, in other examples, the reply text can also be further generated in combination with the interactive text of the historical round of voice dialogue between the user and the intelligent dialogue robot before the current round of voice dialogue.

[0121] Thus, in this embodiment, the associated information of the online human-machine voice dialogue in 104 includes interactive text generated by at least one round of voice dialogue between the user and the intelligent dialogue robot. The interactive text includes the dialogue text corresponding to the user in the current round of voice dialogue, and can also include historical interactive text (including dialogue text content corresponding to the user input historical dialogue speech and reply content (historical reply text) replied by the intelligent dialogue robot to the dialogue text) of the user and the intelligent dialogue robot in the historical round (such as the previous round). Accordingly, the "the preset model generates reply text based on the associated information of the online human-machine voice dialogue and the speech features" in 104 can include:

[0122] 1046, generate the reply text based on the interactive text generated by the at least one round of voice dialogue and the speech features.

[0123] The interactive text generated by the at least one round of voice dialogue can be used to help understand the dialogue context of the current round of voice dialogue, improve the quality of the reply, and thus better ensure the response to the user.

[0124] Further, the associated information of the online dialogue can also include other information, such as the relevant information of the user.

[0125] Exemplarily, in the medical health scenario, the relevant information of the user can include but is not limited to at least one of the following: historical diagnosis records, medical records, allergy history, family medical history, etc. Combining these information can improve the accuracy of diagnosis, so that the medical advice given in the generated conversation text can be more in line with the actual situation of the user (such as avoiding recommending drugs containing ingredients that the user is allergic to), and the user experience can be enhanced.

[0126] In addition, in the e-commerce customer service scenario, the relevant information of the user can include but is not limited to: identity information (such as service level, such as VIP level), portrait information (such as order information, purchase records, historical behavior (such as consultation, complaint, abnormal return record, geographical location). Among them, when generating the conversation text, the conversation generated in combination with the identity information of the user (such as VIP level), historical behavior (such as historical complaint record, return preference) can make the user feel "understood", "valued" and the like; the conversation generated in combination with the historical behavior of the user (such as frequent complaints, abnormal returns) and the geographical location (such as high-risk areas) will be more cautious in content, which can avoid over-promising, or make the conversation in the conversation give more appropriate solutions to the user's problems, and the like.

[0127] That is, one specific implementation of the above step 1046 can include:

[0128] 10462, generating the conversation text based on the interaction text generated by the at least one round of voice dialogue, the relevant information of the user, and the voice feature.

[0129] The foregoing related step content of "generating conversation text based on the relevant information of the online human-computer voice dialogue and the voice feature" in 104 is realized by calling the STT model in the preset model. In addition, the preset model further includes a voice recognition module, which is used to convert the dialogue voice of the user into corresponding dialogue text. Therefore, based on this, in a specific implementation scheme, the "preset model performs voice recognition on the dialogue voice to obtain dialogue text" included in the above 104 can include:

[0130] 1048, the preset model uses the built-in voice recognition module to perform voice recognition on the dialogue voice to obtain the dialogue text.

[0131] In addition, "generating conversation text based on the relevant information of the online human-computer voice dialogue and the voice feature" included in the above 104 can include:

[0132] 10410、The preset model inputs the associated information and the speech feature into a multi-modal dialogue model built in itself, and executes the multi-modal dialogue model to output the reply text.

[0133] The speech recognition module is an ASR module, and the multi-modal dialogue model can be an SST model as shown in FIG. Figure 2B The STT model is built based on a language model (LLM).

[0134] In summary, the scheme provided in the embodiment not only combines the text content of the dialogue speech, but also combines the speech feature of the dialogue speech when generating the reply to the dialogue speech of the user, which can effectively improve the quality of the reply, especially in the medical health scenario, the speech feature of the user is an important diagnostic basis, and combining the speech feature of the user can improve the accuracy and reliability of medical diagnosis, so as to effectively ensure that the medical advice given in the reply is more suitable for the actual situation of the user. The obtaining of the dialogue text and the speech feature corresponding to the dialogue speech and the generation of the reply text are realized through a preset model, and the speech recognition module and the STT model (used for generating the reply text) are integrated in the preset model. The design of such a preset model makes the entire online human-computer voice dialogue interaction realized in two stages, which can reduce the processing delay, improve the response speed, and enhance the user experience. Specifically, the preset model (a speech dialogue model) is designed by integrating the speech recognition module and the STT model, which can reduce the serial execution time of the speech recognition (ASR) module, the text dialogue model (such as LLM), and the speech synthesis (TTS) model in the entire online human-computer voice dialogue interaction, effectively shortening the time for the user to wait for a response, thereby effectively improving the overall voice dialogue interaction experience. In particular, under the background of increasing model parameter quantity, the scheme can provide a more smooth and natural voice interaction experience while ensuring performance.

[0135] The speech synthesis (TTS) model is used to convert the reply text output by the preset model into corresponding reply speech to report to the user.

[0136] That is, in 106, the server converts the reply text into reply speech through the TTS model, and then sends the reply speech to the client to report to the user. In specific implementation, the TTS model can use the speech decoder in it to convert the reply text into corresponding dialogue speech. At the same time when the reply speech is sent to the client, the reply text can also be sent to the client to be displayed on the client interface. Thus, in a specific implementation scheme, the 106 "outputs the reply text to the user" can include:

[0137] 1062, performing speech synthesis on the reply text to generate reply speech;

[0138] 1064、send the reply voice and the reply text to the client, and display the reply text while playing the reply voice by the client.

[0139] In the above, considering that in some fields, such as the medical health field, there are often some special terms, if the reply voice contains these special terms, simply playing the reply voice may cause the user to have difficulty understanding. Therefore, the corresponding reply text is displayed while playing the reply voice, which can facilitate the user to understand the reply voice through text vision. The reply text may be displayed on the interface 22 of the online human-computer voice dialogue interaction shown in Figure 3

[0140] The present specification also provides another online human-computer dialogue method, and the execution subject of the method is the client in the foregoing system. Specifically, as shown in Figure 5

[0141] 202、collect dialogue voice input by the user in the online human-computer voice dialogue;

[0142] 204、send the dialogue voice to the server to trigger the server to input the dialogue voice into a preset model, and the preset model performs: extracting features of the dialogue voice to obtain voice features of the dialogue voice; converting the dialogue voice into dialogue text; and generating the reply text based on the dialogue text and the voice features;

[0143] 206、receive the reply voice returned by the server; wherein the reply voice is generated by performing voice synthesis on the reply text.

[0144] 208、play the reply voice to the user.

[0145] The specific implementation of each step of the foregoing embodiment can be referred to the related content in other embodiments, and will not be repeated here. In addition, the method provided in the embodiment can also include some steps disclosed in other embodiments, and can also be referred to the related content in other embodiments, and will not be repeated here.

[0146] The present specification also provides another online human-computer dialogue method, and the execution subject of the method is the client in the foregoing system. Specifically, as shown in Figure 6 the online human-computer dialogue method includes the following steps:​​

[0147] 302、displaying an interface for health service;

[0148] 304、in response to a voice dialogue triggered through the interface, starting an online man-machine voice dialogue;

[0149] 306、in the online man-machine voice dialogue, collecting a dialogue voice input by a user;

[0150] 308、inputting the dialogue voice into a preset model, and performing the following by the preset model: extracting features of the dialogue voice to obtain voice features of the dialogue voice; and generating a response text based on associated information of the online man-machine voice dialogue and the voice features; wherein the associated information includes interactive text generated by at least one round of voice dialogue between a user and an intelligent dialogue robot, and the interactive text includes dialogue text obtained by performing voice recognition on dialogue voice of the user in the current round of voice dialogue by the preset model;

[0151] 310、outputting the response text to the user.

[0152] In the above, the interface for health service can be an interface displayed when a user opens a health-related intelligent dialogue robot, such as the service interface 21 shown in the embodiment. Figure 3

[0153] For specific implementation of each step in the above embodiment, refer to related content in other embodiments, which will not be repeated here. In addition, the method provided in the embodiment can also include some steps disclosed in other embodiments, and refer to related content in other embodiments, which will not be repeated here.

[0154] The present specification also provides a method for training an intelligent dialogue robot, and the architecture of the intelligent dialogue robot is deployed on a server, so that the training method is implemented by the server. The intelligent dialogue robot includes a preset model and a speech synthesis model; the preset model includes a speech recognition module, a word segmenter, and a trained multi-modal dialogue model. For specific details of each module / function included in the preset model, refer to related content in other embodiments. Specifically, refer to the preset model 110 shown in the embodiment. Figure 7 The training method includes the following steps:

[0155] 400、parameter freezing of the trained multi-modal dialogue model

[0156] 402、using a first training sample set to train the word segmenter to obtain the word segmenter after the first training; the first training sample set includes a plurality of first sample voices and text content corresponding to the first sample voices;

[0157] ​404. Using the second training sample set, the word segmenter after the first training is further trained to obtain the word segmenter after the second training; the second sample set contains multiple second sample speech and the speech features corresponding to the second sample speech;

[0158] 406. Using a third training sample set, the speech synthesis model is trained to obtain the trained speech synthesis model; the third training sample set contains multiple sample texts and the speech corresponding to the sample texts.

[0159] 408. Using the fourth training sample set, the speech recognition module, the word segmenter, the multimodal dialogue model, and the speech synthesis model are jointly fine-tuned and trained to obtain a trained intelligent dialogue robot.

[0160] The fourth training sample set includes multiple third sample speech and fourth sample speech used to respond to the third sample speech. During the joint fine-tuning training process, the speech recognition module is used to convert the third sample speech into corresponding text content, the word segmenter is used to extract the speech features of the third sample speech, and the multimodal dialogue model generates echo text based on the speech features output by the word segmenter and the text content output by the speech recognition module. The speech synthesis model is used to convert the echo text output by the multimodal dialogue model into echo speech. Based on the loss value of the echo speech and the corresponding fourth sample speech, the parameters of the speech recognition module, the word segmenter, the multimodal dialogue model, and the speech synthesis model are jointly fine-tuned.

[0161] In section 402 above, the first training sample set is used to train the word segmenter for speech-text alignment. The goal of training the word segmenter is to extract structured speech features from the first sample speech in the first training sample set and align them with the corresponding text in the first training sample set, so that the word segmenter can learn speech features from speech without text supervision.

[0162] For example, if the word segmenter includes a first word segmenter (a language word segmenter), the training process for the first word segmenter using the first training sample set could be as follows: input the first sample speech from the first training sample set into the first word segmenter; the first word segmenter outputs the language features extracted from the first sample speech (including phoneme-level features (phoneme sequences), speech rule features, etc.); then analyze the differences between the language features of the first sample speech output by the first word segmenter and the language features in the text content corresponding to the first sample speech in the first training sample set. For example, analyze the differences between the phoneme sequence of the first sample speech output by the first word segmenter and the phoneme sequence in the text content corresponding to the first sample speech in the first training sample set, and optimize the parameters of the word segmenter based on these differences.

[0163] It can be seen from the above that the training of the word segmenter aims to make the speech features output by the word segmenter consistent with the speech features in the corresponding text, for example, to make the phoneme sequence output by the first word segmenter consistent with the phoneme sequence in the corresponding text.

[0164] In the above 404, when the word segmenter is trained using the second training sample set, the training process can be, for example, inputting the second sample speech included in the second training sample set into the word segmenter to obtain the speech features of the second sample speech output by the word segmenter, then calculating the loss value between the speech features output by the word segmenter and the corresponding speech features in the second training sample set, and optimizing the parameters of the word segmenter based on the loss value. The second sample speech included in the second training sample set is the data of the ASR task, and the main purpose of this training of the word segmenter is to enable the word segmenter to learn the coding ability (i.e., the speech feature extraction ability) of the speech.

[0165] In the above 406, when the speech synthesis (TTS) model is trained using the third training sample set, the training process can be, for example, inputting the sample text in the third training sample set into the speech synthesis model, executing the speech synthesis model to output the speech corresponding to the sample text, then calculating the loss value between the speech output by the speech synthesis model and the speech corresponding to the third sample text in the third training sample set, and optimizing the parameters of the speech synthesis model according to the loss value.

[0166] In the above 408, the fourth training sample set is online voice dialogue interaction data. Using the fourth training sample set, the speech recognition module, the word segmenter, the multi-modal dialogue module, and the speech synthesis model in the intelligent dialogue robot are trained end-to-end for full-link joint fine-tuning, which can enable the parameters of each module / model in the entire process from speech input to speech output to be optimized cooperatively, and improve the overall response performance of the intelligent dialogue robot.

[0167] The specific implementation of each step described above in this embodiment can be referred to related content in other embodiments, which will not be described in detail here. In addition, the method provided in this embodiment can also include some steps disclosed in other embodiments, which can also be referred to related content in other embodiments, which will not be described here.

[0168] The above is described in combination with Figure 4~7Certain embodiments are described in this specification. It is understood that other embodiments can be practiced, which are within the scope of the following claims, and that logical and / or practical limitations can be used in some circumstances. In some instances, the acts or steps can be performed in an order other than the order in which they are described. Additionally, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0169] The device embodiments corresponding to the method embodiments provided in this specification are described below.

[0170] Figure 8 A structural schematic diagram of an online human-computer dialogue device provided by an example embodiment in this specification is shown. As shown in the figure, Figure 8 The device includes an acquisition module 52, an execution module 54, and an output module 56. Among them,

[0171] The acquisition module 52 is configured to acquire dialogue speech input by a user in an online human-computer voice dialogue.

[0172] The execution module 54 is configured to input the dialogue speech into a preset model, and execute the following by the preset model: performing feature extraction on the dialogue speech to obtain speech features of the dialogue speech; and generating a reply text based on associated information of the online human-computer voice dialogue and the speech features; wherein the associated information includes dialogue text obtained by the preset model performing speech recognition on the dialogue speech.

[0173] The output module is configured to output the reply text to the user.

[0174] In an implementation, the execution module 54, when used to perform feature extraction on the dialogue speech to obtain speech features of the dialogue speech, is specifically configured to: extract language features from the dialogue speech; the language features can reflect the pronunciation structure and language rules of the dialogue speech; and extract acoustic features from the dialogue speech; the acoustic features can reflect the auditory perception attributes of the dialogue speech.

[0175] In an implementation, the preset model includes a first word segmenter and a second word segmenter. When the execution module 54 is used to extract language features from the dialogue speech, it is specifically configured to: input the dialogue speech into the first word segmenter, and execute the first word segmenter to output the language features. When the execution module 54 is used to extract acoustic features from the dialogue speech, it is specifically configured to: input the dialogue speech into the second word segmenter, and execute the second word segmenter to output the acoustic features.

[0176] In an embodiment, the phoneme-level feature included in the language feature reflects pronunciation structure of the dialogue voice. The acoustic feature includes at least one of intonation, speech rate, emotion, timbre, prosody of the dialogue voice.

[0177] In an embodiment, the association information includes interactive text generated by at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactive text includes the dialogue text corresponding to the user in the current round of voice dialogue. The execution module 54 is configured to generate the reply text based on the interactive text generated by the at least one round of voice dialogue and the voice feature.

[0178] In an embodiment, the association information further includes relevant information of the user. The execution module 54 is configured to generate the reply text based on the interactive text generated by the at least one round of voice dialogue, the relevant information of the user, and the voice feature.

[0179] In an embodiment, the preset model includes a speech recognition module and a multi-modal dialogue model. The execution module 54 is configured to perform speech recognition on the dialogue voice to obtain dialogue text by using the speech recognition module. The execution module 54 is configured to input the association information and the voice feature into the multi-modal dialogue model to output the reply text, wherein the multi-modal dialogue model is constructed based on a language model.

[0180] Figure 9 A structural schematic diagram of an online human-computer dialogue device provided by another example embodiment of the present specification is shown. As shown in FIG. 4, the online human-computer dialogue device includes a voice dialogue module 41, an association information acquisition module 42, a voice feature extraction module 43, an execution module 44, and a multi-modal dialogue model 45. Figure 9As shown, the device comprises: a collection module 62, a sending module 64, a receiving module 66, and a broadcast module 68. The collection module 62 is configured to collect dialogue voice input by a user in an online man-machine voice dialogue. The sending module 64 is configured to send the dialogue voice to a server to trigger the server to input the dialogue voice into a preset model, and the preset model is configured to perform the following operations: extracting features of the dialogue voice to obtain voice features of the dialogue voice; converting the dialogue voice into dialogue text; and generating the reply text based on the dialogue text and the voice features. The receiving module 66 is configured to receive reply voice returned by the server, wherein the reply voice is generated by performing voice synthesis on the reply text. The broadcast module 68 is configured to broadcast the reply voice to the user.

[0181] Figure 10 As shown in the structural schematic diagram of an online man-machine dialogue device provided by another example embodiment of the present specification. As shown in the structural schematic diagram of an online man-machine dialogue device provided by another example embodiment of the present specification. Figure 10 As shown, the device comprises: a display module 72, a starting module 74, a collection module 76, an execution module 78, and an output module 710. The display module 72 is configured to display an interface for health services. The starting module 74 is configured to start an online man-machine voice dialogue in response to a voice dialogue starting operation triggered through the interface. The collection module 76 is configured to collect dialogue voice input by a user in the online man-machine voice dialogue. The execution module 78 is configured to input the dialogue voice into a preset model, and the preset model is configured to perform the following operations: extracting features of the dialogue voice to obtain voice features of the dialogue voice; and generating the reply text based on associated information of the online man-machine voice dialogue and the voice features, wherein the associated information comprises interactive text generated in at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactive text comprises dialogue text obtained by the preset model performing voice recognition on dialogue voice of the user in the current round of voice dialogue. The output module 710 is configured to output the reply text to the user.

[0182] Figure 11 As shown in the structural schematic diagram of an intelligent dialogue robot training device provided by another example embodiment of the present specification. The intelligent dialogue robot comprises a preset model and a voice synthesis model; the preset model comprises a voice recognition module, a word segmenter, and a trained multi-modal dialogue model. As shown in the structural schematic diagram of an intelligent dialogue robot training device provided by another example embodiment of the present specification. Figure 11As shown, the training apparatus comprises: a freezing module 82, a training module 84. The freezing module 82 is configured to freeze parameters of the trained multi-modal dialogue model. The training module is configured to train the word segmenter based on a first training sample set to obtain the word segmenter after first training; the first training sample set comprises a plurality of first sample speeches and text content corresponding to the first sample speeches; continue training the word segmenter after the first training based on a second training sample set to obtain the word segmenter after the second training; the second training sample set comprises a plurality of second sample speeches and speech features corresponding to the second sample speeches; train the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set comprises a plurality of sample texts and speech corresponding to the sample texts; jointly fine-tune train the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model based on a fourth training sample set to obtain the trained intelligent dialogue robot; wherein the fourth training sample set comprises a plurality of third sample speeches and fourth sample speeches for responding to the third sample speeches; in the process of joint fine-tuning training, the speech recognition module is configured to convert the third sample speeches into corresponding text content, the word segmenter is configured to extract speech features of the third sample speeches, the multi-modal dialogue model is configured to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is configured to convert the response text output by the multi-modal dialogue model into response speech; and parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model are jointly fine-tuned according to a loss value of the response speech and the corresponding fourth sample speech.

[0183] It should be noted that the above-described devices can implement the technical solutions described in the above corresponding method embodiments, and the principles of the implementation of the above modules or units can be referred to the related content in the above corresponding method embodiments, which will not be described in detail here. In addition, for the convenience of description, the above devices are described in various modules or units. Of course, when implementing one or more of the present specification, the functions of each module or unit can be implemented in the same or more software and / or hardware, or the modules implementing the same function can be combined or implemented by a combination of multiple sub-modules or sub-units. The above-described device embodiments are only illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0184] In addition, the embodiment of the present specification also provides an electronic device. As shown in Figure 12As shown, the electronic device 900 includes a memory 91 and a processor 92.

[0185] The memory 91 can be implemented by any type of at least one volatile or nonvolatile memory device or a combination thereof, such as a Static Random-Access Memory (SRAM), an Electrically Erasable Programmable Read Only Memory (EEPROM), an Erasable Programmable Read Only Memory (EPROM), a Programmable Read-Only Memory (PROM), a Read-Only Memory (ROM), a magnetic memory, a flash memory, a magnetic disc or a optical disc. Also, the whole or part of the memory can be integrated with the processor. The memory can include removable and non-removable components.

[0186] The processor 92 can include one or more general processors and / or dedicated processors.

[0187] Also, the memory 91 can include a non-transitory computer readable medium having stored therein executable program instructions 912 (e.g., compiled or non-compiled program logic and / or machine code). The processor 92 is capable of executing the program instructions 912 stored in the memory to implement any of the methods, processes or functions disclosed in the specification and / or drawings. Further, execution of the program instructions 912 by the processor 92 can cause the processor to utilize corresponding data 911.

[0188] For example, the program instructions 912 can include an operating system 9122 (e.g., operating system kernel, device drivers, and / or other modules) and one or more application programs 9121 (e.g., a browser, a social application, or a game application) installed on the electronic device 900. Similarly, the data 911 can include operating system data 9112 and application data 9111. The operating system data 9112 is primarily accessible to the operating system 9122, while the application data 9111 is primarily accessible to the one or more application programs 9121. The application data 9111 can be located in a file system that is visible or hidden to a user of the electronic device 900.

[0189] The application programs 9121 can communicate with the operating system 9122 through one or more application programming interfaces (APIs). These APIs can help the application programs 9121 read and / or write application data, communicate with the communication components, receive or display information on the user interface, etc. Among some terms, the application programs 9121 can be referred to simply as "apps." In addition, the application programs 9121 can be downloaded to the electronic device through one or more online application stores or application markets. However, the application programs 9121 can also be installed on the electronic device 400 through other means, such as through a web browser or a physical interface (e.g., a USB port) on the electronic device 900

[0190] Further, as shown in Figure 12 , the electronic device further includes a communication component 93, a display 94, a power component 95, an audio component 96, a user interface 99, and other components. Figure 12 Some components are only shown schematically in the electronic device 900, and it does not mean that the electronic device 900 only includes Figure 12 the components shown. In addition, Figure 12 the components within the dashed boxes are optional components, not mandatory components, and the specific components can be determined according to the product form of the electronic device 900. The electronic device 900 of the present embodiment can be implemented as a terminal device such as a desktop computer, a notebook computer, a smart phone, or an IOT device, or as a server device such as a general server, a cloud server, or a server array, or as an integrated device of a terminal device and a server device, etc. If the electronic device 900 of the present embodiment is implemented as a terminal device such as a desktop computer, a notebook computer, a smart phone, etc., it can include the components within the dashed boxes in Figure 12 ; if the electronic device 900 of the present embodiment is implemented as a server device such as a general server, a cloud server, or a server array, etc., it can not include the components within the dashed boxes in Figure 12 .

[0191] The communication component 93 is configured to facilitate wired or wireless communication between the device on which the communication component 93 is located and other devices. The device on which the communication component 93 is located can access a wireless network based on a communication standard, such as a 2G, 3G, 4G / LTE, 5G, or the like, or a combination thereof. In an example embodiment, the communication component 93 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In implementations, the communication component 93 includes a communication interface that enables the electronic device 900 to communicate with other devices, access networks, and / or transmission networks through analog or digital modulation. For example, the communication interface can include a chipset and antenna for wirelessly communicating with a radio access network or access point. In addition, the communication interface can be a wired interface, such as an Ethernet, token ring, or USB port, or a wireless interface, such as a Wifi, Bluetooth, Global Positioning System (GPS), or wide area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface can support other forms of physical layer interfaces and standards or proprietary communication protocols. The communication interface can also include multiple physical communication interfaces, such as a Wifi interface, a Bluetooth interface, and a wide area wireless interface.

[0192] The display 94 includes a screen, which can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touch or a slide action, but also detect duration and pressure related to the touch or slide operation.

[0193] The power supply component 95 provides power to the various components of the device on which the power supply component 95 is located. The power supply component 95 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device on which the power supply component 95 is located.

[0194] The audio component 96 can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive an external audio signal when the device on which the audio component is located is in an operational mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in memory or transmitted via the communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0195] The user interface 97 includes receiving user input and providing output to the user. Thus, the user interface 97 can include input components such as a keypad, keyboard, touch- sensitive or presence-sensitive panel, a computer mouse, trackball, joystick, microphone, still camera, and video camera, among others, and output components such as a display screen (which can be combined with a touch-sensitive panel), a CRT, LCD, LED, display using DLP technology, a printer, other kinds of output devices, or any other known or future developed devices for providing output to a user. The user interface 97 can also generate auditory output, such as through a loudspeaker, a speaker jack, audio output ports, audio output devices, headphones, and other known or future developed devices for generating auditory output. In some embodiments, the user interface 97 can include software, circuitry, or other forms of logic that enables the electronic device 900 to transmit and receive data to and from external user input / output devices. Additionally or alternatively, the electronic device 900 can support remote access from other devices, such as through a communications interface or another physical interface (not shown). The user interface 97 can be configured to receive user input, the location and movement of which can be indicated by a pointer or cursor as described herein. The user interface 97 can also be configured as a display device for rendering or displaying a text segment.

[0196] Accordingly, the embodiments of the present specification also provide a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor is enabled to implement each step in the above-mentioned method embodiments. Wherein, the computer readable storage medium includes volatile or non-volatile or their combination, and can be removable or non-removable. Examples of the computer readable storage medium include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store data and be accessed by a computer

[0197] In addition, the embodiments of the present specification also provide a computer readable storage medium, which stores a computer program, wherein when the computer program is executed in a computer, the computer is enabled to execute the method described above. Figure 3 to Figure 5 In addition, the embodiments of the present specification also provide a computer readable storage medium, which stores a computer program, wherein when the computer program is executed in a computer, the computer is enabled to execute the method described above.

[0198] The embodiments of the present disclosure further provide a computer program product comprising computer programs / instructions which, when executed by a processor, implement the method as Figure 3 to Figure 5 described.

[0199] Those skilled in the art should be aware that, in the above one or more examples, the functions described in the embodiments disclosed in the present specification can be implemented in hardware, software, firmware or any combination thereof. When implemented in software, the functions can be stored in a computer readable medium or transmitted as one or more instructions or codes on a computer readable medium.

[0200] The above detailed description of the specific implementation is further detailed for the purpose of the embodiments disclosed in the present specification, technical solutions and beneficial effects, and it should be understood that the above detailed description is only for the specific implementation of the embodiments disclosed in the present specification, and is not used to limit the protection scope of the embodiments disclosed in the present specification. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the embodiments disclosed in the present specification shall be included in the protection scope of the embodiments disclosed in the present specification.

Claims

1. An online human-to-computer dialog method, characterized by, The application is applied to an intelligent dialogue robot comprising a preset model; The preset model comprises a word segmenter, a trained multi-modal dialogue model and a speech recognition module; The method comprises: obtaining dialogue speech input by a user in an online human-computer voice dialogue; inputting the dialogue speech into the preset model, and performing the following by the preset model: extracting features of the dialogue speech by using the word segmenter to obtain speech features of the dialogue speech; inputting associated information of the online human-computer voice dialogue and the speech features into the multi-modal dialogue model to generate response text by the multi-modal dialogue model; wherein the associated information comprises dialogue text obtained by performing speech recognition on the dialogue speech by using the speech recognition module; outputting the response text to the user; The intelligent dialogue robot further comprises a speech synthesis model; The training of the intelligent dialogue robot comprises: parameter freezing of the trained multi-modal dialogue model; training the word segmenter based on a first training sample set to obtain the word segmenter after the first training; the first training sample set comprises a plurality of first sample speeches and text content corresponding to the first sample speeches; continuing to train the word segmenter after the first training based on a second training sample set to obtain the word segmenter after the second training; the second training sample set comprises a plurality of second sample speeches and speech features corresponding to the second sample speeches; training the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set comprises a plurality of sample texts and speech corresponding to the sample texts; joint fine-tuning training of the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model based on a fourth training sample set to obtain the trained intelligent dialogue robot; The fourth training sample set comprises a plurality of third sample speeches and fourth sample speeches for responding to the third sample speeches; in the joint fine-tuning training process, the speech recognition module is used to convert the third sample speeches into corresponding text content, the word segmenter is used to extract speech features of the third sample speeches, the multi-modal dialogue model is used to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is used to convert the response text output by the multi-modal dialogue model into response speech; and the parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model are fine-tuned according to the loss value of the response speech and the corresponding fourth sample speech.

2. The method of claim 1, wherein, extracting features of the dialogue speech to obtain speech features of the dialogue speech comprises: extracting language features from the dialogue speech; the language features can reflect the pronunciation structure and language rules of the dialogue speech; extracting acoustic features from the dialogue speech; the acoustic features can reflect the auditory perception attributes of the dialogue speech.

3. The method of claim 2, wherein, The preset model comprises a first word segmenter and a second word segmenter; and extracting language features from the dialogue voice, including: inputting the dialogue voice into the first word segmenter, and executing the first word segmenter to output the language features; extracting acoustic features from the dialogue voice, including: inputting the dialogue voice into the second word segmenter, and executing the second word segmenter to output the acoustic features.

4. The method of claim 2, wherein, The phoneme-level features contained in the language features reflect the pronunciation structure of the dialogue voice. The acoustic features include at least one of intonation, speech rate, emotion, timbre, and prosody of the dialogue voice.

5. The method according to any one of claims 1 to 4, characterized in that, The association information includes interactive text generated in at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactive text contains the dialogue text corresponding to the user in the current round of voice dialogue. and, based on the association information of the online human-computer voice dialogue and the voice features, generating a reply text, including: based on the interactive text generated in the at least one round of voice dialogue and the voice features, generating the reply text.

6. The method of claim 5, wherein, The association information further includes relevant information of the user. and, based on the interactive text generated in the at least one round of voice dialogue and the voice features, generating the reply text, including: based on the interactive text generated in the at least one round of voice dialogue, the relevant information of the user, and the voice features, generating the reply text.

7. The method according to any one of claims 1 to 4, characterized in that, The multi-modal dialogue model is constructed based on a language model.

8. An online human-to-computer dialog method, characterized by, including: collecting dialogue voice input by the user in online human-computer voice dialogue; sending the dialogue voice to a server to trigger the server to input the dialogue voice into a preset model including a word segmenter, a trained multi-modal dialogue model, and a speech recognition module, so that the preset model performs: using the word segmenter to extract features from the dialogue voice to obtain voice features of the dialogue voice; using the speech recognition module to convert the dialogue voice into dialogue text; inputting the dialogue text and the voice features into the multi-modal dialogue model to execute the multi-modal dialogue model to generate a reply text; receiving the reply voice returned by the server; wherein the reply voice is generated by performing voice synthesis on the reply text using a voice synthesis model; playing the reply voice to the user; wherein the server has an intelligent dialogue robot deployed thereon, and the intelligent dialogue robot includes the preset model and the voice synthesis model; and, the training of the intelligent dialogue robot includes: parameter freezing of the trained multi-modal dialogue model; training the word segmenter based on a first training sample set to obtain the first trained word segmenter; the first training sample set includes a plurality of first sample voices and text content corresponding to the first sample voices; continuing to train the first trained word segmenter based on a second training sample set to obtain the second trained word segmenter; the second training sample set includes a plurality of second sample voices and voice features corresponding to the second sample voices; training the speech synthesis model based on a third training sample set, to obtain the trained speech synthesis model; the third training sample set contains multiple sample texts and speech corresponding to the sample texts; based on a fourth training sample set, jointly fine-tuning training the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model, to obtain the trained intelligent dialogue robot; wherein the fourth training sample set includes multiple third sample speeches and fourth sample speeches for responding to the third sample speeches; during the joint fine-tuning training process, the speech recognition module is used to convert the third sample speeches into corresponding text content, the word segmenter is used to extract the speech features of the third sample speeches, the multi-modal dialogue model is used to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is used to convert the response text output by the multi-modal dialogue model into response speech; and the parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model are jointly fine-tuned according to the loss value of the response speech and the corresponding fourth sample speech.

9. An online human dialog method, characterized by, comprising: displaying an interface for health services; in response to a voice dialogue starting operation triggered through the interface, starting an online man-machine voice dialogue; in the online man-machine voice dialogue, collecting a dialogue speech input by a user; inputting the dialogue speech into a preset model containing a word segmenter, a trained multi-modal dialogue model and a speech recognition module, and performing the following by the preset model: extracting features of the dialogue speech using the word segmenter to obtain speech features of the dialogue speech; inputting associated information of the online man-machine voice dialogue and the speech features into the multi-modal dialogue model to generate response text by the multi-modal dialogue model; wherein the associated information contains interactive text generated in at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactive text contains dialogue text obtained by performing speech recognition on the dialogue speech input by the user in the current round of voice dialogue using the speech recognition module; outputting the response text to the user; wherein the intelligent dialogue robot contains the preset model and a speech synthesis model; and the training of the intelligent dialogue robot comprises: parameter freezing of the trained multi-modal dialogue model; training the word segmenter based on a first training sample set to obtain the first trained word segmenter; the first training sample set contains multiple first sample speeches and text content corresponding to the first sample speeches; continuing to train the first trained word segmenter based on a second training sample set to obtain the second trained word segmenter; the second training sample set contains multiple second sample speeches and speech features corresponding to the second sample speeches; training the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set contains multiple sample texts and speech corresponding to the sample texts; based on a fourth training sample set, the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model are jointly fine-tuned and trained to obtain the trained intelligent dialogue robot; wherein the fourth training sample set includes multiple third sample speeches and fourth sample speeches for responding to the third sample speeches; in the joint fine-tuning training process, the speech recognition module is used to convert the third sample speeches into corresponding text content, the word segmenter is used to extract the speech features of the third sample speeches, the multi-modal dialogue model is used to generate a response text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is used to convert the response text output by the multi-modal dialogue model into a response speech; according to the loss value of the response speech and the corresponding fourth sample speech, the parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model are jointly fine-tuned. 10.A method of training an intelligent conversational robot, the method comprising: The intelligent dialogue robot includes a preset model and a speech synthesis model; the preset model contains a speech recognition module, a word segmenter, and a trained multi-modal dialogue model; The method comprises: freezing the parameters of the trained multi-modal dialogue model; training the word segmenter based on a first training sample set to obtain the first trained word segmenter; the first training sample set contains multiple first sample speeches and text content corresponding to the first sample speeches; continuing to train the first trained word segmenter based on a second training sample set to obtain the second trained word segmenter; the second training sample set contains multiple second sample speeches and speech features corresponding to the second sample speeches; training the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set contains multiple sample texts and speech corresponding to the sample texts; based on a fourth training sample set, the speech recognition module, the word segmenter, the multi-modal dialogue model and the speech synthesis model are jointly fine-tuned and trained to obtain the trained intelligent dialogue robot; The fourth training sample set includes a plurality of third sample voices and fourth sample voices for responding to the third sample voices; in the joint fine-tuning training process, the speech recognition module is configured to convert the third sample voices into corresponding text content, the word segmenter is configured to extract speech features of the third sample voices, and the multi-modal dialogue model is configured to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is configured to convert the response text output by the multi-modal dialogue model into response voice; and parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model are fine-tuned based on loss values of the response voice and corresponding fourth sample voices.

11. An online human-to-computer dialog system, characterized by The method comprises the following steps: The client collects dialogue voice input by a user in an online human-computer voice dialogue, and sends the dialogue voice to a server; The server is deployed with an intelligent dialogue robot, and the intelligent dialogue robot comprises a preset model and a speech synthesis model, wherein the preset model comprises a word segmenter, a trained multi-modal dialogue model, and a speech recognition module; The server inputs the dialogue voice into the preset model, and the preset model performs the following operations: the word segmenter extracts features of the dialogue voice to obtain speech features of the dialogue voice; the multi-modal dialogue model generates response text based on associated information of the online human-computer voice dialogue and the speech features; and the associated information comprises dialogue text obtained by performing speech recognition on the dialogue voice by using the speech recognition module; The server converts the response text into response voice by using the speech synthesis model, and sends the response voice to the client; The client plays back the response voice to the user; The training of the intelligent dialogue robot comprises the following steps: Parameters of the trained multi-modal dialogue model are frozen; The word segmenter is trained based on a first training sample set to obtain the word segmenter after first training; the first training sample set comprises a plurality of first sample voices and text content corresponding to the first sample voices; The word segmenter after first training is further trained based on a second training sample set to obtain the word segmenter after second training; the second training sample set comprises a plurality of second sample voices and speech features corresponding to the second sample voices; The speech synthesis model is trained based on a third training sample set to obtain the trained speech synthesis model; the third training sample set comprises a plurality of sample texts and voice corresponding to the sample texts; The speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model are jointly fine-tuned based on a fourth training sample set to obtain the trained intelligent dialogue robot; and the fourth training sample set comprises a plurality of third sample voices and fourth sample voices for responding to the third sample voices. The fourth training sample set includes a plurality of third sample voices and fourth sample voices for responding to the third sample voices; in the joint fine-tuning training process, the speech recognition module is configured to convert the third sample voices into corresponding text content, the word segmenter is configured to extract speech features of the third sample voices, and the multi-modal dialogue model is configured to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is configured to convert the response text output by the multi-modal dialogue model into response speech; and parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model are fine-tuned jointly according to a loss value of the response speech and the corresponding fourth sample voice.

12. An on-line man-machine dialog device, characterized by The method comprises: an acquisition module configured to acquire dialogue speech input by a user in an online human-computer voice dialogue; an execution module configured to input the dialogue speech into a preset model comprising a word segmenter, a trained multi-modal dialogue model, and a speech recognition module, so that the preset model performs the following operations: extracting features of the dialogue speech by using the word segmenter to obtain speech features of the dialogue speech; and inputting associated information of the online human-computer voice dialogue and the speech features into the multi-modal dialogue model to generate response text by using the multi-modal dialogue model; wherein the associated information comprises dialogue text obtained by performing speech recognition on the dialogue speech by using the speech recognition module; an output module configured to output the response text to the user; wherein the intelligent dialogue robot comprises the preset model and a speech synthesis model; and training of the intelligent dialogue robot comprises: freezing parameters of the trained multi-modal dialogue model; training the word segmenter based on a first training sample set to obtain the word segmenter after first training; the first training sample set comprises a plurality of first sample voices and text content corresponding to the first sample voices; continuing to train the word segmenter after first training based on a second training sample set to obtain the word segmenter after second training; the second training sample set comprises a plurality of second sample voices and speech features corresponding to the second sample voices; training the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set comprises a plurality of sample texts and speech corresponding to the sample texts; jointly fine-tuning the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model based on a fourth training sample set to obtain the trained intelligent dialogue robot; and The fourth training sample set includes a plurality of third sample voices and fourth sample voices for responding to the third sample voices; in the joint fine-tuning training process, the speech recognition module is configured to convert the third sample voices into corresponding text content, the word segmenter is configured to extract speech features of the third sample voices, and the multi-modal dialogue model is configured to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is configured to convert the response text output by the multi-modal dialogue model into response speech; and parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model are fine-tuned jointly according to a loss value of the response speech and the corresponding fourth sample voice.

13. An on-line man-machine dialog device, characterized by Comprise: The acquisition module is configured to acquire dialogue voice input by a user in an online man-machine voice dialogue; The sending module is configured to send the dialogue voice to a server to trigger the server to input the dialogue voice into a preset model including a word segmenter, a trained multi-modal dialogue model, and a speech recognition module, so that the preset model performs: feature extraction on the dialogue voice by using the word segmenter to obtain speech features of the dialogue voice; and converts the dialogue voice into dialogue text by using the speech recognition module; The dialogue text and the speech features are input into the multi-modal dialogue model to generate response text by using the multi-modal dialogue model; The receiving module is configured to receive response speech returned by the server; wherein the response speech is generated by performing speech synthesis on the response text by using a speech synthesis model; The broadcasting module is configured to broadcast the response speech to the user; The server is deployed with an intelligent dialogue robot, and the intelligent dialogue robot includes the preset model and the speech synthesis model; The training of the intelligent dialogue robot comprises: Parameter freezing is performed on the trained multi-modal dialogue model; The word segmenter is trained based on a first training sample set to obtain the word segmenter after first training; the first training sample set includes a plurality of first sample voices and text content corresponding to the first sample voices; The word segmenter after first training is further trained based on a second training sample set to obtain the word segmenter after second training; the second training sample set includes a plurality of second sample voices and speech features corresponding to the second sample voices; The speech synthesis model is trained based on a third training sample set to obtain the trained speech synthesis model; the third training sample set includes a plurality of sample texts and speech corresponding to the sample texts; The speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model are jointly fine-tuned based on a fourth training sample set to obtain a trained intelligent dialogue robot; The fourth training sample set includes a plurality of third sample voices and fourth sample voices for responding to the third sample voices; in the joint fine-tuning training process, the speech recognition module is configured to convert the third sample voices into corresponding text content, the word segmenter is configured to extract speech features of the third sample voices, and the multi-modal dialogue model is configured to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module; the speech synthesis model is configured to convert the response text output by the multi-modal dialogue model into response speech; and parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model are fine-tuned jointly according to a loss value of the response speech and the corresponding fourth sample voices.

14. An on-line man-machine dialog device, characterized by Comprise: a display module configured to display an interface for health services; a starting module configured to start an online human-computer voice dialogue in response to a voice dialogue starting operation triggered through the interface; a collection module configured to collect dialogue voice input by a user in the online human-computer voice dialogue; an execution module configured to input the dialogue voice into a preset model containing a word segmenter, a trained multi-modal dialogue model, and a speech recognition module, and configured to perform the following operations on the preset model: extracting features of the dialogue voice by using the word segmenter to obtain speech features of the dialogue voice; inputting associated information of the online human-computer voice dialogue and the speech features into the multi-modal dialogue model to generate response text by using the multi-modal dialogue model; wherein the associated information contains interactive text generated in at least one round of voice dialogue between a user and an intelligent dialogue robot, and the interactive text contains dialogue text obtained by performing speech recognition on dialogue voice input by the user in the current round of voice dialogue by using the speech recognition module; an output module configured to output the response text to the user; wherein the intelligent dialogue robot contains the preset model and a speech synthesis model; and training of the intelligent dialogue robot comprises: freezing parameters of the trained multi-modal dialogue model; training the word segmenter based on a first training sample set to obtain the word segmenter after first training; the first training sample set contains a plurality of first sample voices and text content corresponding to the first sample voices; continuing to train the word segmenter after first training based on a second training sample set to obtain the word segmenter after second training; the second training sample set contains a plurality of second sample voices and speech features corresponding to the second sample voices; training the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set contains a plurality of sample texts and speech corresponding to the sample texts; jointly fine-tuning training the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model based on a fourth training sample set to obtain the trained intelligent dialogue robot; The fourth training sample set includes a plurality of third sample voices and fourth sample voices for responding to the third sample voices; in the joint fine-tuning training process, the speech recognition module is configured to convert the third sample voices into corresponding text content, the word segmenter is configured to extract speech features of the third sample voices, the multi-modal dialogue model is configured to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module, and the speech synthesis model is configured to convert the response text output by the multi-modal dialogue model into response speech; and parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model are jointly fine-tuned according to a loss value of the response speech and the corresponding fourth sample voices.

15. An intelligent dialog robot training apparatus, characterized by, The intelligent dialogue robot includes a preset model and a speech synthesis model; the preset model includes a speech recognition module, a word segmenter, and a trained multi-modal dialogue model; The device includes: a freezing module configured to freeze parameters of the trained multi-modal dialogue model; a training module configured to: train the word segmenter based on a first training sample set to obtain the word segmenter after first training; the first training sample set includes a plurality of first sample voices and text content corresponding to the first sample voices; continue training the word segmenter after first training based on a second training sample set to obtain the word segmenter after second training; the second training sample set includes a plurality of second sample voices and speech features corresponding to the second sample voices; train the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set includes a plurality of sample texts and speech corresponding to the sample texts; and perform joint fine-tuning training on the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model based on a fourth training sample set to obtain the trained intelligent dialogue robot; wherein the fourth training sample set includes a plurality of third sample voices and fourth sample voices for responding to the third sample voices; in the joint fine-tuning training process, the speech recognition module is configured to convert the third sample voices into corresponding text content, the word segmenter is configured to extract speech features of the third sample voices, the multi-modal dialogue model is configured to generate response text based on the speech features output by the word segmenter and the text content output by the speech recognition module, and the speech synthesis model is configured to convert the response text output by the multi-modal dialogue model into response speech; and parameters of the speech recognition module, the word segmenter, the multi-modal dialogue model, and the speech synthesis model are jointly fine-tuned according to a loss value of the response speech and the corresponding fourth sample voices.

16. An electronic device, comprising: The device includes a memory and a processor, wherein the memory stores executable program instructions, and the processor executes the program instructions to implement the method of any one of claims 1-10. The device includes a memory and a processor, wherein the memory stores executable program instructions, and the processor executes the program instructions to implement the method of any one of claims 1-10.

17. A computer readable storage medium characterized in that, The computer program is stored in the storage medium and, when executed in the computer, causes the computer to perform the method of any one of claims 1-10.

18. A computer program product, characterised in that, The computer program product comprises the computer program or instructions, which, when executed by a processor, implement the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Voice conversation method and device, and computer readable storage medium

    CN111986675A

  • Dialogue method and device based on voice recognition, terminal equipment and storage medium

    CN118447841A