Online man-machine conversation method, system and device, electronic equipment, storage medium and program product

By integrating the two-stage online human-computer voice dialogue interaction architecture with the automatic speech recognition module and the multimodal dialogue model, the problems of reply accuracy and long response time of intelligent dialogue robots are solved, achieving more efficient user interaction and medical diagnosis accuracy.

CN120723892AActive Publication Date: 2025-09-30ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511140707.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-09-30
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

In existing intelligent conversational robots' voice dialogue interactions, the accuracy and reliability of responses are low, especially in the field of medical diagnosis. In addition, the response time is long, resulting in a poor user experience.

Method used

It adopts a two-stage online human-computer voice dialogue interaction architecture, integrating the automatic speech recognition module and the multimodal dialogue model to retain the sound information in the user's voice, reduce processing delays, and improve response speed.

Benefits of technology

It improves the accuracy and reliability of responses, shortens response time, and enhances user experience. Especially in medical and health scenarios, it can provide medical advice that is more tailored to the user's actual situation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723892A_ABST
    Figure CN120723892A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an online man-machine conversation method, system and device, electronic equipment, a storage medium and a program product. According to the scheme provided by the embodiment of the invention, after the conversation voice input by the user in the online man-machine voice conversation is obtained, the conversation voice is input to the preset model, so that the preset model executes the steps of extracting voice features of the conversation voice, performing voice recognition on the conversation voice to obtain a corresponding conversation text, and outputting the conversation text to the user. And generating a reply text based on the associated information (including the dialogue text corresponding to the dialogue voice of the user of the current round of voice dialogue) of the online man-machine voice dialogue and the voice features of the dialogue voice. Furthermore, the generated reply text can be output to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to an online human-computer dialogue method, system, device, electronic device, storage medium, and program product. Background Art

[0002] Intelligent conversational robots (often also called intelligent conversational assistants) can interact with users, with the goal of helping them solve problems or simply chatting. They have been widely used in various fields, including healthcare. Currently, intelligent conversational robots primarily rely on text-based interactions. However, with the advancement of voice technology, an increasing number of intelligent conversational robots are adding voice-based interaction capabilities. Voice-based interaction is more aligned with user conversational habits, providing a better experience, especially for those unfamiliar with typing. However, current voice-based interactions rely solely on text responses to user voice, resulting in relatively low accuracy and reliability. This shortcoming is particularly pronounced in the field of medical diagnosis.

[0003] Therefore, there is an urgent need to provide an improved solution for the voice dialogue interaction of intelligent dialogue robots. Summary of the Invention

[0004] The various embodiments of this specification provide a method, system, device, electronic device, storage medium, and program product for online human-computer dialogue, which implements voice dialogue interaction for intelligent dialogue robots and can respond to users based on the voice features in the user's voice, thereby improving the accuracy and reliability of responses. In the first embodiment, this specification provides an online human-computer dialogue method. The method includes: Obtain the conversation voice input by the user in the online human-computer voice dialogue; Inputting the conversational speech into a preset model, the preset model then performs the following steps: extracting features from the conversational speech to obtain speech features of the conversational speech; generating a response text based on the associated information of the online human-computer voice conversation and the speech features; wherein the associated information includes the conversation text obtained by performing speech recognition on the conversational speech by the preset model; The reply text is output to the user.

[0005] In a second embodiment, this specification also provides an online human-computer dialogue method. The method includes: Collecting the conversational voice input by the user in the online human-computer voice dialogue; The conversation speech is sent to a server to trigger the server to input the conversation speech into a preset model, and the preset model performs the following operations: extracting features from the conversation speech to obtain speech features of the conversation speech; converting the conversation speech into conversation text; and generating the reply text based on the conversation text and the speech features; Receive the reply voice returned by the server; wherein the reply voice is generated by performing speech synthesis on the reply text.

[0006] The reply voice is broadcast to the user.

[0007] In a third embodiment, this specification also provides an online human-computer dialogue method. The method includes: Displays an interface for health services; In response to a voice dialogue initiation operation triggered through the interface, initiating an online human-computer voice dialogue; In the online human-computer voice dialogue, collecting the dialogue voice input by the user; The conversational speech is input into a preset model, and the preset model performs the following steps: extracting features from the conversational speech to obtain speech features of the conversational speech; and generating the response text based on the associated information of the online human-machine voice dialogue and the speech features; wherein the associated information includes interactions generated by at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactions include the conversation text obtained by performing voice recognition on the user's conversational speech in the current round of voice dialogue by the preset model; The reply text is output to the user.

[0008] In a fourth embodiment, this specification also provides a method for training an intelligent conversational robot. The intelligent conversational robot includes a preset model and a speech synthesis model; the preset model includes a speech recognition module, a word segmenter, and a trained multimodal conversation model. The method includes: Freezing parameters of the trained multimodal dialogue model; Training the word segmenter based on a first training sample set to obtain the word segmenter after the first training; the first training sample set includes a plurality of first sample speech and text content corresponding to the first sample speech; Based on a second training sample set, the word segmenter after the first training is further trained to obtain the word segmenter after the second training; the second training sample set includes a plurality of second sample speech and speech features corresponding to the second sample speech; Training the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set includes a plurality of sample texts and speech corresponding to the sample texts; Based on the fourth training sample set, jointly fine-tune the speech recognition module, the word segmenter, the multimodal dialogue model, and the speech synthesis model to obtain a trained intelligent dialogue robot; In which, the fourth training sample set includes multiple third sample voices and fourth sample voices used to respond to the third sample voices; in the joint fine-tuning training process, the speech recognition module is used to convert the third sample voices into corresponding text content, the word segmenter is used to extract the speech features of the third sample voices, and the multimodal dialogue model uses the speech features output by the word segmenter and the text content output by the speech recognition module to generate reply text; the speech synthesis model is used to convert the reply text output by the multimodal dialogue model into reply speech; according to the loss value of the reply speech and the corresponding fourth sample voice, the parameters of the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model are jointly fine-tuned.

[0009] In a fifth embodiment, this specification provides an online human-computer dialogue system. The system includes: The client is used to collect the conversation voice input by the user in the online human-computer voice dialogue and send the conversation voice to the server; The server is deployed with a preset model, and is configured to input the conversational speech into the preset model, and have the preset model perform the following operations: extracting features from the conversational speech to obtain speech features of the conversational speech; and generating the response text based on the associated information of the online human-computer voice conversation and the speech features; wherein the associated information includes the conversation text obtained by performing speech recognition on the conversational speech by the preset model; The server is further configured to convert the reply text into reply voice and send the reply voice to the client; The client is also used to broadcast the reply voice to the user.

[0010] In a sixth embodiment, this specification provides an online human-computer dialogue device. The device includes: An acquisition module is used to acquire the conversation voice input by the user in the online human-computer voice dialogue; an execution module configured to input the conversational speech into a preset model, and have the preset model perform the following operations: extracting features from the conversational speech to obtain speech features of the conversational speech; and generating the response text based on the associated information of the online human-computer voice conversation and the speech features; wherein the associated information includes the conversation text obtained by performing speech recognition on the conversational speech by the preset model; The output module is used to output the reply text to the user.

[0011] In a seventh embodiment, this specification also provides an online human-computer dialogue device. The device includes: The acquisition module is used to collect the conversation voice input by the user in the online human-computer voice dialogue; a sending trigger module, configured to send the conversation speech to a server to trigger the server to input the conversation speech into a preset model, so that the preset model performs the following operations: extracting features from the conversation speech to obtain speech features of the conversation speech; converting the conversation speech into conversation text; and generating the reply text based on the conversation text and the speech features; A receiving module, configured to receive the reply voice returned by the server; wherein the reply voice is generated by performing speech synthesis on the reply text; The broadcast module is used to broadcast the reply voice to the user.

[0012] In an eighth embodiment, this specification also provides an online human-computer dialogue device. The device includes: A display module, used for displaying an interface for health services; A starting module, configured to respond to a voice dialogue starting operation triggered through the interface and start an online human-computer voice dialogue; A collection module, used to collect the conversation voice input by the user in the online human-computer voice dialogue; An execution module is configured to input the conversational speech into a preset model, and the preset model executes the following: feature extraction of the conversational speech to obtain speech features of the conversational speech; and generation of the response text based on the associated information of the online human-machine voice dialogue and the speech features; wherein the associated information includes interactions generated by at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactions include the conversation text obtained by the preset model performing voice recognition on the user's conversational speech in the current round of voice dialogue; The output module is used to output the reply text to the user.

[0013] In a ninth embodiment, this specification also provides a device for training an intelligent conversational robot. The intelligent conversational robot includes a preset model and a speech synthesis model; the preset model includes a speech recognition module, a word segmenter, and a trained multimodal conversation model. The device includes: A freezing module, configured to freeze parameters of the trained multimodal dialogue model; The training module is used to: train the word segmenter based on the first training sample set to obtain the word segmenter after the first training; the first training sample set includes a plurality of first sample speech and text content corresponding to the first sample speech; based on the second training sample set, continue to train the word segmenter after the first training to obtain the word segmenter after the second training; the second training sample set includes a plurality of second sample speech and speech features corresponding to the second sample speech; based on the third training sample set, train the speech synthesis model to obtain the trained speech synthesis model; the third training sample set includes a plurality of sample texts and speech corresponding to the sample texts; based on the fourth training sample set, link the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model Combined fine-tuning training is performed to obtain a trained intelligent dialogue robot; wherein the fourth training sample set includes multiple third sample voices and fourth sample voices used to respond to the third sample voices; in the joint fine-tuning training process, the speech recognition module is used to convert the third sample voices into corresponding text content, the word segmenter is used to extract the speech features of the third sample voices, and the multimodal dialogue model uses the speech features output by the word segmenter and the text content output by the speech recognition module to generate a reply text; the speech synthesis model is used to convert the reply text output by the multimodal dialogue model into a reply speech; according to the loss value of the reply speech and the corresponding fourth sample voice, the parameters of the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model are jointly fine-tuned.

[0014] In a tenth embodiment, this specification provides an electronic device, including a memory and a processor, wherein the memory stores executable program instructions, and when the processor executes the program instructions, the methods provided in the first to fourth embodiments are implemented.

[0015] In an eleventh embodiment, this specification provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer is caused to execute the methods provided in the first to fourth embodiments.

[0016] In a twelfth embodiment, this specification also provides a computer program product, including a computer program / instruction, which implements the methods provided in the first to fourth embodiments when executed by a processor.

[0017] The solutions provided in the aforementioned embodiments of this specification, after acquiring the user's voice input in an online human-machine voice dialogue, input the voice into a preset model. The preset model then performs the following operations: extracting the voice features of the voice input, performing voice recognition on the voice input to obtain the corresponding dialogue text, and then generating a reply text based on the associated information of the online human-machine voice dialogue (including the dialogue text corresponding to the user's voice input in the current voice dialogue) and the voice features of the voice input. As can be seen, when generating the reply text for the user's voice input, not only the text content of the voice input is incorporated, but also the voice features of the voice input are incorporated. This can effectively improve the quality of the reply text. In particular, in healthcare scenarios, the user's voice features are an important diagnostic basis. Incorporating the user's voice features can improve the accuracy and reliability of medical diagnoses, thereby effectively ensuring that the medical advice provided in the reply text is more tailored to the user's actual situation. The acquisition of the conversational text corresponding to the aforementioned conversational speech, the acquisition of speech features, and the generation of the reply text are all achieved through a preset model that integrates a speech recognition module and a multimodal conversation model (for reply text generation). This preset model design enables the entire online human-computer voice conversation interaction to be implemented in a two-stage manner, which can significantly reduce processing delays, improve response speed, and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings: Figure 1 A schematic diagram of the technical architecture of a traditional online human-computer voice dialogue provided by an exemplary embodiment; Figure 2A and Figure 2B A schematic diagram of the technical architecture on which each method in this specification is implemented is provided as an exemplary embodiment; Figure 3 A schematic diagram of the structure of an online human-computer dialogue system provided by an exemplary embodiment of this specification; Figure 4 、 Figure 5 and Figure 6 A flowchart of an online human-computer dialogue method provided by an exemplary embodiment of this specification; Figure 7 A flowchart of an intelligent conversational robot training method provided by an exemplary embodiment of this specification; Figure 8 、 Figure 9 and Figure 10 A schematic diagram of the structure of an online human-computer dialogue device provided by an exemplary embodiment of this specification; Figure 11 A schematic diagram of a flow chart of an intelligent conversational robot training device provided as an exemplary embodiment of this specification; Figure 12 This is a schematic structural diagram of an electronic device provided as an exemplary embodiment of this specification. DETAILED DESCRIPTION

[0019] With the rapid development of artificial intelligence and mobile internet technologies, intelligent conversational robots (AIs) have become a crucial tool for human-computer interaction in various applications. AIs can interact with users, helping them solve problems or simply chatting. They have been widely used in fields such as healthcare. For example, AIs for health (such as the health managers offered in some apps) are dedicated to helping users resolve various issues before, during, and after medical treatment. They cover a wide range of scenarios, including health consultations, initial disease screenings, medical advice, medication reminders, and rehabilitation guidance, significantly improving users' health management efficiency and overall well-being. Currently, AIs for health and other AIs (such as customer service bots and chatbots) primarily support text-based interaction. For example, users can enter a text description of their health condition, and AIs for health can generate a corresponding medical diagnosis using their language models. With the advancement of voice technology, an increasing number of AIs have added voice interaction capabilities in recent years, allowing users to interact with AIs via voice. Voice interaction is more aligned with user conversational habits, providing a better experience, especially for those unfamiliar with typing. Then, in the current voice dialogue interaction, responses to users are mainly based on the text content of the user's voice dialogue.

[0020] Figure 1 The following is an example of a technical architecture diagram for realizing online human-computer voice dialogue interaction. Figure 1 As shown in the figure, existing online human-computer voice dialogue interactions are mainly implemented based on a three-stage pipeline. Specifically, when a user inputs a voice, Automatic Speech Recognition (ASR) is first performed on the user's input voice to obtain a text representation of the user's voice. Then, based on the text representation of the user's voice, a text dialogue is performed to obtain a reply from the intelligent dialogue robot. Finally, the text content of the intelligent dialogue robot's reply is synthesized through text-to-speech (TTS) to obtain the corresponding reply voice, and the reply voice is broadcast to the user.

[0021] above Figure 1 The online human-computer voice dialogue interaction method shown may have relatively good interactive effects in some application scenarios, but the user experience is not very good in medical scenarios. This is mainly reflected in the following: First, after the user voice is converted into text representation through automatic speech recognition, sound information other than the text content is lost, such as the user's hoarse voice and the user's excitement. This sound information is an important basis for diagnosis; second, automatic speech recognition (ASR), intelligent text dialogue (for reply text generation), and text-to-speech synthesis (TTS) need to be executed serially. When the number of model parameters increases (for example, the intelligent dialogue part starts to use the voice model (LLM) to implement), the execution time will be longer (for example, more than 3 seconds), which leads to long response times, causing users to wait for a long time and a relatively poor experience.

[0022] To address the above issues, the embodiments described in this specification provide a solution, proposing a two-stage online human-computer voice dialogue interaction. This two-stage online human-computer voice dialogue interaction integrates an automatic speech recognition (ASR) module with a text-based dialogue model (such as a language model (LLM)) to form a new voice dialogue model. This approach not only preserves the acoustic information (such as timbre and emotion) in the user's voice, but also significantly reduces processing delays, shortens response times, and improves the user experience.

[0023] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this specification.

[0024] It should be noted that, for ease of description, only the parts related to the relevant technical solutions are shown in the accompanying drawings. In the absence of conflict, the embodiments in this specification and the features in the embodiments may be combined with each other. In addition, the words "first", "second", "third" and the like in the embodiments of this specification are only used for information distinction and do not play any limiting role. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, product or device. In the absence of further restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or device including the elements. In addition, in this specification, unless explicitly stated, "receiving and sending data" does not necessarily mean direct receiving and sending, but may be indirect receiving and sending. For example, when A receives data sent by B, it can be understood that A receives the data sent by B directly, or it can be understood that A receives the data sent by B indirectly through other entities such as C. Similarly, when B sends data to A, it can be understood that B sends the data directly to A, or it can be understood that B sends the data indirectly through other entities such as C. Here, C can be one entity, or two or more entities.

[0025] Furthermore, it should be noted that this specification uses specific words to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different places in this specification does not necessarily refer to the same embodiment. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory. Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it can be understood that the order of steps listed in the embodiments or flowcharts is only one way of executing the steps among many, and does not represent the only execution order. Therefore, when the claims involve method steps, changes and adjustments to the order of such steps, or parallelism between steps are also within the scope of protection of the claims.

[0026] Furthermore, it should be noted that the user data obtained in this manual is authorized by the user and does not involve user privacy.

[0027] The following describes and illustrates various embodiments provided in this specification in conjunction with the accompanying drawings.

[0028] First, the terms used in the embodiments of this specification are explained. It should be understood that this explanation is for a clearer understanding of the embodiments of this specification and does not necessarily constitute a limitation on the embodiments of this specification.

[0029] Automatic Speech Recognition (ASR): It is used to recognize human speech into text, that is, to convert human speech signals into text. It involves multiple processes including speech signal acquisition, feature extraction, acoustic model, language model and decoding.

[0030] Text-to-Speech (TTS): used to convert text information into natural-sounding speech output.

[0031] The preset model is a trained artificial intelligence model. In the specification, the preset model is a voice dialogue model, which integrates the automatic speech recognition (ASR) function and the text dialogue function. The text dialogue function can be implemented through a language model (LLM). That is to say, in one example, the voice dialogue model is mainly implemented by combining automatic speech recognition (ASR) and a language model (LLM). The language model (LLM) is an artificial intelligence model and a key component of natural language processing (NLP) technology. LLM is based on the transformer architecture and can understand and generate high-quality human language text. It is used for dialogue content generation in the dialogue interaction system. It is worth noting here that the embodiments of this specification do not limit the number of parameters supported by the preset model, and the goal is to meet actual application needs.

[0032] Linguistic Tokenizer: This is used to extract high-level structured representations from speech. Specifically, it extracts linguistically structured information from speech, such as phonemes, syllables, and tones.

[0033] Semantic Tokenizer: It is designed to encode semantic and coarse-grained acoustic features in speech.

[0034] Speech Decoder: In text-to-speech synthesis (TTS), it converts text encoding (such as the features output by the language model) into a corresponding speech waveform. For example: the text "Hello" -> encoded into speech features -> the speech decoder generates a speech waveform.

[0035] The technical solutions provided in the following embodiments of this specification are based on Figure 2A and Figure 2B The technical architecture shown in is implemented. Figure 2A and Figure 2B As shown in the figure, this technical architecture is a two-stage online human-computer voice dialogue architecture. One stage is to use the voice dialogue model to process the user's voice to generate the corresponding reply text, and the other stage is to convert the reply text into reply voice through the TTS module and broadcast it to the user.

[0036] Specifically, the voice dialogue model is obtained by integrating the ASR module and the text dialogue model. The input of the voice dialogue model is the user's conversation voice, and the output is the reply text used to respond to the user's conversation voice.

[0037] The text dialogue model within the fused speech dialogue model must combine the speech features of the user's input speech to output a text dialogue result (specifically, the reply text). Therefore, in the following description, the text dialogue model is referred to as a speech-to-text (SST) model. The SST model's input has multiple modal data. In specific implementations, the SST model's input includes at least two modal data types: speech modal data and text modal data. Speech modal data refers to the speech features of the user's speech, and text modal data refers to the text content of the user's speech. Therefore, the SST model is actually a multimodal dialogue model. The specific inputs to the SST model are detailed below and will not be elaborated on here.

[0038] In this solution, the speech dialogue model includes not only the ASR module and the SST model (e.g., based on the LLM model), but also two tokenizers. These tokenizers are the Linguistic Tokenizer and the Semantic Tokenizer. The Linguistic Tokenizer extracts structured linguistic features (high-level representations) from speech, while the Semantic Tokenizer encodes semantic and coarse-grained acoustic features in speech. For a detailed description of these two tokenizers, please refer to the relevant content in the Glossary section above.

[0039] See Figure 2B As mentioned above, the input of the STT model in the speech dialogue model includes the following two parts: 1) Interaction text features are input in text form. This interaction text includes the text corresponding to the user's speech input in the current voice conversation, as well as the historical interaction text generated by previous voice conversations (including the text corresponding to the user's historical speech input and the text of the historical responses from the intelligent conversational robot). The ASR module converts the user's speech input into the corresponding text.

[0040] 2) Speech features. Speech features are obtained by extracting features from the user's conversational speech using a language segmenter and a semantic analyzer.

[0041] Furthermore, the above-mentioned interactive text (specifically, the text features of the interactive text) and speech features are combined and input into the STT model. Executing the STT model will generate the encoding features corresponding to the reply text required to answer the user's conversational speech. Then, the encoding features corresponding to the reply text output by the SST model are input into the TTS model, and the corresponding reply speech can be generated by the TTS model. In specific implementation, the TTS model can call the speech decoder therein to process the encoding features of the reply text, thereby generating the corresponding reply speech. The speech decoder described here is a key component in the TTS model, which can convert the encoding features of the text (such as the features output by the language model) into the corresponding speech waveform.

[0042] Among them, the overall training process of this technical architecture can be, for example: first, based on a trained language model (LLM), on the basis of freezing the parameters of the language model (LLM), use speech-text alignment training to train the tokenizer required for speech feature extraction; then, use the data of the ASR / TTS task to continue training the tokenizer (Tokenizer) and speech decoder (Speech Decoder). For example, the data of the ASR task can be used to train the tokenizer so that the tokenizer can learn the ability to encode speech (that is, feature extraction ability), and the data of the TTS task can be used to train the speech decoder so that the speech decoder can learn the ability to decode text (that is, the ability to convert text encoding into speech waveforms); finally, use voice conversation data (such as user voice input + system voice response) to fine-tune the entire link.

[0043] In this technical architecture, a voice dialogue model is implemented by combining the ASR module with a multimodal dialogue model. This model can learn responses based on different user timbres, emotions, and ambient sounds. This allows the model to recognize and empathize with user details, thereby improving the user experience.

[0044] Although, for Figure 1The problem of losing a large amount of audio details in the three-stage online human-computer voice dialogue interaction shown can also be solved by other methods, such as by customizing the voice feature extraction module to perform emotion recognition, environmental recognition, etc. However, this method has the problem of accuracy loss due to the introduction of more models. On the other hand, it requires additional models or rules to reflect the impact of extracted features on subsequent text generation, which is difficult to be data-driven and long-term optimization effect is relatively difficult. In addition to the above method, online human-computer voice dialogue interaction can also be achieved through a fully end-to-end voice model. This fully end-to-end voice model can couple the three modules of ASR, LLM and TTS to achieve the direct conversion from input voice to the voice of the dialogue response. This voice model coupled with ASR, LLM and TTS is theoretically feasible, but it requires a large amount of supervised training, and the number of LLM parameters will also slow down the speed of TTS, which can easily lead to high response delay and difficult to apply well.

[0045] In summary, the technical architecture presented in this specification implements a voice conversation model by integrating an ASR module and a multimodal conversation model. This technical architecture offers excellent scalability and optimizability, and through continuous data training and model iteration, it can continuously improve the performance and effectiveness of the architecture. It has a wide range of applications and a strong competitive advantage in the market.

[0046] The technical architecture mentioned above is based on the server and client implementation. Figure 1 As shown, the voice conversation model and TTS functionality in the technical architecture are deployed on the server. The reply voice broadcasting function and user voice collection function are deployed on the client. The server can be a server, server cluster, virtual server, or cloud-based system. The client can be, but is not limited to, a smartphone, smart wearable device, tablet, laptop, desktop computer, etc. The server provides corresponding functional services to the client, such as intelligent conversation. Users can initiate an online human-machine voice conversation through a browser, application (app), web application H5 (HyperText Markup Language 5, the fifth generation of HTML, Hypertext Markup Language), light application (also known as mini-program, a lightweight application), or cloud application on the client. After the online human-machine voice conversation is connected to the intelligent conversation robot on the server, an online human-machine voice conversation interaction is established between the client and the server.

[0047] thus, Figure 3 The online human-computer dialogue system (also called service system) provided by an embodiment of this specification is also shown, which includes a client 200 and a server 100. The client 200 is used to collect the conversation voice input by the user in the online human-computer voice dialogue and send the conversation voice to the server; The server 100 is deployed with a preset model (i.e., the aforementioned voice dialogue model), and is configured to input the conversational speech into the preset model, causing the preset model to perform the following operations: extracting features from the conversational speech to obtain voice features of the conversational speech; and generating the response text based on the associated information of the online human-computer voice dialogue and the voice features; wherein the associated information includes the conversation text obtained by performing voice recognition on the conversational speech by the preset model; The server 100 is further configured to convert the reply text into reply voice and send the reply voice to the client; The client 200 is further configured to broadcast the reply voice to the user.

[0048] The online human-machine voice dialogue interaction functionality described above can be integrated into intelligent conversational robots in any field. The intelligent conversational robot is deployed on the server side. Users can access the service interface for human-machine interaction provided by the intelligent conversational robot through the client side and initiate online human-machine voice dialogue interaction through the service interface. The aforementioned preset model (i.e., the voice dialogue model) and TTS model are the core components of the intelligent conversational robot architecture.

[0049] The specific implementation of the functions of the server 100 and the client 200, as well as the initiation of online human-computer voice dialogue interaction, will be described in detail in the following method embodiments and will not be described in detail here.

[0050] The technical solutions provided in this specification will be described below in the form of method embodiments.

[0051] Figure 4 The flowchart of an online human-computer dialogue method provided by an embodiment of this specification is shown. The execution subject of this method is the server in the above system. Figure 4 As shown, the online human-computer dialogue method includes the following steps: 102. Obtaining the conversation voice input by the user in the online human-computer voice dialogue; 104. Input the conversational speech into a preset model, and have the preset model perform the following steps: extracting features from the conversational speech to obtain speech features of the conversational speech; and generating the response text based on the associated information of the online human-computer voice conversation and the speech features; wherein the associated information includes the conversation text obtained by performing speech recognition on the conversational speech by the preset model; 106. Output the reply text to the user.

[0052] In this embodiment, online human-computer voice dialogue is initiated by the user through a service interface for human-computer interaction displayed on the client. The triggering method can be, but is not limited to, triggering by operating corresponding controls, inputting voice, inputting text, etc. The service interface is provided by the corresponding intelligent dialogue robot. Specifically, this service interface can be a traditional human-computer interaction interface that supports text dialogue or graphic dialogue. Therefore, in this specification, the intelligent dialogue robot supports online human-computer voice dialogue interaction, and the intelligent dialogue robot can be an intelligent dialogue assistant (also called an intelligent dialogue system) in any field, such as health services (including medical and health services) and e-commerce customer service.

[0053] For example, taking the smart health manager (a smart conversation assistant that provides health services) as an example, the user enters the human-computer interaction service interface provided by the smart health manager through the client 200. Generally, the default interaction mode supported by this service interface is text conversation or graphic conversation. Figure 3 As shown, the user can trigger the start of online human-computer voice dialogue interaction by operating the "call" control 210 on the service interface 21; the user can also trigger the start of online human-computer voice dialogue interaction by inputting a voice dialogue start instruction (such as please start voice dialogue) in the input box 212 provided in the service interface 21. Of course, in other instances, the voice dialogue start instruction can also be input in other ways, such as voice input, such as operating the "voice input" control 213 to input the voice of "please start voice dialogue". Among them, the interface of online human-computer voice dialogue interaction is as follows Figure 3 After the online human-computer voice dialogue interaction is started, the user and the smart health manager will conduct an online human-computer dialogue interaction in the form of a voice call. During the interaction, the client 200 will use the sound pickup device on it to respectively pick up the user's input dialogue voice and send it to the server 100 for processing.

[0054] Based on the above content, the above step 102 of "obtaining the dialogue voice input by the user in the online human-computer voice dialogue" includes: 1021. Receive the conversation voice sent by the client; The conversation voice is collected by the client through a sound pickup device, which may be, but is not limited to, a microphone.

[0055] After receiving the conversational speech initiated by the client, the server invokes the pre-configured model deployed on it for processing. This pre-configured model includes speech recognition, speech feature extraction, and text generation capabilities. The speech feature extraction function primarily extracts the corresponding linguistic and acoustic features from the received conversational speech. The speech recognition function performs speech recognition on the received conversational speech to obtain the corresponding conversational text content. The text generation function generates the corresponding response text based on the received conversational speech.

[0056] Based on the above content, in a specific implementation scheme, the above step 104 of “the preset model performs feature extraction on the conversation speech to obtain the speech features of the conversation speech” may specifically include: 1042. Extracting language features from the conversational speech, where the language features can reflect the pronunciation structure and language rules of the conversational speech; 1044. Extract acoustic features from the conversational speech, where the acoustic features can reflect auditory perception attributes of the conversational speech.

[0057] In specific implementation, the preset model includes a first word segmenter and a second word segmenter. The first word segmenter is, for example, Figure 2B The language segmenter shown in the second segmenter is Figure 2B The semantic word segmenter shown in FIG. The first word segmenter can be used to extract corresponding language features from the conversation speech. The second word segmenter can be used to extract corresponding acoustic features from the conversation speech.

[0058] That is, the implementation of "extracting language features from the conversation speech" in step 1042 may include: The conversation speech is input into the first word segmenter, and the first word segmenter is executed to output the language features.

[0059] The implementation of “extracting acoustic features from the conversation speech” in step 1044 may include: The conversation speech is input into the second word segmenter, and the second word segmenter is executed to output the acoustic features.

[0060] In the above description, language features include phonemic features and linguistic features. Phonemes are the smallest units of pronunciation in speech. Therefore, in this specification, phonemic features directly reflect the pronunciation structure of conversational speech, such as / p / , / b / , and / t / . Linguistic features reflect the phonetic rules of conversational speech. For example, linguistic features include syntax, part of speech, stress, word boundaries, word pronunciation duration, and pause duration. These features directly reflect the linguistic rules of conversational speech.

[0061] Furthermore, the aforementioned acoustic features include those that reflect the auditory perceptual attributes of conversational speech, such as emotion, timbre (e.g., hoarseness), prosody, intonation, speaking rate, and tone. Furthermore, in other instances, acoustic features may also include physical attributes that reflect conversational speech. These physical attributes include the spectral and energy characteristics of conversational speech, such as fundamental frequency, amplitude, short-term energy, zero-crossing rate, spectrum, Mel-spectrogram, and MFCC.

[0062] Based on the speech features of the aforementioned conversational speech and the corresponding conversational text, the preset model can generate corresponding reply text for responding to the user. Of course, in other examples, the reply text can also be generated by further combining the interactive text of previous voice conversations between the user and the intelligent conversational robot before the current voice conversation.

[0063] Therefore, in this embodiment, the associated information of the online human-computer voice dialogue described in 104 above includes the interactive text generated by at least one round of voice dialogue between the user and the intelligent dialogue robot. This interactive text includes the dialogue text corresponding to the user in the current round of voice dialogue, and may also include the historical interactive text between the user and the intelligent dialogue robot in previous rounds (such as previous rounds) (including the dialogue text content corresponding to the historical dialogue voice input by the user and the reply content (historical reply text) of the intelligent dialogue robot to this dialogue text). And, accordingly, the "the preset model generates the reply text based on the associated information of the online human-computer voice dialogue and the voice features" in 104 above may include: 1046. Generate the reply text based on the interactive text generated by the at least one round of voice dialogue interaction and the voice features.

[0064] The interactive text generated by at least one round of voice dialogue here can be used to help understand the conversation context of the current round of voice dialogue, improve the quality of the reply, and thus better ensure that the response is targeted to the user.

[0065] Furthermore, the associated information of the online conversation may also include other information, such as user-related information.

[0066] For example, in a medical and health scenario, the user's relevant information may include but is not limited to at least one of the following: historical diagnostic records, medical history, allergy history, family medical history, etc. Combining this information can improve the accuracy of the diagnosis, so that the medical advice given in the generated reply text can be more in line with the user's actual situation (such as avoiding recommending drugs containing ingredients that the user is allergic to), which can enhance the user experience.

[0067] Furthermore, in e-commerce customer service scenarios, user-related information may include, but is not limited to: identity information (such as service level, e.g., VIP level), profile information (such as order information, purchase history, historical behavior (such as consultation, complaint, abnormal return, etc. records, and geographic location). When generating a reply text, the reply generated by combining the user's service level (such as VIP level) and other identity information and historical behavior (such as historical complaint records and return preferences) can make the user feel "understood" and "valued". The reply generated by combining the user's historical behavior such as frequent complaints and abnormal returns and geographic location (such as high-risk areas) will make the reply content more cautious, avoid over-promises, or provide more appropriate solutions to user problems in the reply, etc.

[0068] That is, a specific implementation of the above step 1046 may include: 10462. Generate the reply text based on the interactive text generated by the at least one round of voice dialogue, the relevant information of the user and the voice features.

[0069] The aforementioned steps related to "generating a response text based on the associated information of the online human-computer voice dialogue and the voice features" in step 104 are implemented by the preset model by invoking the STT model within it. Furthermore, the preset model also includes a speech recognition module, which is used to convert the user's conversational speech into corresponding conversational text. Therefore, in one specific implementation, the step of "the preset model performing speech recognition on the conversational speech to obtain the conversational text" in step 104 may include: 1048. The preset model uses a built-in speech recognition module to perform speech recognition on the conversation speech to obtain the conversation text.

[0070] Furthermore, the step 104 of “generating a reply text based on the associated information of the online human-computer voice dialogue and the voice features” may include: 10410. The preset model inputs the associated information and the voice features into its own built-in multimodal dialogue model, and executes the multimodal dialogue model to output the reply text.

[0071] The above-mentioned speech recognition module is an ASR module, and the multimodal dialogue model can be Figure 2B The SST model shown in ,the STT model is built based on the language model (LLM).

[0072] In summary, the solution provided in this embodiment not only incorporates the text content of the user's conversational speech but also its speech features when generating a response to the user's conversational speech. This effectively improves the quality of the response. This is particularly true in healthcare scenarios, where user speech features are an important diagnostic basis. Incorporating these features can enhance the accuracy and reliability of medical diagnoses, effectively ensuring that the medical advice provided in the response is more tailored to the user's specific situation. The acquisition of the conversational text corresponding to the conversational speech, the acquisition of speech features, and the generation of the response text are all achieved through a pre-defined model that integrates a speech recognition module and an STT model (for response text generation). This pre-defined model design enables a two-stage online human-computer voice interaction, significantly reducing processing delays, improving response speed, and enhancing the user experience. Specifically, this solution integrates the speech recognition module and the STT model to obtain a preset model (a speech dialogue model). In the entire online human-computer speech dialogue interaction, it can reduce the serial execution time of the speech recognition (ASR) module, the text dialogue model (such as LLM) and the speech synthesis (TTS) model, effectively shortening the time users wait for responses, thereby effectively improving the overall speech dialogue interaction experience. Especially in the context of the increasing number of model parameters, this solution can provide a more streamlined and natural speech interaction experience while ensuring performance.

[0073] The above-mentioned text-to-speech (TTS) model is used to convert the reply text output by the preset model into the corresponding reply voice to be broadcast to the user.

[0074] Specifically, in step 106 above, the server converts the reply text into a spoken voice using the TTS model, and then sends the spoken voice to the client for playback to the user. In a specific implementation, the TTS model can utilize its internal speech decoder to convert the reply text into the corresponding conversational voice. While the spoken voice is being sent to the client, the reply text can also be sent to the client for display on the client interface. Therefore, in one specific implementation, step 106 "outputting the reply text to the user" may include: 1062. Perform speech synthesis on the reply text to generate a reply speech; 1064. Send the reply voice and the reply text to the client, and the client displays the reply text while broadcasting the reply voice.

[0075] In the above, considering that in some fields, such as the medical and health field, there are often some special terms, if the reply voice contains these special terms, simply broadcasting this reply voice may cause the user to have problems understanding it. For this reason, here, while broadcasting the reply voice, the corresponding reply text will also be displayed synchronously, which can facilitate users to understand the reply voice through text vision. Among them, the reply text can be, for example, Figure 3 It is displayed on the online human-computer voice dialogue interaction interface 22 shown in FIG.

[0076] This specification also provides another online human-computer dialogue method, the execution subject of which is the client in the above system. Figure 5 As shown, the online human-computer dialogue method includes the following steps: 202. Collecting the conversation voice input by the user in the online human-computer voice dialogue; 204. Send the conversation speech to the server to trigger the server to input the conversation speech into a preset model, and the preset model executes the following steps: extracting features from the conversation speech to obtain speech features of the conversation speech; converting the conversation speech into conversation text; and generating the reply text based on the conversation text and the speech features. 206. Receive the reply voice returned by the server; wherein the reply voice is generated by performing speech synthesis on the reply text.

[0077] 208. Broadcast the reply voice to the user.

[0078] The specific implementation of the above steps provided in this embodiment can refer to the relevant content in other embodiments, and will not be described in detail here. In addition, the method provided in this embodiment can also include some steps disclosed in other embodiments, and can also refer to the relevant content in other embodiments, and will not be described in detail here.

[0079] This specification also provides another online human-computer dialogue method, in which some steps (such as 302, 304, and 306) are executed by the client, and the remaining steps (such as 308 and 310) are executed by the server, that is, the method is implemented by the client and the server in a collaborative manner. In addition, the scenario in which this method is applied is a medical and health scenario. For details, see Figure 6 As shown, the online human-computer dialogue method includes the following steps: 302. Displaying an interface for health services; 304. In response to the voice dialogue initiation operation triggered through the interface, initiating an online human-computer voice dialogue; 306. In the online human-computer voice dialogue, collecting the dialogue voice input by the user; 308. Input the conversational speech into a preset model, and the preset model executes: feature extraction on the conversational speech to obtain speech features of the conversational speech; and generates the response text based on the associated information of the online human-machine voice dialogue and the speech features; wherein the associated information includes interactive text generated by at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactive text includes the conversation text obtained by performing voice recognition on the user's conversational speech in the current round of voice dialogue by the preset model; 310. Output the reply text to the user.

[0080] In the above, the interface for health services can be triggered when the user opens the health-related intelligent dialogue robot, such as Figure 3 The service interface 21 shown in FIG.

[0081] The specific implementation of the above steps provided in this embodiment can refer to the relevant content in other embodiments, and will not be described in detail here. In addition, the method provided in this embodiment can also include some steps disclosed in other embodiments, and can also refer to the relevant content in other embodiments, and will not be described in detail here.

[0082] This specification also provides a method for training an intelligent conversation robot. The architecture of the intelligent conversation robot is deployed on the server, so that the training method is implemented by the server. The intelligent conversation robot includes a preset model and a speech synthesis model; the preset model contains a speech recognition module, a word segmenter, and a trained multimodal conversation model. For detailed descriptions of each module / function included in the preset model, please refer to the relevant content in other embodiments. Specifically, see Figure 7 As shown, the training method includes the following steps: 400. Freeze parameters of the trained multimodal dialogue model 402. Train the word segmenter using a first training sample set to obtain the word segmenter after first training; the first training sample set includes a plurality of first sample speech sounds and text contents corresponding to the first sample speech sounds; 404. Continuing to train the word segmenter after the first training using a second training sample set to obtain the word segmenter after the second training; the second sample set includes a plurality of second sample speech and speech features corresponding to the second sample speech; 406. Train the speech synthesis model using a third training sample set to obtain the trained speech synthesis model; the third training sample set includes a plurality of sample texts and speech corresponding to the sample texts; 408. Using a fourth training sample set, perform joint fine-tuning training on the speech recognition module, the word segmenter, the multimodal dialogue model, and the speech synthesis model to obtain a trained intelligent dialogue robot. In which, the fourth training sample set includes multiple third sample voices and fourth sample voices used to respond to the third sample voices; in the joint fine-tuning training process, the speech recognition module is used to convert the third sample voices into corresponding text content, the word segmenter is used to extract the speech features of the third sample voices, and the multimodal dialogue model uses the speech features output by the word segmenter and the text content output by the speech recognition module to generate reply text; the speech synthesis model is used to convert the reply text output by the multimodal dialogue model into reply speech; according to the loss value of the reply speech and the corresponding fourth sample voice, the parameters of the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model are jointly fine-tuned.

[0083] In the above 402, the first training sample set is used to implement speech-text alignment training for the word segmenter. The goal of this training of the word segmenter is to extract structured speech features from the first sample speech in the first training sample set and align them with the corresponding text in the first training sample set, so that the word segmenter can learn speech features from speech without text supervision.

[0084] For example, the word segmenter includes a first word segmenter (a language word segmenter), and the training process of the first word segmenter using the first training sample set can be, for example: the first sample speech in the first training sample set is input into the first word segmenter, and the first word segmenter outputs the language features extracted from the first sample speech (including phoneme-level features (phoneme sequences), speech rule features, etc.), and then analyzes the difference between the language features of the first sample speech output by the first word segmenter and the language features in the text content corresponding to the corresponding first sample speech in the first training sample set. For example, the difference between the phoneme sequence of the first sample speech output by the first word segmenter and the phoneme sequence in the text content corresponding to the corresponding first sample speech in the first training sample set is analyzed to optimize the parameters of the word segmenter based on this difference.

[0085] As can be seen from the above, the main goal of training the word segmenter this time is to make the speech features output by the word segmenter consistent with the speech features in the corresponding text, for example, to make the phoneme sequence output by the first word segmenter consistent with the phoneme sequence in the corresponding text.

[0086] In step 404 above, when training the word segmenter using the second training sample set, the training process may include, for example, inputting the second sample speech contained in the second training sample set into the word segmenter to obtain speech features of the second sample speech output by the word segmenter; then, calculating the loss between the speech features output by the word segmenter and the corresponding speech features in the second training sample set, and optimizing the parameters of the word segmenter based on this loss. The second sample speech contained in the second training sample set is data for the ASR task, and the main purpose of training the word segmenter is to enable the word segmenter to learn speech encoding capabilities (i.e., speech feature extraction capabilities).

[0087] In the above 406, when the speech synthesis (TTS) model is trained using the third training sample set, the training process may be, for example: inputting the sample text in the third training sample set into the speech synthesis model, executing the speech synthesis model to output the speech corresponding to the sample text, and then calculating the loss value between the speech corresponding to the sample text output by the speech synthesis model and the speech corresponding to the third sample text in the third training set, thereby optimizing the parameters of the speech synthesis model based on this loss value.

[0088] In step 408 above, the fourth training sample set is online voice dialogue interaction data. This fourth training sample set is used to perform end-to-end, full-link joint fine-tuning training on the intelligent conversational robot's speech recognition module, word segmenter, multimodal conversation module, and speech synthesis model. This allows for the coordinated optimization of module / model parameters throughout the entire process from voice input to voice output, improving the overall responsiveness of the intelligent conversational robot.

[0089] The specific implementation of the above steps provided in this embodiment can refer to the relevant content in other embodiments, and will not be described in detail here. In addition, the method provided in this embodiment can also include some steps disclosed in other embodiments, and can also refer to the relevant content in other embodiments, and will not be described in detail here.

[0090] Combined with the above Figures 4-7 This specification describes specific embodiments. It should be noted that for each specific embodiment described, other embodiments are within the scope of the appended claims, and that, in some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0091] The following introduces the device embodiments corresponding to the various method embodiments provided in this specification.

[0092] Figure 8 FIG. 1 shows a schematic diagram of the structure of an online human-computer dialogue device provided by an exemplary embodiment of this specification. Figure 8 As shown, the device includes: an acquisition module 52, an execution module 54, and an output module 56. An acquisition module 52 is used to acquire the conversation voice input by the user in the online human-computer voice dialogue; An execution module 54 is configured to input the conversational speech into a preset model, and have the preset model perform the following operations: extracting features from the conversational speech to obtain speech features of the conversational speech; and generating a response text based on the associated information of the online human-computer voice conversation and the speech features; wherein the associated information includes the conversation text obtained by performing speech recognition on the conversational speech by the preset model; The output module is used to output the reply text to the user.

[0093] In one embodiment, the above-mentioned execution module 54, when used to extract features of the conversational speech and obtain the speech features of the conversational speech, is specifically used to: extract language features from the conversational speech; the language features can reflect the pronunciation structure and language rules of the conversational speech; extract acoustic features from the conversational speech; the acoustic features can reflect the auditory perception attributes of the conversational speech.

[0094] In one embodiment, the preset model includes a first word segmenter and a second word segmenter. Furthermore, the execution module 54, when used to extract language features from the conversational speech, is specifically configured to: input the conversational speech into the first word segmenter, execute the first word segmenter to output the language features. When used to extract acoustic features from the conversational speech, the execution module 54 is specifically configured to: input the conversational speech into the second word segmenter, execute the second word segmenter to output the acoustic features.

[0095] In one embodiment, the phoneme-level features included in the language features reflect the pronunciation structure of the conversational speech. The acoustic features include at least one of the intonation, speech rate, emotion, timbre, and rhythm of the conversational speech.

[0096] In one embodiment, the associated information includes interactive text generated from at least one round of voice dialogue between a user and an intelligent conversational robot, wherein the interactive text includes the conversation text corresponding to the user in the current round of voice dialogue. Furthermore, the execution module 54, when configured to generate a reply text based on the associated information of the online human-machine voice dialogue and the voice features, is specifically configured to generate the reply text based on the interactive text generated from the at least one round of voice dialogue and the voice features.

[0097] In one embodiment, the associated information further includes user-related information. Furthermore, the execution module 54, when configured to generate the reply text based on the interactive text generated by the at least one round of voice conversation and the voice features, is specifically configured to generate the reply text based on the interactive text generated by the at least one round of voice conversation, the user-related information, and the voice features.

[0098] In one embodiment, the preset model includes a speech recognition module and a multimodal dialogue model. Furthermore, the execution module 54, when used to perform speech recognition on the conversation speech to obtain the dialogue text, is specifically configured to: utilize the speech recognition module to perform speech recognition on the conversation speech to obtain the dialogue text. When used to generate a response text based on the associated information of the online human-computer voice dialogue and the speech features, the execution module 54 is specifically configured to: input the associated information and the speech features into the multimodal dialogue model, and output the response text; wherein the multimodal dialogue model is constructed based on a language model.

[0099] Figure 9 FIG. 1 shows a schematic diagram of the structure of an online human-computer dialogue device provided by another exemplary embodiment of this specification. Figure 9 As shown, the device includes: a collection module 62, a sending module 64, a receiving module 66, and a broadcasting module 68. The collection module 62 is used to collect the conversational voice input by the user in the online human-computer voice dialogue. The sending module 64 is used to send the conversational voice to the server to trigger the server to input the conversational voice into a preset model, and the preset model performs the following operations: feature extraction of the conversational voice to obtain the voice features of the conversational voice; conversion of the conversational voice into conversational text; and generation of the reply text based on the conversational text and the voice features. The receiving module 66 is used to receive the reply voice returned by the server; wherein the reply voice is generated by performing speech synthesis on the reply text. The broadcasting module 68 is used to broadcast the reply voice to the user.

[0100] Figure 10 FIG. 1 shows a schematic diagram of the structure of an online human-computer dialogue device provided by another exemplary embodiment of this specification. Figure 10As shown, the device includes: a display module 72, a startup module 74, a collection module 76, an execution module 78, and an output module 710. The display module 72 is used to display an interface for health services. The startup module 74 is used to initiate an online human-machine voice dialogue in response to a voice dialogue startup operation triggered through the interface. The collection module 76 is used to collect the conversational speech input by the user during the online human-machine voice dialogue. The execution module 78 is used to input the conversational speech into a preset model, which then performs the following operations: feature extraction on the conversational speech to obtain speech features of the conversational speech; and generation of the reply text based on the associated information of the online human-machine voice dialogue and the speech features. The associated information includes the interactive text generated by at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactive text includes the conversational text obtained by the preset model performing voice recognition on the user's conversational speech in the current round of voice dialogue. The output module 710 is used to output the reply text to the user.

[0101] Figure 11 The diagram shows the structure of an intelligent dialogue robot training device provided by another exemplary embodiment of this specification. The intelligent dialogue robot includes a preset model and a speech synthesis model; the preset model contains a speech recognition module, a word segmenter, and a trained multimodal dialogue model. Figure 11As shown, the training device includes: a freezing module 82 and a training module 84. The freezing module 82 is used to freeze the parameters of the multimodal dialogue model that has been trained. The training module is used to train the word segmenter based on the first training sample set to obtain the word segmenter after the first training; the first training sample set contains a plurality of first sample speech and text content corresponding to the first sample speech; based on the second training sample set, continue to train the word segmenter after the first training to obtain the word segmenter after the second training; the second training sample set contains a plurality of second sample speech and speech features corresponding to the second sample speech; based on the third training sample set, train the speech synthesis model to obtain the trained speech synthesis model; the third training sample set contains a plurality of sample texts and speech corresponding to the sample texts; based on the fourth training sample set, link the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model. Combined fine-tuning training is performed to obtain a trained intelligent dialogue robot; wherein the fourth training sample set includes multiple third sample voices and fourth sample voices used to respond to the third sample voices; in the joint fine-tuning training process, the speech recognition module is used to convert the third sample voices into corresponding text content, the word segmenter is used to extract the speech features of the third sample voices, and the multimodal dialogue model uses the speech features output by the word segmenter and the text content output by the speech recognition module to generate a reply text; the speech synthesis model is used to convert the reply text output by the multimodal dialogue model into a reply speech; according to the loss value of the reply speech and the corresponding fourth sample voice, the parameters of the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model are jointly fine-tuned.

[0102] What needs to be explained here about each of the above-mentioned devices is that: each of the devices provided above can implement the technical solutions described in the corresponding method embodiments above. The specific implementation principles of each of the above-mentioned modules or units can refer to the relevant content in the corresponding method embodiments above, and will not be described in detail here. In addition, for the convenience of description, the above devices are described by being divided into various modules or units according to their functions. Of course, when implementing one or more of the present specifications, the functions of each module or unit can be implemented in the same or more software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0103] In addition, the embodiments of this specification also provide an electronic device. Figure 12 As shown, the electronic device 900 includes a memory 91 and a processor 92 .

[0104] The memory 91 can be implemented by at least one volatile or nonvolatile memory device of any type, or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. Furthermore, all or part of the memory can be integrated with the processor. The memory can include both removable and non-removable components.

[0105] The processor 92 may include one or more general-purpose processors and / or special-purpose processors.

[0106] Furthermore, the memory 91 may include a non-transitory computer-readable medium having executable program instructions 912 (e.g., compiled or non-compiled program logic and / or machine code) stored therein. The processor 92 is capable of executing the program instructions 912 stored in the memory to implement any method, process, or function disclosed in this specification and / or the accompanying drawings. Furthermore, the execution of the program instructions 912 by the processor 92 may cause the processor to use corresponding data 911.

[0107] For example, the program instructions 912 may include an operating system 9122 (e.g., an operating system kernel, device drivers, and / or other modules) and one or more applications 9121 (e.g., a browser, a social networking application, or a gaming application) installed on the electronic device 900. Similarly, the data 911 may include operating system data 9112 and application data 9111. The operating system data 9112 is primarily accessible to the operating system 9122, while the application data 9111 is primarily accessible to one or more applications 9121. The application data 9111 may be located in a file system visible to or hidden from the user of the electronic device 900.

[0108] The application 9121 can communicate with the operating system 9122 through one or more application programming interfaces (APIs). These APIs help the application 9122 read and / or write application data, transmit or receive information via communication components, receive or display information on a user interface, etc. In some terms, the application 9121 can be simply referred to as an "app". In addition, the application 9121 can be downloaded to the electronic device through one or more online application stores or application markets. However, the application 9121 can also be installed on the electronic device 400 through other means, such as through a web browser or a physical interface on the electronic device 900 (e.g., a USB port). Further, if Figure 12 As shown, the electronic device also includes: a communication component 93, a display 94, a power component 95, an audio component 96, a user interface 99 and other components. Figure 12 Only some components are shown schematically, which does not mean that the electronic device 900 only includes Figure 12 In addition, Figure 12 The components in the dotted box are optional components, not mandatory components, and the specific components depend on the product form of the electronic device 900. The electronic device 900 of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT device, or a server device such as a conventional server, a cloud server or a server array, or an integrated device of a terminal device and a server device. If the electronic device 900 of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it can include Figure 12 If the electronic device 900 of this embodiment is implemented as a conventional server, cloud server or server array and other server devices, it may not include Figure 12 Components within the dotted box.

[0109] The communication component 93 is configured to facilitate wired or wireless communication between the device in which the communication component resides and other devices. The device in which the communication component 93 resides can access a wireless network based on a communication standard, such as a 2G, 3G, 4G / LTE, 5G, or other mobile communication network, or a combination thereof. In one exemplary embodiment, the communication component 93 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In a specific implementation, the communication component 93 includes a communication interface that enables the electronic device 900 to communicate with other devices, access networks, and transmission networks using analog or digital modulation. For example, the communication interface may include a chipset and an antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface may be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, a Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface may also support other physical layer interfaces and standard or proprietary communication protocols. The communication interface may also include multiple physical communication interfaces, such as a Wi-Fi interface, a Bluetooth interface, and a wide-area wireless interface.

[0110] The display 94 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can detect not only the boundaries of a touch or slide action, but also the duration and pressure associated with the touch or slide action.

[0111] The power supply assembly 95 provides power to various components of the device in which it is located. The power supply assembly 95 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.

[0112] The audio component 96 can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in a memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0113] The user interface 97 described above includes both receiving user input and providing output to the user. Thus, the user interface 97 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensing panel, computer mouse, trackball, joystick, microphone, still camera, and video camera. It may also include output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, or other similar devices known or developed in the future. The user interface 97 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, or other similar devices known or developed in the future. In certain embodiments, the user interface 97 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, the electronic device 900 may support remote access from other devices, such as via a communication interface or another physical interface (not shown). The user interface 97 may be configured to receive user input, and its position and movement may be indicated by an indicator or cursor as described herein. The user interface 97 may also be configured as a display device for rendering or displaying text snippets.

[0114] Accordingly, an embodiment of the present specification also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement each step in the above method embodiment. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission medium. Furthermore, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer is caused to execute the following Figures 3 to 5 Described method.

[0115] The embodiments of this specification also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the following Figures 3 to 5 Described method.

[0116] Those skilled in the art will appreciate that, in one or more of the above examples, the functions described in the various embodiments disclosed in this specification may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0117] The specific implementation methods described above further illustrate in detail the purposes, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above description is only the specific implementation methods of the multiple embodiments disclosed in this specification, and is not intended to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.

Claims

1. An online human-computer dialogue method, characterized in that: include: Obtain the conversation voice input by the user in the online human-computer voice dialogue; Inputting the conversational speech into a preset model, the preset model then performs the following steps: extracting features from the conversational speech to obtain speech features of the conversational speech; generating a response text based on the associated information of the online human-computer voice conversation and the speech features; wherein the associated information includes the conversation text obtained by performing speech recognition on the conversational speech by the preset model; The reply text is output to the user.

2. The method according to claim 1, characterized in that Extracting features of the conversational speech to obtain speech features of the conversational speech includes: Extracting language features from the conversational speech; the language features can reflect the pronunciation structure and language rules of the conversational speech; Acoustic features are extracted from the conversational speech; the acoustic features can reflect the auditory perception attributes of the conversational speech.

3. The method according to claim 2, characterized in that The preset model includes a first word segmenter and a second word segmenter; And, extracting language features from the conversation speech, including: Inputting the conversation speech into the first word segmenter, executing the first word segmenter to output the language features; Extracting acoustic features from the conversation speech includes: The conversation speech is input into the second word segmenter, and the second word segmenter is executed to output the acoustic features.

4. The method according to claim 2, characterized in that The phoneme-level features included in the language features reflect the pronunciation structure of the conversational speech; The acoustic feature includes at least one of the intonation, speaking speed, emotion, timbre, and rhythm of the conversational speech.

5. The method according to any one of claims 1 to 4, characterized in that The associated information includes the interactive text generated by at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactive text includes the dialogue text corresponding to the user in the current round of voice dialogue; as well as, Generating a reply text based on the associated information of the online human-computer voice dialogue and the voice features, including: The reply text is generated based on the interactive text generated by the at least one round of voice dialogue and the voice features.

6. The method according to claim 5, characterized in that The associated information also includes relevant information of the user; And, generating the reply text based on the interactive text generated by the at least one round of voice dialogue interaction and the voice features, including: The reply text is generated based on the interaction text generated by the at least one round of voice dialogue interaction, the relevant information of the user and the voice features.

7. The method according to any one of claims 1 to 4, characterized in that The preset model includes a speech recognition module and a multimodal dialogue model; And, performing speech recognition on the conversation speech to obtain the conversation text, including: Using the speech recognition module, performing speech recognition on the conversation speech to obtain the conversation text; Generating a reply text based on the associated information of the online human-computer voice dialogue and the voice features, including: The associated information and the speech features are input into the multimodal dialogue model, and the reply text is output; wherein the multimodal dialogue model is constructed based on a language model.

8. An online human-computer dialogue method, characterized in that: include: Collecting the conversational voice input by the user in the online human-computer voice dialogue; The conversation voice is sent to the server to trigger the server to input the conversation voice into a preset model, and the preset model performs: feature extraction on the conversation voice to obtain voice features of the conversation voice; Converting the conversation speech into conversation text; generating a reply text based on the conversation text and the voice features; Receiving a reply voice returned by the server; wherein the reply voice is generated by performing speech synthesis on the reply text; The reply voice is broadcast to the user.

9. An online human-computer dialogue method, characterized in that: include: Displays an interface for health services; In response to a voice dialogue initiation operation triggered through the interface, initiating an online human-computer voice dialogue; In the online human-computer voice dialogue, collecting the dialogue voice input by the user; The conversational speech is input into a preset model, and the preset model performs the following steps: extracting features from the conversational speech to obtain speech features of the conversational speech; and generating a response text based on the associated information of the online human-machine speech dialogue and the speech features; wherein the associated information includes the interactive text generated by at least one round of speech dialogue between the user and the intelligent dialogue robot, and the interactive text includes the conversation text obtained by performing speech recognition on the user's conversational speech in the current round of speech dialogue by the preset model; The reply text is output to the user.

10. A method for training an intelligent conversational robot, characterized in that: The intelligent dialogue robot includes a preset model and a speech synthesis model; the preset model contains a speech recognition module, a word segmenter, and a trained multimodal dialogue model; The method comprises: Freezing parameters of the trained multimodal dialogue model; Training the word segmenter based on a first training sample set to obtain the word segmenter after the first training; the first training sample set includes a plurality of first sample speech and text content corresponding to the first sample speech; Based on a second training sample set, the word segmenter after the first training is further trained to obtain the word segmenter after the second training; the second training sample set includes a plurality of second sample speech and speech features corresponding to the second sample speech; Training the speech synthesis model based on a third training sample set to obtain the trained speech synthesis model; the third training sample set includes a plurality of sample texts and speech corresponding to the sample texts; Based on the fourth training sample set, jointly fine-tune the speech recognition module, the word segmenter, the multimodal dialogue model, and the speech synthesis model to obtain a trained intelligent dialogue robot; In which, the fourth training sample set includes multiple third sample voices and fourth sample voices used to respond to the third sample voices; in the joint fine-tuning training process, the speech recognition module is used to convert the third sample voices into corresponding text content, the word segmenter is used to extract the speech features of the third sample voices, and the multimodal dialogue model uses the speech features output by the word segmenter and the text content output by the speech recognition module to generate reply text; the speech synthesis model is used to convert the reply text output by the multimodal dialogue model into reply speech; according to the loss value of the reply speech and the corresponding fourth sample voice, the parameters of the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model are jointly fine-tuned.

11. An online human-computer dialogue system, characterized in that: include: The client is used to collect the conversation voice input by the user in the online human-computer voice dialogue and send the conversation voice to the server; The server is deployed with a preset model, configured to input the conversational speech into the preset model, and have the preset model perform the following operations: extracting features from the conversational speech to obtain speech features of the conversational speech; and generating a response text based on the associated information of the online human-computer voice conversation and the speech features; wherein the associated information includes the conversation text obtained by performing speech recognition on the conversational speech by the preset model; The server is further configured to convert the reply text into reply voice and send the reply voice to the client; The client is also used to broadcast the reply voice to the user.

12. An online human-computer dialogue device, characterized in that: include: An acquisition module is used to acquire the conversation voice input by the user in the online human-computer voice dialogue; an execution module configured to input the conversational speech into a preset model, and have the preset model perform the following operations: extracting features from the conversational speech to obtain speech features of the conversational speech; and generating a response text based on the associated information of the online human-computer voice conversation and the speech features; wherein the associated information includes the conversation text obtained by performing speech recognition on the conversational speech by the preset model; The output module is used to output the reply text to the user.

13. An online human-computer dialogue device, characterized in that: include: The acquisition module is used to collect the conversation voice input by the user in the online human-computer voice dialogue; a sending module, configured to send the conversational speech to a server, so as to trigger the server to input the conversational speech into a preset model, and for the preset model to perform feature extraction on the conversational speech to obtain speech features of the conversational speech; Converting the conversation speech into conversation text; generating a reply text based on the conversation text and the voice features; A receiving module, configured to receive the reply voice returned by the server; wherein the reply voice is generated by performing speech synthesis on the reply text; The broadcast module is used to broadcast the reply voice to the user.

14. An online human-computer dialogue device, characterized in that: include: A display module, used for displaying an interface for health services; A starting module, configured to respond to a voice dialogue starting operation triggered through the interface and start an online human-computer voice dialogue; A collection module, used to collect the conversation voice input by the user in the online human-computer voice dialogue; An execution module is configured to input the conversational speech into a preset model, and the preset model executes the following: feature extraction of the conversational speech to obtain speech features of the conversational speech; and generation of a response text based on the associated information of the online human-machine voice dialogue and the speech features; wherein the associated information includes interactions generated by at least one round of voice dialogue between the user and the intelligent dialogue robot, and the interactions include the conversation text obtained by performing voice recognition of the user's conversational speech in the current round of voice dialogue by the preset model; The output module is used to output the reply text to the user.

15. An intelligent dialogue robot training device, characterized in that: The intelligent dialogue robot includes a preset model and a speech synthesis model; the preset model contains a speech recognition module, a word segmenter, and a trained multimodal dialogue model; The device comprises: A freezing module, configured to freeze parameters of the trained multimodal dialogue model; The training module is used to: train the word segmenter based on the first training sample set to obtain the word segmenter after the first training; the first training sample set includes a plurality of first sample speech and text content corresponding to the first sample speech; based on the second training sample set, continue to train the word segmenter after the first training to obtain the word segmenter after the second training; the second training sample set includes a plurality of second sample speech and speech features corresponding to the second sample speech; based on the third training sample set, train the speech synthesis model to obtain the trained speech synthesis model; the third training sample set includes a plurality of sample texts and speech corresponding to the sample texts; based on the fourth training sample set, link the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model Combined fine-tuning training is performed to obtain a trained intelligent dialogue robot; wherein the fourth training sample set includes multiple third sample voices and fourth sample voices used to respond to the third sample voices; in the joint fine-tuning training process, the speech recognition module is used to convert the third sample voices into corresponding text content, the word segmenter is used to extract the speech features of the third sample voices, and the multimodal dialogue model uses the speech features output by the word segmenter and the text content output by the speech recognition module to generate a reply text; the speech synthesis model is used to convert the reply text output by the multimodal dialogue model into a reply speech; according to the loss value of the reply speech and the corresponding fourth sample voice, the parameters of the speech recognition module, the word segmenter, the multimodal dialogue model and the speech synthesis model are jointly fine-tuned.

16. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores executable program instructions, and when the processor executes the program instructions, the method according to any one of claims 1 to 10 is implemented.

17. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 10.

18. A computer program product, characterized in that The computer program product includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Voice conversation method and device, and computer readable storage medium

    CN111986675A

  • Spoken language understanding method and device combined with voice information, equipment and storage medium

    CN114021582A

  • Dialogue text processing method and device, electronic equipment and storage medium

    CN117076648A

  • Dialogue method and device based on voice recognition, terminal equipment and storage medium

    CN118447841A

  • Voice recognition-based interaction method, apparatus and device, and storage medium

    CN120199252A