Question-answering method and apparatus, and device and storage medium

By recognizing the question text and identifying matching auxiliary information on the client side, and generating response text or response voice, the problems of response accuracy and data security in AI question-answering applications are solved, achieving more efficient question-answering services and data protection.

WO2026002003A1PCT designated stage Publication Date: 2026-01-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/103293
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

In AI-powered question-answering applications, the service capabilities of models still need improvement, especially during the interaction process, where existing technologies struggle to effectively enhance the accuracy of responses and data security.

Method used

By acquiring the question's voice and recognizing the question's text on the client side, identifying auxiliary information that matches the question's text, and sending it to the server to generate a response text or voice, the system avoids storing auxiliary information on the server or in the cloud, improving response accuracy and ensuring data security.

Benefits of technology

It improves service capabilities in question-and-answer scenarios, ensures the accuracy of responses, and protects the data security of auxiliary information by avoiding storage on the server or in the cloud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025103293_02012026_PF_FP_ABST
    Figure CN2025103293_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A question-answering method and apparatus, and a device and a storage medium. The method comprises: in response to question speech of a user having been received, acquiring question text identified from the question speech (410); on the basis of the question text, determining assistance information which matches the question text (420); sending the assistance information to a serving end (430); and receiving from the serving end at least one of response text or response speech for the question speech (440), wherein the response speech corresponds to the response text, and the response text is determined on the basis of the question text and the assistance information. In this way, a user-friendly service capability in a human-computer interaction scenario can be improved, and data security can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, apparatus and storage medium for question answering

[0001] The present application claims priority to the Chinese patent application No. 202410869789.8, filed on June 28, 2024, entitled "Method, device, apparatus and storage medium for question answering", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The example embodiments of the present disclosure generally relate to the field of computer, and in particular, to a method, device, apparatus and computer readable storage medium for question answering. BACKGROUND

[0003] With the development of artificial intelligence technology, various types of application forms have emerged, including but not limited to text question answering, voice question answering, and text-to-image applications. For the analysis of such applications, it is found that the service capability of the model in the interaction process still needs to be improved. SUMMARY

[0004] In a first aspect of the present disclosure, a method for question answering is provided. The method comprises: in response to receiving a question voice of a user, obtaining a question text recognized from the question voice; determining auxiliary information matched with the question text based on the question text; sending the auxiliary information to a server; and receiving at least one of a response text or a response voice for the question voice from the server, the response voice corresponding to the response text, and the response text being determined based on the question text and the auxiliary information.

[0005] In a second aspect of the present disclosure, a method for question answering is provided. The method comprises: in response to receiving a question voice of a user from a client, sending a question text recognized from the question voice to the client; receiving auxiliary information matched with the question text recognized from the question voice from the client; determining a response text based on the question text and the auxiliary information; and feeding back at least one of the response text or a response voice to the client, the response voice corresponding to the response text.

[0006] In a third aspect of the present disclosure, a device for question answering is provided. The device comprises: a question text obtaining module configured to, in response to receiving a question voice of a user, obtain a question text recognized from the question voice; an auxiliary information determining module configured to determine auxiliary information matched with the question text based on the question text; an auxiliary information sending module configured to send the auxiliary information to a server; and a response receiving module configured to receive at least one of a response text or a response voice for the question voice from the server, the response voice corresponding to the response text, and the response text being determined based on the question text and the auxiliary information.

[0007] In a fourth aspect of the present disclosure, an apparatus for question answering is provided. The apparatus includes: a question text sending module configured to send, to a client, question text recognized from a question voice of a user in response to receiving the question voice from the client; an auxiliary information receiving module configured to receive, from the client, auxiliary information matched with the question text recognized from the question voice; an answer text determining module configured to determine answer text based on the question text and the auxiliary information; and an answer feedback module configured to feed back, to the client, at least one of the answer text or answer voice corresponding to the answer text.

[0008] In a fifth aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the electronic device to perform the method of the first aspect or the method of the second aspect.

[0009] In a sixth aspect of the present disclosure, a computer-readable storage medium is provided. The medium has stored thereon computer-executable instructions that, when executed by a processor, implement the method of the first aspect or the method of the second aspect.

[0010] In a seventh aspect of the present disclosure, a computer program product is provided. The computer program product includes computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method of the first aspect or the method of the second aspect according to the present disclosure.

[0011] It should be understood that the description in this section is not intended to define key or essential features of embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0012] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which:

[0013] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0014] FIG. 2 shows a flow diagram of a signaling flow for question answering according to some embodiments of the present disclosure;

[0015] FIG. 3 shows a schematic diagram of an example architecture for question answering according to some embodiments of the present disclosure;

[0016] FIG. 4 shows a flow diagram of a process for question answering according to some embodiments of the present disclosure;

[0017] FIG. 5 shows a flowchart of a process for question answering, according to some embodiments of the present disclosure;

[0018] FIG. 6 shows an exemplary structural block diagram of an apparatus for question answering, according to some embodiments of the present disclosure;

[0019] FIG. 7 shows an exemplary structural block diagram of an apparatus for question answering, according to some embodiments of the present disclosure; and

[0020] FIG. 8 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It will be appreciated that the drawings of the present disclosure and the embodiments are only for exemplary purposes and should not be construed as limiting the scope of protection of the present disclosure.

[0022] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meanings of "including but not limited to", i.e., "comprising but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.

[0023] In this document, unless explicitly stated otherwise, performing a step "in response to" A does not mean that the step is performed immediately after A, but can include one or more intermediate steps.

[0024] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, obtaining, using, storing or deleting of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0025] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant users and the authorization of the relevant users should be obtained through appropriate means, wherein the relevant users can include any type of right subject, such as individuals, enterprises and groups.

[0026] For example, in response to receiving an active request of a user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require obtaining and using information of the relevant user, so that the relevant user can autonomously select whether to provide information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solution of the present disclosure according to the prompt information.

[0027] As an optional but non-limiting implementation manner, in response to receiving an active request of a relevant user, the manner of sending a prompt information to the relevant user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.

[0028] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0029] As used herein, the term "model" can learn an association between respective inputs and outputs from training data, such that a corresponding output can be generated for a given input after training is completed. The generation of a model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is one example of a model based on deep learning. In this document, a "model" can also be referred to as a "machine learning model", a "learning model", a "machine learning network", or a "learning network", which terms are used interchangeably herein.

[0030] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, which typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence, such that the output of a previous layer is provided as input to a subsequent layer, with the input layer receiving the input to the neural network and the output of the output layer as the final output of the neural network. Each layer of a neural network includes one or more nodes (also referred to as processing nodes or neurons), each of which processes input from the previous layer.

[0031] Generally, machine learning can include three stages, namely a training stage, a testing stage, and an application stage (also referred to as an inference stage). In the training stage, a given model can be trained using a large amount of training data, iteratively updating parameter values until the model is able to derive consistent inferences from the training data that satisfy an intended objective. Through training, the model can be considered to have learned an association (also referred to as a mapping) from input to output from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to determine whether the model is able to provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs based on the trained parameter values to determine corresponding outputs.

[0032] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example environment 100, an application 120 is installed in a client 110. A user 140 can interact with the application 120 via the client 110 and / or an attached device of the client 110. For example, the application 120 can capture a voice 145 of the user 140 via a voice capturing device (e.g., a microphone) of the client 110.

[0033] In embodiments of the present disclosure, the application 120 can be any suitable application with a question-answering function. For example, the application 120 can provide question-answering with a digital assistant. The digital assistant supports text question-answering services, voice question-answering services, and content question-answering in other modalities with the user 140. In some embodiments, the application 120 or the digital assistant therein can utilize a machine learning model 160 (which can include one or more machine learning models, e.g., can include a machine learning model 160-1, a machine learning model 160-2, …, a machine learning model 160-N, etc., where N is a positive integer) to support interactions with the user 140. For example, the application 120 or the digital assistant therein can utilize one or more machine learning models 160 to provide question-answering services to the user 140.

[0034] In the environment 100, the client 110 can present a user interface 150 of the application 120 if the application 120 is in an active state. The user interface 150 can include various pages that the application 120 is capable of providing, such as a question-answering page with the digital assistant, etc. In some embodiments, the client 110 can play a voice 152 and present text 154 in the user interface 150. The voice 152 can include, for example, the voice 145 from the user 140 or a voice of a response to the voice 145.

[0035] The machine learning models 160 can be different types of models. In some embodiments, one or more machine learning models 160 can be built based on a language model (LM). The machine learning models used are content generative models, capable of generating corresponding outputs based on model inputs. In some embodiments, the language model based machine learning models are capable of textual modality model inputs (e.g., natural language and / or machine language) and / or non-textual modality model inputs (e.g., images, speech, videos, etc.), and are capable of generating desired outputs from the model inputs and a prompt word. The prompt word here is used to guide the machine learning model to generate an output that is capable of addressing a user need indicated by the model inputs. In an application scenario for supporting user question answering, the input of the user 140 can be provided as at least a portion of the model inputs (other portions can include the prompt word) to the machine learning model 160. The input of the user is considered as a question. Based on the model output, a corresponding answer can be generated to be provided to the user 140.

[0036] In some embodiments, one or more machine learning models 160 can be voice related models, including an automatic speech recognition (ASR) model and a text-to-speech (TTS) model. The input of the ASR model is speech, and the output is text. The input of the TTS model is text, and the output is corresponding speech.

[0037] In some embodiments, the client 110 communicates with the server 130 to implement provisioning of services of the application 120. As shown in FIG. 1, the server 130 can invoke the machine learning models 160 to support question answering functionality between the application 120 and the user 140 based on outputs of the machine learning models 160. The client 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including an accessory or peripheral device thereof or any combination thereof. In some embodiments, the client 110 can also support any type of interface for the user (such as “wearable” circuitry, etc.). The server 130 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, etc. The server 130 can be implemented based on a cloud environment, for example.

[0038] It should be understood that the structure and functionality of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0039] As mentioned above, with the development of artificial intelligence technology, various types of application forms have emerged, including but not limited to text question answering, voice question answering, and text-to-image applications. For the analysis of such applications, it is found that the service capability of the model in the interaction process still needs to be improved.

[0040] Therefore, according to an embodiment of the present disclosure, an improved solution for question answering is provided. According to the solution of the embodiment of the present disclosure, in response to receiving a question voice of a user, a question text recognized from the question voice is obtained; based on the question text, auxiliary information matched with the question text is determined; the auxiliary information is sent to a server; and at least one of a response text or a response voice for the question voice is received from the server, the response voice corresponds to the response text, and the response text is determined based on the question text and the auxiliary information.

[0041] In this way, in the question answering scenario, by determining the auxiliary information matched with the question text based on the question text by the client and sending the auxiliary information to the server, the response text and the response voice can be obtained for the question text and the auxiliary information, which can improve the accuracy of the response content, thereby improving the service capability in the question answering scenario, and avoiding saving the auxiliary information in the server or the cloud, which is conducive to protecting the data security of the auxiliary information.

[0042] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0043] FIG. 2 shows a flowchart of a signaling flow 200 for question answering according to some embodiments of the present disclosure. The signaling flow 200 involves the client 110 and the server 130. FIG. 3 shows a schematic diagram of an example architecture 300 for question answering according to some embodiments of the present disclosure. For ease of discussion, the signaling flow 200 will be described with reference to the environment of FIG. 1 and in conjunction with the example architecture 300 shown in FIG. 3.

[0044] In an embodiment of the present disclosure, as shown in the signaling flow 200, in response to receiving a question voice of a user 140, the client 110 sends (202) the question voice to the server 130. The server 130 receives (204) the question voice of the user sent by the client 110, and recognizes a question text from the question voice. The server 130 sends (208) the recognized question text to the client 110, and the client 110 receives (210) the question text from the server 130.

[0045] As shown in FIG. 3, the client 110 can include a business module 310. The business module 310 can correspond to an application that supports voice question answering. The business module 310 includes an application voice module 311 and an application message module 312. The server 130 can include an access layer 340, a voice module 350, and a question answering module 360. The access layer 340 includes a network interface 341 and a network interface 342. A first network connection can be built between the network interface 341 of the client 110 and the server 130. A second network connection can be built between the network interface 342 of the client 110 and the server 130. Alternatively or additionally, the first network connection and / or the second network connection can be a long connection in compliance with a transmission control protocol (TCP). Of course, the first network connection and the second network connection are not limited to the long connection in compliance with the TCP protocol, and the first network connection and the second network connection can also be network connections in compliance with other communication protocols. In a specific implementation, a selection configuration can be made according to actual needs.

[0046] The application voice module 311 can receive a voice (e.g., the voice 145) from a user (e.g., the user 140). In some embodiments, the application voice module 311 can receive a voice (i.e., “user voice”, also referred to as “question voice” in this document) from the user 140. The application voice module 311 can send the user voice to the network interface 341 of the server 130 via the first network connection, and send the user voice to the voice module 350 of the server 130 through the network interface 341.

[0047] The voice module 350 can perform an ASR function and / or a TTS function by means of an ASR model 371 and / or a TTS model 372 in a voice service 370. The ASR model 371 and / or the TTS model 372 are deployed on the server 130, and can also be referred to as the ASR model 371 and / or the TTS model 372 being deployed on the cloud. Exemplarily, the voice module 350 can send the user voice to the ASR model 371, and obtain an ASR text (i.e., a question text converted from the question voice) corresponding to the user voice from the ASR model 371.

[0048] The voice module 350 can send the obtained ASR text to the network interface 341 of the access layer 340. The access layer 340 can in turn send the obtained ASR text to the application voice module 311 of the client 110. Alternatively or additionally, the access layer 340 can send the ASR text to the client 110 via the long connection between the network interface 341 and the client 110.

[0049] In some embodiments, the application speech module 311 of the client 110 can also provide the user speech to an ASR model 321 (also referred to as a“first machine learning model” herein sometimes) in the local model capability 320 of the client 110, perform speech recognition on the user speech with the ASR model 321, and generate ASR text corresponding to the user speech. In this way, the speed of obtaining the ASR text can be improved. When the network communication capability is low, it can be ensured that the client 110 can effectively obtain the ASR text corresponding to the user speech, and thus ensure that the question and answer can proceed normally.

[0050] In some embodiments, the application speech module 311 can send the ASR text to the application message module 312 in response to receiving the ASR text from the server 130 or receiving the ASR text from the ASR model 321. The application message module 312 can present the ASR text to the user 140. For example, the application message module 312 can present the ASR text (i.e., present the first question text converted from the first question speech) in a user interface.

[0051] In embodiments of the present disclosure, as shown in the signaling flow 200, the client 110 determines (212) the auxiliary information matching the question text based on the question text recognized from the question speech. The client 110 sends (214) the auxiliary information to the server 130, and the server 130 receives (216) the auxiliary information from the client 110.

[0052] It can be understood that the question text in the auxiliary information determined by the client 110 to match the question text can be the ASR text received by the client 110 from the server 130, or the ASR text obtained by the client 110 from the ASR model 321 in the local model capability 320.

[0053] In some embodiments, the client 110 can be locally deployed with an auxiliary information library 323, and the auxiliary information library 323 can store at least one category of auxiliary information. The client 110 can determine the auxiliary information matching the question text from the auxiliary information library 323 based on the question text. The auxiliary information herein can include various information related to the current user, which can assist in improving the intention understanding of the question and answer model 380 to the user question and improving the question and answer capability and the question and answer satisfaction in the question and answer scenario. In some embodiments, the auxiliary information can be information provided by the user in advance. For example, the at least one category of auxiliary information can include but is not limited to a profile file, schedule information, a knowledge base of a field of interest, and the like provided by the user in advance. In some embodiments, the auxiliary information can be summarized and recorded from historical question and answer records, and the like. With the auxiliary information, the model for question and answer can generate a response that is more satisfactory and more in line with the user's expectation.

[0054] It should be noted that the auxiliary information in the embodiments of the present disclosure is obtained, stored and used with the authorization of the user. In the embodiments of the present disclosure, various types of auxiliary information for the current user are stored locally on the client, and one or more auxiliary information is provided to the server when needed, which can ensure data security.

[0055] In some embodiments, the client 110 provides the question text to an information determination model 322 (also referred to as a “second machine learning model” herein in some cases) in the local model capability 320, and determines the auxiliary information matching the ASR text by using the information determination model 322. The application voice module 311 can receive the auxiliary information fed back by the information determination model 322. In some embodiments, the information determination model 322 can be a text matching model to perform text matching between the ASR text corresponding to the question voice and the auxiliary information, so as to determine the most matching auxiliary information.

[0056] Alternatively or additionally, the information determination model 322 can determine the auxiliary information matching the ASR text from the auxiliary information library 323 for the user 140 based on the ASR text. Alternatively or additionally, the ASR model 321 or the ASR model 371 can be configured to perform text recognition on the user voice in a streaming manner, sequentially generate at least one text sequence of the ASR text, and the ASR model 321 or the ASR model 371 can provide the at least one text sequence of the ASR text to the application voice module 311 in the generation order. The application voice module 311 can provide the at least one text sequence of the ASR text to the information determination model 322 in the generation order, and the information determination model 322 can determine the matching auxiliary information based on all or part of the text sequences in the at least one text sequence.

[0057] Exemplarily, when adding the auxiliary information to the auxiliary information library 323, the category corresponding to the auxiliary information can be determined, the auxiliary information is added to the auxiliary information library 323, and the label for labeling the category corresponding to the auxiliary information is also added to the auxiliary information library 323. The information determination model 322 can determine the category corresponding to the ASR text based on the received at least one text sequence. Then, the corresponding auxiliary information can be retrieved from the auxiliary information library 323 based on the determined category corresponding to the ASR text.

[0058] It should be understood that the information determination model 322 is not limited to determining the auxiliary information matching the ASR text from the auxiliary information library 323, but also determines the auxiliary information matching the ASR text from the network or the remote device. For example, the information determination model 322 can also obtain various information such as real-time information, news events, weather information or historical knowledge as auxiliary information matching the ASR text.

[0059] In some embodiments, after the application speech module 311 of the client 110 obtains the auxiliary information, the application speech module 311 can send the auxiliary information to the network interface 341 of the server via the first network connection, and the network interface 341 can provide the auxiliary information to the speech module 350.

[0060] In embodiments of the present disclosure, as shown in the signaling flow 200, the server 130 determines (218) the response text based on the question text and the auxiliary information. The server 130 sends (220) the response text to the client 110, and the client 110 receives the response text from the server 130.

[0061] In some embodiments, the server 130 can generate a model input for a question model, i.e., a question and answer model 380 (sometimes also referred to as a “third machine learning model” herein) based on the question text and the auxiliary information. The server 130 can invoke the question and answer model 380 to generate the response text based on the model input.

[0062] Alternatively or additionally, the speech module 350 in the server 130 can also construct a sending message based on the obtained ASR text, send the sending message and the auxiliary information to the question and answer module 360. The question and answer module 360 can generate a model input based on the sending message and the auxiliary information, and the question and answer module 360 can invoke the question and answer model 380 to generate the response text based on the model input. The question and answer module 360 can receive the response text from the question and answer model 380, and the question and answer module 380 can send the response text to the access layer 340, and the network interface 342 of the access layer 340 can send the response text to the application message module 312 of the client 110 via the second network connection. In this way, the ASR text and the auxiliary message can be passed between the speech module 350 and the question and answer module 360 of the server 130, which is conducive to simplifying the interaction process between the server 130 and the client 110.

[0063] For example, assume that the auxiliary information library 323 can store historical question and answer records of the user. The historical question and answer records include question and answer record A "What is the weather today?", and the context record of the question and answer record A includes question and answer record B "What clothes are suitable to wear?". Assume that the ASR text includes "What is the weather today?", the information determination model 322 can take the question and answer record A, the question and answer record B, and the context relationship between the question and answer record A and the question and answer record B as auxiliary information. The client 110 can upload the auxiliary information to the server 130, and the question and answer model 380 can splice the ASR text "What is the weather today?" and the auxiliary information "What is the weather today?", "What clothes are suitable to wear?" to form a model input. The response text generated by the question and answer model 380 based on the model input can include "It is sunny today, the temperature is XX, and it is suitable to wear light clothes", or "It is cloudy today, the temperature is XX, and it is suitable to wear warm clothes", and the like. In this way, the question and answer model 380 more accurately understands the user's question intention, can reduce the number of user questions, and can improve the question and answer efficiency.

[0064] For another example, assume that the auxiliary information library 323 can include schedule information provided by the user, and the user adds schedule information reminders in the application 120, which can include XXXX-XX-XX, and go on a business trip from city A to city B. Assume that the user 140 asks (i.e., the ASR text) "What is the weather today?" on XXXX-XX-XX, and the information determination model 322 can take the schedule information reminder as auxiliary information for the user. The question and answer model 380 can splice the ASR text "What is the weather today?" and the auxiliary information "XXXX-XX-XX, and go on a business trip from city A to city B" as a model input. The response text generated by the question and answer model 380 based on the model input can include "The weather in city A today is XXXXXX, and the weather in city B today is XXXXXX", and the like. In this way, although the user 140's question is very general, the question and answer model 380 accurately understands the user 140's question intention, and the provided response text or response voice can be more accurate and more suitable for the actual needs of the user 140.

[0065] Alternatively or additionally, the speech module 350 can also determine a context of the ASR text based on the ASR text, construct a send message based on the ASR text, send the send message, the auxiliary information, and the context of the ASR text to the question-answering module 360. The question-answering module 360 can generate a model input of the question-answering model 380 based on the send message, the auxiliary information, and the context of the ASR text. Then, the question-answering module 360 can provide the model input to the question-answering model 380, trigger the question-answering model 380 to generate a response text based on the model input. The question-answering module 360 can receive the response text fed back by the question-answering model 380, send the response text to the access layer 340, and send the response text to the application message module 312 of the client 110 through the network interface 342. By adding the context of the ASR text, the accuracy of the response text can be improved.

[0066] In some embodiments, the server 130 can generate a first model input for the question model 380 (i.e., the third machine learning model) based on the ASR text before receiving the auxiliary information from the client 110. The server 130 can provide the first model input to the question-answering model 380 to make the question-answering model 380 generate a response text for the first model input. Alternatively or additionally, the speech module 350 can receive the ASR text from the ASR model 371, assuming that the speech module 350 has not received the auxiliary information yet. The speech module 350 can send the ASR text to the question-answering module 360. The question-answering module 360 can generate the first model input based on the ASR text, provide the first model input to the question-answering model 380, and trigger the question-answering model 380 to generate a response text for the first model input. In this way, in the case where valid auxiliary information cannot be obtained, the response text for the first model input can still be fed back to the client 110, ensuring the normal progress of the question-answering. In the case where the auxiliary information is delayed for a long time, the response text for the first model input can still be fed back to the client 110 in a timely manner, ensuring the timeliness of the response text feedback.

[0067] In some embodiments, the speech module 350 of the server 130 can send the auxiliary information to the question-answering module 360 after receiving the auxiliary information from the client 110. The question-answering module 360 can generate a second model input for the question-answering model 380 based on the ASR text and the auxiliary information. The question-answering module 360 can provide the second model input to the question-answering model 380, interrupt the question-answering model 380 from generating the response text for the first model input, and trigger the question-answering model 380 to generate a response text for the second model input. The question-answering model 380 can feed back the response text for the second model input to the application message module 312 of the client 110. In this way, in the case where the auxiliary information can be obtained in a timely manner, a more accurate response text for the user 140 can be obtained, improving the quality of the question-answering.

[0068] In embodiments of the present disclosure, as shown in signaling flow 200, the server 130 generates the response speech corresponding to the response text, and the server 130 sends (226) the response speech to the client 110, and the client 110 receives (228) the response speech from the server 130.

[0069] The response text here can be the response text input for the first model, or the response text input for the second model. In some embodiments, the question-answering module 360 can receive the response text from the question-answering model 380, and send the response text to the speech module 350, and the speech module 350 can provide the response text to the TTS model 372, and trigger the TTS model to generate the response speech corresponding to the response text. The response speech here is the speech converted from the response text, which can also be referred to as TTS speech.

[0070] In some embodiments, after receiving the TTS speech, the speech module 350 can send the TTS speech to the network interface 341 of the access layer 340, and send the TTS speech to the application speech module 311 of the client 110 via the first network connection. It can be understood that in actual application, the server 130 can feed back the response text to the client 110 according to the question-answering mode of the application 120, or feed back the response speech to the client 110, or feed back the response text and the response speech to the client 110 respectively. For example, in the case where the application 120 is configured to perform text question-answering with the user 140, the server 130 can feed back the response text to the client 110. In the case where the application 120 is configured to perform speech question-answering with the user 140, the server 130 can feed back the TTS speech to the client 110. In the case where the application 120 is configured to perform text and speech question-answering with the user 140, the server 130 can feed back the response text and the TTS speech to the client 110 respectively.

[0071] In embodiments of the present disclosure, as shown in signaling flow 200, the client 110 presents (230) the response text while playing (230) the response speech. The playing of the response speech and the presenting of the response text here are synchronized.

[0072] In embodiments of the present disclosure, the “synchronization” of the presenting can refer to that the time when the response speech is played and the time when the response text is presented are the same or almost the same. It should be understood that depending on the time accuracy, processing performance, etc., the synchronization in time can allow certain errors. For example, in the case where the time difference between the time when the response speech is played and the time when the response text is presented is less than a certain threshold, it can be considered that the playing of the response speech and the presenting of the response text are synchronized. In addition, the use of “synchronization” in other embodiments of the present disclosure also follows similar principles.

[0073] Exemplarily, the client 110 can play the TTS voice by utilizing the application voice module 311, and present the response text in the user interface 150 synchronously with the playing of the TTS voice by utilizing the application message module 312. Alternatively or additionally, the application message module 312 can present the response text in the user interface 150 synchronously with the playing of the TTS voice based on the correspondence between the TTS voice and the response text. Here the correspondence between the TTS voice and the response text can just be the corresponding voice part of each text character or text string in the response text.

[0074] In summary, according to the embodiments of the present disclosure, in the question-answering scenario, the response text and the response voice can be generated based on the question text and the auxiliary information matched with the question text, the service capability in the question-answering scenario can be improved, and the auxiliary information can be avoided to be saved in the server or the cloud, which is beneficial to guarantee the data security of the auxiliary information.

[0075] FIG. 4 shows a flowchart of a process 400 for question-answering according to some embodiments of the present disclosure. The process 400 can be implemented at the client 110.

[0076] At block 410, the client 110 acquires the question text recognized from the question voice in response to receiving the question voice of the user.

[0077] At block 420, the client 110 determines the auxiliary information matched with the question text based on the question text.

[0078] At block 430, the client 110 sends the auxiliary information to the server 130.

[0079] At block 440, the client 110 receives at least one of the response text or the response voice for the question voice from the server 130, the response voice corresponds to the response text, and the response text is determined based on the question text and the auxiliary information.

[0080] In some embodiments, the process 400 is further configured to receive the question text recognized from the question voice from the server 130.

[0081] In some embodiments, the process 400 is further configured to provide the question voice to the first machine learning model to make the first machine learning model perform voice recognition on the question voice, and receive the question text fed back by the first machine learning model.

[0082] In some embodiments, at least one category of auxiliary information is stored in the auxiliary information library of the client 110 for the user, and the process 400 is further configured to determine the auxiliary information matched with the question text from the auxiliary information library based on the question text.

[0083] In some embodiments, the process 400 is further configured to provide the question text to a second machine learning model to cause the second machine learning model to determine the auxiliary information matching the question text; and receive the auxiliary information fed back by the second machine learning model.

[0084] In some embodiments, the process 400 is further configured to receive the answer speech from the server 130 via a first network connection between the client 110 and the server 130; and / or receive the answer text from the server 130 via a second network connection between the client 110 and the server 130.

[0085] In some embodiments, the process 400 is further configured to present the answer text while playing the answer speech at the client 110, the playing of the answer speech being synchronized with the presenting of the answer text.

[0086] FIG. 5 illustrates a flowchart of a process 500 for question answering, according to some embodiments of the present disclosure. The process 500 can be implemented at the server 130.

[0087] At block 510, the server 130 sends, to the client 110, a question text recognized from the question speech in response to receiving the question speech from the client 110.

[0088] At block 520, the server 130 receives, from the client 110, auxiliary information matching the question text recognized from the question speech.

[0089] At block 530, the server 130 determines an answer text based on the question text and the auxiliary information.

[0090] At block 540, the server 130 feeds back at least one of the answer text or an answer speech corresponding to the answer text to the client 110.

[0091] In some embodiments, the process 500 is further configured to generate a model input for a third machine learning model based on the question text and the auxiliary information; and invoke the third machine learning model to generate the answer text based on the model input.

[0092] In some embodiments, the process 500 is further configured to generate a first model input for the third machine learning model based on the question text before receiving the auxiliary information from the client 110; and provide the first model input to the third machine learning model to cause the third machine learning model to generate an answer text for the first model input.

[0093] In some embodiments, the process 500 is further configured to: after receiving the auxiliary information from the client 110, generate a second model input for a third machine learning model based on the question text and the auxiliary information; and provide the second model input to the third machine learning model, interrupt the third machine learning model from generating the response text for the first model input, trigger the third machine learning model to generate the response text for the second model input.

[0094] In some embodiments, the process 500 is further configured to: feedback the response speech to the client 110 via the first network connection between the server 130 and the client 110; and / or feedback the response text to the client 110 via the second network connection between the server 130 and the client 110.

[0095] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 6 shows an exemplary structural block diagram of an apparatus 600 for question answering, according to some embodiments of the present disclosure. The apparatus 600 can be implemented as or included in the client 110. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.

[0096] As shown in FIG. 6, the apparatus 600 includes a question text obtaining module 610, an auxiliary information determining module 620, an auxiliary information sending module 630, and a response receiving module 640. The question text obtaining module 610 is configured to, in response to receiving a question speech of a user, obtain question text recognized from the question speech. The auxiliary information determining module 620 is configured to determine auxiliary information matching the question text based on the question text. The auxiliary information sending module 630 is configured to send the auxiliary information to a server. The response receiving module 640 is configured to receive at least one of a response text or a response speech for the question speech from the server, the response speech corresponding to the response text, and the response text being determined based on the question text and the auxiliary information.

[0097] In some embodiments, the question text obtaining module 610 is further configured to: receive the question text recognized from the question speech from the server.

[0098] In some embodiments, the question text obtaining module 610 is further configured to: provide the question speech to a first machine learning model to make the first machine learning model perform speech recognition on the question speech; and receive the question text fed back by the first machine learning model.

[0099] In some embodiments, the user has at least one category of auxiliary information stored in an auxiliary information library of the client, and the auxiliary information determining module 620 is further configured to: determine the auxiliary information matching the question text from the auxiliary information library based on the question text.

[0100] In some embodiments, the auxiliary information determination module 620 is further configured to: provide the question text to a second machine learning model to cause the second machine learning model to determine the auxiliary information matching the question text; and receive the auxiliary information fed back by the second machine learning model.

[0101] In some embodiments, the answer receiving module 640 is further configured to: receive the answer speech from the server via a first network connection between the client and the server; and / or receive the answer text from the server via a second network connection between the client and the server.

[0102] In some embodiments, the apparatus 600 further includes a presentation playing module configured to present the answer text while playing the answer speech at the client, the playing of the answer speech being synchronized with the presentation of the answer text.

[0103] FIG. 7 shows an exemplary structural block diagram of an apparatus 700 for question answering, according to some embodiments of the present disclosure. The apparatus 700 can be implemented as or included in the server 130. The various modules / components in the apparatus 700 can be implemented by hardware, software, firmware, or any combination thereof.

[0104] As shown in FIG. 7, the apparatus 700 includes a question text sending module 710, an auxiliary information receiving module 720, an answer text determination module 730, and an answer feedback module 740. The question text sending module 710 is configured to send, to the client, a question text recognized from a question speech of a user in response to receiving the question speech from the client. The auxiliary information receiving module 720 is configured to receive, from the client, auxiliary information matching the question text recognized from the question speech. The answer text determination module 730 is configured to determine an answer text based on the question text and the auxiliary information. The answer feedback module 740 is configured to feed back, to the client, at least one of the answer text or an answer speech corresponding to the answer text.

[0105] In some embodiments, the answer text determination module 730 is further configured to: generate a model input for a third machine learning model based on the question text and the auxiliary information; and invoke the third machine learning model to generate the answer text based on the model input.

[0106] In some embodiments, the answer text determination module 730 is further configured to: generate a first model input for the third machine learning model based on the question text before receiving the auxiliary information from the client; and provide the first model input to the third machine learning model to cause the third machine learning model to generate an answer text for the first model input.

[0107] In some embodiments, the response text determination module 730 is further configured to: after receiving the auxiliary information from the client, generate a second model input for the third machine learning model based on the question text and the auxiliary information; and provide the second model input to the third machine learning model, interrupting the third machine learning model from generating the response text for the first model input, triggering the third machine learning model to generate the response text for the second model input.

[0108] In some embodiments, the response feedback module 740 is further configured to: feed back the response speech to the client via the first network connection between the server and the client; and / or feed back the response text to the client via the second network connection between the server and the client.

[0109] The units and / or modules included in the apparatus 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine-executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 700 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.

[0110] It should be understood that one or more steps in the above methods can be performed by an appropriate electronic device or combination of electronic devices. Such electronic devices or combinations of electronic devices can include, for example, the client 110 in FIG. 1.

[0111] FIG. 8 shows a block diagram of an electronic device 800 in which one or more embodiments of the present disclosure can be implemented. It should be understood that the electronic device 800 shown in FIG. 8 is merely an example and should not be construed to limit the functionality and scope of the embodiments described herein. The electronic device 800 shown in FIG. 8 can be used to implement the client 110 or the server 130 in FIG. 1, or to implement the apparatus 600 in FIG. 6 or the apparatus 700 in FIG. 7.

[0112] As shown in FIG. 8, electronic device 800 is in the form of a general-purpose electronic device. Components of electronic device 800 can include, but are not limited to, one or more processors or processing units 810, memory 820, storage 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit(s) 810 can be actual or virtual processors and capable of executing various processing in accordance with programs stored in memory 820. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of electronic device 800.

[0113] Electronic device 800 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or nonvolatile media. Memory 820 can be volatile (such as, for example, registers, cache, random access memory (RAM)), non-volatile (such as, for example, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage 830 can be removable or non-removable and can include machine-readable media, such as, for example, flash drives, disks, or any other media capable of storing information and / or data and accessible by electronic device 800.

[0114] Electronic device 800 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 8, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard drive"), and a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM or other optical medium). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 820 can include a computer program product 825 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0115] Communication unit(s) 840 enable communication with other electronic devices via communication media. Additionally, functionality of components of electronic device 800 can be implemented in a single computing cluster or a plurality of computer machines capable of communication through a communication connection. Thus, electronic device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.

[0116] Input device 850 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 860 can be one or more output devices, such as a display, a speaker, a printer, etc. Electronic device 800 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through communication unit 840, as desired, in order to communicate with a user in order to interact with electronic device 800, or to communicate with any device (e.g., a network card, a modem, etc.) that enables electronic device 800 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0117] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0118] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0119] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0120] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0121] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The machine-readable medium can be a single medium, or multiple media, of the same or different type. The computer program product can be one or more computer program components embodied in medium and / or transmission signals. The computer program product can have one or more computer readable and / or computer executable components embodied in medium and / or transmission signals. The computer program product can be one or more computer readable and / or computer executable components embodied in medium and / or transmission signals. The computer program product can also be at least one or any combination of circuitry, hardware, and / or executable instructions related to the operation of the present disclosure as claimed in the appended claims. It should be noted that a software module could reside in a computer program product. The computer program product could be a computer- readable medium having computer readable program code embodied therein. The computer readable program code can be executed by a computer or computer processor to perform various functions as described herein. The computer program product could be a memory, a memory device, a memory chip, a memory package, or any combination thereof. The computer program product could also be a hardware logic circuit such as a logic chip, a logic device, a logic package, or any combination thereof. The computer program product could also be a hardware logic circuit such as a logic chip, a logic device, a logic package, or any combination thereof.

[0122] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although being shown and described in terms of certain implementations, it is apparent that with the advance of technological age, one of ordinary skill in the art can devise certain modifications and alterations of the described implementations. Therefore, the disclosures and descriptions of the implementations are illustrative and not restrictive. Changes can be made by one of ordinary skill in the art, which accomplishes the same employment, function, and / or result of the related implementations as described. Modifications, additions, or omissions can be made to the implementations described herein without departing from the spirit or essential characteristics of the implementations. Changes can be made to the implementations described herein without departing from the spirit or essential characteristics of the implementations. The scope of the disclosure should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. The claims are intended to cover all modifications and alterations of the described implementations.

Claims

1. A method for question answering, applied to a client, the method comprising: in response to receiving a question speech of a user, obtaining a question text recognized from the question speech; based on the question text, determining auxiliary information matching the question text; sending the auxiliary information to a server; and receiving at least one of a response text or a response speech for the question speech from the server, the response speech corresponding to the response text, and the response text being determined based on the question text and the auxiliary information. 2.The method of claim 1, wherein obtaining a question text recognized from the question speech comprises: receiving the question text recognized from the question speech from the server. 3.The method of claim 1, wherein obtaining a question text recognized from the question speech comprises: providing the question speech to a first machine learning model to cause the first machine learning model to perform speech recognition on the question speech; and receiving the question text fed back by the first machine learning model. 4.The method of claim 1, wherein a library of auxiliary information of at least one category is stored in the client, and wherein determining auxiliary information matching the question text comprises: based on the question text, determining auxiliary information matching the question text from the library of auxiliary information. 5.The method of claim 1, wherein determining auxiliary information matching the question text comprises: providing the question text to a second machine learning model to cause the second machine learning model to determine the auxiliary information matching the question text; and receiving the auxiliary information fed back by the second machine learning model. 6.The method of claim 1, wherein receiving at least one of a response text or a response speech for the question speech from the server comprises: receiving the response speech from the server via a first network connection between the client and the server; and / or receiving the response text from the server via a second network connection between the client and the server. 7.The method of claim 6, further comprising: presenting the response text while playing the response speech at the client, the playing of the response speech being synchronized with the presenting of the response text. 8.A method for question answering, applied to a server, the method comprising: in response to receiving a question speech of a user from a client, sending a question text recognized from the question speech to the client; receiving auxiliary information matching the question text recognized from the question speech from the client; based on the question text and the auxiliary information, determining a response text; and feeding back at least one of the response text or a response speech to the client, the response speech corresponding to the response text. 9.The method of claim 8, wherein determining a response text based on the question text and the auxiliary information comprises: ​ ​ ​ generating a model input for a third machine learning model based on the question text and the auxiliary information; and calling the third machine learning model to generate the answer text based on the model input.

10. The method of claim 8, further comprising: generating a first model input for a third machine learning model based on the question text before receiving the auxiliary information from the client; and providing the first model input to the third machine learning model to cause the third machine learning model to generate an answer text for the first model input.

11. The method of claim 10, wherein determining an answer text based on the question text and the auxiliary information comprises: generating a second model input for the third machine learning model based on the question text and the auxiliary information after receiving the auxiliary information from the client; and providing the second model input to the third machine learning model to interrupt the third machine learning model from generating an answer text for the first model input, triggering the third machine learning model to generate an answer text for the second model input.

12. The method of claim 8, wherein feeding back at least one of the answer text or an answer speech to the client comprises: feeding back the answer speech to the client via a first network connection between the server and the client; and / or feeding back the answer text to the client via a second network connection between the server and the client.

13. An apparatus for question answering, comprising: a question text obtaining module configured to obtain a question text recognized from a question speech of a user in response to receiving the question speech; an auxiliary information determining module configured to determine auxiliary information matching the question text based on the question text; an auxiliary information sending module configured to send the auxiliary information to a server; and an answer receiving module configured to receive at least one of an answer text or an answer speech for the question speech from the server, the answer speech corresponding to the answer text, and the answer text being determined based on the question text and the auxiliary information.

14. An apparatus for question answering, comprising: a question text sending module configured to send a question text recognized from a question speech of a user to a client in response to receiving the question speech from the client; an auxiliary information receiving module configured to receive auxiliary information matching the question text recognized from the question speech from the client; an answer text determining module configured to determine an answer text based on the question text and the auxiliary information; and an answer feeding back module configured to feed back at least one of the answer text or an answer speech to the client, the answer speech corresponding to the answer text.

15. An electronic device, comprising: at least one processor; and ​ at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-7 or 8-12.

16. A computer-readable storage medium having stored thereon computer- executable instructions executable by a processor to implement a method according to any one of claims 1-7 or 8-12.

17. A computer program product comprising a computer program, wherein the computer- executable instructions, when executed by a processor, implement a method according to any one of claims 1-7 or 8-12.

Citation Information

Patent Citations

  • Deep learning-based intelligent response system

    CN106202301A

  • Voice question and answer interaction method and device, computer device and storage medium

    CN109670088A

  • Intelligent voice interaction robot and voice interaction method thereof

    CN110019683A

  • Voice recognition-based interrogation dialogue method, interrogation dialogue device and storage medium

    CN110335595A

  • Question answering method and system based on voice interaction, electronic equipment and storage medium

    CN114005440A