Method and apparatus for inputting data in dialogue, and device, medium and program product
By using a fast but less accurate speech recognition model to initially present text data, and then replacing it with more accurate text data, combined with a machine learning model to generate responses, the problem of low efficiency in traditional speech recognition is solved, and user experience and recognition accuracy are improved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-28
AI Technical Summary
Traditional speech recognition models are inefficient at determining the question text, causing users to wait a long time for the speech recognition results, which affects the user experience.
Two speech recognition models with different recognition speeds and accuracies are used to quickly present preliminary text data, and then replace it with more accurate text data after confirmation. At the same time, a machine learning model is used to generate responses.
It improves the speed of feedback from voice data to text data, reduces user waiting time, and enhances recognition accuracy and response efficiency.
Smart Images

Figure CN2024134276_28052026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, media, and procedures for inputting data in a dialogue. Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, electronic devices, computer-readable storage media, and computer program products for inputting data in a dialogue. Background Technology
[0002] With the development of information technology, various electronic devices can provide people with a variety of services in work and life. For example, electronic devices can be equipped with applications that provide dialogue services. Applications with dialogue capabilities can output corresponding responses based on user-input questions. Electronic devices or applications can use trained machine learning models to determine the corresponding response to a question and provide that response to the user. Both the user's question and the response to that question can be presented on a specific page of the application, which could be a dialogue page between the user and a digital assistant. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for inputting data in a dialogue is provided. The method includes: acquiring first text data associated with voice data, the voice data being obtained from a user in the dialogue, and the first text data being acquired by recognizing the voice data using a first speech recognition model; presenting the first text data in the dialogue; acquiring second text data associated with the voice data, the second text data being acquired by recognizing the voice data using a second speech recognition model; and, in response to determining that the second text data is different from the first text data, replacing the first text data with the second text data in the dialogue.
[0004] In a second aspect of this disclosure, an apparatus for inputting data in a dialogue is provided. The apparatus includes: a first text acquisition module configured to acquire first text data associated with voice data, the voice data being from a user in the dialogue, and the first text data being acquired by recognizing the voice data using a first speech recognition model; a first text presentation module configured to present the first text data in the dialogue; a second text acquisition module configured to acquire second text data associated with the voice data, the second text data being acquired by recognizing the voice data using a second speech recognition model; and a second text presentation module configured to replace the first text data with the second text data in the dialogue in response to determining that the second text data is different from the first text data.
[0005] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to perform the method according to a first aspect of this disclosure.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;
[0011] Figure 2 illustrates an example architecture for inputting data in a dialogue according to some embodiments of the present disclosure;
[0012] Figures 3A to 3F illustrate examples of dialog pages according to some embodiments of the present disclosure;
[0013] Figure 4 shows a flowchart of a method for inputting data in a dialogue according to some embodiments of the present disclosure;
[0014] Figure 5 shows a schematic structural block diagram of a device for inputting data in a dialogue according to some embodiments of the present disclosure; and
[0015] Figure 6 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. Detailed Implementation
[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below. In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after "A", but may include one or more intermediate steps.
[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0020] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.
[0021] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.
[0022] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 120 is installed on an electronic device 110. A user 130 can interact with the application 120 via the electronic device 110 and / or an attached device of the electronic device 110. For example, the application 120 can capture the voice 135 of the user 130 via a voice capture device (e.g., a microphone) of the electronic device 110.
[0023] In embodiments of this disclosure, application 120 can be any suitable application with human-computer dialogue capabilities. For example, application 120 can provide a digital assistant for human-computer dialogue. This digital assistant supports text-based dialogue services, voice interaction services, and content dialogue in other modalities with user 130.
[0024] In environment 100, if application 120 is active, electronic device 110 can provide page 150 of application 120. Page 150 may include various types of pages that application 120 can provide, such as a user-digital assistant dialogue page (where the current and historical dialogues, including text dialogue content, may be presented), and so on. In some embodiments, electronic device 110 may play voice 152 on page 150. Voice 152 may, for example, include voice 135 from user 130 or voice response to voice 145.
[0025] In some embodiments, application 120 or its digital assistant may utilize machine learning model 140 (which may include one or more machine learning models, such as machine learning model 140-1, machine learning model 140-2, ..., machine learning model 140-N, etc., where N is a positive integer. For ease of description, the one or more machine learning models are collectively referred to as machine learning model 140 herein) to support interaction with user 130. For example, application 120 or its digital assistant may utilize one or more machine learning models 140 to provide question-and-answer services to user 130. In audio dialogue scenarios, the question in the question-and-answer process is audio input by the user, and the response may be provided to the user in audio form, text, or other modalities.
[0026] Machine learning model 140 can be of different types. In some embodiments, one or more machine learning models 140 may be built based on a language model (LM). The machine learning model used is a content-generative model, capable of generating corresponding outputs based on model inputs. In some embodiments, the language model-based machine learning model can handle textual modal model inputs (e.g., natural language and / or machine language) and / or non-textual modal model inputs (e.g., images, speech, video, etc.), and can generate the desired output based on the model inputs and prompt words. Here, prompt words are used to guide the machine learning model to generate outputs that address the user needs indicated by the model inputs. In application scenarios supporting user dialogue, user 130's input can be provided to machine learning model 140 as at least a part of the model inputs (other parts may include prompt words). This user input is considered a user question. Based on the model output, a corresponding response can be generated and provided to user 130.
[0027] In Figure 1, electronic device 110 can be any type of computing device, including terminal devices or server devices. Terminal devices can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Server devices can include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.
[0028] It should be noted that if electronic device 110 is a terminal device, it can directly present page 150 to user 130 via its own display screen. If electronic device 110 is a server device, it can send page 150 to the terminal device corresponding to user 130, so that the terminal device can present page 150 to user 130.
[0029] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0030] As mentioned earlier, applications with dialogue capabilities can output corresponding responses based on user-input questions. Electronic devices or applications can use trained machine learning models to determine the response to a question and provide it to the user. Both the user's question and the response to that question can be presented on a specific page within the application, which could be a dialogue page between the user and a digital assistant. If the question is a voice-based question, it can be referred to as a voice question. Electronic devices can use Automatic Speech Recognition (ASR) models to determine the text corresponding to the voice question and present the text on the aforementioned dialogue page.
[0031] Traditionally, electronic devices typically respond by acquiring the complete spoken question and using a high-performance, large-scale speech recognition model to determine the corresponding text. While this ensures the accuracy of the text, it results in a longer time required to determine the text, making it less efficient at providing the user with the text. This can lead to users waiting for extended periods for the speech recognition results.
[0032] In view of this, according to embodiments of the present disclosure, an improved scheme for inputting data in a dialogue is provided. According to this scheme, first text data associated with voice data, the voice data being from a user in the dialogue, is obtained, and the first text data is obtained by recognizing the voice data using a first speech recognition model. The first text data is presented in the dialogue. Second text data associated with the voice data is obtained, the second text data being obtained by recognizing the voice data using a second speech recognition model. In response to determining that the second text data is different from the first text data, the first text data is replaced with the second text data in the dialogue.
[0033] Therefore, first text data can be quickly provided to the user, and in response to obtaining more accurate second text data, the first text data can be replaced with the second text data. Alternatively and / or additionally, a time range can be specified (e.g., referred to as a first time range, such as 1.5 seconds or other values), and if the second text data received within the specified time range differs from the first text data, a replacement is performed. This can improve feedback to user input, quickly provide the user with text data corresponding to the voice data, increase the speed at which the user obtains information, and reduce user waiting time. Some exemplary embodiments of this disclosure will continue to be described below with reference to the accompanying drawings.
[0034] Figure 2 illustrates a schematic diagram of an example architecture 200 for audio dialogue according to some embodiments of the present disclosure. The example architecture 200 can be implemented at an electronic device 110. For ease of discussion, the example architecture 200 will be described with reference to the environment 100 of Figure 1. It should be noted that the operations performed by the aforementioned electronic device 110, and the operations performed by the electronic device 110 as described below, may specifically be performed by relevant applications (e.g., application 120) installed on the electronic device 110. In some embodiments, when the electronic device 110 is a terminal device, the operations performed on the electronic device 110 can be completed with the assistance of other devices (e.g., a server).
[0035] Example architecture 200 includes at least speech recognition model 210 and speech recognition model 220. In some embodiments, the recognition speed of speech recognition model 210 is higher than that of speech recognition model 220. Speech recognition model 210 can be referred to as a first speech recognition model, and speech recognition model 220 as a second speech recognition model. As an example, the model size of speech recognition model 210 can be smaller than that of speech recognition model 220. The model size of each machine learning model is related to the parameter size, structural complexity, etc. of that machine learning model. Generally, the larger the parameter size or the more complex the structure of a machine learning model, the larger its model size. Machine learning models with larger model sizes require greater resource overhead, including but not limited to overhead related to computing resources, memory resources, time resources, etc.
[0036] In summary, when the model size of speech recognition model 210 is smaller than that of speech recognition model 220, the recognition accuracy of speech recognition model 210 may be lower than that of speech recognition model 220, and the recognition speed of speech recognition model 210 is usually higher than that of speech recognition model 220.
[0037] Electronic device 110 can acquire voice data 201 from a user (e.g., user 130) during a conversation (which may correspond, for example, to voice 135 in FIG1, and may be referred to as a question voice). For example, electronic device 110 can acquire user input on a conversation page between the user and a digital assistant, and if the user input is audio, determine the user input as voice data 201.
[0038] Figures 3A to 3F illustrate examples 300A and 300B of a dialog page according to some embodiments of the present disclosure. Further details regarding the reception of voice data are described with reference to Figures 3A and 3B. As shown in example 300A, the dialog page may display historical messages between the user and a digital assistant (e.g., message 301 from the user and message 302 from the digital assistant, as shown in the figures). The dialog page may also include an input box 310. The electronic device 110 may receive user input from the user via the input box 310. In some embodiments, the electronic device 110 may display example 300B in response to receiving a trigger on the input box 310 in example 300A. As shown in example 300, the electronic device 110 may receive voice data 201 from the user in response to receiving a trigger on the input box 310. The electronic device 110 may also display a prompt message 312 in example 300B, which may be used to indicate to the user that voice data 201 is currently being received. The prompt message 312 can be any appropriate type of prompt message, such as text, image, or icon. For example, the prompt message 312 can include the prompt text "Release to send, move up to cancel" as shown in the figure, etc.
[0039] Electronic device 110 can provide speech data 201 to speech recognition model 210. As an example, electronic device 110 can create a prompt word input for speech recognition model 210 based at least on speech data 201 (e.g., also based on a prompt word template), and provide the prompt word input to speech recognition model 210. Speech recognition model 210 can determine text data 215 for speech data 201 based on this prompt word input. The text data 215 determined by the first speech recognition model (i.e., speech recognition model 210) can be referred to as first text data.
[0040] Electronic device 110 can acquire text data 215 from speech recognition model 210 and present it in a dialogue (e.g., on a dialogue page). Text data 215 can be presented on the dialogue page, for example, as a dialogue message from the user. It should be noted that text data 215 and speech data 201 can be presented in the dialogue individually or in conjunction. For example, a visual representation corresponding to speech data 201 can be presented on the page, which the user can click to play the speech data 201, and further, text data 215 can be presented below this representation (or elsewhere).
[0041] Since the speech recognition model 210 requires a certain amount of time to determine the text data 215 based on the speech data 201, in some embodiments, the electronic device 110 may, in response to determining that the text data 215 has not been received within a specified time range (which may be referred to as a second time range, and which may be any suitable time range, such as 0.3 seconds), present a visual representation in the dialogue for waiting for the text data 215. This visual representation may be any suitable visual representation such as text, symbols, or charts; as an example, it may be an ellipsis "...". The electronic device 110 may then, in response to acquiring the text data 215, replace the visual representation for waiting for the text data 215 with the text data 215 in the dialogue.
[0042] Referring to Figures 3C and 3D, which illustrate examples 300C and 300D of a dialogue page according to some embodiments of the present disclosure. As shown in example 300C, electronic device 110 may present a visual representation 303 in the dialogue page in response to receiving voice data 201 but not text data 215. Electronic device 110 may then switch to presenting example 300D in response to receiving text data 215. As shown in example 300D, electronic device 110 may present a message 304 corresponding to text data 215 in the dialogue page. Message 304 may include the text of text data 215, “What happy today?”. Thus, timely feedback to the user regarding the receipt of voice data can be provided, reducing user waiting time.
[0043] In some embodiments, after presenting text data 215 in a dialogue, the electronic device 110 may also present a visual representation in the dialogue for waiting for a response to the voice data. This visual representation can also be any suitable visual representation such as text, symbols, or charts; for example, it could be an ellipsis "...". As shown in Example 300D, the electronic device 110 may present a visual representation 305 in the dialogue. This provides the user with feedback that a response is being determined, reducing user waiting time and thus reducing user anxiety while waiting for a response.
[0044] In some embodiments, the electronic device 110 may also, in response to not receiving or being unable to receive text data 215 within another specified time range (which may be longer than the second time range), present a visual representation in the dialog indicating that text data 215 cannot be received. As an example, the electronic device 110 may present a prompt message on the dialog page indicating that text data 215 has not been received. This prompt message may include any suitable type of message such as text, symbols, or icons. For example, the electronic device 110 may present a message including an exclamation mark on the dialog page, presented as a message from the user, indicating that the electronic device 110 has not received text data 215. This informs the user that text confirmation failed and instructs the user to resubmit voice data, allowing the user to take appropriate action promptly.
[0045] In some embodiments, the electronic device 110 may also provide voice data 201 to the speech recognition model 220. The electronic device 110 may, for example, simultaneously provide voice data 201 to both the speech recognition model 210 and the speech recognition model 220. Similarly, the electronic device 110 may create a prompt input for the speech recognition model 220 based at least on the voice data 201 (e.g., also based on a prompt word template), and provide the prompt input to the speech recognition model 220. The speech recognition model 220 may determine text data 225 for the voice data 201 based on this prompt input. The text data 225 determined by the second speech recognition model (i.e., speech recognition model 220) may be referred to as second text data.
[0046] Electronic device 110 can acquire text data 225 from speech recognition model 220. It is understood that speech recognition model 210 determines text data 215 faster than speech recognition model 220 determines text data 225; therefore, electronic device 110 acquires text data 215 first, and then text data 225. In response to acquiring text data 225, electronic device 110 can determine whether text data 225 is the same as text data 215.
[0047] Electronic device 110 may, for example, replace text data 215 with text data 225 in a dialogue in response to text data 225 being different from text data 215. Referring to Figures 3D and 3E, Figure 3E illustrates an example 300E of a dialogue page according to some embodiments of the present disclosure. Electronic device 110 may, in response to receiving text data 225 and determining that text data 225 is different from text data 215, present message 306 in the dialogue page and no longer present message 304. Message 306 may include the text “What happened today?” from text data 225. Here, the incorrect “happy” in message 304 is corrected to “happened” in message 306. In this way, recognition accuracy can be improved.
[0048] Because the recognition accuracy of speech recognition model 210 is lower than that of speech recognition model 220, there are instances where speech recognition model 210 misses or makes recognition errors (i.e., text data 215 has missing or incorrect words). When text data 225 differs from text data 215, the number of text units included in text data 225 may differ from the number of text units included in text data 215. In this case, electronic device 110 may, for example, present text data 225 according to its data volume during a dialogue.
[0049] Referring to Figures 3D and 3F, Figure 3F illustrates an example 300F of a dialog page according to some embodiments of the present disclosure. When the amount of text data 225 is greater than that of text data 215 (e.g., text data 225 includes the text "What happened today? Any news?", and text data 215 includes the text "What happy today?"), the electronic device 110 may, for example, present text data 215 as a single line in the dialog page shown in example 300D, and may present text data 225 as a double line in the dialog page shown in example 300F. In this way, the most recently identified and accurate text can be presented on the page.
[0050] It is understandable that if text data 225 is the same as text data 215, electronic device 110 can continue to present text data 215 in the conversation, or it can use text data 225 to replace text data 215 in the conversation.
[0051] In some embodiments, the electronic device 110 may acquire a specified time range (which may be referred to as a third time range) for the text data 225. The third time range can be any suitable time range; for example, it can be 0.5 seconds. For instance, in response to not acquiring the text data 225 within the third time range, the electronic device 110 may re-invoke the speech recognition model 220 to recognize the speech data 201. That is, in response to not acquiring the text data 225 within the third time range, the electronic device 110 may resend the speech data to the speech recognition model 220 to utilize the speech recognition model 220 to determine the text data 225.
[0052] In some embodiments, the electronic device 110 may also acquire another specified time range for the text data 225 (which may be referred to as a first time range). The first time range can also be any suitable time range, which may, for example, be greater than a third time range; for example, it may be 1.5 seconds. The electronic device 110 may, in response to acquiring the text data 225 within the first time range, replace the text data 215 with the text data 225 in the conversation. In some embodiments, the electronic device 110 may also, in response to not acquiring the text data 225 within the first time range, continue to present the text data 215 in the conversation. In this case, even if the electronic device 110 receives the text data 225 outside the first time range, the electronic device 110 may still continue to present the text data 215 without replacing it with the text data 225. Alternatively and / or additionally, the text data 215 may also be replaced with the text data 225 received outside the first time range.
[0053] In some embodiments, example architecture 200 may further include machine learning model 230. Machine learning model 230 may be based on a language model (LM). A language model, by learning from a large corpus, is capable of question-answering. Machine learning models may also be based on other suitable models. Specific configuration areas are provided during the feature creation process to support user-provided prompts, and the configuration of prompts can be done using natural language. This allows users to easily constrain the model's output and configure diverse digital assistants. In some embodiments, machine learning model 230 may also be based on any suitable model architecture, including but not limited to Transformer models, convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep neural networks (DNNs), and so on.
[0054] Electronic device 110 can send text data to machine learning model 230 to determine a response 235 for the text data (which is also a response to voice data 201). As an example, electronic device 110 can send text data 225 to machine learning model 230 in response to acquiring text data 225 within a first time frame, and can send text data 215 to machine learning model 230 in response to not acquiring text data 225 within the first time frame. Thus, more accurate second text data can be prioritized for sending to the machine learning model while ensuring low latency, improving the accuracy of the response. If second text data is not received for a longer period, first text data can be sent to the machine learning model, ensuring the efficiency of the response.
[0055] In some embodiments, electronic device 110 may, in response to determining that text data has been sent to machine learning model 230, present a visual representation in the dialogue indicating that the text data has been sent. This visual representation may also include any suitable visual representation; for example, it may be an underline at the location of the text data presented in the dialogue. Continuing to refer to FIG3E, electronic device 110 may, in response to text data 225 being sent to machine learning model 230, add an underline to the text shown in message 306 to indicate that text data 225 has been sent to machine learning model 230.
[0056] In some embodiments, the electronic device 110 may also receive an interaction request for a visual representation of which text data has been sent. The electronic device 110 may receive the interaction request in any suitable manner. For example, the electronic device 110 may determine that an interaction request for a visual representation has been received in response to receiving a trigger on the visual representation (including but not limited to a click operation, long press operation, hover operation, swipe operation, etc. on the visual representation).
[0057] Electronic device 110 can edit text data based on an interaction request received for a visual representation. Referring to FIG3E, electronic device 110 can determine that an interaction request for text data 225 has been received in response to a triggering of message 306 including an underline, and then edit text data 225 based on the interaction request. The interaction request may instruct the addition, deletion, or modification of text in text data 225. Electronic device 110 can then send the edited text data to machine learning model 230 in response to receiving a confirmation request for the edited text data. For example, referring to FIG3E, electronic device 110 can send the edited text data 225 to machine learning model 230.
[0058] In this scenario, machine learning model 230 can determine the corresponding response based on the edited text data. It can be understood that if machine learning model 230 is determining a response based on the unedited text data, it can stop determining the response based on the unedited text data and determine the corresponding response based on the edited text data. This allows users to edit the text data, making it more aligned with the latest user needs and improving the accuracy of responses in the conversation.
[0059] Electronic device 110 may present response 235 in response to receiving response 235 at machine learning model 230. In some embodiments, where electronic device 110 has already presented a visual representation in the dialogue for waiting for a response to voice data, electronic device 110 may replace the visual representation in the dialogue with response 235 in response to receiving response 235.
[0060] In some embodiments, the electronic device 110 may also, in response to not receiving a response 235 within a certain time range, present a visual representation in the dialogue indicating that a response 235 could not be received. As an example, the electronic device 110 may present a prompt message on the dialogue page indicating that a response 235 was not received. This prompt message may include any suitable type of message such as text, symbols, or icons. For example, the electronic device 110 may present a message including an exclamation mark on the dialogue page, presented as a message from the digital assistant, indicating that the electronic device 110 has not received a response 235. This allows the user to be informed that the response has failed, and may instruct the user to resubmit voice data or resend text data to the machine learning model, enabling the user to take appropriate action promptly.
[0061] In summary, according to the embodiments of this disclosure, first text data can be quickly provided to the user, and in response to obtaining more accurate second text data, the first text data can be replaced with the second text data. This can improve feedback to user input, quickly provide the user with text data corresponding to the voice data, increase the speed at which the user obtains information, and reduce the user's waiting time.
[0062] Figure 4 shows a flowchart of a method 400 for inputting data in a dialogue according to some embodiments of the present disclosure. Method 400 can be implemented at the electronic device 110 of Figure 1. Method 400 will be described with reference to the environment 100 of Figure 1.
[0063] In box 410, electronic device 110 acquires first text data associated with voice data, the voice data being from a user in a conversation, and the first text data being acquired by recognizing the voice data using a first speech recognition model.
[0064] In box 420, electronic device 110 presents the first text data in the dialogue.
[0065] In box 430, electronic device 110 acquires second text data associated with voice data, the second text data being acquired by recognizing the voice data using a second speech recognition model.
[0066] In box 440, in response to determining that the second text data is different from the first text data, the electronic device 110 uses the second text data to replace the first text data in the dialogue.
[0067] In some embodiments, method 400 further includes: sending text data to a machine learning model, the text data including first text data or second text data; and presenting a response from the machine learning model to the text data in a dialogue.
[0068] In some embodiments, sending text data to a machine learning model includes at least one of the following: sending first text data to a machine learning model in response to determining that second text data has not been received within a first time range; or sending second text data to a machine learning model in response to determining that second text data has been received within a first time range.
[0069] In some embodiments, method 400 further includes: in response to determining that text data has been sent to a machine learning model, presenting a visual representation in the dialogue that the text data has been sent.
[0070] In some embodiments, method 400 further includes: in response to receiving an interaction request for a visual representation, editing text data based on the interaction request; and in response to receiving an acknowledgment request for the edited text data, sending the edited text data to a machine learning model.
[0071] In some embodiments, method 400 further includes: in response to determining that first text data has not been received within a second time range, presenting a visual representation in the dialogue for waiting for the first text data.
[0072] In some embodiments, obtaining the second text data includes: in response to determining that the second text data has not been received within a third time range, invoking a second speech recognition model to recognize the speech data.
[0073] In some embodiments, method 400 further includes: after presenting first text data in a dialogue, presenting a visual representation in the dialogue for waiting for a response to the voice data.
[0074] In some embodiments, replacing first text data with second text data in a conversation includes: in response to acquiring second text data within a first time range, replacing first text data with second text data in a conversation.
[0075] In some embodiments, the first recognition speed of the first speech recognition model is higher than the second recognition speed of the second speech recognition model.
[0076] In some embodiments, replacing first text data with second text data in a dialogue includes: presenting second text data in the dialogue according to the amount of second text data.
[0077] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 5 shows an exemplary structural block diagram of an apparatus 500 for inputting data in a dialog according to some embodiments of this disclosure. The apparatus 500 may be implemented as or included in the electronic device 110. The various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0078] As shown in Figure 5, the device 500 includes a first text acquisition module 510, configured to acquire first text data associated with voice data, the voice data being from a user in a conversation, and the first text data being acquired by recognizing the voice data using a first speech recognition model. The device 500 also includes a first text presentation module 520, configured to present the first text data in a conversation. The device 500 further includes a second text acquisition module 530, configured to acquire second text data associated with the voice data, the second text data being acquired by recognizing the voice data using a second speech recognition model. The device 500 also includes a second text presentation module 540, configured to replace the first text data with the second text data in a conversation in response to determining that the second text data is different from the first text data.
[0079] In some embodiments, the apparatus 500 further includes: a text data sending module configured to send text data to a machine learning model, the text data including first text data or second text data; and a response presentation module configured to present a response from the machine learning model to the text data in a dialogue.
[0080] In some embodiments, the text data sending module is further configured to: send first text data to the machine learning model in response to determining that second text data has not been received within a first time range; or send second text data to the machine learning model in response to determining that second text data has been received within a first time range.
[0081] In some embodiments, the apparatus 500 further includes: a first visual representation rendering module configured to render a visual representation in a dialogue that the text data has been sent in response to determining that text data has been sent to a machine learning model.
[0082] In some embodiments, the apparatus 500 further includes: a text data editing module configured to edit text data based on an interaction request for a visual representation in response to receiving such an interaction request; and an editing confirmation module configured to send the edited text data to a machine learning model in response to receiving a confirmation request for the edited text data.
[0083] In some embodiments, the apparatus 500 further includes: a second visual representation presentation module configured to present a visual representation in a dialogue for waiting for the first text data in response to determining that the first text data has not been received within a second time range.
[0084] In some embodiments, the second text acquisition module 530 is further configured to: in response to determining that no second text data has been received within a third time range, invoke a second speech recognition model to recognize speech data.
[0085] In some embodiments, the apparatus 500 further includes: a third visual representation presentation module configured to present a visual representation for waiting for a response to voice data in a dialogue after presenting first text data in the dialogue.
[0086] In some embodiments, the second text rendering module 540 is further configured to: in response to acquiring second text data within a first time range, replace the first text data with the second text data in a conversation.
[0087] In some embodiments, the first recognition speed of the first speech recognition model is higher than the second recognition speed of the second speech recognition model.
[0088] In some embodiments, the second text presentation module 540 is further configured to present the second text data in the dialogue according to the amount of the second text data.
[0089] The units and / or modules included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.
[0090] It should be understood that one or more steps in the above methods can be performed by a suitable electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices may, for example, include the electronic device 110 in FIG1.
[0091] Figure 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 600 shown in Figure 6 can be used to implement the electronic device 110 of Figure 1 or the device 500 of Figure 5.
[0092] As shown in Figure 6, the electronic device 600 is in the form of a general-purpose electronic device. Components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0093] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.
[0094] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0095] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0096] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0097] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0098] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0099] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0100] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some, as newer, implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0102] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for inputting data in a dialogue, comprising: First text data associated with the voice data is obtained, the voice data being from the user of the conversation, and the first text data is obtained by recognizing the voice data using a first speech recognition model; The first text data is presented in the dialogue; Acquire second text data associated with the voice data, wherein the second text data is obtained by recognizing the voice data using a second speech recognition model; as well as In response to determining that the second text data is different from the first text data, the second text data is used to replace the first text data in the dialogue.
2. The method according to claim 1, further comprising: Send text data to a machine learning model, wherein the text data includes either the first text data or the second text data; as well as The dialogue presents responses from the machine learning model to the text data.
3. The method of claim 2, wherein sending the text data to the machine learning model comprises at least one of the following: In response to determining that the second text data has not been received within a first time range, the first text data is sent to the machine learning model; or In response to determining that the second text data has been received within a first time range, the second text data is sent to the machine learning model.
4. The method of claim 3, further comprising: In response to determining that the text data has been sent to the machine learning model, a visual representation of the text data being sent is presented in the dialogue.
5. The method of claim 4, further comprising: In response to receiving an interaction request for the visual representation, the text data is edited based on the interaction request; as well as In response to receiving a confirmation request for the edited text data, the edited text data is sent to the machine learning model.
6. The method of claim 1, further comprising: In response to determining that the first text data has not been received within a second time range, a visual representation for waiting for the first text data is presented in the dialogue.
7. The method according to claim 1, wherein obtaining the second text data comprises: In response to determining that the second text data has not been received within a third time range, the second speech recognition model is invoked to recognize the speech data.
8. The method of claim 1, further comprising: After the first text data is presented in the dialogue, a visual representation for waiting for a response to the voice data is presented in the dialogue.
9. The method of claim 1, wherein replacing the first text data with the second text data in the dialogue comprises: In response to acquiring the second text data within a first time frame, the second text data is used to replace the first text data in the dialogue.
10. The method according to claim 1, wherein the first recognition speed of the first speech recognition model is higher than the second recognition speed of the second speech recognition model.
11. The method of claim 1, wherein replacing the first text data with the second text data in the dialogue comprises: The second text data is presented in the dialogue according to the amount of the second text data.
12. An apparatus for inputting data in a dialogue, comprising: The first text acquisition module is configured to acquire first text data associated with the voice data, the voice data being from the user of the conversation, and the first text data being acquired by recognizing the voice data using a first speech recognition model; A first text rendering module is configured to render the first text data in the dialogue; The second text acquisition module is configured to acquire second text data associated with the voice data, wherein the second text data is acquired by recognizing the voice data using a second speech recognition model; as well as A second text rendering module is configured to replace the first text data with the second text data in the dialogue in response to determining that the second text data is different from the first text data.
13. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 11.
14. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 11.
15. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Voice recognition method and device
CN110610697A
Speech recognition method, device, equipment, medium and program product
CN114187913A
Voice processing method and system, computer readable storage medium and program product
CN114678029A
Speech recognition method and device, equipment and storage medium
CN115240685A
Speech recognition method, device and equipment and computer readable storage medium
CN116805490A