System for providing natural pronunciation by voice assistant and method thereof
By introducing technologies such as automatic speech recognition, natural language understanding, communication classification and natural pronunciation analysis in the voice assistant system, we can identify users' interruptions and context changes, and provide real-time and intuitive hybrid responses, solving the problem of unnatural response of existing voice assistants and realizing continuous and natural dialogue between users and voice assistants.
Patent Information
- Application Number
- CN202380056681.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-07-25
- Filing Date
- 2023-06-29
- Publication Date
- 2025-05-06
AI Technical Summary
The response provided by existing voice assistants during real-time conversations with users is usually mechanical and unnatural, unable to recognize user interruptions and context changes, resulting in the user being forced to re-initiate conversations and affecting the user experience.
The user's unsegmented voice input is converted into text format through the automatic speech recognition module. The natural language understanding module extracts the user's information and intentions. The communication classification unit classifies the user's input. The natural pronunciation analyzer unit calculates the context's natural pronunciation interval. The context sorter unit prioritizes and sorts the response. The intelligent activity engine provides real-time intuitive hybrid response based on response, context pause and user's activities.
It realizes continuous and natural dialogue between voice assistants and users, overcomes the mechanical response problems provided by existing voice assistants, and improves the naturalness of user experience and dialogue.
Smart Images

Figure CN119948559A_ABST
Abstract
Description
Technical Field
[0001] A system and method for providing natural pronunciation by a voice assistant are provided. The present disclosure particularly relates to a system and method for facilitating natural conversations between a user and a voice assistant, wherein the voice assistant recognizes interruptions by the user and provides one or more real-time intuitive mixed responses based on responses, contextual pauses, and ongoing activities of the user. Background Art
[0002] The use of voice assistants in our daily lives has grown exponentially over the past few years, where we rely on input from voice assistants for various activities such as time, weather, recipes during cooking, distance to a particular destination of interest, etc. Although the use of voice assistants has proven to be helpful in our daily affairs, existing voice assistants often provide mechanical and unnatural responses during real-time conversations with users. In some cases, users are unable to provide immediate responses to voice assistants due to factors such as motion-based events (running, climbing stairs, driving), long-term instruction-based activities (cooking, exercising), and the user's natural state (coughing, sneezing, shaking).
[0003] In this case, the voice assistant loses the context of the ongoing conversation, so the user is forced to re-initiate the conversation with the voice assistant, which inconveniences the user. In addition, existing voice assistants lack the intelligence to provide contextual responses even during interference or interruptions, which prevents users from establishing a continuous and natural conversation with the voice assistant.
[0004] The above shortcomings are illustrated by the following example. Consider a user requesting a voice assistant to play a song, where the user says "Can you play this song..." and coughs before completing the request. The voice assistant responds with a list of the top 5 songs, rather than waiting for the user to complete the request after coughing. This example shows that the voice assistant is unable to recognize the real-time context and acknowledge the user's interruption.
[0005] In another example, a user requested a voice assistant to calculate the distance between his current location and his office. However, because the user made this request while running, the user did not get a response from the voice assistant. In this case, the voice assistant was not smart enough to wait for the user to stop running before responding to his question, resulting in the user repeating his question multiple times.
[0006] Therefore, there is a need for a voice assistant that facilitates a continuous and natural conversation with a user despite user interruptions and other distractions. Summary of the invention
[0007] Solution to the problem
[0008] According to aspects of the present disclosure, a method includes the following steps: an automatic speech recognition module converts an unsegmented speech input received from a user into a text format, wherein a natural language understanding module extracts the user's information and intention from the converted text input. In addition, a communication classification unit classifies the data extracted from the user input into one or more predefined categories, wherein instructions are predefined categories of the communication classification unit, which refer to long-term communications between a user and a voice assistant and include a set of contextual responses. Similarly, requests and commands are predefined categories of the communication classification unit, which refer to instantaneous communications that include prompt responses from the voice assistant.
[0009] After classifying the user input, the natural pronunciation analyzer unit calculates the contextual natural pronunciation intervals during the real-time conversation between the user and the voice assistant. In addition, the context sorter unit prioritizes and sorts the responses to the static input and dynamic input provided by the user, wherein the contextual responses are obtained from the virtual server. Thus, the intelligent activity engine provides one or more real-time intuitive hybrid responses based on the responses, contextual pauses, and ongoing activities.
[0010] In addition, the present disclosure discloses a system for providing natural pronunciation by a voice assistant, wherein the system includes an automatic speech recognition module for converting one or more unsegmented voice inputs provided by the user after saying a wake-up word into a text format in real time. In addition, the natural language understanding module in the system extracts the user's information and intentions from the converted text input, wherein the natural language understanding module includes a communication classification unit for classifying the user input into one or more predefined categories. In addition, the system also includes a processing module for analyzing and processing inputs from the natural language understanding module and the activity recognition module.
[0011] Therefore, the present disclosure provides a system and method for making existing voice assistants more intelligent and capable of having a continuous and natural conversation with a user, thereby overcoming the mechanical and monotonous responses provided by existing voice assistants. The processing module provided in the system calculates context pauses and provides one or more real-time intuitive mixed responses based on the response, context pauses, and the user's ongoing activities. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The foregoing and other features of certain embodiments of the present disclosure will become more apparent when the following detailed description is read in conjunction with the accompanying drawings, in which like reference numerals refer to like elements, and in which:
[0013] Figure 1 A method for providing natural pronunciation by a voice assistant is shown;
[0014] Figure 2A system for providing natural pronunciation by a voice assistant is shown;
[0015] Figure 3 A block diagram representation of a natural language understanding module having a communication classification unit is shown;
[0016] Figure 4 shows a block diagram representation of an activity recognition module;
[0017] Figure 5 A block diagram representation of a processing module is shown;
[0018] Figure 6 shows a block diagram representation of a natural pronunciation analyzer unit;
[0019] Figure 7 A block diagram representation of a context sequencer unit is shown;
[0020] Figure 8 A block diagram representation of a smart activity engine is shown;
[0021] Fig. 9 A flow chart showing a first use case of the present disclosure;
[0022] Fig.10 A flow chart showing a second use case of the present disclosure;
[0023] Fig.11 A flowchart illustrating a third use case of the present disclosure; and
[0024] Fig.12 A flowchart showing a fourth use case of the present disclosure is shown. DETAILED DESCRIPTION
[0025] Reference will now be made in detail to the description of this subject matter, one or more examples of which are illustrated in the accompanying drawings. Each example is provided to illustrate the subject matter, and is not restrictive. Various changes and modifications that are obvious to those skilled in the art to which the present disclosure belongs are considered to be within the spirit, scope, and envision of the present disclosure.
[0026] Figure 1 A method for providing natural pronunciation by a voice assistant is shown, wherein the method 200 includes starting with the automatic speech recognition module 101 converting the unsegmented speech input received from the user into a text format in operation 201, wherein one or more queries from the user are received in the form of unsegmented speech input after the pronunciation of the wake-up word detected by the wake-up word engine 102. The natural language understanding module 103 extracts the user's information and intent from the converted text input.
[0027] In operation 202, the communication classification unit 104 classifies the extracted data related to the user's information and intention obtained from the user input into one or more predefined categories. In one embodiment, requests and commands are predefined categories of the communication classification unit 104, which refer to instantaneous communications that include prompt responses from the voice assistant. In addition, instructions are predefined categories of the communication classification unit 104, which refer to long-term communications between the user and the voice assistant and include a set of contextual responses.
[0028] In addition, after the communication classification unit 104 classifies the user data, the activity recognition module 105 tracks the user's activities and recognizes the user's environment in real time. Subsequently, in operation 203, the natural pronunciation analyzer unit 109 calculates the contextual natural pronunciation intervals during the real-time conversation between the user and the voice assistant. In addition, in operation 204, the context sorter unit 110 prioritizes and sorts the responses to the static input and dynamic input provided by the user, wherein the contextual responses are obtained from the virtual server 108.
[0029] After prioritizing and sorting responses to static and dynamic inputs provided by the user, the intelligent activity engine 111 identifies the user's active listening state and performs subsequent tasks based on the identified user's listening state and contextual natural pronunciation intervals. In operation 205, after identifying the user's active listening state and contextual natural pronunciation intervals, the method 200 provides one or more real-time intuitive mixed responses. In addition, artificial intelligence and deep neural networks are used to learn user behavior and identify the effectiveness of natural pronunciation.
[0030] Figure 2A system for providing natural pronunciation by a voice assistant is shown, wherein the system 100 includes an automatic speech recognition module 101 (e.g., an automatic speech recognition code executed by a processor, or a dedicated circuit designed to implement an automatic speech recognizer), which is used to convert one or more unsegmented voice inputs provided by a user after saying a wake-up word into a text format in real time, wherein after the wake-up word engine 102 (e.g., a detection code executed by a processor and configured to detect a wake-up word in a sound received by a microphone, or a dedicated circuit designed to implement a wake-up word analyzer) detects the pronunciation of the wake-up word, one or more queries from the user are received in the form of unsegmented voice input. In addition, the system 100 includes a natural language understanding module 103 (e.g., a natural language understanding code executed by a processor, or a dedicated circuit designed to implement a natural language understanding system), which is used to extract the user's information and intention from the converted text input, wherein the natural language understanding module 103 includes a communication classification unit 104 (e.g., a communication classification code executed by a processor, or a dedicated circuit designed to implement a communication classifier), which is used to classify the user input into one or more predefined categories such as requests, commands, and instructions.
[0031] In addition, the system 100 includes a processing module 106 (e.g., a processing code executed by a processor, or a dedicated circuit) for analyzing and processing inputs from the natural language understanding module 103 and the activity recognition module 105 (e.g., an activity recognition code executed by a processor, or a dedicated circuit designed to implement an activity recognizer), wherein the processing module 106 includes a response engine 107 (e.g., a response code executed by a processor, or a dedicated circuit designed to implement a response engine) for obtaining contextual responses to questions and requests provided by the user from the virtual server 108 based on the user's intent and entity. In addition, the processing module 106 includes a natural pronunciation analyzer unit 109 (e.g., a natural pronunciation analyzer code executed by a processor, or a dedicated circuit designed to implement a natural pronunciation analysis system), a context sorter unit 110 (e.g., a context sorter code executed by a processor, or a dedicated circuit designed to implement a context sorter), and an intelligent activity engine 111 (e.g., an intelligent activity code executed by a processor, or a dedicated circuit designed to implement an intelligent activity system), wherein the natural pronunciation analyzer unit 109 calculates the contextual natural pronunciation intervals during a real-time conversation between a user and a voice assistant. In addition, the context sorter unit 110 prioritizes and sorts responses to static inputs and dynamic inputs provided by the user, wherein the contextual responses are obtained from the virtual server 108. In addition, the intelligent activity engine 111 identifies the active listening state of the user and performs subsequent tasks based on the identified listening state of the user. Therefore, the text-to-speech conversion module 112 (e.g., a text-to-speech conversion code executed by a processor, or a dedicated circuit designed to implement a text-to-speech converter) converts the contextual text-based response determined by the processing module 106 into a speech format.
[0032] Figure 3A block diagram representation of a natural language understanding module with a communication classification unit is shown, wherein the natural language understanding module 103 including the communication classification unit 104 further includes a semantic parsing unit 113 (e.g., semantic parsing code executed by a processor, or a dedicated circuit designed to implement a semantic parser) for converting natural language in a text format into a machine-readable logical format required for further analysis and processing. In addition, the output of the semantic parsing unit 113 is provided to a sentiment analyzer unit 114 (e.g., sentiment analysis code executed by a processor, or a dedicated circuit designed to implement a sentiment analyzer) for accepting parsed data from the semantic parsing unit 113 and detecting positive, negative, and neutral sentiments present in the input provided by the user. In addition, an intent classification unit 115 (e.g., intent classification code executed by a processor, or a dedicated circuit designed to implement an intent classifier) and an entity recognition unit 116 (e.g., entity recognition code executed by a processor, or a dedicated circuit designed to implement an entity recognizer) are provided, wherein the intent classification unit 115 classifies the user intent based on the emotions detected by the emotion analyzer unit 114, and the entity recognition unit 116 determines one or more objects present in the input provided by the user.
[0033] In addition, the communication classification unit 104 in the natural language understanding module 103 includes a term tagger unit 117 (e.g., a term tagging code executed by a processor, or a dedicated circuit designed to implement a term tagger) for splitting the user input into one or more tokens, wherein each token represents a word in the input provided by the user. In addition, the communication classification unit 104 includes a natural language processing (NLP) model 118 (e.g., an NLP model code executed by a processor, or a dedicated circuit designed to implement an NLP model) for classifying the tokens into one or more predefined categories. In one embodiment, the natural language processing (NLP) model 118 can be, but is not limited to, a Bidirectional Encoder Representation from Transformer (BERT) variant. In addition, the communication classification unit 104 includes a relationship extraction unit 119 (e.g., a relationship extraction code executed by a processor, or a dedicated circuit designed to implement a relationship extractor) and a classification unit 120 (e.g., a classification code executed by a processor, or a dedicated circuit designed to implement a classifier), wherein the relationship extraction unit 119 determines the context and relationship between the classified tokens, and the classification unit 120 determines the type of conversation based on data provided at the output of the relationship extraction unit 119.
[0034] Figure 4A block diagram representation of an activity recognition module for tracking a user's activities and identifying the user's environment in real time is shown, wherein the activity recognition module 105 includes a data stream analyzer unit 121 (e.g., a data stream analysis code executed by a processor, or a dedicated circuit designed to implement a data stream analyzer) for organizing unorganized raw data such as audio, video, sensor data, etc. received from multiple input devices that are connected within a predefined proximity (e.g., connected to the processor). In addition, an object motion tracking unit 122 in the activity recognition module 105 (e.g., an object motion tracking code executed by a processor, or a dedicated circuit designed to implement an object motion tracker) is provided for extracting information related to a primary object present in the organized data received from the data stream analyzer unit 121, and determining an expected action of the identified object. In one embodiment, a convolutional neural network (CNN) may be used to extract information related to the primary object present in the organized data.
[0035] In addition, the activity recognition module 105 includes an environment classifier unit 123 (e.g., an environment classifier code executed by a processor, or a dedicated circuit designed to implement an environment classifier) for determining the environment of the vicinity of the user based on the organized data received from the data stream analyzer unit 121. In one embodiment, the environment classifier unit 123 can use K-means clustering to identify the environment near the user. In addition, the activity recognition module 105 includes an object location tracking unit 124 (e.g., an object location tracking code executed by a processor, or a dedicated circuit designed to implement an object location tracker) for receiving data related to the user's environment from the environment classifier unit 123 and determining the user's location. In one embodiment, the user's location can be determined using a triangulation technique that takes into account all peripheral devices present near the user, such as a microphone, a camera, a global positioning system (GPS), etc.
[0036] In addition, the activity recognition module 105 includes an activity classification unit 125 (e.g., an activity classification code executed by a processor, or a dedicated circuit designed to implement an activity classifier) for classifying the user activity data into activity type, activity state, and real-time environment based on the inputs received from the object tracking unit 122, the environment classifier unit 123, and the object position tracking unit 124. In one embodiment, the activity classification unit 125 can classify the user activity data using a K-nearest neighbor (K-NN) algorithm. In one example, if the inputs of the activity classification unit 125 are driving, outdoor, turning, speed bump, and GPS location, the activity classification unit 125 classifies the inputs as follows: activity type: motion; activity state: ongoing; environment: outdoor, turning, speed bump.
[0037] Figure 5 A block diagram representation of a processing module for analyzing and processing inputs from a natural language understanding module 103 and an activity recognition module 105 is shown, wherein the processing module 106 includes a response engine 107 for obtaining contextual responses to user-provided questions and requests from a virtual server 108 based on user intent and entities. In addition, the processing module 106 includes a natural pronunciation analyzer unit 109, a context sorter unit 110, and an intelligent activity engine 111, wherein the natural pronunciation analyzer unit 109 calculates contextual natural pronunciation intervals during a real-time conversation between a user and a voice assistant. In addition, the context sorter unit 110 prioritizes and sorts responses to static and dynamic inputs provided by the user, wherein the contextual responses are obtained from the virtual server 108. In addition, the intelligent activity engine 111 identifies the active listening state of the user and performs subsequent tasks based on the identified listening state of the user. Therefore, the text-to-speech conversion module 112 converts the contextual text-based responses determined by the processing module 106 into a speech format.
[0038] Figure 6 A block diagram representation of a natural pronunciation analyzer unit is shown, wherein the natural pronunciation analyzer unit 109 includes a context pause analyzer unit 126 (e.g., a context pause analyzer code executed by a processor, or a dedicated circuit designed to implement a context pause analyzer) for accepting inputs from the natural language understanding module 103, the activity recognition module 105, and the response engine 107, and using artificial intelligence and machine learning techniques to train the received data for future analysis to introduce context pauses based on the time required to complete a specific task requested by the user. In one embodiment, training the received data may include the following steps: token embedding, sentence embedding, and transformer position embedding. In addition, a modified BERT variant may be used to introduce context pauses based on the time required to complete a specific task requested by the user.
[0039] In one example, consider the following input from the natural language understanding module 103: intent: cooking; entity: noodles; communication type: instruction. In addition, consider the following input from the response engine 107: step 1: boil water for 5 minutes; step 2: add noodles to the boiling water; step 3: cook for 2 minutes; step 4: add seasoning and stir; step 5: serve noodles. The context pause analyzer unit 126 processes the input received from the natural language understanding module 103 and the response engine 107 and provides the following output: step 1: 5 minute pause; step 2: 0 minute pause; step 3: 2 minute pause; step 4: 1 minute pause; step 5: 0 minute pause.
[0040] In addition, the natural pronunciation analyzer unit 109 includes a real-time analyzer engine 127 (e.g., a real-time analysis code executed by a processor, or a dedicated circuit designed to implement a real-time analyzer) for accepting input from the activity recognition module 105 and calculating the contextual time delay required to complete real-time intervention during an ongoing conversation between the user and the voice assistant. In one embodiment, the real-time analyzer engine 127 can use CNN to calculate the contextual time delay. In one example, the input from the activity recognition module 105 to the real-time analyzer engine 127 is considered as follows: activity type: conversation; activity state: ongoing; environment: living room, news sound. In addition, consider that the activity recognition module 105 also captures the following sounds: ringtones / phone conversations / discussions. The real-time analysis engine 127 processes the input from the activity recognition module 105 and determines the output when the user receives a call during it, and introduces intelligent pauses accordingly based on the real-time analysis.
[0041] In addition, the natural pronunciation analyzer unit 109 includes a probability engine 128 (e.g., a probability calculation code executed by a processor, or a dedicated circuit designed to implement a probability engine) for accepting inputs from the context pause analyzer unit 126 and the real-time analyzer engine 127, and determining the probability of a time delay for a predefined event taking into account the probability of occurrence of similar events that have occurred, wherein the estimated time delay calculated by the probability engine 128 for the predefined event is based on the user's lifestyle, habits, and data related to the user's whereabouts. In one embodiment, the probability of a time delay for a predefined event can be estimated using Bayes' theorem.
[0042] Figure 7 A block diagram representation of a context sequencer unit is shown, wherein the context sequencer unit 110 determines whether the context state of a user input is static or dynamic by calculating the similarity between the user input and the appropriate response obtained from the response engine 107. In one embodiment, an N×N correlation matrix may be used to calculate the similarity between the user input and the appropriate response obtained from the response engine 107. The context sequencer unit 110 includes a pre-processing unit 129 (e.g., a pre-processing code executed by a processor, or a dedicated circuit designed to perform pre-processing), a priority calculation unit 130 (e.g., a priority calculation code executed by a processor, or a dedicated circuit designed to implement a priority calculator), and a command queuing unit 131 (e.g., a command queuing code executed by a processor, or a dedicated circuit designed to implement a command queue), wherein a plurality of tasks obtained from the response engine 107 are assigned to the context sequencer unit 110 in parallel, and the weight calculation and prioritization subsequent stages performed by the pre-processing unit 129 and the priority calculation unit 130, respectively, are performed simultaneously.
[0043] The pre-processing unit 129 is used to calculate the weights of the multiple tasks obtained from the response engine 107, wherein the weight of each task is estimated based on the static / dynamic context, the waiting time and the response length. When calculating the weights of the multiple tasks, the priority calculation unit 130 determines the priority of each task based on the weight calculated for the corresponding task by the pre-processing unit 129. Subsequently, a command queuing unit 131 is provided for sorting the tasks and the associated time delays of the tasks based on the priorities estimated by the priority calculation unit 130.
[0044] Figure 8 A block diagram representation of a smart activity engine is shown, wherein the smart activity engine 111 includes a smart reply unit 132 (e.g., a smart reply code executed by a processor, or a dedicated circuit designed to implement a smart reply system), which accepts inputs from the response engine 107, the context sequencer unit 110, and the probability engine 128, and is used to provide one or more real-time intuitive mixed responses based on the response, context pause, and the ongoing activity. In addition, the smart activity engine 111 includes a smart reply selection unit 133 (e.g., a smart reply selection code executed by a processor, or a dedicated circuit designed to implement a smart reply selector), such as but not limited to a long short-term memory (LSTM), for selecting the most appropriate and intelligent reply from multiple replies provided by the smart reply unit 132, wherein the reply selected by the smart reply selection unit 133 in the smart activity engine 111 is provided to the text-to-speech conversion module 112 through an actuator, which cascades the smart reply with the calculated context pause.
[0045] Fig. 9A flowchart of a first exemplary use case of the present disclosure is shown, which involves an unexpected natural pause during a conversation between a user and a voice assistant. Consider the following user input "Can you play a song?", in which the user coughs before completing the question. The natural language understanding module 103 provided in the system 100 identifies the event (i.e., cough) and the user's surrounding environment (i.e., home). The user's biological state includes, but is not limited to, coughing, hiccups, sneezes, yawns, etc. In addition, the communication classification unit 104 classifies the user's input as a command. Subsequently, the activity recognition module 105 analyzes the user's input and confirms the user's biological state, i.e., coughing. In addition, the activity recognition module 105 identifies the activity type as a conversation and identifies the activity state as "ongoing". In addition, based on the input from the natural language understanding module 103 and the activity recognition module 105, the context sorter unit 110 in the processing module 106 determines that the probability of natural pronunciation is high, and then sorts the tasks and the associated time delays of the tasks based on the estimated priority (i.e., high). In the first use case, since there is a single task, the context sorter unit 110 provides the following response to the user, "Which song?" and records the time at which the response is provided to the user. In addition, the voice assistant responds immediately to the user's input, so the waiting time is recorded as zero. Based on the priority and sorting of the tasks calculated by the context sorter unit 110, the intelligent activity engine 111 determines the response status, i.e., waiting is required because the user is coughing. In addition, since the user is coughing, the pause time is infinite, and subsequent input is received only after the user coughs. Therefore, based on the above estimation, the voice assistant responds: "It seems like you are coughing. I'm waiting for your reply." Therefore, it is obvious that the system 100 and the method 200 overcome the challenges of the prior art by providing a real-time intuitive hybrid response based on the response, the context pause, and the user's ongoing activity.
[0046] Fig.10A flowchart of a second exemplary use case of the present disclosure is shown, which involves user movement during a conversation between a user and a voice assistant. Consider the following user input "What is the distance between my location and the office", where the user asks the question while driving. The natural language understanding module 103 provided in the system 100 identifies the event (i.e., driving and the user's surroundings), which is identified by sharp turns or frequent braking of the user's vehicle. Motion-based events include, but are not limited to, running, driving, using stairs, taking elevators, etc. In addition, the communication classification unit 104 classifies the user's input as a command. Subsequently, the activity recognition module 105 analyzes the user's input and confirms the user's current driving condition. In addition, the activity recognition module 105 identifies the activity type as a strenuous exercise activity, and identifies the activity state as "ongoing". In addition, based on the input from the natural language understanding module 103 and the activity recognition module 105, the context sorter unit 110 in the processing module 106 determines that the probability of natural pronunciation is high, and then sorts the tasks and the associated time delays of the tasks based on the estimated priority (i.e., high). In this use case, since there is a single task, the context sorter unit 110 provides the following response to the user, "The distance between your current location and the office is 20Km", and records the time when the response is provided to the user. In addition, the voice assistant responds to the user's input immediately, so the waiting time is recorded as zero. Based on the priority and ranking of the tasks calculated by the context sorter unit 110, the intelligent activity engine 111 determines the response state, that is, waiting until the user concentrates and listens. Therefore, based on the above estimation, the voice assistant waits until the user concentrates and listens, instead of providing a response when the user cannot hear the voice assistant, which is counterproductive. Therefore, it is obvious that the system 100 and the method 200 overcome the challenges of current systems by providing real-time intuitive hybrid responses based on responses, context pauses, and the user's ongoing activities.
[0047] Fig.11A flowchart of a third exemplary use case of the present disclosure is shown, which relates to the ability of a voice assistant to maintain the context of a conversation between a user and a voice assistant during a long activity such as exercise, cooking, etc. (which is interrupted by input outside the context). Consider the following user input "how to make noodles", wherein the natural language understanding module 103 provided in the system 100 recognizes the event (i.e., cooking and the user's surroundings), which is recognized by background noise. In addition, the communication classification unit 104 classifies the user's input as an instruction. Subsequently, the activity recognition module 105 analyzes the user's input and confirms the user's current status. In addition, the activity recognition module 105 analyzes the instruction and calculates that a pause of 10 minutes is required. In addition, the activity recognition module 105 identifies the activity type as a long activity and identifies the activity state as "ongoing". In addition, based on the input from the natural language understanding module 103 and the activity recognition module 105, the context sorter unit 110 in the processing module 106 determines that the probability of natural pronunciation is high, and then sorts the tasks and the associated time delays of the tasks based on the estimated priority (i.e., high). In the third use case, since there are two tasks, the context sequencer unit 110 provides the following response to the user, "Boil water for 5 minutes", and records the time when the response is provided to the user. In addition, the voice assistant prompts that the waiting time is five minutes, and gives a second response after five minutes, namely, "Add noodles to boiling water". Consider an out-of-context question asked by the user in a long cooking activity, where there is no waiting time associated with the question. Based on the priority and sorting of the tasks calculated by the context sequencer unit 110, the intelligent activity engine 111 determines the response state, namely, waiting for 5 minutes between the task of boiling water for 5 minutes and the task of adding noodles to boiling water. Therefore, based on the above estimation, the voice assistant waits for 5 minutes after the first instruction to boil water for 5 minutes. Since the voice assistant does not have to provide the next step before the 5 minutes are completed, the voice assistant responds to the out-of-context question. After the 5 minutes are completed, a second instruction related to adding noodles to boiling water is provided. Therefore, it is obvious that the system 100 and the method 200 overcome the challenges of the current system by providing real-time intuitive hybrid responses based on responses, context pauses, and the user's ongoing activities.
[0048] Fig.12A flowchart of a fourth exemplary use case of the present disclosure is shown, which involves secondary activity interference during a conversation between a user and a voice assistant. Consider the following situation: the voice assistant is playing news for the user, and the user is listening to the news. In addition, consider that the user is interrupted by a call. The natural language understanding module 103 provided in the system 100 identifies the event (i.e., the call) and the user's surroundings (i.e., the office). In addition, the communication classification unit 104 classifies the user's input as an instruction. Subsequently, the activity recognition module 105 analyzes the instruction and the user's current call status. In addition, the activity recognition module 105 identifies the activity type as a conversation and identifies the activity status as "ongoing". In addition, based on the input from the natural language understanding module 103 and the activity recognition module 105, the context sorter unit 110 in the processing module 106 determines that the probability of natural pronunciation is high, and then sorts the tasks and the associated time delays of the tasks based on the estimated priority (i.e., high). In this use case, since there is a single task, the contextual sequencer unit 110 provides the following response to the user, "Finance Minister Nirmala Sitharaman said the budget stands for continuity, stability and predictability in taxation," and records the time when the response was provided to the user. Based on the priority and ranking of the tasks calculated by the contextual sequencer unit 110, the intelligent activity engine 111 determines the response state, i.e., wait until the user disconnects the call. Therefore, based on the above estimation, the voice assistant waits until the user disconnects the call instead of providing a response when the user answers the call. Therefore, it is apparent that the system 100 and method 200 overcome the challenges of current systems by providing real-time intuitive hybrid responses based on responses, contextual pauses, and the user's ongoing activities.
[0049] The present disclosure provides a system 100 and a method 200 for making existing voice assistants more intelligent and capable of having a continuous and natural conversation with a user, thereby overcoming the mechanical and monotonous responses provided by existing voice assistants. The processing module 106 provided in the system calculates context pauses and provides one or more real-time intuitive mixed responses based on the response, context pauses, and the user's ongoing activities.
[0050] At least one of the multiple modules described herein may be implemented by an AI model. Functions associated with AI may be performed by non-volatile memory, volatile memory, and a processor. The processor may include one or more processors. At this time, the one or more processors may be a general-purpose processor (e.g., a central processing unit (CPU), an application processor (AP), etc.), a graphics-specific processing unit (e.g., a graphics processing unit (GPU), a visual processing unit (VPU)), and / or an AI-specific processor (e.g., a neural processing unit (NPU)).
[0051] The one or more processors control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are provided by training or learning. Here, providing by learning means: by applying a learning algorithm to multiple learning data, a predefined operating rule or AI model with desired characteristics is formulated. The learning can be performed in the device itself that executes the AI according to the embodiment, and / or can be implemented by a separate server / system.
[0052] The AI model can be composed of multiple neural network layers. Each layer has multiple weight values, and the layer operation is performed by the calculation of the previous layer and the operation of multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recursive deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks. A learning algorithm is a method for training a predetermined target device (e.g., a robot) using multiple learning data to enable, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0053] The device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory storage medium" simply means that it is a tangible device and does not contain signals (e.g., electromagnetic waves). The term does not distinguish between a case where data is semi-permanently stored in a storage medium and a case where data is temporarily stored in a storage medium. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0054] According to an embodiment, the method according to various embodiments of the present disclosure may be included and provided in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)) or through an application store (e.g., Play Store TM ) online distribution (e.g., downloading or uploading), or directly between two user devices (e.g., smartphones). If distributed online, at least a part of the computer program product (e.g., a downloadable application) may be temporarily generated or at least temporarily stored in a machine-readable storage medium (e.g., a manufacturer's server, an application store's server, or a memory of a relay server).
[0055] While at least one exemplary embodiment has been presented in the foregoing detailed description, it should be appreciated that a vast number of variations exist.
Claims
1. A method for providing natural pronunciation by a voice assistant, the method comprising: converting at least one unsegmented speech input received from a user into a text format, extracting user information from the converted text format of the at least one unsegmented speech input; categorizing the extracted user information into one or more predefined categories; calculating at least one contextual natural pronunciation interval during a real-time conversation between the user and the voice assistant; prioritizing and ranking at least one contextual response to the at least one unsegmented speech input received from the user, wherein the at least one contextual response is obtained from a virtual server; as well as One or more real-time responses are provided based on at least one or more of: the at least one contextual response, at least one contextual pause, and ongoing ambient activity detected in proximity to the user.
2. The method according to claim 1, further comprising: Detect the pronunciation of the wake-up word, Wherein, the at least one unsegmented voice input received from the user is received after the wake-up word is detected.
3. The method according to claim 1, in, The one or more predefined categories include requests and commands, and Wherein, the request predefined category and the command predefined category each include an instant communication, and the instant communication includes at least one prompt response from the voice assistant.
4. The method according to claim 1, in, The one or more predefined categories include instructions that include lengthy communications between the user and the voice assistant and also include multiple contextual responses.
5. The method according to claim 1, further comprising: After the user data is classified, the user's activities are tracked in real time and the user's environment is identified in real time.
6. The method according to claim 1, further comprising: After prioritizing and sorting the at least one contextual response to the at least one unsegmented speech input received from the user, identifying an active listening state of the user and performing at least one task based on the identified listening state of the user and the at least one contextual natural pronunciation interval.
7. The method according to claim 1, wherein: Providing the one or more real-time responses also includes using a machine pre-trained deep neural network.
8. A system for providing natural pronunciation by a voice assistant, the system comprising: at least one memory configured to store computer program code; at least one processor in communication with the at least one memory and configured to execute the at least one instruction to: Recognize the pronunciation of the wake-up word, Based on the recognized pronunciation of the wake-up word, convert at least one unsegmented voice input provided by the user into a text format in real time, extracting user information from the converted text format of the at least one unsegmented speech input, classifying the extracted user information into one or more predefined categories, calculating at least one contextual natural pronunciation interval during a real-time conversation between the user and the voice assistant, prioritizing and sorting at least one contextual response to at least one unsegmented speech input provided by the user, wherein the at least one contextual response is obtained from a virtual server, An active listening state of the user is identified, and at least one task is performed based on the identified listening state of the user.
9. The system according to claim 8, in, The one or more predefined categories include requests and commands, and Wherein, the request predefined category and the command predefined category each include instant communication, and the instant communication triggers a prompt response from the voice assistant.
10. The system according to claim 8, wherein: The one or more predefined categories include instructions that include lengthy communications between the user and the voice assistant and also include multiple contextual responses.
11. The system according to claim 8, wherein: The at least one processor is further configured to execute the at least one instruction to: The at least one context response is obtained from the virtual server based on the extracted user information.
12. The system according to claim 8, wherein: The at least one processor is further configured to execute the at least one instruction to: The at least one contextual response is converted from a text format to a speech format.
13. The system according to claim 8, wherein: The at least one processor is further configured to execute the at least one instruction to: converting the at least one unsegmented speech input provided by the user into a text format in real time by detecting positive, negative and neutral emotions present in the at least one unsegmented speech input provided by the user, and The user information is extracted from the converted text format of the at least one unsegmented speech input by classifying the user's intent based on an emotion detected in the at least one unsegmented speech input provided by the user, and determining one or more objects present in the at least one unsegmented speech input provided by the user.
14. The system according to claim 8, wherein: The at least one processor is further configured to execute the at least one instruction to classify the extracted user information into one or more predefined categories by: splitting the converted text format of the at least one unsegmented speech input provided by the user into one or more tokens, wherein each token represents a word in the converted text format of the at least one unsegmented speech input provided by the user, classifying the tags into one or more predefined categories, determining the context and relationships between the classified tokens, and A type of conversation is determined based on the determined context and the determined relationships between the classified tokens.
15. The system according to claim 8, wherein: The at least one processor is further configured to execute the at least one instruction to: Receive and organize unorganized raw data from multiple input devices located within a predefined proximity, extracting information related to the user from the organized data, determining an intended action of the user based on the extracted information, determining an environment near the user based on the organized data, receiving data related to the user's environment and determining the user's location based on the received data related to the user's environment, and Data related to the user is classified into activity type, activity state, and real-time environment based on at least one of: information extracted from the organized data, a determination of an environment near the user, and a determined location of the user.